search

MiaoYin Studio cover platform is now online

visibility 40chat_bubble_outline 0favorite 6
Updated: 26/08/16Published: 26/08/16
MiaoYin Studio cover platform is now online

Inference Comparison

After Inference

Not provided

Model Introduction

This week, we added a new feature: the MiaoYin Studio cover table.

It can create online cover features for all existing models, and convert the model's tone in real time. Before that, you should understand the features and product pricing.

Image [1]-Miaoyin-RVC Sound Model Workshop

Step 1: Choose available timbres

The "Available Voice Library" only displays the models your current account is authorized to use, including:

  • A free model available to all users
  • Models you purchased separately
  • A model available to regular alchemist members
  • Models available to permanent alchemist members
  • A model that acquires usage rights through other authorization methods

For example, even if you are not a member, as long as you have purchased a certain sound individually, that tone will still appear in your voice library.

After logging in, both regular and permanent alchemists will automatically display the available models based on their respective member benefits. You don't need to manually verify permissions, nor will you see models that your current account is not authorized to use.

After selecting a tone, you can continue uploading the audio and set inference parameters.

"Free," "Purchased," or "Membership-Available" means you have platform access to the model, but does not mean you have simultaneously obtained song copyright, sound authorization, or commercial licenses.

Step 2: Upload audio

Click the upload area on the timeline and select the audio file you want to convert.

For better cover effects, it is recommended to upload:

  • Audio with clear vocals and minimal noise
  • Stable volume with no severe clipping
  • Dry sound with less reverberation
  • WAV, FLAC, or high-bitrate MP3 files
  • Audio not exceeding 6 minutes in length and no larger than 50 MB

After uploading, the original audio will appear on the first timeline. You can listen to it, drag playback progress, or re-upload other audio files.

When the same audio is used again within the validity period, the system will try to reuse the uploaded files to avoid duplicate uploads and waiting.

Step 3: Set inference parameters

MiaoYin Studio retains the inference parameters commonly used by RVC. If this is your first time using it, you can keep the default value directly.

The two parameters you need to understand most are as follows.

Tone sandhi (half tone).

Pitch sandhi is used to adjust the pitch difference between the input vocals and the target timbre.

  • Left adjustment: Lower the pitch
  • Adjust to the right: Raise the pitch
  • 0: Maintains the original pitch

If the original sound is low but the target tone is high, the pitch change can be appropriately increased; If the original sound is higher and the target tone is lower, you can lower the pitch change appropriately.

Common reference values:

  • Convert male voice to higher-pitched female voice: You can start from +6 to +12
  • Convert female voice to a lower male voice: start from -6 to -12
  • Sounds with similar ranges: it is recommended to first use 0

These values are not fixed rules. The optimal settings may vary depending on the song and timbre model. It is recommended to generate the default values first and then fine-tune based on the listening results.

Volume envelope

Volume envelope controls how much the conversion results follow the volume changes of the original audio or model output.

It's not just a simple "volume control":

  • Lower value: preserves more variation in the original audio's intensity
  • Higher values: More uses the volume envelope of the converted audio
  • Default: Suitable for most audio

If the volume fluctuations in the generated results are unnatural, you can slightly adjust this setting. Don't adjust too many at once, or you may experience fluctuating volume or lost details.

Index strength

Index strength determines the extent to which model feature indexes are used during reasoning.

Higher values usually bring results closer to the target timbre, but setting them too high may also increase noise, articulation, or unnatural timbre characteristics.

When using it for the first time, it is recommended to keep the default values.

Consonant protection

Consonant protection is used to reduce damage to consonants, breath sounds, and voiceless sounds during conversion.

If the generated results show blurred enunciation, loss of consonants, or abnormal sibilance, adjustments can be made accordingly. It is recommended to keep the default value when using it for the first time.

Voices separated

If you upload a complete song that already mixes vocals and accompaniment, you can enable "vocal separation."

Once enabled, the system will first separate the audio into vocals and accompaniment, converting only the separated vocals, and finally remixing the converted vocals and accompaniment.

The timeline will display the following five tracks:

  1. Original audio
  2. Accompaniment
  3. Separated by human voices
  4. Vocals after conversion
  5. Conversion results

All five tracks can be previewed and downloaded separately, making it convenient for post-processing in other audio software.

Accompaniment volume

Control the volume of the accompaniment in the final mix.

If the accompaniment drowns out the vocals, it can be lowered appropriately; If the accompaniment is too small, it can be appropriately increased.

Vocal volume

Controls the volume of vocals after conversion during the final mix.

If vocals are not prominent enough, they can be appropriately increased; If there is a pop or vocals that are too strong, it is recommended to lower it.

Voices echo

Once enabled, the system adds a sense of space to the transformed vocals, making the fusion of vocals and accompaniment more natural.

If the original audio already has obvious reverberation, it is recommended to enable it cautiously to avoid excessive reverberation in the final result.

Notes on vocal separation

Currently, vocal separation is mainly used to extract vocals and accompaniment, and cannot guarantee the complete removal of reverb, harmony, background vocals, or other complex sounds from the original audio.

If the original song has obvious reverberation, the separated vocals may still contain some spatial remnants. For better results, it is recommended to use professional audio separation or de-reverberation software in advance, and upload only after obtaining high-quality dry sound. At this point, you can disable MiaoYin Studio's vocal separation function to directly convert the processed vocals.

Vocal separation requires additional computational resources, so the generation cost is higher than ordinary conversion. Before activation, the estimated cost will be automatically displayed on the right panel.

It should be noted that no vocal separation technique can guarantee complete losslessness. The final result will be affected by the song type, reverberation level, overlap between vocals and accompaniment, and the quality of the original audio.

Step 4: Check the price and start generating

it

After selecting the timbre, uploading audio, and setting parameters, the estimated cost for this task will be displayed on the right panel.

Fees will be deducted from the bonus duration first; If the giveaway period is insufficient, use your account balance according to the current billing rules. Please confirm that your gift limit or account balance is sufficient before submitting.

After confirming everything is correct, click "Start generating cover."

If the selected model is used on a GPU server for the first time, the system will automatically complete the following steps:

  1. Prepare and deploy the timbre model
  2. After deployment, the cover task is automatically submitted
  3. Waiting for the quest to be queued
  4. Perform vocal separation (if enabled).
  5. Perform tone conversion
  6. Mixing vocals and accompaniment
  7. Generate the final result

No repeated button clicks are needed during model preparation. After deployment is complete, the system will automatically continue the generation operation you just submitted.

Deployed models are stored in the server cache. The next time you use the same model, you usually don't need to wait for another deployment.

Step 5: Listen and download

Once generation, you can listen directly to different tracks on the timeline.

When vocal separation is enabled, the original audio, accompaniment, separated vocals, converted vocals, and conversion results all have separate download buttons.

If vocal separation is not enabled, you can download the original audio and the final conversion result.

The system also creates a history of inference. You can find it on the right side under "Reasoning History":

  • Listen to the historical generation results
  • Download historical generated results
  • Check the timbres used and generation times
  • Delete unnecessary historical records

File retention and privacy

We value the privacy of every user.

User-uploaded audio, intermediate processing files, and final generated results are retained by default for 7 days. During the retention period, you can listen to or download at any time.

After the retention period expires, the system will automatically clean up the relevant files. Once a file is deleted, it cannot be recovered, so please promptly download the result that needs long-term preservation.

You can also manually delete the reasoning history. After deletion, related history and corresponding audio files on the server will be simultaneously cleared and cannot be recovered.

The platform will not disclose your audio traceability information on the page, nor will it display your uploaded content to other users.

Generation failures and refunds

MiaoYin Studio is currently in public beta. During model deployment, task queuing, audio upload, or inference, factors such as network status, GPU load, and audio format may be affected.

If the task is not successfully completed:

  • Tasks that are not actually deducted will not incur any fees
  • The
  • deducted balance will be automatically refunded
  • The
  • deducted bonus time will be automatically released
  • Failed tasks are not charged as successful ones

If you encounter temporary network errors, please wait a moment or refresh the page to check the task status. It is not recommended to repeatedly submit the same task.

Tips for better results

  • Try to use clear, clean vocals with minimal reverberation.
  • When the original range differs significantly from the target timbre, adjust the key change appropriately.
  • When using a model for the first time, keep the default parameters first.
  • Do not adjust multiple parameters at once, otherwise it will be difficult to determine which parameter affects the results.
  • When using a full song, vocal separation can be enabled; If there is already a dry sound, it is recommended to turn it off.
  • After generation, listen separately to "separated vocals," "converted vocals," and "conversion results," allowing faster identification of issues caused by separation, conversion, or final mixing.
  • Please download important works within 7 days and back them up yourself.

AI cover price list

User type Normal conversion Activate vocal separation
Regular users ¥0.12/minute ¥0.25/minute
An ordinary alchemist ¥0.08/minute ¥0.16/minute
Permanent Alchemist ¥0.05/minute ¥0.10/minute

Once vocal separation is enabled, you will be charged at the unit price of "vocal separation," and there will be no additional fees added to the regular conversion price.

Giving away time

User type Bonus credit limit Validity period
Newly registered users 3 minutes One-time gift
An ordinary alchemist 50 minutes per month Valid for the current month and reissued the following month
Permanent Alchemist The first 100 minutes are included Long-term effectiveness
Permanent Alchemist After that, 10 minutes per month Valid for the current month and reissued the following month
The

bonus time will be prioritized. The bonus will only be deducted from your account balance after the bonus period is used up.

Common duration price references

The following prices do not include bonus credit deductions.

Normal conversion

Audio duration Regular users An ordinary alchemist Permanent Alchemist
1 minute ¥0.12 ¥0.10* ¥0.10*
3 minutes ¥0.36 ¥0.24 ¥0.15
5 minutes ¥0.60 ¥0.40 ¥0.25
6 minutes ¥0.72 ¥0.48 ¥0.30

Activate vocal separation

Audio duration Regular users An ordinary alchemist Permanent Alchemist
1 minute ¥0.25 ¥0.16 ¥0.10
3 minutes ¥0.75 ¥0.48 ¥0.30
5 minutes ¥1.25 ¥0.80 ¥0.50
6 minutes ¥1.50 ¥0.96 ¥0.60

* For a single balance deduction, the minimum deduction amount is ¥0.10.

Billing rules

  • Billing is based on the actual audio duration, with one billing unit every 10 seconds.
  • Fractions less than 10 seconds are counted as 10 seconds.
  • The duration of the gift is prioritized over the account balance.
  • If the reward duration is sufficient to cover this task, your account balance will not be deducted.
  • If the bonus time only covers part of the audio, the remaining time will be charged according to the unit price corresponding to your current account.
  • For a single balance deduction, the minimum deduction amount is ¥0.10.
  • After enabling vocal separation, the entire audio segment is calculated based on the unit price of vocal separation.
  • The
  • estimated cost of the task will be displayed above the submit button. After confirming, you can submit.

For example, a 3-minute 22-second audio segment is counted as 3 minutes and 30 seconds. If the account still has 2 minutes of bonus time, the system will use those 2 minutes first, and the remaining 1 minute 30 seconds will be charged according to the current user level's unit price.

Refund if you fail

After submitting a task, the system will temporarily lock the required gift duration and account balance.

If the task is successfully completed, the system will confirm the charge; If upload, model deployment, voice separation, or inference fail, the system will automatically:

  • Refund the withheld account balance
  • Restores the bonus duration used this
  • time
  • Mark the transaction record as refunded

Failed tasks will not be charged as successful ones.

Scope of use

The AI cover feature only provides technical services such as audio processing, voice conversion, and vocal separation. Before using this feature, please ensure that you have legal rights to use the uploaded audio, songs, accompaniment, vocals, and selected tone models.

Applicable range

  • Handling original or legally authorized songs, accompaniment, and recordings by myself.
  • Used for personal study, technical testing, private auditions, and non-commercial creation.
  • After obtaining full authorization, it is used for commercial scenarios such as video scores, gaming, advertising, live streaming, and music distribution.
  • Use the corresponding model according to the authorization scope marked on the timbre model page.

Tone model authorization instructions

"Free", "Purchased", "Available for Ordinary Alchemists", or "Available for Permanent Alchemists" only means the current account has the right to call the model for reasoning, and does not automatically grant the following rights:

  • Copyright of song lyrics and composition
  • Copyrights for accompaniment and recording products
  • The
  • voice and personality rights of the original singer and dubbing artist
  • Model characters, names, images, and related intellectual property
  • Commercial distribution, advertising endorsements, or public communication authorization

Purchasing models, signing up for memberships, or paying generation fees are purchasing model usage qualifications and computing services, which do not equate to obtaining ownership of model files, copyright, or commercial authorization.

Prohibited Scope

Do not use this function:

  • Unauthorized imitation, transformation, or public dissemination of a voice identifiable to a natural person.
  • Creating false performances, false endorsements, fraud, impersonation, or content that may mislead the public.
  • Uploading or handling pirated songs, unauthorized accompaniment, dry voices, recordings, or other infringing content.
  • Creating illegal, pornographic, violent, insulting, defamatory, harassing, or violating others' privacy.
  • Extracting, reselling, sharing, cracking, or otherwise obtaining the platform's model files and download addresses.
  • Bypassing membership, billing, access control, signature verification, or other security restrictions.
  • Describing the generated results as actual performances, official works, or authorized works by the original singer, voice actor, or related rights holders.

 

 

 

Carrier

Carrier

inventory_20person_add0

Version Details

OriginalNo

Recommended Parameters

Threshold0.25
Pitch0
Index Rate0.75
Volume Factor1.25
Sample Length192
Harvest Processes2
Fade Length100
Extra Inference Time500
infoFor reference only, adjust according to your input!

Related Models

No related models

Comments

No comments yet