This week, we added a new feature: the MiaoYin Studio cover table.
It can create online cover features for all existing models, and convert the model's tone in real time. Before that, you should understand the features and product pricing.
![Image [1]-Miaoyin-RVC Sound Model Workshop](https://apis.klrvc.com/wp-content/uploads/2026/08/da54f3143820260816113302-1024x498.png)
Step 1: Choose available timbres
The "Available Voice Library" only displays the models your current account is authorized to use, including:
- A free model available to all users
- Models you purchased separately
- A model available to regular alchemist members
- Models available to permanent alchemist members
- A model that acquires usage rights through other authorization methods
For example, even if you are not a member, as long as you have purchased a certain sound individually, that tone will still appear in your voice library.
After logging in, both regular and permanent alchemists will automatically display the available models based on their respective member benefits. You don't need to manually verify permissions, nor will you see models that your current account is not authorized to use.
After selecting a tone, you can continue uploading the audio and set inference parameters.
"Free," "Purchased," or "Membership-Available" means you have platform access to the model, but does not mean you have simultaneously obtained song copyright, sound authorization, or commercial licenses.
Step 2: Upload audio
Click the upload area on the timeline and select the audio file you want to convert.
For better cover effects, it is recommended to upload:
- Audio with clear vocals and minimal noise
- Stable volume with no severe clipping
- Dry sound with less reverberation
- WAV, FLAC, or high-bitrate MP3 files
- Audio not exceeding 6 minutes in length and no larger than 50 MB
After uploading, the original audio will appear on the first timeline. You can listen to it, drag playback progress, or re-upload other audio files.
When the same audio is used again within the validity period, the system will try to reuse the uploaded files to avoid duplicate uploads and waiting.
Step 3: Set inference parameters
MiaoYin Studio retains the inference parameters commonly used by RVC. If this is your first time using it, you can keep the default value directly.
The two parameters you need to understand most are as follows.
Tone sandhi (half tone).
Pitch sandhi is used to adjust the pitch difference between the input vocals and the target timbre.
- Left adjustment: Lower the pitch
- Adjust to the right: Raise the pitch
0: Maintains the original pitch
If the original sound is low but the target tone is high, the pitch change can be appropriately increased; If the original sound is higher and the target tone is lower, you can lower the pitch change appropriately.
Common reference values:
- Convert male voice to higher-pitched female voice: You can start from
+6 to +12
- Convert female voice to a lower male voice: start from
-6 to -12
- Sounds with similar ranges: it is recommended to first use
0
These values are not fixed rules. The optimal settings may vary depending on the song and timbre model. It is recommended to generate the default values first and then fine-tune based on the listening results.
Volume envelope
Volume envelope controls how much the conversion results follow the volume changes of the original audio or model output.
It's not just a simple "volume control":
- Lower value: preserves more variation in the original audio's intensity
- Higher values: More uses the volume envelope of the converted audio
- Default: Suitable for most audio
If the volume fluctuations in the generated results are unnatural, you can slightly adjust this setting. Don't adjust too many at once, or you may experience fluctuating volume or lost details.
Index strength
Index strength determines the extent to which model feature indexes are used during reasoning.
Higher values usually bring results closer to the target timbre, but setting them too high may also increase noise, articulation, or unnatural timbre characteristics.
When using it for the first time, it is recommended to keep the default values.
Consonant protection
Consonant protection is used to reduce damage to consonants, breath sounds, and voiceless sounds during conversion.
If the generated results show blurred enunciation, loss of consonants, or abnormal sibilance, adjustments can be made accordingly. It is recommended to keep the default value when using it for the first time.
Voices separated
If you upload a complete song that already mixes vocals and accompaniment, you can enable "vocal separation."
Once enabled, the system will first separate the audio into vocals and accompaniment, converting only the separated vocals, and finally remixing the converted vocals and accompaniment.
The timeline will display the following five tracks:
- Original audio
- Accompaniment
- Separated by human voices
- Vocals after conversion
- Conversion results
All five tracks can be previewed and downloaded separately, making it convenient for post-processing in other audio software.
Accompaniment volume
Control the volume of the accompaniment in the final mix.
If the accompaniment drowns out the vocals, it can be lowered appropriately; If the accompaniment is too small, it can be appropriately increased.
Vocal volume
Controls the volume of vocals after conversion during the final mix.
If vocals are not prominent enough, they can be appropriately increased; If there is a pop or vocals that are too strong, it is recommended to lower it.
Voices echo
Once enabled, the system adds a sense of space to the transformed vocals, making the fusion of vocals and accompaniment more natural.
If the original audio already has obvious reverberation, it is recommended to enable it cautiously to avoid excessive reverberation in the final result.
Notes on vocal separation
Currently, vocal separation is mainly used to extract vocals and accompaniment, and cannot guarantee the complete removal of reverb, harmony, background vocals, or other complex sounds from the original audio.
If the original song has obvious reverberation, the separated vocals may still contain some spatial remnants. For better results, it is recommended to use professional audio separation or de-reverberation software in advance, and upload only after obtaining high-quality dry sound. At this point, you can disable MiaoYin Studio's vocal separation function to directly convert the processed vocals.
Vocal separation requires additional computational resources, so the generation cost is higher than ordinary conversion. Before activation, the estimated cost will be automatically displayed on the right panel.
It should be noted that no vocal separation technique can guarantee complete losslessness. The final result will be affected by the song type, reverberation level, overlap between vocals and accompaniment, and the quality of the original audio.
Step 4: Check the price and start generating
it
After selecting the timbre, uploading audio, and setting parameters, the estimated cost for this task will be displayed on the right panel.
Fees will be deducted from the bonus duration first; If the giveaway period is insufficient, use your account balance according to the current billing rules. Please confirm that your gift limit or account balance is sufficient before submitting.
After confirming everything is correct, click "Start generating cover."
If the selected model is used on a GPU server for the first time, the system will automatically complete the following steps:
- Prepare and deploy the timbre model
- After deployment, the cover task is automatically submitted
- Waiting for the quest to be queued
- Perform vocal separation (if enabled).
- Perform tone conversion
- Mixing vocals and accompaniment
- Generate the final result
No repeated button clicks are needed during model preparation. After deployment is complete, the system will automatically continue the generation operation you just submitted.
Deployed models are stored in the server cache. The next time you use the same model, you usually don't need to wait for another deployment.
Step 5: Listen and download
Once generation, you can listen directly to different tracks on the timeline.
When vocal separation is enabled, the original audio, accompaniment, separated vocals, converted vocals, and conversion results all have separate download buttons.
If vocal separation is not enabled, you can download the original audio and the final conversion result.
The system also creates a history of inference. You can find it on the right side under "Reasoning History":
- Listen to the historical generation results
- Download historical generated results
- Check the timbres used and generation times
- Delete unnecessary historical records
File retention and privacy
We value the privacy of every user.
User-uploaded audio, intermediate processing files, and final generated results are retained by default for 7 days. During the retention period, you can listen to or download at any time.
After the retention period expires, the system will automatically clean up the relevant files. Once a file is deleted, it cannot be recovered, so please promptly download the result that needs long-term preservation.
You can also manually delete the reasoning history. After deletion, related history and corresponding audio files on the server will be simultaneously cleared and cannot be recovered.
The platform will not disclose your audio traceability information on the page, nor will it display your uploaded content to other users.
Generation failures and refunds
MiaoYin Studio is currently in public beta. During model deployment, task queuing, audio upload, or inference, factors such as network status, GPU load, and audio format may be affected.
If the task is not successfully completed:
- Tasks that are not actually deducted will not incur any fees
The - deducted balance will be automatically refunded
The - deducted bonus time will be automatically released
- Failed tasks are not charged as successful ones
If you encounter temporary network errors, please wait a moment or refresh the page to check the task status. It is not recommended to repeatedly submit the same task.
Tips for better results
- Try to use clear, clean vocals with minimal reverberation.
- When the original range differs significantly from the target timbre, adjust the key change appropriately.
- When using a model for the first time, keep the default parameters first.
- Do not adjust multiple parameters at once, otherwise it will be difficult to determine which parameter affects the results.
- When using a full song, vocal separation can be enabled; If there is already a dry sound, it is recommended to turn it off.
- After generation, listen separately to "separated vocals," "converted vocals," and "conversion results," allowing faster identification of issues caused by separation, conversion, or final mixing.
- Please download important works within 7 days and back them up yourself.
AI cover price list
| User type |
Normal conversion |
Activate vocal separation |
| Regular users |
¥0.12/minute |
¥0.25/minute |
| An ordinary alchemist |
¥0.08/minute |
¥0.16/minute |
| Permanent Alchemist |
¥0.05/minute |
¥0.10/minute |
Once vocal separation is enabled, you will be charged at the unit price of "vocal separation," and there will be no additional fees added to the regular conversion price.
Giving away time
| User type |
Bonus credit limit |
Validity period |
| Newly registered users |
3 minutes |
One-time gift |
| An ordinary alchemist |
50 minutes per month |
Valid for the current month and reissued the following month |
| Permanent Alchemist |
The first 100 minutes are included |
Long-term effectiveness |
| Permanent Alchemist |
After that, 10 minutes per month |
Valid for the current month and reissued the following month |
The
bonus time will be prioritized. The bonus will only be deducted from your account balance after the bonus period is used up.
Common duration price references
The following prices do not include bonus credit deductions.
Normal conversion
| Audio duration |
Regular users |
An ordinary alchemist |
Permanent Alchemist |
| 1 minute |
¥0.12 |
¥0.10* |
¥0.10* |
| 3 minutes |
¥0.36 |
¥0.24 |
¥0.15 |
| 5 minutes |
¥0.60 |
¥0.40 |
¥0.25 |
| 6 minutes |
¥0.72 |
¥0.48 |
¥0.30 |
Activate vocal separation
| Audio duration |
Regular users |
An ordinary alchemist |
Permanent Alchemist |
| 1 minute |
¥0.25 |
¥0.16 |
¥0.10 |
| 3 minutes |
¥0.75 |
¥0.48 |
¥0.30 |
| 5 minutes |
¥1.25 |
¥0.80 |
¥0.50 |
| 6 minutes |
¥1.50 |
¥0.96 |
¥0.60 |
* For a single balance deduction, the minimum deduction amount is ¥0.10.
Billing rules
- Billing is based on the actual audio duration, with one billing unit every 10 seconds.
- Fractions less than 10 seconds are counted as 10 seconds.
- The duration of the gift is prioritized over the account balance.
- If the reward duration is sufficient to cover this task, your account balance will not be deducted.
- If the bonus time only covers part of the audio, the remaining time will be charged according to the unit price corresponding to your current account.
- For a single balance deduction, the minimum deduction amount is ¥0.10.
- After enabling vocal separation, the entire audio segment is calculated based on the unit price of vocal separation.
The - estimated cost of the task will be displayed above the submit button. After confirming, you can submit.
For example, a 3-minute 22-second audio segment is counted as 3 minutes and 30 seconds. If the account still has 2 minutes of bonus time, the system will use those 2 minutes first, and the remaining 1 minute 30 seconds will be charged according to the current user level's unit price.
Refund if you fail
After submitting a task, the system will temporarily lock the required gift duration and account balance.
If the task is successfully completed, the system will confirm the charge; If upload, model deployment, voice separation, or inference fail, the system will automatically:
- Refund the withheld account balance
- Restores the bonus duration used this
time
- Mark the transaction record as refunded
Failed tasks will not be charged as successful ones.
Scope of use
The AI cover feature only provides technical services such as audio processing, voice conversion, and vocal separation. Before using this feature, please ensure that you have legal rights to use the uploaded audio, songs, accompaniment, vocals, and selected tone models.
Applicable range
- Handling original or legally authorized songs, accompaniment, and recordings by myself.
- Used for personal study, technical testing, private auditions, and non-commercial creation.
- After obtaining full authorization, it is used for commercial scenarios such as video scores, gaming, advertising, live streaming, and music distribution.
- Use the corresponding model according to the authorization scope marked on the timbre model page.
Tone model authorization instructions
"Free", "Purchased", "Available for Ordinary Alchemists", or "Available for Permanent Alchemists" only means the current account has the right to call the model for reasoning, and does not automatically grant the following rights:
- Copyright of song lyrics and composition
- Copyrights for accompaniment and recording products
The - voice and personality rights of the original singer and dubbing artist
- Model characters, names, images, and related intellectual property
- Commercial distribution, advertising endorsements, or public communication authorization
Purchasing models, signing up for memberships, or paying generation fees are purchasing model usage qualifications and computing services, which do not equate to obtaining ownership of model files, copyright, or commercial authorization.
Prohibited Scope
Do not use this function:
- Unauthorized imitation, transformation, or public dissemination of a voice identifiable to a natural person.
- Creating false performances, false endorsements, fraud, impersonation, or content that may mislead the public.
- Uploading or handling pirated songs, unauthorized accompaniment, dry voices, recordings, or other infringing content.
- Creating illegal, pornographic, violent, insulting, defamatory, harassing, or violating others' privacy.
- Extracting, reselling, sharing, cracking, or otherwise obtaining the platform's model files and download addresses.
- Bypassing membership, billing, access control, signature verification, or other security restrictions.
- Describing the generated results as actual performances, official works, or authorized works by the original singer, voice actor, or related rights holders.