search

[Original] RVC Fine-tuning Bottom Mold - Yue f048K

visibility 1.6kchat_bubble_outline 0favorite 13
Updated: 25/04/30Published: 25/04/02
[Original] RVC Fine-tuning Bottom Mold - Yue f048K

Inference Comparison

After Inference

0:000:00

Model Introduction

We are pleased to announce today that we are releasing a new base model fine-tuned by Miaoyin, which is fine-tuned based on the RVC2 base model. Below is an introduction to some of the language we trained.

Added: Japanese ≈ 10 hours of high-quality voice including multi-person speaking datasets and some anime datasets.

Added: 15 hours of high-quality studio sound quality in Chinese ≈.

Added: 2-hour Chinese song dataset.

Currently, in version 1.0, we have temporarily added these two voice lines. As for whether the model trained at the bottom of the newly fine-tuned version can reach a good level, we are still testing, but we have already achieved successful inference training using only a small portion of the data. This is a very good start.

Because the fine-tuning model data was quite large, this was especially tight on our GPU, so we only stepped 30,000 steps.

471863056420250402165041

If you don't understand this curve, you can check out a previous article in the user help section for detailed explanations. Our loss values have always fluctuated—at least from the G model's loss values, I can't judge them with minimal training. But after testing these inference models, their performance is quite good, since the base model is excellent in RVC_v2.

Next, I will use several audio clips (the bottom model inferes directly, since the original dataset has not been trained). )

Original audio

<!--[if lt IE 9]>document.createElement('audio');<![endif]-->

Base model reasoning

Without adding original audio data to the bottom model, the inference effect of the bottom model can reach 6-7 points of similarity.

Next, the second paragraph

Original audio

Base model reasoning

If you listen carefully, although the original audio is not in the base model, the output timbre is very similar. In subsequent inference model training, you only need to use our base model + speaker (possibly a small portion of speech) + emotional statements to obtain a good reasoning model.

Of course, we will continue to improve this model in the future, and permanent alchemists should prioritize downloading and enjoying it.

Below is the usage tutorial:

Placing the model into the native RVC project:

assets\pretrained_v2

Directly extract the two pth files starting with D and G to the directory of this folder. Do not change the usage names.

You need to pay attention to these parameters during training.

Image [2]-Miaoyin-RVC Tone Model Workshop

The target sampling rate + treble + version must be consistent, since our fine-tuning model is fine-tuned under f048K, so it must be consistent here.

Image[3]-Miaoyin-RVC Tone Model Workshop

After modifying the model path, if you select the correct target sampling rate, you only need to

G Model Path:

assets/pretrained_v2/G_mygfyue48K.pth

D Model Path:

assets/pretrained_v2/D_mygfyue48K.pth

Before training, pay attention to the path of the training model.

Since we haven't compared it with RVC_V2 base molds, we'll do more in later spare time.

D_mygfyue48K(model_hash):b3640c5ac8fdc81b3934be1173d4f2a6ec2d815742f477632f9a727a442a7034

G_mygfyue48K(model_hash):087e08d6e3992de28deba1cd87c7a56eac5b257b8b72d97c8fc83c94f2a6e99b

Subsequent plans include further fine-tuning of other voice lines and tweaking the singing parts.

 

Miaoyin model

Miaoyin model

inventory_20person_add0

Version Details

OriginalNo

Recommended Parameters

Threshold0.25
Pitch0
Index Rate0.75
Volume Factor1.25
Sample Length192
Harvest Processes2
Fade Length100
Extra Inference Time500
infoFor reference only, adjust according to your input!

Related Models

No related models

Keywords

#rvc#Download#Free of charge#Miaoyin Workshop#Undermembrane#Bottom model#Fine-tuning#Model

Comments

No comments yet