Ai Qiuqiu 24000 steps Wenyu RVC model
Inference Comparison
After Inference
Other Inference
Model Introduction
This was an exciting model training session, and at this moment, we were using the latest fine-tuned pre-trained model.
The model dataset is about 15 minutes long and includes raw dialogue and tone performance, with a total of 24,000 steps

(partial training loss function curve).
In reasoning tests, enunciation is better than the default low-mode model, and the optimization of songs is also relatively good.
However, the only thing to note about the model is that it is not ideal for breathing, possibly due to insufficient loudness during processing of breathing in the original data or loss of breathing during separation. For the second training, I tried adding corresponding breathing sounds to slightly relieve it, but if I needed to use the sound card to suppress the original breathing sound or try to keep it as low as possible, I wanted to use it.
In terms of tone testing, for example, "laughing softly while talking" is fine, but if you laugh too loudly, electric currents can appear, which is a current drawback of this model.
Of course, everyone's device and input may have different effects.
zhangzhang
Version Details
Recommended Parameters
Related Models
Keywords
Comments
No comments yet

