48K Ai Rimo: A very clean and girly voice RVC model
Inference Comparison
After Inference
Other Inference
Model Introduction
As the title suggests, this model handles the noise of this raw dataset the most "exhausting" of all currently available models.
Many alchemists use AI for one-click denoising when training models, which is currently the most widely used method, but the result is that the model output is often poor, with various types of noise.
I used Ai+ purely handcrafted, even the breathing was processed and deduced most of the original sounds, and the results were very good.
However!
Currently, I still can't train the laughing part. Although I tried processing laughter with similar tone and putting it into "alchemy," the effect still didn't work. I then threw the laughter from the original dataset into the inference, which was problematic. So I concluded that if a real model needs to produce a model of true "laughter," it needs a large amount of expression in the original dataset.
And there are many kinds of laughter. In the original dataset, some of the currently popular "laughter" are thrown in, such as: "Hmph!" Hmph! ~" "hhhhhhh" and so on, the output effect is still decent, but that kind of spontaneous smile is hard to support in this model.
This requires me to collect enough and diverse expressions of tone inthe later stages to achieve this.
For AI covers, this model is already sufficient; even if I reasoned without accompaniment, the effect was still very good.
The recommended tone for the model is roughly between -3 and 5, which can be adjusted according to your own needs.
Breathing sounds and such are fine, with no "zi" sound, and the song's reasoning is also well presented.
Effects vary from person to person; it doesn't mean your output will be exactly what I demonstrate.
I uploaded a copy from Google Cloud, and you can also download it from Google.
zhangzhang
Version Details
Recommended Parameters
Related Models
Keywords
Comments
No comments yet

