A high-quality model with a 48k sampling rate, a mature Yin Shasha model
Inference Comparison
After Inference
Other Inference
Model Introduction
Please pay attention to the Cn_hubert version, strictly follow the https://klrvc.com/zh/tutorial/3714 tutorial to deploy the model.
This model uses CN HuBERT(cn_hubert) for feature extraction. Compared to raw EN HuBERT(en_hubert), CN HuBERT usually delivers better < in Chinese speech scenarios>strong data-start="95" data-end="110" clarity and pronunciation stability, and is better for common "ambiguous, sticky words." These issues have some improvement effect.
data-end="247" Please make sure you can < strong data-start="151" data-end="170"> strictly follow the tutorial/file replacement instructions above. If you still use the original en_hubert or fail to replace the required files, it may cause feature distribution mismatches, resulting in issues such as data-end="243", which outputs nonsense, abnormal pronunciation, or instability.
data-end="353" The reference audio on the right demonstrates the model's performance under "soft" and "deep" style inputs. It should be noted that the actual effect is affected by factors such as timbre, speech speed, enunciation, recording quality, and emotional state > the original input
Recommended pitch parameter set to 13. This is not a fixed value; please fine-tune it based on your fundamental frequency and target listening experience, prioritizing "natural, comfortable, no distortion, and no obvious distortion." Because it uses cn_hubert, it is not friendly to other languages, but it supports Cantonese well.
zhangzhang
Version Details
Recommended Parameters
Related Models
Keywords
Comments
No comments yet

