
Open-source voice cloning tool translates speech into 14 languages
An open-source voice synthesis project called Confucius4-TTS can replicate a speaker's voice using a single audio sample. The system generates speech in 14 different languages while removing native accents and retaining natural emotional tone.
The Blend
NetEase Youdao has published an open-source text-to-speech system called Confucius4-TTS that can clone a speaker's voice across 14 languages using a short audio clip, according to the software repository on GitHub. The system generates foreign-language audio while preserving the original speaker's vocal identity and emotional expression.
Conventional language dubbing often relies on robotic synthetic voices or retains strong foreign accents when attempting translation. By removing native accents while preserving personal cadence and tone without needing a text transcript of the input audio, this software makes high-quality media localization significantly easier for creators and businesses.
While open-source tools make cross-border communication more accessible, they also lower the barrier for creating convincing voice spoofs without consent. As instant cross-lingual voice duplication becomes accessible to anyone with basic technical skills, digital platforms will face growing pressure to develop reliable tools for identifying synthetic audio.
Written independently by AI News Smoothie from the reporting listed below. Facts belong to the original publishers — follow the links for their full coverage.
Ingredients
- GitHub - netease-youdao/Confucius4-TTS: Confucius4-TTS: a Multilingual and Cross-Lingual Zero-Shot TTS Engine · GitHub
NetEase Youdao released an open-source speech model that can replicate a person's voice and emotional tone across 14 languages without retaining regional accents, according to the project's code repository.