OpenDub takes a video in one language and gives it back in another, and the new voice is a clone of the speaker's own. You do not record a sample or train anything. The voice comes out of the video you are dubbing. This is how that works, and what you can choose.
Where the cloned voice comes from
Before anything is cloned, Demucs splits the voice from the music and effects, so every later step hears a clean voice. Silero VAD then marks where someone is talking, and the words are transcribed.
The voice is cloned from the densest 6–12 seconds of whole sentences in the original. Every line of the translation is then spoken in it. When the dub is done, the app gives you the clip the voice was cloned from along with the dubbed MP4, the dub audio as WAV, and subtitles in both languages as SRT, so you can hear exactly what the clone was made of.
Keeping the tone, not only the voice
A cloned voice that reads every line the same way does not sound like the speaker. So each line's loudness and pace, measured against the speaker's average, go with its words to a model acting as a voice director. It picks an emotion, how expressive to be, and whether to whisper or shout. The line is then spoken that way.
Speaking speed is never a setting. It is chosen per line, so each one starts and ends with the speaker.
Checking every line
Each new clip is transcribed back and generated again if it matches the text less than 72%. A line more than 15% longer or shorter than the original gets up to three more tries: reworded twice, then spoken faster or slower. What is left is stretched between 0.75× and 1.3×.
In the app on your computer you can also fix what you hear. Edit a translation, change a line's tone, or drag its block on the timeline, then apply. Only the lines you touched are spoken again, in about ten seconds.
17 languages to dub into
The cloned voice can speak 17 languages: English, 简体中文, 繁體中文, 日本語, 한국어, Español, Français, Deutsch, Português, Italiano, Русский, हिन्दी, Bahasa Indonesia, Bahasa Melayu, Tiếng Việt, ไทย, العربية. The language of the original video is detected automatically, so you only choose the one to dub into.
Which voice models clone the voice
OpenDub itself is open source under AGPL-3.0 and free to use. The cloning is done by the voice model you choose:
- OmniVoice. The free voice. It runs in the OpenDub app on your computer, at about 15–35 seconds a line on a laptop CPU. Its model weights are licensed for non-commercial use only.
- Higgs Audio. Transcribes, translates, reads the tone and clones the voice. You use your own API key and pay Boson AI directly.
- ElevenLabs. Scribe and instant voice cloning. You use your own API key and pay ElevenLabs directly. A temporary cloned voice is deleted from your ElevenLabs account when the dub finishes.
With your own Higgs Audio or ElevenLabs key, the dub runs in the card at the top of opendub.app, for videos up to five minutes, with nothing to install. The app on your computer has no length limit.
What is sent to clone the voice
The video file never leaves your device. What goes out is the speech from your video, to transcribe it, and a clip of it to clone the voice, plus the text of each line. It goes only to the voice model you choose, with your own API key. The key is used for that one dub, sent only to that provider, and never stored. The privacy page has the details.
Only clone voices you have the rights to use.
For the whole pipeline step by step, see How OpenDub dubs a video in the speaker's voice.