OpenDub takes a video in one language and gives it back in another. The new voice is a clone of the speaker's own. Each line starts and ends with their lips, and subtitles in the new language are burned in. It is open source under AGPL-3.0, and most of the work runs on your own device. This is what it does, step by step.
Eight steps
- Separate. Demucs splits the voice from the music and effects. The soundtrack survives the dub, and every later step hears a clean voice.
- Find speech. Silero VAD marks where someone is talking. Those regions decide how the audio is cut up for transcription.
- Transcribe. Higgs STT writes down the words. You do not have to say which language they are in; it is detected. A local Whisper model supplies only the timing, aligned word by word.
- Translate to fit. Whole sentences are translated, one line out for each line in, and the count is checked. Each line gets a character budget sized to how long the speaker took to say it.
- Read the tone. Each line's loudness and pace, measured against the speaker's average, go with its words to a model acting as a voice director. It picks an emotion, how expressive to be, and whether to whisper or shout.
- Speak. The voice is cloned from the densest 6–12 seconds of whole sentences in the original. Every line is then spoken in it.
- Check and fit. Each new clip is transcribed back and generated again if it matches the text less than 72%. A line more than 15% longer or shorter than the original gets up to three more tries: reworded twice, then spoken faster or slower. What is left is stretched between 0.75× and 1.3×.
- Mix and render. The new voice goes over the music, matched to the original's loudness. Subtitles are timed to the dubbed voice and burned in.
Speaking speed is never a setting. It is chosen per line, so each one fits the moment it replaces.
What leaves your device
The video file never does. Separation, speech detection, word timing, mixing and rendering run locally. What goes out is speech clips and text, and only to the voice model you choose, with your own API key: Higgs Audio from Boson AI, or ElevenLabs. On opendub.app the key is used for that one dub, sent only to that provider, and never stored.
If you want nothing to leave at all, the app on your computer can dub with free engines instead: Whisper, a local language model and OmniVoice. OmniVoice is slower, about 15–35 seconds a line on a laptop CPU, and its weights are for non-commercial use only.
Three ways to run it
- In the browser. The card at the top of opendub.app runs the whole dub in your tab, for videos up to five minutes.
- The app on your computer. No length limit, and the voice is removed faster. The Mac (Apple Silicon) and Windows downloads carry Python, ffmpeg and everything else inside. On macOS and Linux there is also a one-line install of about 2 GB.
- Headless. From a clone of the repository,
.venv/bin/python -m opendub.pipeline video.mp4 --to jadubs a file with no page at all, once./run.shhas set it up.
The app also lets you fix what you hear. Edit a translation, change a line's tone, or drag its block on the timeline, then apply. Only the lines you touched are spoken again, in about ten seconds.
What to expect
On an M1 Max, the 74-second demo video takes about four minutes to dub with a key limited to about one request per second. That dub makes about 70 Higgs calls. There are 17 languages to dub into. The app gives you the dubbed MP4, the dub audio as WAV, subtitles in both languages as SRT, and the clip the voice was cloned from.
Who it is for
Anyone with a talk, a lesson or a product video who wants it heard in another language in their own voice, without uploading the video to get there. Only clone voices you have the rights to use.