Blog

How to translate a video into another language

Translating a video can mean two things: subtitles in another language, or the speech itself in another language. OpenDub does both in one pass. It writes the translation as subtitles and speaks it in a clone of the speaker's own voice. That is what AI video translation means in this post: translated speech and translated subtitles together. This post walks through the steps and what you get at the end.

What translating a video means here

OpenDub is a video translator that dubs. It separates the speech from the music, transcribes it, translates each sentence to fit its moment, and speaks it back in a clone of the original voice. Subtitles in the new language can be burned in.

The translation is sized to the speech it replaces. Whole sentences are translated, one line out for each line in, and each line gets a character budget sized to how long the speaker took to say it. So each dubbed line is timed to last as long as the speaker took to say it.

How to translate a video, step by step

  1. Choose the video. MP4, MOV or WebM, from your own device. The speech can be in any language, and you do not have to say which: the language of the original video is detected automatically.
  2. Choose the language to dub into. There are 17, listed below. To translate a video into English, choose English.
  3. Choose a voice model. Higgs Audio and ElevenLabs use your own API key. OmniVoice is the free voice and runs in the OpenDub app on your computer. The differences are under "Where it runs and what it costs" below.
  4. Start the dub. OpenDub transcribes the speech, translates it, clones the voice from the video and speaks every line. Each new clip is transcribed back and generated again if it matches the text less than 72%.

You do not record a voice sample or train anything. The voice is cloned from the densest 6–12 seconds of whole sentences in the video you are translating.

What you get

  • The dubbed MP4.
  • The dub audio as WAV.
  • Subtitles in both languages as SRT.
  • The clip the voice was cloned from.

The picture stays the same

OpenDub replaces the voice and leaves the picture as it is. It does not change the lips in the picture. If you choose burned-in subtitles, they are added to the picture.

The music stays too

There are two ways to treat the original sound:

  • Replace the voice keeps the music and effects and swaps only the speech.
  • Voice-over leaves the original faintly underneath.

In the browser, the card has a checkbox, Remove the original voice on this device, which keeps the music and swaps only the speech. More on this in Keep background music when dubbing a video.

Languages you can translate a video into

17 languages: English, 简体中文, 繁體中文, 日本語, 한국어, Español, Français, Deutsch, Português, Italiano, Русский, हिन्दी, Bahasa Indonesia, Bahasa Melayu, Tiếng Việt, ไทย, العربية.

Where it runs and what it costs

OpenDub itself is open source under AGPL-3.0 and free to use. It needs no account. What you pay, and whether you install anything, depends on the voice model:

  • Higgs Audio or ElevenLabs, in the browser. You use your own API key and pay Boson AI or ElevenLabs directly. The dub runs in the card at the top of opendub.app, with nothing to install, for videos up to five minutes.
  • OmniVoice, in the app on your computer. OmniVoice is the free voice. It needs the OpenDub app on your computer and runs at about 15–35 seconds a line on a laptop CPU. Its model weights are licensed for non-commercial use only.

The app on your computer has no length limit. There are downloads for Mac (Apple Silicon) and Windows, and macOS and Linux also have a one-line install.

Is it a free video translator?

The software is free and open source. Translating a video for free means using the OmniVoice voice, which runs in the OpenDub app on your computer, for non-commercial use. Translating a video in the browser with nothing to install means using your own Higgs Audio or ElevenLabs key and paying that provider directly. Paid OpenDub credits, which need no key, are listed as coming soon.

Is the video uploaded?

No. The video stays on your device. Separation, speech detection, timing, mixing and rendering run there. Only speech clips and text go to the voice model you choose. An API key is kept in memory for that one dub, sent only to that provider, and never stored. The privacy page has the details.

How long it takes

One measured case: on an M1 Max, the 74-second demo video takes about four minutes to dub with a key limited to about one request per second. That dub makes about 70 Higgs calls.

Only clone voices you have the rights to use.

Related reading

Dub a video on opendub.app · All posts