The question used to be simple: subtitle it or dub it? By 2020 there were three answers, once AI dubbing became a real line item. In 2026 there are four, because live speech-to-speech shipped inside the meeting tools you already pay for, and it is now a delivery option rather than a demo.
Four modes sound harder to choose between than two. They are not, because they sort cleanly on two axes: how much delay you can accept, and what an error costs once it is out. This post defines the four, then runs them through three lenses — the content, the audience, and the budget.
[Average read time: 8 minutes]
The four modes, defined
Subtitling. The spoken dialogue is translated and shown as timed text on screen while the original audio plays untouched. Generally the cheapest option, and usually the fastest to turn around.
Human dubbing. The original dialogue is replaced by voice actors performing in the target language, timed to the lip movements of the people on screen. It costs more and takes longer, because you are casting, directing, recording, and mixing a new performance.
AI dubbing. Software generates a target-language voice track that replaces the original audio. Two variables are worth separating: the voice can be fully synthetic or a clone of the original speaker carried into the new language, and the on-screen mouth can be left alone or re-rendered to form the new words. That second option, visual lip-sync, is a distinct capability rather than a setting, and it is what removes the mouth-mismatch you get from audio-only dubbing.
Live speech-to-speech. A target-language voice generated straight from source speech as it happens, sometimes carrying the speaker's own vocal identity, with no written script step in between.
There is also a lighter option that keeps its place: UN-style voice-over, where a translated voice reads over the original audio, which stays faintly audible underneath, the way you hear it in a UN briefing or a news interview. It favors accuracy over polish, and for a lot of factual content it is the correct answer rather than the cheap one.
Where AI dubbing is strong is easy to state: it is fast, the cost per minute is low, and it scales across many languages at once. Where it struggles is performance nuance, emotional range, comedic timing, song, and accent. And accountability, which is not a quality problem: when a voice has been cloned, someone has to own consent for that clone, and someone has to stand behind the version a network airs.
Latency is the first axis, and it decides more than people expect
Real-time is not a switch. It is a range, and the trade-offs move as you slide along it.
- Live: sub-second to a few seconds. A webinar interpreted as it happens.
- Near-live: minutes to hours. A recorded all-hands dubbed overnight for the regional offices.
- Batch: the finished-asset workflow. The file is the product, and there is time to review it.
Underneath, even live output runs through stages: speech recognition turns sound into text, machine translation converts it, then text-to-speech or voice conversion speaks it back, sometimes with live lip-sync on top. Each stage adds latency and each stage adds its own errors. Microsoft's Teams Interpreter is explicitly cascaded this way (Source: Microsoft Learn, 2026-06-03). Cascaded systems chain those stages; end-to-end models try to compress them into one. Either way the engineering tension is the same: lower latency costs quality, higher quality costs latency.
Two practical notes on the live tier before you plan around it.
Language coverage is smaller than the marketing suggests, and it is per-tool. Teams Interpreter supports 9 languages for speaking and listening: Chinese (Mandarin), English, French, German, Italian, Japanese, Korean, Portuguese, and Spanish (Source: Microsoft Learn, 2026-06-03). Google Meet's Speech Translation launched with English and Spanish, with Italian, German, and Portuguese following (Source: 9to5Google, 2025-05-20). The larger aggregate figures you will see quoted, "90+ languages" and similar, usually describe offline dubbing or a preview path. Do not plan around them as live coverage.
Single latency numbers are not useful. Production live interpretation runs roughly one to a few seconds end to end, and first-chunk latency, meaning how long until you hear the first words, is the metric a listener actually experiences. Sub-second is achievable in some systems but is not universal, and published benchmarks mix vendor claims with research figures (Source: Forasoft, 2026).
Then there is the structural ceiling, and it is the reason latency belongs on the same axis as stakes. Real-time output is draft quality by construction. There is no review pass between the moment the system generates a line and the moment it reaches the listener. Whatever it gets wrong, it ships.
One more maturity caution: voice-preserving dubbing for broadcast is not as settled as the meeting features. ElevenLabs' Dubbing v2 was still in alpha as of 2025. Separate what has shipped and is stable, meaning live meeting interpretation, from what is still maturing, meaning polished voice-matched video dubbing, before you promise either to a stakeholder.

Lens 1: the content
Look at the content before the budget. It narrows the method on its own.
A rough map. High-volume corporate video, e-learning, explainers, and social clips are reasonable AI dubbing candidates. Flagship film, scripted TV, and drama want human dubbing or subtitles. Anything culturally specific or carried by performance leans toward subtitling or a human dub.
To get specific: complex animated graphics with off-screen narration usually do better with replaced on-screen titles and a laid-down voice-over than with a frame buried in subtitles. For a series or a film, lip-sync dubbing is generally preferred, because viewers settle into a story faster when the mouths roughly match the words.
Culturally specific content is where we slow down. Consider the scene in Parasite that turns on the song "Dokdo is Our Land," a reference loaded with Korean political meaning. Subtitling preserves that; a dub, AI or human, tends to flatten it.
For on-screen speakers there is one more test, and it is visual rather than linguistic. AI lip-sync has real strengths for front-facing, well-lit talking-head footage: HeyGen offers video translation with lip-sync across a long list of languages, Meta added an in-platform translate, dub, and optional lip-sync feature for Reels in late 2025 (Source: Meta, 2025), and Deepdub and others work the same territory. Iteration is easy, too — a script edit or a new language can be regenerated without bringing talent back. But small artifacts in the mouth or in the prosody read as off, and they show up most on close-ups and tight shots, which is exactly where high-stakes material lives. Watch the close-ups before you sign off. The gap is narrowing, so treat that as a current observation rather than a permanent verdict.
Where a human performer still leads is content that carries intent. A CEO address, sensitive HR messaging, brand storytelling: these live on pacing, emphasis, and emotional nuance, and a voice actor working with a director makes hundreds of small choices about where to land a pause or lean on a word.
Lens 2: the audience
Who is watching changes the answer as much as what they are watching.
Cinephiles, film students, and many streaming subscribers prefer subtitles, because they want the original performance. Dubbing scripts have to deviate from a literal translation to match lip movement, so something is always traded away. For young children who cannot read fast yet, and for older viewers dealing with small text or reading fatigue, dubbing is generally the easier watch. Treat that as an industry observation rather than settled science: for children specifically there is peer-reviewed evidence pointing the other way, since same-language subtitles appear to help literacy.
Geography matters, and this is where the older version of this argument needed correcting.
France, Germany, Italy, and Spain are strong dubbing markets, for historical reasons. That is preference, mostly, not law. The one real legal wrinkle is France's Toubon Law (1994), which requires French subtitling or dubbing for broadcast audiovisual content but specifically exempts original-version films (Source: Wikipedia, Toubon Law). Germany and Italy did have compulsory dubbing under pre-war regimes; that is history, not current regulation. So France carries a narrow legal requirement and the others are market preference.
On the other side are the traditional subtitling markets: high-English-proficiency parts of Northern and Western Europe. The Netherlands is among them and is not Scandinavian — Scandinavia is Denmark, Norway, and Sweden. On the 2025 EF English Proficiency Index the Netherlands ranks first, for the seventh year running, ahead of Croatia, Austria, and Germany. The Nordic countries sit high but not at the very top: Norway fifth, Denmark seventh, Sweden eighth.
Streaming nudged some of this. Platforms now default to and prioritize dubbing in preference markets such as France, Germany, and Japan, while always shipping subtitles and keeping the original audio available. Netflix reported that when it streamed dubbed and subtitled cuts of the French drama Marseille to two viewer groups, the group watching the dubbed version was far more likely to finish (Source: The National, 2020). That is a platform result rather than peer-reviewed work, so read it as directional; the dubbed cut of the German series Dark is reported to have drawn most of its English-language audience, which points the same way.
Asian markets are sometimes cited as shifting toward dubbing. That reading traces largely to one industry executive, and later survey data has a majority of adults in China and South Korea still preferring subtitles. The honest framing is that studios are offering more dubs, not that consumers have flipped.
Lens 3: the budget, and what finishing actually costs
With content and audience narrowing the field, cost sorts the rest.
The ranking is steady. Subtitling is cheapest and fastest. Human dubbing is the most expensive and slowest. AI dubbing sits in between, closer to subtitling on price and speed. Live sits outside the ranking, because you are usually buying it as part of a platform subscription you already have rather than as a per-minute service.
The temptation is to read the AI slot as a free upgrade over subtitles. It is not, quite. It trades the lowest cost for a quality and risk profile you have to look at language by language and asset by asset.
Two costs are easy to leave out of a comparison and expensive to discover later.
Broadcast finishing. A network or regulated channel has delivery specs, and a person has to mix, check, and certify that the final file meets them: loudness, captioning to accessibility standards, file format. A tool gives you a draft. Bringing it up to a channel's requirements is craft work with pass-or-fail criteria.
Consent and disclosure. Re-rendering a real person's mouth alters their likeness; carrying their voice into another language clones it. Both carry obligations a tool cannot sign off on, and in the EU those obligations bind the party that publishes rather than the vendor that generated the file. Our posts on voice cloning and on disclosure and provenance cover the mechanics.
How to sort your catalogue
One question does most of the work: what does an error cost once it is published?
- Informational, internal, small error survivable — live or near-live is a reasonable tool. Webinars and conference Q&A, internal e-learning and onboarding, live captions on a town hall. Captions are the lowest-risk real-time output you can pick, because text on screen carries fewer ways to go wrong than a synthetic voice speaking your translation.
- High volume, low stakes, finished asset — AI dubbing, with a review pass sized to the risk.
- Flagship, performance-driven, or culturally loaded — human dubbing or subtitles, depending on the audience.
- Anything with a broadcaster's delivery specs, a real person's identity, or legal weight — a named person in the loop before it ships, whatever mode you picked.
Most real projects fall across several rows, and that is normal rather than a compromise. You might carry a whole training library on AI dubbing and reserve the studio for the executive introduction that opens it.
Our white paper, Localizing in the Age of Generative AI, goes deeper on where AI fits in that workflow and where it does not. If you want help drawing the line for a specific set of videos, we are glad to talk it through.
