Machine translation, neural voices, and AI lip-sync now do parts of video localization that used to need a person, and they do them well enough that buyers are asking whether they still need a vendor at all. The backlash we see is not against AI. It is against AI with nobody on the hook.
So the question that decides a localization purchase in 2026 is not which model a vendor runs. It is narrower and much easier to ask: which named person decides that the output is good enough, and who answers when it is not.
This post is the long version of that question. We walk the workflow stage by stage, show what it does to your invoice, give three documented cases of what happens when nobody is steering, and end with the five questions that tell you whether a job needs paid human review at all.
[Average read time: 9 minutes]
Where the line sits
Before the stages, name the line.
AI does the volume work. First-draft translation, rough subtitle timing, draft text-to-speech, bulk processing across many files and languages. Speed and consistency matter here, and machines are good at both. These are the commodity stages of the job, and automating them is the right call.
A person does the rest: transcreation (rewriting a message so it lands in another culture, not just another language), voice casting, consent and licensing, native-speaker checks, on-screen text and motion-graphics finishing, and the final sign-off. Direction and finishing, in other words, plus every call where being wrong is expensive.
We are not anti-AI. Our job is to be the party that decides when AI is enough and when it is not, then to put a name against that decision. That referee role is part of what you are buying.
The workflow, stage by stage
Stage 1: LLM pre-translation, the fast first draft
A large language model or neural machine translation engine produces a first pass across many languages quickly. Fed with the right inputs it gets better: your past approved translations and your glossaries keep terminology consistent from one video to the next.
Where it helps is volume and speed. Where it fails without supervision is everything that needs a reader who knows the language. Tone and register slip. False friends sneak through. Models invent text that was never said, get gender wrong, and miss cultural references.
So the human here is a post-editor, not a proofreader. The linguist treats the model output as raw material and reworks it, a step the industry calls machine translation post-editing. Every line gets read, corrected, and rewritten before it moves on.
Stage 2: media adaptation, fitting language to picture
This is the step pure-AI pipelines tend to skip or do crudely. Translated text has to fit the picture, not just the meaning.
Subtitles have reading-speed limits and character counts. Dubbed lines have to match the time the speaker's mouth is moving. Languages do not occupy the same space: a line that fits an English frame often overruns once it becomes German or French. A person resolves that against the actual footage, trimming and rephrasing so the words fit the shot.
On-screen text and motion graphics are their own problem. A title card, a lower third, a label baked into the animation: these are not audio, and a voice tool will not touch them. Adapting them is in-house craft, and most automated pipelines leave it on the table.
Stage 3: voice, real or approved-cloned, never unlicensed
There are three honest ways to produce a localized voice. Record a real voice actor. Use a neural text-to-speech voice licensed for commercial use. Or clone a specific person's voice, with that person's documented consent and a fee for the use. The tools are good; capability is not the question. Provenance is.
The line is simple: no voice goes into a deliverable without rights behind it.
One common misunderstanding is worth clearing up here. SAG-AFTRA has strong AI terms now, and its 2025 Interactive Media Agreement, ratified on 9 July 2025 to end the 2024–2025 video game strike, requires separate written and reasonably specific consent before a performer's digital replica is created or used, with that consent void if the use exceeds the described scope (Source: SAG-AFTRA, 2025). Those protections apply to union members under signatory contracts. Most corporate, e-learning, and industrial narration is non-union and sits outside SAG-AFTRA's reach. If you want those terms on a non-union job, someone has to write them into the contract. We do that ourselves rather than imply a union covers work it does not. The mechanics are in our post on voice cloning, consent, and licensing.
Stage 4: visual dubbing and lip-sync, where it earns its place
AI lip-sync has come a long way. For front-facing, well-lit talking-head footage, the kind common in e-learning and corporate explainers, it holds up. For broadcast work, extreme camera angles, and high-emotion delivery, it still falls short of a skilled human dubbing team. We use it where it earns its place and finish by hand where it does not. Our decision grid post covers how to sort content across subtitling, human dubbing, AI dubbing, and live.
Stage 5: native QA, finishing, and disclosure
A native speaker reviews the whole video in context, the on-screen text and graphics are finished, and a named person signs off that it is ready to air.
Disclosure belongs here, and in Europe it is now a legal matter rather than a courtesy: the EU AI Act's transparency duties apply from 2 August 2026, and they fall on the deployer who publishes, meaning the brand or studio, not only the AI vendor. Our EU AI Act and provenance post has the detail, including the machine-readable marking duty, which is a separate obligation on a different party.

The drafting line on the invoice shrinks. The review line does not.
What this does to pricing and roles
If the machine writes the first draft, does the work simply get cheaper? It gets cheaper in one place and not in another, and the split is the whole story.
Post-editing is faster than translating from a blank page, though the saving is narrower than the headlines suggest, and how much narrower depends heavily on the language pair and the subject matter. On hard pairs or specialized content the gap shrinks further. So the drafting line on an invoice shrinks. The review line does not, because the review is now the work.
The failure mode changed too, and that is what raises the cost of review rather than lowering it. Older tools produced output that read as broken, so a mistake announced itself. Modern models produce output that reads smoothly and is sometimes confidently wrong. A clean, fluent sentence that says the opposite of the source is harder to catch than a garbled one.
Some errors are old and still here. Translating from a language with no gendered third-person pronoun, such as Turkish or Uzbek, into one that forces a choice, the model has to guess, and it guesses from training data, so "the doctor" drifts male and "the nurse" drifts female. Idiom stays brittle. Google Translate is a useful marker of where the tooling sits: as of December 2025 it runs a hybrid setup, with a large language model behind consumer text translation for slang and context, and the older NMT engine kept in place for low-latency real-time work (Source: Google, 2025-12-12).
The linguist's role shifts with all this. Less time goes to producing words, more to deciding which machine words to keep, which to cut, and which carry a risk that no fluency can paper over.
It also helps to be honest about the money around the industry, because the figures get quoted as if settled and they are not. Mordor Intelligence puts the language services market near 75 billion dollars for 2026; Nimdzi has measured it closer to 72 billion dollars for 2024 (Source: Mordor Intelligence; Nimdzi). They disagree because they count different things. Both put it above 70 billion and rising, with spending shifting toward review and finishing rather than draft production.
There is one number from our own records that we think earns its place next to those. Measured in June 2026, our production archive held 230,579 files across 42 languages and 21 years. Of the 12,579 video deliverables carrying explicit version tags, 85 percent reached a version 2 or higher, 29 percent reached a version 3, and one ran to 22 (Source: JBI archive study, June 2026). A single-pass AI draft is version one of something that usually needs several.

Fluent output is harder to audit than broken output.
Three ways it breaks when nobody is steering
The argument for review is easier to make with cases than with principle, so here are three, all documented, all public.
Meaning reverses, and it reads perfectly. In October 2017, Facebook's translation system rendered a Palestinian man's Arabic post. He had written "good morning." The system produced "attack them" in Hebrew and "hurt them" in English. No Arabic-speaking officer reviewed the post; he was arrested before the error was caught, then released after questioning (Source: Quartz, October 2017). Part of why this is structurally hard to catch is that the person signing off on a translation often cannot read the target language. They are trusting how fluent the output sounds rather than what it says.
The error rate that matters is not the average. A study of emergency-department discharge instructions run through Google Translate found the Spanish output around 92 percent accurate and the Chinese output around 81 percent. Among the errors that occurred, roughly 2 percent in Spanish and 8 percent in Chinese carried the potential to cause real clinical harm (Source: JAMA Internal Medicine via ScienceDaily, February 2019). Newer systems have narrowed that gap. In medicine, law, financial figures, and public-health messaging, a small share of harmful flips is still not a rounding error.
The public version costs more than the desk version. In November 2018, the Spanish Ministry of Industry published a press release that ran an official's name through automatic translation. Dolores del Campo came out in English as "It is pain of field" (Source: The Guardian, November 2018). The story traveled far further than the announcement ever would have.
Software teams have known a version of this asymmetry for decades, though the usual citation deserves a caveat: the "100 times more expensive in production" multiplier traces back to undated internal training material rather than a published study, and it has been picked apart since (Source: The Register, 2021). The direction holds even though the multiplier does not. A flagged sentence at review is a quick edit; the same sentence on air is a different problem.
Five questions before you buy the human version
"Human-certified" is a phrase you will see in proposals, and it currently tells you almost nothing. There is no official badge for localized video as of 2026, and no standards body issues one. The closest real things are process certifications: ISO 17100, which covers translation by qualified humans plus a separate revision step, and ISO 18587, which covers human post-editing of machine-translation output (Source: ISO, 2026). Both certify a workflow. Neither is a stamp you can put on a finished dubbed video.
So the buyer has to define the term. Run the job through five questions instead of a gut call:
- Stakes. What does a public error actually cost you here, in legal, reputational, or regulatory terms?
- Audience. Is this headed for broadcast or regulated distribution, or is it internal and low-reach?
- Rights. Is a real person's voice, face, or performance involved at any point?
- Permanence. Will this live for years and shape how the brand is seen, or is it disposable?
- Verifiability. Can the vendor name who signed off, or only tell you that "a human reviewed it"?
Answers clustering toward high stakes, broad audience, real people, long life, and a vendor who can name names give you a case for paying. Clustering the other way, you probably do not have one.
What a credible claim contains
If the job clears the checklist, here is how to separate a real commitment from a rubber stamp.
Named accountability, not anonymous "human-in-the-loop" language. Someone should be willing to be the person who decided.
A defined scope, in writing: what the human actually decided, and what was auto-generated. "Reviewed" with no boundaries is theater.
Consent and rights documentation for any synthetic voice or likeness. In the US, the NO FAKES Act, aimed at unauthorized AI replicas of voice and likeness, advanced out of the Senate Judiciary Committee in mid-2026 but is still pending and has not become law (Source: US Senate Judiciary Committee, 2026). You want a vendor whose paperwork would survive that kind of rule whether or not it passes.
Videos, not tool access. A real service hands you the finished file to spec. If the deliverable is software you have to run, you are doing the work and you are the one accountable for it.
An audit trail a broadcaster or legal team can inspect. FCC caption rules in the US set outcome standards for accuracy, timing, and completeness rather than literally requiring human verification, but meeting those standards on pre-recorded programming effectively takes human quality control (Source: FCC, 47 CFR Part 79). Ask whether the audit trail records who performed that control and when.
When you should not pay for it
Often, honestly. A weekly internal update, a batch of product how-to clips, a training video with a six-week shelf life: none of these need a named person signing a certificate. Paying for the label here buys reassurance you will not use.
Industry estimates put AI dubbing in the rough range of a few dollars per minute against several hundred for traditional studio work, with claimed savings cited as high as 90 percent or more (Source: Research and Markets, 2026). Treat those as vendor and analyst figures with wide spread rather than audited facts, but the direction is real, and it is real precisely because most video does not need the studio path.
If a vendor tells you everything needs human certification, that is a sales position, not advice. The same goes for "AI does it all, you will never tell the difference." Both sell one answer to a question that depends on the job.
Where the stakes are real, ask what the human review covers, who is accountable, and whether the deliverable is a finished video or a tool you have to operate. Where they are not, save the money. Our white paper, Localizing in the Age of Generative AI, sets out how we draw that line in practice, and our insight report, What 13 Terabytes of Human Localization Reveal About AI's Real Limits, has the archive analysis behind it.
