AI Studio
Dub Studio · Internal production tool
Your video. Any language. Same voices.
We localize finished video from start to finish, keeping the original speakers, the music and ambience, and the timing. Our own dubbing model does the voices, and a director signs off on every line.



Our point of view
We didn’t rent a model.
We built our own.
Most AI dubbing is a thin layer over a general-purpose voice API, the same engine everyone else calls with a logo on top. Ours is different. It is an internal model we built for one job: carrying a performance into another language. It is a first version, and it gets better every month. It already handles things the general tools miss, because we trained and tuned it for dubbing rather than for a demo.
The rest of the box
Our own model for the voice. The best tools for the rest.
-23 LUFS · EBU R128 · CONFORMED
The dubbing voice is our own internal model, and the tooling we build on top: adaptation that fits a line to its slot, scoring that picks the take, and a director who signs it off. For the work the model does not cover — AI video and avatars — we use the best of what exists. On top of all of it sits the part nobody sells you: direction, retakes, and quality control.
Speech technology
technology partner
Avatar video
FOCA certified
Three ways to be heard
Whose voice speaks the new language.
S01 → S49 · ALWAYS-ON
Every line carries a stage direction (tone, emotion, pacing, subtext) that the voice actually performs. Our team writes those directions, and rewrites them when a read lands wrong. It is closer to directing an actor than configuring a setting.
The pipeline
Eight steps. A human can interrupt any of them.
This is a studio process that happens to be automated, not an automation that happens to involve a studio. The pipeline runs end to end on its own — and at every stage a human validates the output and corrects what the model got wrong.
ACT I · PREP
Any source video. Audio extracted at production quality, 48 kHz.
Word-level transcription, with each speaker tracked through the whole programme. Interviews, two-host demos and panels are handled as what they are: several people, not one averaged voice.
Music and effects are isolated from speech. The original score and sound design survive the dub untouched. We replace the talking, nothing else.
ACT II · STUDIO
Each line is adapted to land in the exact time slot of the original, at the speaker’s natural density. A subtitle can run long and nobody notices. A dub cannot.
Cloned with written consent, designed from a brief, or drawn from a voice library you already own.
Every line is generated in multiple takes, up to a dozen on the ones that carry the scene. Each take is scored three ways: does it say exactly the translated words; does its energy follow the original delivery, so a shouted line stays shouted and a whisper stays a whisper; and is it still recognisably the same voice.
ACT III · DELIVERY
Our team auditions the picked takes line by line, edits any translation by hand, re-directs a delivery (warmer, slower, a smile in the voice), regenerates, or overrides the pick. This is the step the machine does not get a vote in.
Takes are placed back on the timeline over the original music bed, micro-timed so the correction stays inaudible, mixed, and rendered onto the original picture. Your delivery specs are the deliverable.
human can step in
The figures and lines shown here are an example, not a delivered job. What the automation buys you is speed; what you pay for is the ability to stop at any step and change the result — and those corrections feed back, so the system gets the next one closer.
Listen
Range you can hear
Pick a voice, then an emotion. Same voice, different intent — that is the part a model usually flattens.
Synthetic voices, produced in our AI studio.
Coverage
Twenty-six languages, at three depths. We say which is which.
English and German run through a model fine-tuned for voice acting: the direction is part of the generation. The base model carries a dozen more natively, without that fine-tune. Beyond it, a second engine covers the rest and we add the emotion.
- English
- German
- Spanish
- Chinese
- Arabic
- Portuguese
- Russian
- Japanese
- French
- Korean
- Italian
- Turkish
- Polish
- Hindi
- Dutch
- + 11 more
Top row: directed emotion, native. Middle: the base model, no acting fine-tune. Dashed: a second engine, emotion added and edited. Same delivery standard throughout — different roads to it.
Our own systems
The tool that does the work.
This is the studio pipeline we built, running a real animation from English into German, French and Spanish. It can dub a video from start to finish on its own, and a person can step in on any line: the translation, the delivery, the take. Full-auto when you want speed, hands-on when you want it exact.


speakers tracked
segments
languages out
multitrack master
Where we draw the line
The rules we don’t bend for a deadline.
We design voices, we do not copy people
Cloning copies a person. Casting finds the right voice for the message — we do it synthetically, and we direct it. A cast voice is built to be non-identifiable: not a real person with the name filed off.
A voice is directed, not generated
Emotion is a direction, given take by take and judged by a human ear before it is kept. If nobody in the room can hear the difference between two takes, we have not finished working.
Cloning happens in writing, or not at all
When a project genuinely calls for a real voice, we clone it only with that speaker’s written consent, scoped to the project it was given for. No exceptions. Anything not made for delivery is marked as such, in the file and in the brief.
Processing inside an agreed perimeter
Where a project requires it, the pipeline is deployed on controlled infrastructure. Tell us the constraint before we scope and we will tell you exactly how it runs for you.
What you are buying is a voice chosen for the message and directed until it lands — not a person’s voice borrowed at speed. The rules above exist because we also clone when a project asks for it, and that is the part where a signature has to be on file.
FAQ · Ask us anything
Asked, and answered.
Our own, at the core, built with a partner. The dubbing voice comes from an internal model we developed with Voxist and keep training. It is a first version that improves month over month, tuned for carrying a performance across languages rather than for a general demo. Around it runs the rest of our own work: the adaptation that fits a translated line into its time slot, the scoring that decides which take survives, and the studio interface our directors use. When a job needs AI video or avatars, we use the best third-party tools for that, Synthesia among them. The voice and its direction stay ours. We do not publish the recipe. What you buy is that model, the parts we built around it, and the director at the end of the line.
Anywhere. Our team can rewrite a translation by hand, re-direct a line (warmer, slower, a smile in the voice), regenerate it, or throw out the automatic pick and choose another take. The pipeline proposes options and a person makes the call.
Only with the speaker’s written consent, scoped to the project it was given for. The alternative is a designed voice that imitates no living person and carries no identity-rights baggage. For a lot of brand work, that is the better option anyway.
The engine can be deployed on controlled infrastructure, so for work under an NDA that names sub-processors we can keep processing inside an agreed perimeter. Tell us the constraint before we scope and we will tell you exactly how the pipeline is deployed for your project.
English, German and French today, from any source language we can transcribe. More are in validation. We add a language when it clears the same bar as the rest of the studio, not when a model claims it can.
Keep it. The question is who signs off on the languages you cannot read. We are the direction and the finish: adaptation that fits the mouth, a director on the retakes, mix, compliance, and a linguist’s approval before anything ships.
Cheaper per pass, not per result. Professional first versions get revised as a matter of course: a client, a reviewer or a native speaker looks at good work and asks for better. A workflow that ships version 1 unreviewed is betting your brand on nobody looking. You pay us to be the ones who look.
Yes. Dub Studio is an internal tool we run on selected projects, so we take on work we can watch closely. Send us a few minutes of something you have already localized the traditional way and we will dub it, so you are comparing against a result you already trust.
Ready when you are
Send us a video, and we’ll tell you if it fits.
Tell us what you need localized and we will tell you honestly whether this is the right tool for it. A 30-minute call.