The brief
A clinical-study app developer had an instructional video library produced in English and needed it to work in other markets. Their English production vendor was already in place; what was missing was everything that comes after — voice-over, captioning, and the on-screen text that carries a good part of the instruction.
We took the full localization lifecycle, sitting between that English vendor and the target markets, so that voice, captions and on-screen text arrive aligned rather than as three separate deliveries someone else has to reconcile.
Five languages
Spanish (Spain), Hebrew, French (Canada), German and Italian.
What made it difficult
Instructional content leaves no room to drift
The point of these videos is that a viewer follows them. Localized audio and the visuals it refers to have to stay locked together — a step named in the voice-over while the screen shows another one is worse than no localization at all. Final QC checks voice-over accuracy, audio-to-visual sync, text alignment against the translations, and mix balance.
On-screen text without the source files
Where the original project files exist, the on-screen text is replaced in them. Where they do not, we mask or cover the existing English and build the localized assets directly on top. Either way the result is a video whose graphics read as though they had been made in that language.
Running voice and graphics in parallel
Voice-over recording and on-screen text creation run at the same time rather than in sequence, with timing and visual cues checked against the translated audio as both progress. That is what holds the schedule: up to an hour of localized content — 15 to 20 short videos — inside two to three weeks, with several languages running in parallel and minimal impact on the overall timeline.
One review cycle for three kinds of asset
Reviewers work in Frame.io and comment in context, directly within the frame, which is the only practical way to give precise feedback on voice-over, captioning and on-screen text at once. Side-by-side comparison of revisions confirms each requested correction landed, and takes the version-control question off the table.
The result
One point of contact instead of three: voice-over, captioning and on-screen text consolidated into a single vendor relationship. Localized audio and visual cues land precisely against each other, so the videos keep the instructional intent of the original. And the pipeline is repeatable — a consistent production line for high volumes of short-form instructional content on tight deadlines.
