ElevenLabs launched Dubbing v2 in May and updated its announcement in July. The company says the model conditions directly on the original performance to carry tone, pacing, delivery, and emotional intent into more than 90 languages. It also uses sync-aware translation to align starts, stops, and pacing with the source.

Those are ElevenLabs’ documented claims. This is my practical reading of what they mean for post-production, not a hands-on test. Better performance transfer could remove a lot of friction from the first dubbing pass. It does not make localization a one-click delivery step.

Preserving performance is the meaningful shift

Translation can be technically accurate and still sound disconnected from the picture. A pause, hesitation, change in energy, or emphasis may be as important as the words. If Dubbing v2 consistently carries more of that original performance across languages, it addresses one of the most visible weaknesses in earlier AI dubbing.

That matters because audiences react to performance before they analyze the technology behind it. A voice can be clean and intelligible but still feel wrong if the emotion, pace, or emphasis no longer matches the face and the edit.

A translated voice can sound convincing and still be wrong for the edit.

Synchronization is not the same as editorial timing

ElevenLabs says its sync-aware translation system automatically aligns the beginning, end, and pacing of translated speech. That is useful, but timing is not only a technical measurement. Editors also use silence, reaction shots, music changes, and cuts to control how an idea lands.

A line that fits the same number of seconds may still put the important word in the wrong place. It may step on a sound cue, weaken a reveal, or feel rushed in a language that needs a different sentence structure. The waveform can line up while the storytelling does not.

Every language creates a new picture pass

Localization also reaches beyond dialogue. On-screen text, lower thirds, supers, captions, legal copy, and calls to action may all need to change. Text expansion can break a motion system that worked perfectly in English. A translated phrase may need more time on screen, a smaller type size, or a different animation rhythm.

This is why motion design and version planning belong in the localization conversation early. Templates should allow for longer copy. Important text should not be baked into footage. Clean plates, organized graphics, and a delivery list by language can prevent a fast audio pass from turning into a slow visual rebuild.

Quality control still needs a local human ear

Current ElevenLabs documentation labels Dubbing v2 as an alpha model. It also says the v2 web workflow is automatic and does not offer content editing, while the editable Dubbing Studio remains tied to the legacy v1 model and is in maintenance mode. The company’s managed Productions service adds human translators, voice casting, pacing adjustments, professional mixing, and quality control.

That division is revealing. Automation can generate the language version, but someone still needs to verify pronunciation, names, idioms, cultural meaning, brand terminology, and legal language. A native speaker who understands the intended audience is not a cosmetic final check. That review is part of the creative and factual accuracy of the piece.

The final mix still matters

A localized track has to live inside the full soundtrack. Dialogue level, room tone, ambience, music, effects, and transitions need to feel like one piece rather than a replacement voice dropped on top. If the original performance changes in length or energy, the surrounding audio may need to move with it.

The same is true across platforms. A broadcast mix, a social version, and a presentation video may have different loudness and delivery requirements. Multiply that by several languages, captions, and aspect ratios, and version control becomes as important as voice quality.

Build localization into the project from the beginning

The strongest response to better AI dubbing is not to remove the post team. It is to design a workflow that lets the technology handle a faster first pass while experienced people concentrate on what needs judgment.

That means keeping approved scripts connected to timecode, separating dialogue from music and effects, preserving clean masters, building flexible graphics, using clear version names, and defining who approves each language. It also means testing the current product limitations before promising a workflow around them. ElevenLabs’ documentation says the Dubbing v2 API is not yet live, so an automated pipeline should not be presented as production-ready until the actual access and controls are available.

AI dubbing is becoming more capable, and that is genuinely useful. The opportunity is not simply to make more versions. It is to make good versions with less mechanical work and more attention available for language, performance, picture, sound, and delivery.

If your next project needs senior editing, motion design, sound, finishing, and organized versions for multiple platforms, let’s talk about the post-production workflow.

Primary sources