ElevenLabs Dubbing v2 Preserves Performance Across 90+ Languages
ElevenLabs has introduced Dubbing v2, an Alpha dubbing model designed to carry a speaker's delivery, emotion, and pacing into translated audio.
ElevenLabs has launched Dubbing v2, an Alpha AI dubbing model designed to preserve a speaker's original performance when audio or video is localized into another language. Rather than depending solely on a transcript, the model conditions directly on the original delivery. The company says this enables translated tracks to retain more of the source speaker's emotion, tone, pacing, emphasis, pitch, and voice identity.
That shift addresses a familiar weakness in automated dubbing. Text can convey the words that were spoken, but it does not fully represent how they were delivered. When translation and synthetic speech are driven principally by text, the result can sound flatter than the source. According to ElevenLabs' official Dubbing v2 announcement, the new model uses the original performance as part of the input so that intonation and emotional intent can carry across languages.
For creators and organizations localizing video, podcasts, campaigns, or other spoken content, the practical goal is more natural multilingual delivery without rebuilding each asset through separate translation, voice, editing, and engineering stages. Dubbing v2 supports more than 90 languages and includes synchronization-aware translation intended to align starts, stops, and pacing with the source material.
What Dubbing v2 changes
The central change is not simply broader language translation. It is the model's performance-centered approach to localization. ElevenLabs positions Dubbing v2 as an end-to-end workflow that can preserve the character of a source performance while generating a translated track.
The company highlights several capabilities:
- Retention of the original speaker's emotional delivery, tone, and emphasis across translations.
- Preservation of voice identity, including pitch and tonality.
- Synchronization-aware translation to better align timing and pacing with the original content.
- Speaker separation for material containing multiple speakers.
- Availability through ElevenCreative and ElevenProductions, with workflows aimed at creators, marketers, and professional production teams.
The distinction matters because language localization has two related but different requirements: translating meaning and recreating delivery. Dubbing v2 is intended to bring those stages closer together. A marketing campaign, for example, may depend on the pace and emotional emphasis of its original narration, not merely the literal wording. Likewise, a creator's recognizable delivery can be part of the audience experience that a conventional translated voice track may lose.
| Area | Prior V1 workflow | Dubbing v2 |
|---|---|---|
| Core approach | More transcript-centered dubbing | Conditions directly on the original performance |
| Delivery preservation | Did not claim the same performance-transfer fidelity | Designed to retain emotion, tone, pacing, and emphasis |
| Self-serve workflow | Dubbing Studio remains associated with V1 | Automatic Dubbing uses the Dubbing v2 Alpha model |
| Workflow status | Dubbing Studio is in maintenance mode | Available in ElevenCreative and ElevenProductions |
The comparison also clarifies an important rollout limitation. Dubbing v2 is not yet a universal replacement across every ElevenLabs dubbing surface. Supporting documentation associates Automatic Dubbing with Dubbing v2 Alpha, while Dubbing Studio remains on V1 for a more granular workflow. Teams that rely on studio-level controls should therefore assess the current product path rather than assume immediate feature parity.
Availability, pricing, and developer implications
Dubbing v2 is available in ElevenCreative and through ElevenProductions.
ElevenLabs also offers a Creator Dubbing Partner Program with discounted access for eligible creators. For initial use, the company describes a seven-day free usage window with upfront quotas of one minute on Free, 15 minutes on Starter, and 30 minutes on Creator+.
Official documentation sets out per-minute Automatic Dubbing pricing and concurrent-job limits for Free, Starter, Creator, Pro, Scale, Business, and Enterprise plans. Those limits matter for teams processing larger libraries because concurrency can affect how much dubbing work can run at once, independently of the per-minute usage cost.
For developers, the key caveat is timing. ElevenLabs says API access is coming soon, rather than generally available at launch. That means Dubbing v2's immediate use is centered on the company's platform workflows, not direct production API integration. Live or real-time dubbing is also not currently available.
The Alpha label is equally significant. ElevenLabs is presenting a substantial improvement in expressiveness, but users should expect the model to continue evolving and may encounter rough edges. Organizations with high-stakes localization requirements should validate output quality across their target languages, speaker combinations, and content formats before standardizing a workflow.
Organizations evaluating performance-aware dubbing alongside their existing content systems can work with Scalevise on AI workflow architecture, automation, and implementation planning, particularly as developer access and production integration options expand.
Frequently Asked Questions
What is ElevenLabs Dubbing v2?
ElevenLabs Dubbing v2 is an Alpha AI dubbing model that conditions on an original speaker's performance to preserve delivery characteristics in translated audio.
How many languages does Dubbing v2 support?
ElevenLabs says Dubbing v2 supports more than 90 languages.
Is ElevenLabs Dubbing v2 available through an API?
No. ElevenLabs says API access is coming soon, so it was not generally available at launch.
Where can users access Dubbing v2 now?
Dubbing v2 is available in ElevenCreative and through ElevenProductions. Automatic Dubbing uses the Dubbing v2 Alpha model.
Does Dubbing v2 replace Dubbing Studio?
Not currently. Dubbing Studio remains associated with the V1 model and is in maintenance mode, while Dubbing v2 is used in Automatic Dubbing.
Conclusion
ElevenLabs Dubbing v2 moves AI localization beyond transcript-led translation by making the source performance part of the dubbing process. Its focus on expressive delivery, timing, and voice continuity could make multilingual tracks more faithful to their originals, although the Alpha status, lack of live dubbing, and pending API access remain important constraints for production teams.