ElevenLabs Dubbing v2 Adds Accent and Audio Controls for Video Localization
ElevenLabs has introduced Dubbing v2, expanding its AI localization workflow with performance preservation, accent controls, multi-speaker support, and background-audio handling.
ElevenLabs has introduced Dubbing v2, a rearchitected AI dubbing model built to retain a speaker's emotion, delivery, and timing while translating content across more than 90 languages. The update expands the company’s localization proposition beyond a basic language replacement: its documented controls cover dialect-specific accents, multi-speaker material, and background audio management for more complex video and audio scenes.
According to ElevenLabs’ Dubbing v2 announcement, the model is designed for creators, marketers, studios, and broadcasters that need to localize video at production scale. It is integrated with ElevenCreative for one-click video localization and ElevenProductions, the company’s professional localization service.
The central goal is preserving the original performance rather than simply generating translated speech. That matters for material in which pacing, vocal emphasis, and emotional delivery are part of the message, including marketing campaigns, creator videos, and professionally produced programming. Dubbing v2 is intended to synchronize translated dialogue with the source speaker’s timing and delivery across its supported languages.
What Dubbing v2 changes for localization workflows
The most practical additions are the controls documented for ElevenLabs' dubbing workflow. They give teams more ways to shape a dub around the source material and the intended audience, particularly when a project includes regional language variation or a mix of dialogue and sound.
| Localization need | Dubbing v2 capability | Documented control or workflow |
|---|---|---|
| Regional language variation | Dialect-specific accent selection | target_accent, marked experimental |
| Scenes with several voices | Multi-speaker dubbing | num_speakers |
| Music, effects, or ambient sound | Background audio management | foreground_audio_file, background_audio_file, and drop_background_audio |
| End-to-end video localization | Integrated production workflows | ElevenCreative and ElevenProductions |
For Spanish-language localization, the accent capability is especially relevant. ElevenLabs highlights sharper locale accuracy for regional accents such as Castilian Spanish and Latin American Spanish. The documented target_accent field provides a mechanism for applying dialect-specific accents, although ElevenLabs labels that control experimental. Teams should therefore treat it as a useful production option that still merits review against their own editorial and brand requirements.
The update also addresses a common limitation in automated dubbing: real source media rarely consists of one clean voice track. Interviews, shows, advertisements, and social content can include multiple speakers, music, effects, and ambient audio. The num_speakers control and separate foreground and background audio inputs indicate that Dubbing v2 is designed to accommodate these more complicated inputs. The drop_background_audio option gives users a documented way to remove background sound when that is appropriate for the output.
Why audio separation matters
Separating foreground speech from background material can improve control over the localized result. A team may want dialogue translated while retaining music and effects, or may decide that background audio should be removed for a particular version. Those are distinct workflow choices, and the documented parameters make them explicit rather than treating every source file as a single undifferentiated audio track.
That does not remove the need for quality assurance. Accent choices, speaker identification, translation timing, and background-sound treatment all affect whether a localized version fits its intended audience. Dubbing v2 provides more control points, but organizations remain responsible for reviewing outputs in the contexts where accuracy, tone, and production quality matter most.
Integration and developer outlook
ElevenLabs positions Dubbing v2 as part of a broader production stack. ElevenCreative offers one-click video localization, while ElevenProductions is aimed at professional localization services. The company also says Dubbing v2 can participate in its Creator Dubbing Partner Program.
For developers, the announcement says API access is coming soon. Supporting dubbing API documentation already describes controls including target_accent, num_speakers, and foreground and background audio fields for the endpoint used to dub a video or audio file. The announcement does not provide a date or detailed rollout terms for broader API access, so teams planning automated integrations should watch ElevenLabs for availability details.
Organizations evaluating how to connect AI dubbing to content operations can work with Scalevise on AI workflow automation, API integration, and localization-ready implementation.
Frequently Asked Questions
What is ElevenLabs Dubbing v2?
ElevenLabs Dubbing v2 is a rearchitected AI dubbing model designed to preserve a source speaker's emotion, delivery, and timing while translating content across more than 90 languages.
Can Dubbing v2 apply regional accents?
Yes. ElevenLabs documents an experimental target_accent field for dialect-specific accents and highlights improved locale accuracy for accents including Castilian and Latin American Spanish.
How does Dubbing v2 handle multiple speakers and background audio?
The documented dubbing controls include num_speakers for multi-speaker material, plus foreground and background audio inputs and a drop_background_audio option for managing sound layers.
Is ElevenLabs Dubbing v2 available through an API?
ElevenLabs says API access is coming soon. Its supporting dubbing API documentation describes the relevant accent, speaker, and audio-management controls, but the announcement does not specify a broader API rollout date.
Conclusion
Dubbing v2 moves ElevenLabs' localization offering toward more production-aware AI dubbing. By combining performance preservation with documented controls for regional accents, multiple speakers, and background sound, the update gives localization teams more flexibility for real-world media while leaving review and rollout decisions firmly in their hands.