ElevenLabs has launched the Dubbing v2 API, a major upgrade to its AI-powered dubbing platform that enables developers and enterprises to translate audio and video into more than 90 languages while preserving the original speaker’s emotion, tone, pacing, timing, and vocal identity. The company says the new API addresses one of the biggest challenges in AI localization—making translated speech sound as if it were naturally spoken by the original speaker rather than generated from a flat text translation.

Unlike conventional dubbing systems that translate transcripts into synthetic speech, Dubbing v2 directly conditions on the original audio performance. This allows the model to recreate subtle vocal characteristics such as emotional delivery, pauses, emphasis, and speaking rhythm across multiple languages. The API is designed for developers building multilingual media, enterprise communication tools, e-learning platforms, gaming experiences, and global content distribution services.

ElevenLabs Introduces Dubbing v2 API

The new API enables organizations to automate multilingual content localization while maintaining the authenticity of the original speaker.

Key capabilities include:

  • AI dubbing in 90+ languages.
  • Preservation of emotion and vocal performance.
  • Automatic voice cloning.
  • Synchronization with the speaker’s original timing and pacing.
  • API integration for developers and enterprise workflows.

Dubbing v2 API Snapshot

FeatureDetails
CompanyElevenLabs
ProductDubbing v2 API
Language Support90+ languages
Key DifferentiatorEmotion-preserving AI dubbing
Primary UsersDevelopers, enterprises, media companies

Performance-Based AI Dubbing

A major advancement in Dubbing v2 is its ability to model the original speaker’s performance rather than relying solely on translated text.

According to ElevenLabs, the system preserves:

  • Emotional delivery.
  • Speaking rhythm.
  • Tone of voice.
  • Timing and pauses.
  • Vocal identity through automatic voice cloning.

This allows translated speech to sound significantly more natural, especially for expressive content such as interviews, educational videos, podcasts, films, and creator content.

Traditional AI Dubbing vs Dubbing v2

Traditional AI DubbingElevenLabs Dubbing v2
Transcript-based speech generationConditions directly on original performance
Limited emotional expressionPreserves emotion and delivery
Generic synthesized voicesAutomatic voice cloning
Less natural pacingMaintains original timing and rhythm

Designed for Global Content Localization

The API is intended to help organizations scale multilingual content production.

Potential use cases include:

  • Video localization.
  • Film and television dubbing.
  • Online education.
  • Corporate training.
  • Podcasts.
  • Marketing campaigns.
  • Creator content.
  • Gaming dialogue.
  • Customer support videos.

Developers can integrate the API into existing workflows to automate translation and dubbing without requiring traditional recording sessions for every language.

Automatic Voice Cloning

One of the platform’s key capabilities is automatic voice cloning.

Instead of requiring manual voice creation, the system:

  • Identifies the original speaker.
  • Recreates vocal characteristics.
  • Maintains pitch and tone.
  • Preserves speaker identity across translated languages.

This enables audiences in different regions to hear localized content while retaining the recognizable voice of the original presenter or actor.

Why It Matters

Demand for multilingual content is growing rapidly across streaming, enterprise communications, online education, and creator platforms.

Emotion-preserving dubbing offers several advantages:

  • Faster localization.
  • Lower production costs.
  • More consistent global branding.
  • Improved viewer engagement.
  • Reduced dependence on traditional dubbing studios.

For creators and enterprises producing content at scale, AI-powered dubbing can significantly shorten production timelines while expanding international reach.

Competition in AI Voice Technology

The launch strengthens ElevenLabs’ position in the rapidly evolving AI voice market, where companies are increasingly competing to deliver more natural, multilingual speech generation.

Industry priorities now include:

  • Emotional speech synthesis.
  • Real-time voice translation.
  • Voice cloning.
  • Low-latency speech generation.
  • Enterprise-scale localization.

Rather than focusing solely on text-to-speech quality, providers are increasingly investing in preserving authentic human performance across languages.

Looking Ahead

The Dubbing v2 API represents a significant step forward in AI-powered content localization by shifting the focus from simple translation to preserving the emotional performance of the original speaker. With support for more than 90 languages, automatic voice cloning, and synchronization of tone, pacing, and delivery, ElevenLabs is positioning its platform as a comprehensive solution for developers, creators, and enterprises seeking to produce high-quality multilingual content at scale.

Looking ahead, demand for AI dubbing is expected to accelerate as streaming platforms, global businesses, educational institutions, and digital creators increasingly localize content for international audiences. As competition intensifies in AI voice technology, advances in emotion preservation, voice authenticity, and production automation are likely to become key differentiators, helping transform multilingual media production from a manual process into an AI-driven workflow.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.