On September 28, 2026, ElevenLabs launched Eleven v4 and Eleven v4 Turbo, a new generation of speech-generation models designed to make AI voices more expressive, context-aware and consistent across longer conversations. The models expand language support from more than 70 languages in the previous generation to more than 90 languages, while introducing improved voice cloning and more granular control over how generated speech should sound.

The launch also targets one of the biggest limitations of AI voice agents: the trade-off between natural-sounding speech and response speed. Eleven v4 is positioned as the higher-quality model for expressive speech, while Eleven v4 Turbo is designed for real-time applications, with ElevenLabs reporting median time to first speech of about 150 milliseconds in its September test, with network delay excluded. That figure describes Turbo over WebSocket, not the full time a user waits for an agent to answer.

Red speech waves pass through a voice chip and spread into multiple languages, illustrating ElevenLabs v4
Concept illustration of voice synthesis across languages; model performance figures below are company reported.

Key takeaways

  • ElevenLabs has launched Eleven v4 and Eleven v4 Turbo.
  • Both models support 90+ languages, compared with 70+ for Eleven v3.
  • V4 provides greater control over emotion, pacing, tone, character and delivery.
  • Users can use inline instructions such as laughter, whispers and sound effects to influence delivery.
  • Instant Voice Clones can be created from just 10 seconds of audio, according to ElevenLabs.
  • V4 is designed for creative and long-form production, while V4 Turbo focuses on real-time voice agents.
  • ElevenLabs says V4 improves speaker consistency in longer generations and multi-speaker conversations.
  • The launch comes as the company rapidly expands its enterprise voice-agent business.

ElevenLabs V4 focuses on how AI speaks, not just what it says

Traditional text-to-speech systems primarily solve a straightforward problem: converting written words into audio. The challenge becomes much harder when the system needs to understand how those words should be delivered.

The same sentence can sound reassuring, threatening, excited, sarcastic or uncertain depending on the speaker, situation and surrounding conversation.

ElevenLabs says V4 has been designed around this problem. Its new architecture is intended to interpret tone, pacing, emotion, character and context, allowing generated speech to change depending on the situation rather than simply reading a script in a fixed manner.

This matters particularly for applications where speech is part of the experience rather than simply an accessibility feature.

An audiobook narrator, game character, advertising voice, virtual assistant and customer-service agent may all use the same underlying speech technology, but they need very different delivery styles.

Eleven v4 is therefore positioned as a model for situations where the emotional quality of the voice is as important as the words themselves.

More detailed control over voice expression

One of the biggest changes in V4 is the amount of control available to users.

As TechCrunch reported on September 28, ElevenLabs previously introduced inline tags that could influence speech delivery. With V4, the company says users can use more detailed instructions and combine multiple directions in sequence.

For example, a script can include instructions for a speaker to laugh, whisper, become angry or change the way a particular phrase is delivered. The system can also interpret instructions involving sound effects.

That allows a creator to provide something closer to a director’s instructions rather than simply entering a block of text.

The company’s examples include tags for actions such as laughter, whispered delivery and environmental sounds. ElevenLabs says V4 follows these directions more accurately than earlier versions.

This could be particularly important for media production.

Instead of generating individual sentences and manually editing the emotion into each recording, creators can increasingly describe the desired performance directly in the prompt or script.

The practical effect is a shift from text-to-speech toward something closer to text-to-performance.

V4 improves consistency for longer audio

Another problem with generative voice systems is consistency.

A short voice sample may sound convincing, but maintaining the same speaker identity over an audiobook, advertisement or extended dialogue can be more difficult. Small differences in pronunciation, tone or vocal characteristics can become noticeable as the amount of generated material increases.

ElevenLabs says V4 introduces improvements designed to preserve speaker identity across longer generations.

The company says the model uses a new approach for capturing speaker identity and can maintain the characteristics of a voice across narration, dialogue, advertisements and conversational agents.

The model is also designed to handle conversations between multiple speakers more naturally.

Rather than treating every sentence as an independent audio-generation task, V4 can use the broader scene or conversational context to determine how one speaker should respond to another.

That could make AI-generated conversations sound less like a collection of separately generated lines and more like an actual exchange.

ElevenLabs model language coverageEleven v3 supported over 70 languages; Eleven v4 and v4 Turbo support over 90, according to ElevenLabs.Language coverage, by model familyCompany-reported minimums; counts are not quality scoresEleven v370+Eleven v490+Sources: ElevenLabs model documentation and Sept 28 announcement
ElevenLabs says the v4 family supports more than 90 languages; coverage alone does not establish equal performance in each language.

V4 expands language support beyond 90 languages

Language coverage is another major part of the release.

ElevenLabs says the V4 model family supports more than 90 languages, compared with more than 70 supported by V3. The company’s documentation lists languages ranging from English, Spanish and French to Hindi, Bengali, Gujarati, Marathi, Tamil, Telugu, Urdu, Japanese, Mandarin, Korean and Cantonese.

The significance is not simply the number of languages.

For voice AI to work effectively across international markets, a system needs to reproduce more than the correct words. Pronunciation, rhythm, accent, pacing and emotional delivery also need to remain natural.

ElevenLabs says V4 improves how speech captures rhythm, emotion and delivery across languages. It also says a voice can speak in another supported language while maintaining the original speaker’s identity and adopting the appropriate native accent.

The company specifically highlighted improvements in Japanese, Brazilian Portuguese, Mandarin and Cantonese.

For businesses operating across multiple markets, this capability could reduce the need to create separate voice identities and production processes for every language.

Ten seconds of audio can create an Instant Voice Clone

ElevenLabs is also improving its voice-cloning capabilities with V4.

The company says its Instant Voice Clones can capture a voice with high fidelity using only 10 seconds of audio. It also says V4 improves speaker similarity and maintains voice identity more reliably across generated material.

The technology could make voice production substantially faster.

A traditional voice production workflow may require a professional actor, recording session, editing and additional recordings whenever a line needs to be changed.

With voice cloning, authorized users can potentially generate new dialogue without bringing the original speaker back into the studio for every modification.

That has obvious applications in games, audiobooks, advertising, localization and digital assistants.

However, the same capability also increases the importance of consent and identity protections around synthetic voices. A voice clone can represent a person’s identity, making authorization and responsible use important considerations as these systems become easier to operate.

Eleven v4 Turbo targets real-time AI agents

The second model introduced alongside V4 is Eleven v4 Turbo.

While V4 is positioned around high-quality expressive speech, Turbo is designed for applications where the AI needs to respond during a live interaction.

ElevenLabs reports median time to first speech of approximately 150 milliseconds for V4 Turbo over WebSocket in its September testing. Its methodology removed network delay, so this figure does not measure the full response time experienced by a caller. The company’s model documentation separately lists roughly 100 milliseconds of median inference latency; the two metrics measure different intervals.

That distinction is important for voice agents.

A person talking to a customer-service agent does not expect the system to generate an entire answer before beginning to speak. Long pauses can make an automated system feel slow or broken.

Lower latency allows the agent to begin responding more quickly, making the interaction feel closer to a human conversation.

ElevenLabs says V4 Turbo has been optimized together with its ElevenAgents platform rather than treating the speech model and agent system as completely separate components.

The company’s documentation identifies customer-support agents, AI assistants and interactive characters as potential applications for V4 Turbo.

What the Turbo latency metric measuresA request passes through network delivery, model inference, speech streaming and playback. ElevenLabs reports about 150 milliseconds from request to first audible speech in a controlled WebSocket comparison with network delay removed.What does “150 ms” measure?ElevenLabs’ September 2026 test of v4 Turbo over WebSocketRequestNetworkSpeech startsFull reply~150 ms: request to first speech, with network delay removedIt is not a measured end-to-end caller response time.
The company’s 150 ms figure excludes network latency and should not be read as the complete time to answer a user.

Why voice agents are becoming strategically important

The V4 launch is taking place at a time when the voice-AI market is moving beyond demonstrations and toward business operations.

Companies are increasingly using AI agents for customer service, sales, appointment scheduling, support and other repetitive conversations.

ElevenLabs said in May that enterprises were deploying its voice agents across customer support, sales, hiring and marketing operations. The company reported that it had surpassed $500 million in annualized recurring revenue during the first four months of 2026, after ending 2025 at $350 million ARR.

By September, the scale had increased further.

Reuters reported that ElevenLabs’ AI agents were handling more than 15 million conversations per week, three times the level reported in February. The company said its agents were being used for activities including refunds, insurance renewals and appointment bookings.

That context helps explain why ElevenLabs is emphasizing both expression and latency.

For consumer-facing voice agents, sounding natural is not enough if the system responds too slowly. Conversely, a fast agent that sounds robotic can make an interaction frustrating.

V4 and V4 Turbo are essentially attempts to address both sides of that problem.

ElevenLabs is also scaling rapidly as a business

The model launch arrives during a major period of growth for ElevenLabs. Our coverage of the ElevenLabs–UMG licensing deal covers a separate music development.

In February, the company announced a $500 million Series D funding round at an $11 billion valuation. It said the round brought its total funding to $781 million and would support expansion of ElevenAgents, research and international operations.

The company’s valuation has since moved substantially higher.

On September 30, ElevenLabs announced a $300 million employee tender offer at a $22 billion valuation, according to Reuters. Unlike a conventional funding round, the transaction primarily provides liquidity by allowing employees and existing shareholders to sell shares rather than simply injecting new capital into the company.

The valuation increase reflects investor expectations around the broader AI-agent market, particularly the potential for AI systems to handle real-world interactions rather than simply generate text or images.

It also raises the competitive stakes for ElevenLabs.

The competition in AI voice is intensifying

ElevenLabs is not operating in an empty market.

The company faces competition from specialized speech startups including Cartesia, Deepgram, Fish Audio, Boson and WellSaid Labs, while larger technology companies such as Google and OpenAI are also developing increasingly capable voice systems.

The competitive environment is changing the definition of a good speech model.

Early AI voice systems were largely judged on whether they sounded understandable and human-like.

Today’s systems are increasingly being evaluated on a broader set of characteristics:

CapabilityWhy it matters
Natural expressionMakes generated speech feel less robotic
Context awarenessHelps voices respond appropriately to conversations
Voice consistencyImportant for long-form audio and characters
Language coverageEnables global deployment
Low latencyCritical for real-time agents
Voice cloningEnables scalable personalized audio
Fine-grained controlGives creators more control over delivery

The combination is becoming particularly important for enterprise applications, where a company may want one AI voice system to operate across customer service, marketing, localization and internal workflows.

India could be an important market for multilingual voice AI

The expansion to more than 90 languages is also significant for markets such as India.

India’s linguistic diversity creates a different challenge from markets dominated by one or two major languages. A voice agent that works well in English but struggles with Hindi, Bengali, Marathi, Tamil, Telugu or other widely used languages has limited reach.

ElevenLabs’ V4 language list includes Hindi, Bengali, Gujarati, Marathi, Tamil, Telugu and Urdu, among many other languages.

The company has also been expanding its presence in India. Its February funding announcement listed Bengaluru among its international locations, while TechCrunch reported that the company had been hiring across markets including India, Europe and Brazil.

For readers tracking related Lapaas Voice coverage, the earlier ElevenLabs Dubbing v2 API report concerns a separate product. For Indian businesses, multilingual AI agents could eventually be useful in customer support, banking, commerce, healthcare, government services and education.

The more difficult challenge will be ensuring that improved language coverage translates into consistently natural regional speech rather than simply adding more languages to a checklist.

What the V4 launch means for creators

For creators, V4 could reduce the gap between traditional voice production and generative audio.

Audiobook producers can use more consistent narration. Game developers can create characters with different emotional responses. Advertisers can produce localized campaigns. Media companies can experiment with dubbing without recording every version manually.

The ability to give detailed performance instructions also gives creators more control over the final output.

Instead of selecting from a small collection of predefined emotional voices, users can increasingly specify how a particular sentence should sound.

That is an important change because the value of synthetic speech is moving away from simply replacing a human recording and toward enabling new types of production workflows.

What the V4 launch does not solve

Despite the improvements, V4 does not eliminate every challenge surrounding AI-generated voices.

Voice cloning still raises questions about consent, impersonation and unauthorized use. Greater realism can make fraudulent or misleading audio more convincing.

There are also technical questions around how consistently the models perform across all 90-plus supported languages and accents. ElevenLabs’ strongest claims about preference and benchmark performance are based on the company’s own testing or cited evaluations, so they should not automatically be treated as universal evidence that V4 is better in every scenario. ElevenLabs says V4 was preferred by about 75% of listeners in its blind head-to-head testing against selected competing systems, while its documentation identifies Artificial Analysis as a source for its ranking claim.

Independent testing across a wider range of languages, accents, use cases and real-world conversations will provide a clearer picture of how large the practical improvement is.

The Bigger Picture

The larger story behind ElevenLabs V4 is that voice AI is becoming an interface layer for AI agents.

Text-based AI can generate an answer, but a voice agent needs to deliver that answer at the right speed, in the right tone and with enough contextual awareness to keep the conversation moving.

That makes speech generation increasingly important to the overall AI-agent stack.

ElevenLabs is trying to position itself around that opportunity by combining expressive speech models, voice cloning, multilingual capabilities and an agent platform rather than competing only as a text-to-speech API provider.

Its recent growth suggests investors and enterprises see voice agents as more than an experimental technology. The company says enterprise adoption is already driving a substantial share of its business, while its agents are handling millions of conversations each week.

Looking Ahead

The next test for ElevenLabs will be whether V4 can turn better speech quality and lower latency into sustained enterprise adoption. The company will need to demonstrate that its models remain reliable across languages, industries and long conversations while maintaining the speed required for real-time interactions.

The competitive race will also move beyond simply making AI voices sound human. As OpenAI, Google and specialized voice startups continue improving their systems, the differentiator may increasingly become the complete agent platform: how quickly an AI can understand a user, decide what to say, respond naturally and complete a task. ElevenLabs V4 gives the company a stronger speech layer for that race, but the broader voice-agent market is still developing.

What independent testing does and does not show

The announcement is dated September 28, rather than October 6. TestingCatalog’s launch report describes v4 as the production-oriented model and Turbo as the streaming option. Its latency comparisons repeat company measurements. They are useful for understanding the product claim, but they are not an independent test of call-center response time.

A narrower independent check came from AI Tool Radar’s September 30 German-language test. The writer used one German voice, one script and three runs per model. The median time to generate a full sentence was 15.8 seconds for v3, 7.9 seconds for v4 and 2.5 seconds for Turbo. All three transcribed without error in that small trial. The tester explicitly cautioned that the setup used REST and included network delay, so those results cannot be compared directly with ElevenLabs’ WebSocket figure. It also does not validate Hindi or other Indian-language quality.

For an Indian customer-support deployment, the relevant pilot would use the actual supported language, local accent, mixed-language phrases, background noise, telephony connection and live agent workflow. A team should measure the full interval from the customer finishing a turn to audible speech, as well as task completion and escalation quality. The model’s language list confirms availability; it does not by itself establish performance on those tasks.

Sources and method

Questions readers may have

When did ElevenLabs launch v4?

ElevenLabs announced Eleven v4 and Eleven v4 Turbo on September 28, 2026. This article was published on October 6 and explains that earlier release.

Does support for Hindi mean it will sound natural in every Indian accent?

No. ElevenLabs lists Hindi and several other Indian languages as supported, but the cited independent hands-on check was in German. Regional accents and code-switching need their own tests.

Is 150 milliseconds the full time to answer a call?

No. ElevenLabs measured median time to first speech for v4 Turbo over WebSocket with network delay removed. A caller’s full wait also depends on transcription, reasoning, telephony and other network steps.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.