ElevenLabs launches v4 speech models with 90 language support and lower latency for voice agents

Voice cloning now works from 10 seconds of audio as the startup targets its growing enterprise calling business

By LineZotpaper
Published
Read Time2 min
ElevenLabs released two new speech models on Monday, ElevenLabs v4 and v4 Turbo, promising more expression control, lower latency for voice agents and support for more than 90 languages. The release is the company's first major model refresh since v3 last year and arrives as more than 55% of its business now comes from large companies.

ElevenLabs launched two new speech models on Monday, saying the v4 generation adopts a new architecture that allows for better control and faster cloning. Users will be able to clone a voice with just 10 seconds of audio. The model was teased at an event in Warsaw earlier this year.

On the creative side, the company said v4 handles voice identity better over longer chunks of text and keeps the context of the text in mind while reading it aloud to change expressions. ElevenLabs introduced inline tags to define expression with v3 and is expanding them in v4, letting users stack multiple tags and have the model follow the sequence.

Language support has grown from 70 to more than 90 languages. The startup said it observed the biggest quality jumps in Japanese, Brazilian Portuguese, Mandarin and Cantonese.

The new model is also pitched at voice agents. ElevenLabs said v4 has lower latency to allow for more fluid conversation and can start generating audio as soon as the LLM behind it starts generating answers. The model can also handle confrontations, escalations and holds differently for better issue resolution.

Competition in speech models has intensified, with startups including Cartesia, Deepgram, Fish Audio, Boson and WellSaid Labs building expressive speech models, while Google and OpenAI have improved their own voice offerings.

On the business side, ElevenLabs raised $500 million from Sequoia earlier this year at an $11 billion valuation, and there are rumors of a follow-up round that would value the company at $22 billion. Its annualized revenue run rate has climbed from roughly $330 million at the start of the year to over $600 million, and headcount has surpassed 800, with aggressive hiring in India, Europe and Brazil. In a recent interview with TechCrunch, co-founder and CEO Mati Staniszewski said the company is aiming for an IPO "in the next years," without committing to a timeline.

§

Analysis

Why This Matters

  • The update targets voice agents, a fast-growing enterprise use case, and could raise the bar for conversational AI audio quality and latency.
  • ElevenLabs has scaled quickly, and the rumored $22 billion valuation makes this release a test of whether the technology keeps pace with its commercial ambitions.
  • Wider language coverage and stacked expression tags extend synthetic speech to more creators and markets, potentially accelerating adoption.

Background

ElevenLabs is a synthetic speech company known for text-to-speech, voice cloning and audio tools used by creators, publishers and businesses. The voice AI field has become crowded, with expressive speech models deployed in audiobooks, dubbing, marketing and increasingly automated customer service. Voice agents are a key battleground because they demand fast, natural responses that can handle difficult conversations. Last year's v3 model introduced inline tags for directing tone and delivery; v4 builds on that approach with a new architecture aimed at both creative work and real-time interaction.

Key Perspectives

ElevenLabs: The company positions v4 as a step forward for creative expression and enterprise voice agents, pointing to lower latency, stacked expression tags and better handling of confrontations, escalations and holds. Enterprise customers: Large companies, now more than 55% of ElevenLabs' business, stand to gain more natural customer-facing voice agents, but may weigh the cost and reliability of upgrading. Competitors and big tech: Startups such as Cartesia, Deepgram, Fish Audio, Boson and WellSaid Labs, along with Google and OpenAI, are all improving voice models, keeping pressure on ElevenLabs to stay ahead. Skeptics: The rumored $22 billion valuation is double the $11 billion round from earlier this year. With a revenue run rate above $600 million, the price assumes continued rapid growth and durable technical leadership in a market where rivals are closing in.

What to Watch

  • Whether the rumored follow-up funding round at a $22 billion valuation is confirmed and who leads it.
  • Adoption of v4 in enterprise voice agent deployments, especially around latency and issue resolution.
  • The pace of quality improvements in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
  • Any concrete IPO timeline, given Staniszewski has said an offering is "in the next years" without committing to a date.

Sources

Zotpaper

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.