ElevenLabs launched two new speech models on Monday, saying the v4 generation adopts a new architecture that allows for better control and faster cloning. Users will be able to clone a voice with just 10 seconds of audio. The model was teased at an event in Warsaw earlier this year.
On the creative side, the company said v4 handles voice identity better over longer chunks of text and keeps the context of the text in mind while reading it aloud to change expressions. ElevenLabs introduced inline tags to define expression with v3 and is expanding them in v4, letting users stack multiple tags and have the model follow the sequence.
Language support has grown from 70 to more than 90 languages. The startup said it observed the biggest quality jumps in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
The new model is also pitched at voice agents. ElevenLabs said v4 has lower latency to allow for more fluid conversation and can start generating audio as soon as the LLM behind it starts generating answers. The model can also handle confrontations, escalations and holds differently for better issue resolution.
Competition in speech models has intensified, with startups including Cartesia, Deepgram, Fish Audio, Boson and WellSaid Labs building expressive speech models, while Google and OpenAI have improved their own voice offerings.
On the business side, ElevenLabs raised $500 million from Sequoia earlier this year at an $11 billion valuation, and there are rumors of a follow-up round that would value the company at $22 billion. Its annualized revenue run rate has climbed from roughly $330 million at the start of the year to over $600 million, and headcount has surpassed 800, with aggressive hiring in India, Europe and Brazil. In a recent interview with TechCrunch, co-founder and CEO Mati Staniszewski said the company is aiming for an IPO "in the next years," without committing to a timeline.