Startups

ElevenLabs Ships v4 Speech Models as Revenue Run Rate Tops $600 Million

ElevenLabs launched v4 and v4 Turbo with 90-language support and 10-second voice cloning, as annualized revenue run rate climbs past $600 million.

By Daniel Okafor

2 min read

Updated

ElevenLabs’ new v4 speech model supports more expression control and 90 languages
ElevenLabs’ new v4 speech model supports more expression control and 90 languagesjenschapter3 / Openverse

What's News

  • ElevenLabs launched v4 and v4 Turbo on Monday, supporting 90+ languages, up from 70
  • Annualized revenue run rate grew from ~$330 million to over $600 million this year
  • The company raised $500 million led by Sequoia at an $11 billion valuation; a rumored follow-up round would value it at $22 billion

ElevenLabs launched two new speech models on Monday, ElevenLabs v4 and v4 Turbo, as its annualized revenue run rate has climbed from roughly $330 million at the start of the year to over $600 million.

The new generation brings more expression control, lower latency for voice agents, and support for more than 90 languages — up from 70 in the previous version. The company said it observed the biggest quality jump in Japanese, Brazilian Portuguese, Mandarin and Cantonese.

ElevenLabs released its v3 model last year and teased the successor at an event in Warsaw earlier this year. For v4, the startup adopted a new architecture that allows for better control and faster cloning. With v4, users can clone a voice from just 10 seconds of audio.

Longer Text, Sharper Expression

On the creative side, the model handles voice identity better over longer chunks of text. It also keeps the context of the text in mind while reading aloud, changing expressions to match the material. ElevenLabs introduced inline tags to define expression with v3 and is expanding them in v4: users can now stack multiple tags, and the model follows the sequence.

Built for Enterprise Calling

The enterprise calling business has scaled rapidly over the last year. More than 55% of ElevenLabs' business now comes from large companies, and the company positioned v4 squarely at that segment.

The new version has lower latency to allow for more fluid conversation. The v4 model can start generating audio as soon as the LLM behind it starts generating answers. The model can also handle confrontations, escalations and holds differently, which the company said enables better issue resolution — a direct pitch at customer-service deployments.

A Crowded Field

Competition in speech models has ramped up. Startups including Cartesia, Deepgram, Fish Audio, Boson and WellSaid Labs have built expressive speech models, while Google and OpenAI have also improved their voice offerings.

ElevenLabs carries substantial financial weight into that fight. The company raised $500 million earlier this year in a round led by Sequoia that valued it at $11 billion. Rumors already circulate of a follow-up fundraising round that would value the company at $22 billion.

The startup has also hired aggressively across markets such as India, Europe and Brazil, growing headcount to over 800 people.

The IPO Question

In a recent interview with TechCrunch, co-founder and CEO Mati Staniszewski said the company is aiming for an IPO "in the next years," but didn't commit to a timeline.

With revenue nearly doubling since January, a doubling of valuation rumored in a potential new round, and v4 aimed squarely at the enterprise calling market that now drives most of its business, ElevenLabs is positioning itself as the scale player in a speech-AI field crowded with both startups and the largest labs.

Original: youtube.com

Share this article:

More from Daniel Okafor

Daniel Okafor

Show full bio

Correspondent covering business strategy at Business Bearings.

276 articles

Related articles

« Previous articleNext article »