ElevenLabs Scales to $600 Million ARR While Defending Margins Against Customer Competition
The voice AI startup now valued at $22 billion is navigating a blurring line between platform and customer, as enterprises it once trained now build rival models.
The $22 Billion Voice Layer
ElevenLabs has reached $600 million in annual recurring revenue, a milestone that underscores how quickly voice synthesis has moved from novelty to infrastructure. The company builds models that convert text into human-sounding speech, technology now embedded in customer service operations at Klarna, Deutsche Telekom, Cisco, and Adobe. Klarna alone routes first-line phone support for 35 million users in the United States through ElevenLabs' platform.
The four-year-old startup is now valued at $22 billion, according to people familiar with recent funding discussions. Co-founder and CEO Mati Staniszewski says the business splits roughly 55 per cent enterprise and 45 per cent small and medium businesses, developers, and creators who use the platform for audiobooks, dubbing, and music production.
At Opentechwire, we have tracked how voice AI has shifted from a research curiosity to a core layer in conversational systems. ElevenLabs' revenue velocity suggests enterprises are no longer experimenting but committing budget at scale. The question is whether that momentum can withstand two pressures: commoditisation and customer defection.
The Commoditisation Clock
Staniszewski appeared at a conference in Toronto in late September 2026 and acknowledged a prediction he made a year earlier. He had said audio models would become commoditised within a couple of years. His revised view is more cautious. Quality differences at the model level remain significant, he said, and the gap may persist for three to five years.
The company's stated ambition is to pass the Turing test for conversational AI, a goal that requires not only intelligible speech but emotional intelligence. Systems need to detect sentiment, adjust pacing, and modulate tone in real time. Staniszewski said that has not yet been achieved by any platform.
Yet the commoditisation thesis is not hypothetical. Open-weight models are closing the gap, particularly in lower-stakes use cases. ElevenLabs itself offers customers a menu of reasoning layers, from proprietary frontier models to open-source alternatives. Staniszewski said the choice depends on the application. For informational queries with no transaction risk, open-source models suffice because the knowledge base defines the experience. For financial services, where authentication and error-free execution are required, frontier models still lead.
Customers as Competitors
A more immediate challenge is the blurring boundary between platform and application. Decagon, a conversational AI platform, initially trained its voice product on ElevenLabs. It now runs queries through its own models, effectively competing with the vendor that helped it get started.
Staniszewski described the dynamic as a structural shift in AI. The lines between model companies, platform companies, and application companies have become unclear. Anthropic, once purely a model provider, now operates as a platform and is expanding into applications. He expects this pattern to continue.
The risk for ElevenLabs is that enterprise customers with sufficient scale and technical capability will build in-house alternatives. The company is betting that breadth of use cases, quality differentiation, and speed of iteration will keep defection rates manageable. But the funding rounds we have followed across the region show that conversational AI startups are raising capital specifically to verticalise and own the full stack.
Government Deployments and Model Choices
ElevenLabs counts the US government and multiple European governments as customers. Staniszewski said each deployment is customised based on regulatory and operational requirements. Some governments use open-weight models; others use closed-source or fine-tuned variants. Data residency and sovereignty are negotiated case by case.
In Poland, the company is working with the public health system to reduce no-show rates for medical appointments, which run at 18 per cent. ElevenLabs deployed agents that call patients with reminders. The system uses models optimised on local healthcare knowledge, integrated while maintaining data residency within Poland.
The presence of Chinese open-weight models in ElevenLabs' menu raises questions about supply chain and security review processes for government contracts. Staniszewski did not detail vetting procedures but said the company applies know-your-customer checks to every enterprise deployment and has safeguards against self-replicating or recurrent agent behaviour.
Disclosure and the Agent Economy
Staniszewski said businesses should disclose when a customer is speaking to an AI agent rather than a human. His reasoning is pragmatic. People are not yet accustomed to synthetic voices in service contexts, and undisclosed automation risks eroding trust.
He expects that norm to shift within five years, when individuals have their own agents acting on their behalf. At that point, he said, people will expect to interact with agents and the disclosure question will become moot.
In the interim, ElevenLabs recommends that customers offer a choice. If the wait time for a human operator is 30 minutes, present the agent as an alternative. Staniszewski said most customers choose the agent and are surprised by the quality of the interaction.
Margins, Training Data, and the IPO Question
Staniszewski declined to provide specific figures on gross margins but said the company is willing to compress them further if it accelerates market share growth. ElevenLabs benefits from in-house research capabilities that allow it to fine-tune and constrain models efficiently, but the priority is proving value to customers. If passing on cost savings strengthens those relationships, the company will do it.
On training data, Staniszewski said ElevenLabs has millions of hours of customer service calls. In some cases, the company has co-developed models with enterprise clients for specific use cases. The core challenge has not been data volume but annotation. ElevenLabs employs thousands of contract annotators who label not only what was said but how it was said, including timing, emotion, and accent. The company brought in voice coaches to improve accent detection accuracy.
A portion of the training data is synthetic, though Staniszewski did not quantify it. The approach mirrors trends we have seen in other model companies, where synthetic data generation is used to fill gaps in labelled corpora.
Asked about IPO timing, Staniszewski said the company is preparing the foundation to go public in the next few years but has not committed to a specific date. Reports have pointed to 2028, but he would not confirm. The decision will depend on market conditions and the company's readiness.
Precautions and Exposure
Staniszewski said there is alignment among AI companies on the need to pace deployment responsibly, though he stopped short of endorsing public slowdowns or specific regulatory interventions. ElevenLabs does not train text models or the intelligence layer of conversational systems, which insulates it from some of the debates around frontier model risk.
He said the company is less exposed than Hugging Face to misuse because ElevenLabs does not support self-replicating or recurrent agent logic. Customers cannot use the platform to create agents that spawn other agents. Cybersecurity risk remains a concern for the wider industry, but Staniszewski said the company has implemented precautions including mandatory KYC for enterprise accounts.
ElevenLabs is betting that voice synthesis remains a defensible layer even as the rest of the conversational AI stack consolidates. The $600 million ARR figure suggests enterprises agree, at least for now. Whether that holds as customers build their own models, and as open-weight alternatives improve, will determine whether the $22 billion valuation was prescient or premature.



