Product timeline
Voice Conversion
Research demo. Early proof the speech-to-speech tech could work.
Voice Conversion
Research demo. Early proof the speech-to-speech tech could work.
“The first AI that can laugh”
Emotion research that became the company's expressiveness thesis.
“The first AI that can laugh”
Emotion research that became the company's expressiveness thesis.
Public launch — Eleven (Monolingual) v1
The launch. First human-like English TTS on a beta web platform. The date ElevenLabs dates itself to.
Public launch — Eleven (Monolingual) v1
The launch. First human-like English TTS on a beta web platform. The date ElevenLabs dates itself to.
Multilingual v1
Added 7 languages (FR, DE, HI, IT, PL, PT, ES). Richer emotion.
Multilingual v1
Added 7 languages (FR, DE, HI, IT, PL, PT, ES). Richer emotion.
Multilingual v2 (out of beta)
~30 languages, highest consistency. Still the long-form / audiobook default in 2026.
Multilingual v2 (out of beta)
~30 languages, highest consistency. Still the long-form / audiobook default in 2026.
PCM output format
Added raw PCM audio output to the API.
PCM output format
Added raw PCM audio output to the API.
Voice translation / Dubbing
First dubbing tool — break language barriers for content.
Voice translation / Dubbing
First dubbing tool — break language barriers for content.
Turbo v2 + Projects
First high-quality + low-latency model. Projects long-form editor (later renamed Studio).
Turbo v2 + Projects
First high-quality + low-latency model. Projects long-form editor (later renamed Studio).
Speech-to-Speech API
Voice changer goes public on the API.
Speech-to-Speech API
Voice changer goes public on the API.
TTS endpoints with timestamps
Character-level timing data for sync use cases.
TTS endpoints with timestamps
Character-level timing data for sync use cases.
AI Sound Effects
Generate SFX from text descriptions.
AI Sound Effects
Generate SFX from text descriptions.
Voice Isolator API
Strip background noise from audio.
Voice Isolator API
Strip background noise from audio.
Turbo v2.5
Faster, cheaper, more languages.
Turbo v2.5
Faster, cheaper, more languages.
+3 languages
Hungarian, Vietnamese, Norwegian.
+3 languages
Hungarian, Vietnamese, Norwegian.
API Key Permissions
Scoped permissions for API keys.
API Key Permissions
Scoped permissions for API keys.
ElevenReader app
Consumer TTS reader app.
ElevenReader app
Consumer TTS reader app.
Studio got bigger & better
Major Studio upgrade for long-form production.
Studio got bigger & better
Major Studio upgrade for long-form production.
Flash (v2 / v2.5)
Ultra-low latency ~75ms, 32 languages, ~50% cheaper. Becomes THE real-time / agents model. eleven_flash_v2_5.
Flash (v2 / v2.5)
Ultra-low latency ~75ms, 32 languages, ~50% cheaper. Becomes THE real-time / agents model. eleven_flash_v2_5.
Conversational AI v1
The agents platform launches: ASR + your LLM + low-latency TTS + turn-taking.
Conversational AI v1
The agents platform launches: ASR + your LLM + low-latency TTS + turn-taking.
Studio available to everyone
Long-form editor opens to all users.
Studio available to everyone
Long-form editor opens to all users.
Scribe (v1) — first STT model
Most accurate transcription; diarization, timestamps, 90+ langs. First move into understanding voice.
Scribe (v1) — first STT model
Most accurate transcription; diarization, timestamps, 90+ langs. First move into understanding voice.
Sound Effects in Studio
SFX integrated into the Studio editor.
Sound Effects in Studio
SFX integrated into the Studio editor.
Agent testing + Multimodal ConvAI
Agent testing tools and multimodal (voice + text) conversations.
Agent testing + Multimodal ConvAI
Agent testing tools and multimodal (voice + text) conversations.
Conversational AI 2.0
First-turn latency <500ms, real-time turn-taking (reads pauses/fillers), auto language switching, interrupt handling, RAG, enterprise readiness.
Conversational AI 2.0
First-turn latency <500ms, real-time turn-taking (reads pauses/fillers), auto language switching, interrupt handling, RAG, enterprise readiness.
Eleven v3 (alpha)
Most expressive TTS: audio tags [laughs]/[whispers], multi-speaker dialogue mode, 70+ langs. Not for real-time at launch.
Eleven v3 (alpha)
Most expressive TTS: audio tags [laughs]/[whispers], multi-speaker dialogue mode, 70+ langs. Not for real-time at launch.
ElevenLabs app (mobile)
Native mobile app.
ElevenLabs app (mobile)
Native mobile app.
Voice Design v3
Generate a brand-new voice from a text prompt.
Voice Design v3
Generate a brand-new voice from a text prompt.
Conversational AI supports WebRTC
Lower-latency real-time transport for agents.
Conversational AI supports WebRTC
Lower-latency real-time transport for agents.
Eleven Music
Studio-grade song generation on licensed data (Merlin + Kobalt deals), cleared for broad commercial use.
Eleven Music
Studio-grade song generation on licensed data (Merlin + Kobalt deals), cleared for broad commercial use.
Eleven Music API
First licensed music API for developers. 50% off in August.
Eleven Music API
First licensed music API for developers. 50% off in August.
Agents support Chat Mode
Text agents, not just voice.
Agents support Chat Mode
Text agents, not just voice.
Eleven v3 (alpha) in the API
Developer access to the expressive flagship.
Eleven v3 (alpha) in the API
Developer access to the expressive flagship.
Renamed → “ElevenLabs Agents”
Conversational AI rebrands. 2M+ agents created, 33M+ conversations YTD. Roadmap: workflow builder, testing, Expressive Mode.
Renamed → “ElevenLabs Agents”
Conversational AI rebrands. 2M+ agents created, 33M+ conversations YTD. Roadmap: workflow builder, testing, Expressive Mode.
Productions
Human-edited, done-for-you content (managed dubbing/localization).
Productions
Human-edited, done-for-you content (managed dubbing/localization).
Scribe v2 Realtime
Streaming STT (~150ms); feeds the turn-taking system.
Scribe v2 Realtime
Streaming STT (~150ms); feeds the turn-taking system.
ElevenLabs Image & Video
Integrates 3rd-party models (Veo/Sora/Kling/Wan/Flux/Seedance) for visual generation.
ElevenLabs Image & Video
Integrates 3rd-party models (Veo/Sora/Kling/Wan/Flux/Seedance) for visual generation.
Templates on the Creative Platform
Pre-built creative workflow presets.
Templates on the Creative Platform
Pre-built creative workflow presets.
Eleven Music — new tools
Explore / edit / produce tooling for music.
Eleven Music — new tools
Explore / edit / produce tooling for music.
Scribe v2 (batch)
Most accurate batch STT; subtitles/captions at scale; now powers Studio. scribe_v2 + scribe_v2_realtime.
Scribe v2 (batch)
Most accurate batch STT; subtitles/captions at scale; now powers Studio. scribe_v2 + scribe_v2_realtime.
Revolut selects ElevenAgents
Revolut adopts ElevenAgents for customer support.
Revolut selects ElevenAgents
Revolut adopts ElevenAgents for customer support.
Eleven v3 → General Availability
v3 leaves alpha across UI + API.
Eleven v3 → General Availability
v3 leaves alpha across UI + API.
$500M Series D at $11B valuation
Led by Sequoia, with a16z + ICONIQ; new: Lightspeed, Evantic, BOND. Telecom flagged as priority vertical.
$500M Series D at $11B valuation
Led by Sequoia, with a16z + ICONIQ; new: Lightspeed, Evantic, BOND. Telecom flagged as priority vertical.
BCG strategic partnership
Consulting-led enterprise CX agent deployments.
BCG strategic partnership
Consulting-led enterprise CX agent deployments.
Expressive Mode for ElevenAgents
Two new models: Eleven v3 Conversational (expressive real-time dialogue) + a new turn-taking system (uses Scribe v2 Realtime to infer emotion).
Expressive Mode for ElevenAgents
Two new models: Eleven v3 Conversational (expressive real-time dialogue) + a new turn-taking system (uses Scribe v2 Realtime to infer emotion).
Agents harden
Versioning, MCP tools as first-class type, content + focus + prompt-injection guardrails, conversation redaction, semantic search, file uploads, LLM-list endpoint. SDKs rebranded → ElevenAgents (v2.36).
Agents harden
Versioning, MCP tools as first-class type, content + focus + prompt-injection guardrails, conversation redaction, semantic search, file uploads, LLM-list endpoint. SDKs rebranded → ElevenAgents (v2.36).
Google Cloud / NVIDIA extension
Multi-year deal: NVIDIA Blackwell GPUs, Gemini in Agents, Veo in Creative, listed on Google Cloud Marketplace.
Google Cloud / NVIDIA extension
Multi-year deal: NVIDIA Blackwell GPUs, Gemini in Agents, Veo in Creative, listed on Google Cloud Marketplace.
Three-pillar brand consolidation
Brand framing into ElevenAgents / ElevenCreative / ElevenAPI.
Three-pillar brand consolidation
Brand framing into ElevenAgents / ElevenCreative / ElevenAPI.
Flows — node-based creative canvas
Chains 35+ image/video models + full audio stack into one visual pipeline; reusable. “Flows API coming soon.”
Flows — node-based creative canvas
Chains 35+ image/video models + full audio stack into one visual pipeline; reusable. “Flows API coming soon.”
Users page GA + SIP dynamic vars
Agents Users page GA; SIP headers as dynamic variables; content threshold guardrail.
Users page GA + SIP dynamic vars
Agents Users page GA; SIP headers as dynamic variables; content threshold guardrail.
Environment Variables + Auth Connections API
Workspace config APIs, KB URL refresh/auto-sync, guardrail retry-with-feedback, Music Marketplace docs, client SDK v1.0.0-rc.
Environment Variables + Auth Connections API
Workspace config APIs, KB URL refresh/auto-sync, guardrail retry-with-feedback, Music Marketplace docs, client SDK v1.0.0-rc.
IBM watsonx + Basic/Full Seats
IBM partnership: TTS/STT in watsonx Orchestrate (10k+ voices, PCI, Zero-Retention/HIPAA). Basic vs Full Seats (20 Basic on paid plans).
IBM watsonx + Basic/Full Seats
IBM partnership: TTS/STT in watsonx Orchestrate (10k+ voices, PCI, Zero-Retention/HIPAA). Basic vs Full Seats (20 Basic on paid plans).
Client SDK v1.0.0 GA
Big breaking changes, ships with a migration skill. MCP tool scoping per node, tool mocking, mTLS, video-to-music endpoint, transcribe-from-URL (YouTube/TikTok).
Client SDK v1.0.0 GA
Big breaking changes, ships with a migration skill. MCP tool scoping per node, tool mocking, mTLS, video-to-music endpoint, transcribe-from-URL (YouTube/TikTok).
Scoped analysis + test folders
Per-agent conversation analysis, test folders, multimodal messages in hooks.
Scoped analysis + test folders
Per-agent conversation analysis, test folders, multimodal messages in hooks.
KB fuzzy search + topic discovery
Knowledge base content search, conversation topic discovery, flexible branch merging, Gemini 3.1 Pro / Qwen LLMs.
KB fuzzy search + topic discovery
Knowledge base content search, conversation topic discovery, flexible branch merging, Gemini 3.1 Pro / Qwen LLMs.
ElevenLabs Devs YouTube channel
Developer channel: tutorials, API walkthroughs, deep dives. Plus workflow node overrides, usage analytics endpoint. ★ Directly relevant to Dev Community Growth.
ElevenLabs Devs YouTube channel
Developer channel: tutorials, API walkthroughs, deep dives. Plus workflow node overrides, usage analytics endpoint. ★ Directly relevant to Dev Community Growth.
Agent trust context + events
agent_response_complete event, pre_tool_speech mode, trust context (low/high), Scribe realtime keyterms.
Agent trust context + events
agent_response_complete event, pre_tool_speech mode, trust context (low/high), Scribe realtime keyterms.
Conversation tags + new LLMs
Conversation tags; claude-opus-4-7, gpt-5.4/5.5 added.
Conversation tags + new LLMs
Conversation tags; claude-opus-4-7, gpt-5.4/5.5 added.
Studio Agents
Agentic help inside Studio.
Studio Agents
Agentic help inside Studio.
ElevenCreative Templates (expanded)
More templates; SIP logs, KB in-place edits, SMS conversations, API analytics, IP allowlisting.
ElevenCreative Templates (expanded)
More templates; SIP logs, KB in-place edits, SMS conversations, API analytics, IP allowlisting.
Flows real-time collaboration + SFX v2
Multi-user Flow editing. SFX v2 model (up to 30s, seamless looping).
Flows real-time collaboration + SFX v2
Multi-user Flow editing. SFX v2 model (up to 30s, seamless looping).
Agent version metadata + procedures
Version metadata, text_only filter, procedure loading/skills, gemini-3.1-flash-lite.
Agent version metadata + procedures
Version metadata, text_only filter, procedure loading/skills, gemini-3.1-flash-lite.
Speech Engine
Add real-time voice to your own LLM. ElevenLabs does STT + turn-taking + TTS + playback; you keep the LLM/logic and stream text over a WebSocket. The BYO-LLM middle ground. Plus Exotel/Intercom/Telegram/Freshdesk.
Speech Engine
Add real-time voice to your own LLM. ElevenLabs does STT + turn-taking + TTS + playback; you keep the LLM/logic and stream text over a WebSocket. The BYO-LLM middle ground. Plus Exotel/Intercom/Telegram/Freshdesk.
Music v2
Better vocals/instrumentation/arrangement, improved multilingual; chunk-based composition plans.
Music v2
Better vocals/instrumentation/arrangement, improved multilingual; chunk-based composition plans.
Dubbing v2
First dubbing model that carries the original speaker's emotion/performance across languages. Creator Dubbing Partner Program; enterprise via ElevenProductions.
Dubbing v2
First dubbing model that carries the original speaker's emotion/performance across languages. Creator Dubbing Partner Program; enterprise via ElevenProductions.
Telephony + workflow transfers
Exotel telephony, workflow-aware agent transfers, agent test repeat runs, all agents versioned, 5GB STT uploads.
Telephony + workflow transfers
Exotel telephony, workflow-aware agent transfers, agent test repeat runs, all agents versioned, 5GB STT uploads.
Flows Agent
Conversational assistant inside the Flows canvas. Describe what you want; it selects models, creates + wires nodes, and runs generations — asking clarifying questions before spending credits.
Flows Agent
Conversational assistant inside the Flows canvas. Describe what you want; it selects models, creates + wires nodes, and runs generations — asking clarifying questions before spending credits.
v1 model deprecation + turn_model
v1 models + scribe_v1 removal notice (Jul 9). Per-agent turn_model (turn_v2/turn_v3); KB Google Drive sync; text-to-dialogue zero-retention mode.
v1 model deprecation + turn_model
v1 models + scribe_v1 removal notice (Jul 9). Per-agent turn_model (turn_v2/turn_v3); KB Google Drive sync; text-to-dialogue zero-retention mode.
Avatars
Talking-head video from script + voice + avatar in one tab. Persistent reusable identities; pairs any voice (incl. clones) with lip-sync. New Avatar node in Flows for batch video.
Avatars
Talking-head video from script + voice + avatar in one tab. Persistent reusable identities; pairs any voice (incl. clones) with lip-sync. New Avatar node in Flows for batch video.
Music v2 in the API
model_id=music_v2 with composition plans. Conversation evaluation rerun, telephony/WhatsApp filters, Unity SDK WebGL bridge. JS + Python SDKs at v2.53.0.
Music v2 in the API
model_id=music_v2 with composition plans. Conversation evaluation rerun, telephony/WhatsApp filters, Unity SDK WebGL bridge. JS + Python SDKs at v2.53.0.
Ads Engine + TELUS Digital
Ads Engine connects Google / Meta / LinkedIn ad accounts; localizes existing creatives across 50+ languages and pushes them back. Same day: TELUS Digital named preferred implementation partner for ElevenAgents.
Ads Engine + TELUS Digital
Ads Engine connects Google / Meta / LinkedIn ad accounts; localizes existing creatives across 50+ languages and pushes them back. Same day: TELUS Digital named preferred implementation partner for ElevenAgents.