
API integration for ElevenLabs
We build your ElevenLabs connector
We build the ElevenLabs connector to the ElevenLabs API to generate your voices in French, clone a voice by the rules and transcribe your audio. With a cache, a spending cap and a replaceable engine.
- Senior product team
- voice integrations in production
- from scoping to monitoring
What does the ElevenLabs API provide and which products should integrate it?
ElevenLabs is a high-quality speech synthesis platform that turns text into natural-sounding audio in over seventy languages. You integrate it into voice agents, audio content creation tools, accessibility applications or voice notification systems. The voice quality is noticeably better than standard text-to-speech, and the platform also allows voice cloning from a recording. The practical outcome: a personalised voice response can be generated in seconds without a studio or human re-recording every time the content changes.
What our clients build on the ElevenLabs API
The voice of a phone voice agent
ElevenLabs is the default speech synthesis provider of Twilio ConversationRelay. It is the French voice of the agent that answers your calls.
Personalised voice notifications
Appointment reminders, job confirmations, on-call alerts that read the actual incident context. A name, an amount, a time slot inside the message.
An audio version of your content
Every publication on your site or in your knowledge base gets its listenable version, regenerated only when the text actually changes.
Transcribing calls and field notes
scribe_v2 indexes your recordings and voice notes, and makes their content searchable from inside your business tool.
What it changes in how you produce content
Engineering in service of a measurable outcome: no more studio, multilingual with no re-recording, a bill under control.
The studio leaves the critical path
Write, record, edit, republish becomes an API call. A text correction propagates in minutes instead of a catch-up recording session.
Multilingual stops being a project
One pipeline produces French, English and Spanish with no re-recording. Markets that studio costs made unprofitable suddenly open up.
Messages that are actually personalised
A name, an amount, a date, a time slot fit inside the sentence. No more artificial phrasing like your order number, beep, is ready.
A bill that follows actual listens
With a cache, identical audio is never generated twice. Without one, spending follows the number of deployments rather than real usage.
How we ship your ElevenLabs connector
Scoping
Which content, what character volume, which voices, which languages. The cloning regime and the consent it requires are settled before coding.
Listening comparison
Your real sentences, with your proper nouns, your numbers and your addresses, played model by model. You approve the voice, not us.
Development
Internal synthesis service, cache addressed by fingerprint, model identifier in configuration, pronunciation lexicon, spending cap.
Monitoring
Character counter under supervision, traceability of every generated audio, and a failover test to the backup engine played regularly.
What the ElevenLabs API allows
- Multilingual speech synthesis
- eleven_v3 covers more than seventy languages, eleven_multilingual_v2 twenty-nine including French, and eleven_flash_v2_5 claims around 75 milliseconds.
- Transcription, including realtime
- scribe_v2 handles more than ninety languages, scribe_v2_realtime claims around 150 milliseconds. ElevenLabs is no longer only a voice provider.
- Voice cloning, two regimes
- Instant cloning needs one to two minutes of clean audio. Professional cloning needs thirty to a hundred and eighty minutes, plus a verification step.
- Voice agent platform
- In-house recognition, your choice of model including your own, native Twilio integration, SIP trunk, batch outbound calls, RAG, tools and automated tests.
The vocabulary of the ElevenLabs API
- Instant Voice Cloning
- Instant cloning from one to two minutes of clean audio, available on every tier, near immediate and with no verification. The absence of any check moves the entire legal responsibility onto you, not the platform.
- Professional Voice Cloning
- Professional cloning, thirty to a hundred and eighty minutes of audio, three to six hours of fine-tuning and a limited number of slots per tier. The documentation is explicit: your own voice only, with verification.
- Cost per character
- Flash and Turbo sit at 0.05 dollar per thousand characters, Multilingual v2 and v3 at 0.10 dollar. That one-to-two ratio is steered in code, not in a commercial negotiation.
- Zero Retention Mode
- An option that prevents processed content from being retained on ElevenLabs servers. Worth requesting explicitly when the synthesised text contains a name, an amount or an address.
- Data residency
- Isolated environments for the EU, India and Singapore, reserved for the Enterprise plan. On lower tiers the sovereignty argument cannot be made to your customer, and that has to be said upfront.
- Concurrency
- The number of simultaneous calls allowed on the agents platform, from four to forty depending on the tier. An unplanned peak turns into burst pricing or refusal, not into gentle degradation.
The real constraints of the ElevenLabs API
Professional cloning requires your voice
The documentation is explicit: a professional clone can only be made of your own voice, and every clone goes through a microphone verification. Cloning an executive's voice means that person does it from their own account.
Versioning is fast and retroactive
eleven_flash_v2 and eleven_turbo_v2_5 are already deprecated. An integration that hard-codes a model identifier eventually breaks, or produces a different output without anyone touching the text.
Claimed latency is not perceived latency
The 75 milliseconds of eleven_flash_v2_5 carry a footnote in the documentation and count neither the network, nor queueing, nor telephony transport. What matters is the time to the first audio byte, under load.
French quality has to be measured
Liaisons, numbers, dates, email addresses, acronyms and proper nouns do not come out the same way from one model to another, nor from one voice to another. Pronouncing a street name is a test case, not a detail.
eleven_flash_v2_5 or eleven_multilingual_v2?
The two models we use most. The choice follows the use case, and it can differ from one piece of content to another inside the same application.
| Criterion | eleven_flash_v2_5Interactive | eleven_multilingual_v2Content |
|---|---|---|
| Claimed latency | Around 75 ms | Not highlighted |
| Languages covered | 32, French included | 29, French included |
| Cost per 1,000 characters | 0.05 dollar | 0.10 dollar |
| The right use | Voice agent, guidance, alert | Article, training, dubbing |
| Effect of caching | Low, the text is often unique | High, content gets regenerated |
| Long-form reading | Good enough for an exchange | More nuance over time |
| Maximum coverage | No, thirty-two languages | eleven_v3 goes beyond 70 |
The two coexist in the same application: Flash for interactive, Multilingual v2 for content people listen to. The identifier stays in configuration because the catalogue moves, as the deprecation of eleven_flash_v2 and eleven_turbo_v2_5 shows.
What we measure on an ElevenLabs integration
The other voice and AI building blocks
ElevenLabs combines more often than it replaces, but these options are worth discussing during scoping.
ElevenLabsWe build your ElevenLabs connectorThis page
GeminiThe model that decides, when voice is only the system's output.
MistralThe sovereign option on text, hosted in France.
HeyGenWe build your HeyGen connectorTranscriptionWe build your transcription connector
Anthropic (Claude)A model that calls your tools, not one more chat
OpenAIThe model writes into your tools, it does not chat
CursorThe agent reaches your internal tools, not just your codeWe combine ElevenLabs with
The stack that surrounds ElevenLabs on our projects.
ElevenLabs: your questions
Three steps. First put an internal synthesis service in front of the API, with the model identifier in configuration and a cache addressed by the fingerprint of the text, voice and model triplet: the same audio is never generated twice, and you can change engine without touching application code. Then maintain a pronunciation lexicon for your brand names, products, cities and acronyms, approved by listening with you. Finally set a spending cap per user and per period, with a character counter exposed in monitoring. The sensitive part is not the API call, it is the cache, the lexicon and the traceability of generated audio.
The public grid is per character: 0.05 dollar per thousand characters for Flash and Turbo, 0.10 dollar for Multilingual v2 and v3. Tiers run from Starter at 6 dollars up to Business at 990 dollars, and the agents platform bills 0.080 dollar per minute beyond what is included, 0.160 dollar at burst pricing. Those one-to-two ratios, between models as well as between standard and burst pricing, are steered in code. At scale the real lever is caching: without it, the bill follows the number of deployments rather than the number of listens.
Not freely. The ElevenLabs documentation is explicit: a professional clone can only be created of your own voice, and every professional clone goes through a verification process designed to confirm the voice belongs to you. Even with their agreement, you cannot create the professional clone of a third party on your account. In practice, cloning an executive's voice means that person runs the verification from their own account, or a contractual and technical arrangement to be scoped. Instant cloning requires no verification at all, which moves the entire legal responsibility onto you: an identifiable person's voice falls under the GDPR as soon as it identifies them, and a written agreement bounded in duration and scope, and revocable, is not a formality.
French is covered by eleven_multilingual_v2, eleven_v3 and eleven_flash_v2_5, but the useful question is not coverage, it is rendition. Liaisons, numbers, dates, email addresses, acronyms and proper nouns do not come out the same way from one model to another, nor from one voice to another. So we do not answer from a catalogue: we play your real sentences, with your product names and your street names, model by model, and you make the call. The pronunciation lexicon that comes out of that session is a deliverable, and it is what separates it works from it is usable in production.
Only under conditions. ElevenLabs documents data residency in isolated environments for the EU, India and Singapore, but it is reserved for the Enterprise plan. On lower tiers the sovereignty argument cannot be made to your own customer, and that has to be said upfront. An optional Zero Retention Mode also prevents processed content from being retained on ElevenLabs servers. In every case we minimise what leaves: synthesised text often contains a name, an amount or an address, and reading out a confirmation does not require the full customer record. When the contract does not allow that frame, we propose another engine.
An ElevenLabs integration project?
Let's talk. 30 minutes to scope your volumes, listen to what the models do with your own sentences and tell you honestly what is feasible.
Discuss my ElevenLabs project