API integration for Transcription
We build your transcription connector
We wire speech-to-text into your business: already recorded calls, field dictation, meetings. Whisper, gpt-transcribe, Deepgram or Voxtral, depending on latency, language and residency. Not a single logo.
- Senior product team
- STT in production
- WER and cost measured on your files
What does a transcription API provide and how do you integrate it into a product?
A transcription API turns an audio recording or phone call into structured text, with speakers identified and keywords extractable. You integrate it into automatic meeting note-taking software, call centres that want to analyse conversations, medical applications that transcribe consultations, or legal platforms that index proceedings. The practical result: spoken information becomes searchable, summaries are generated without human intervention, and data stays in your system without manual transfer.
What our clients have transcribed
Indexing already recorded support calls
Not the live callbot. Full text 3 parcels, 90 days, restricted access. Support no longer listens for 40 minutes.
Field report dictated in the van
The technician speaks, STT first, an LLM next to structure, the ERP receives the file.
SRT captions for internal training
Timestamps: still whisper-1 (or an equivalent pipeline). gpt-transcribe is not the default SRT path.
Diarisation of a three-voice meeting
gpt-4o-transcribe-diarize, or Deepgram diarize. The CRM updates, it is not the same model as a plain file.
What this changes in your operations
Engineering in service of a measurable result: audio becomes data, typing drops, the engine stays replaceable.
Support finds the sentence, not the tape
The customer said 3 parcels is searchable. Forty minutes of listening become a query, bounded access.
Dictation replaces the form
Field, video, meeting: text lands in the ERP. AI structures next. Transcription first, always.
The engine stays replaceable
One product interface. Switching transcription vendor does not break the business journey.
Compliance is set before recording
Notice to people, legal basis, retention, workplace CNIL rules. The transcript is only enforceable if collection is scoped.
How we ship your transcription connector
Scoping
File or live, residency, volume, diarisation. Whisper GPU, Voxtral eu, Deepgram or OpenAI: settled before coding.
Measurement
WER on your French files (noise, 8 kHz, names, SIRET). Not an English marketing percentage. You approve.
Development
Internal interface, two implementations, 25 MB chunking, hash idempotence, transcript stored on your side.
Monitoring
Model, cost, duration, retention. Journal: who transcribed what. Deepgram redaction if needed. No useless PII to the vendor.
What a transcription connector allows
- File batch
- OpenAI POST /v1/audio/transcriptions, gpt-transcribe recommended. Deepgram POST /v1/listen?model=nova-3. Voxtral Mini Transcribe 2.
- Live streaming
- Deepgram WebSocket, Voxtral Realtime, OpenAI Realtime. Different latency, price, contract. Do not sell live with a batch API.
- Diarisation and timestamps
- OpenAI diarisation: another model, diarized_json, auto chunking beyond 30 s. SRT / VTT: still whisper-1. Translation: to English only.
- Self-host Whisper
- MIT, inference on your side, GPU, VAD. Real sovereignty, no vendor SLA. The quote is not a POST.
Transcription connector vocabulary
- gpt-transcribe
- Recommended OpenAI file model. A 2023 article we plug in Whisper is wrong for a new file project, except timestamps, SRT or EN translation.
- 25 MB
- OpenAI file cap. An hour of 16 kHz wav often exceeds it. Chunking or compression first, or Deepgram / self-host Whisper.
- Diarisation
- Who speaks. At OpenAI, another model (gpt-4o-transcribe-diarize), another payload. Deepgram: diarize option on Nova-3. Two paths, two quotes.
- WER
- Word error rate. Measured on your corpus (noise, 8 kHz, proper nouns, SIRET). We do not publish a marketing percentage.
- Batch vs live
- Already recorded file vs WebSocket / Realtime. Different latency, price, contract. Selling live captions with a batch API is a false scoping.
- Self-host Whisper
- MIT model, GPU on your side, no OpenAI API. Real sovereignty, GPU ops, not a POST. The quote is not the same job as a key.
The real constraints of a transcription API
25 MB and another model for diarisation
Two official OpenAI traps. Chunking before the call. gpt-4o-transcribe-diarize does not accept prompts like gpt-transcribe. Two API paths.
whisper-1 is no longer the file default
gpt-transcribe is. whisper-1 stays for SRT, timestamps, translation to English. A 2023 Whisper API connector is rescope, not a rename of the model ID.
Live is not the file
Deepgram WS, Voxtral Realtime, OpenAI Realtime: different latency, different price, different contract. Do not sell streaming with a batch POST, it is not the same product.
A call transcript is a processing activity
Information, legal basis, retention, CNIL (workplace recording). No vendor training by default: contracts to reread, not a generic ZDR slogan.
File transcription or live streaming?
Two STT architectures. The right choice depends on when the audio exists, not on the most cited logo.
| Criterion | File batchAlready recorded | Live streamingDuring the audio |
|---|---|---|
| When | After the call, visio, dictation | During, captions, agent |
| OpenAI | gpt-transcribe, 25 MB | Realtime / gpt-live-transcribe |
| Deepgram | POST /v1/listen Nova-3 | WebSocket /v1/listen |
| Mistral | Voxtral Mini Transcribe 2 | Voxtral Realtime (may leave regional) |
| Whisper GPU | Natural (files) | VAD + chunks pipeline, another project |
| SRT / diarisation | whisper-1 / diarize model | Depends on the live vendor |
| The right case | Support index, dictation, training | Captions, callbot (other page) |
The live callbot is the AI voice agent page. Here: STT connector, often file. Two implementations tested behind transcribe(). WER on your corpus, not a bench.
What we measure on a transcription connector
The other AI building blocks
Transcription is combined more often than it is replaced. These options are discussed at scoping.
GeminiThe model that structures the transcript afterwards, or multimodal audio.
MistralVoxtral and EU inference, when STT must stay in the same lab.
HeyGenWe build your HeyGen connector
ElevenLabsWe build your ElevenLabs connector
Anthropic (Claude)A model that calls your tools, not one more chat
OpenAIThe model writes into your tools, it does not chat
CursorThe agent reaches your internal tools, not just your codeWe combine Transcription with
The stack around an STT connector on our projects.
Transcription: your questions
Four steps. Pick file or live, and residency (Vertex/Deepgram EU, Voxtral api.eu, Whisper on-prem). Put an internal transcribe() interface with two implementations tested. Chunking (25 MB OpenAI), retry, hash idempotence, transcript stored on your side. Measure WER on your French audio. The sensitive part is not the POST, it is the architecture choice, diarisation on another model, and the CNIL frame if these are calls.
Two Whispers. whisper-1 at OpenAI: still useful for SRT, timestamps and translation to English, no longer the default for a new file (that is gpt-transcribe). Open-source Whisper: inference on your side, MIT, GPU, no OpenAI API, no vendor SLA. A Whisper API article often mixes the two. We split them at scoping. High KD on that query: the page states the job (STT connector), it does not promise to rank on OpenAI docs.
gpt-transcribe is the file model recommended today (POST /v1/audio/transcriptions). whisper-1 stays for timestamps, SRT, VTT, and /v1/audio/translations (English only). File streaming exists on gpt-* models, not on whisper-1. Diarisation is a third path (gpt-4o-transcribe-diarize). A connector that only exposes a whisper identifier will break on the first need for speakers or captions. We keep the identifier in configuration.
Under conditions. Informing people, legal basis, retention period, recipients. CNIL doctrine on listening and recording calls at the workplace. Separate operations data from tuning data. No useless PII sent; Deepgram redaction if needed. Vendor contracts (training, retention) to reread: we do not promise generic ZDR. The journal (who transcribed what, model, cost) is a deliverable, not an afterthought.
A first file flow (gpt-transcribe or Deepgram) with chunking and internal storage ships in two to three weeks. Two engines, diarisation, SRT, self-host Whisper and a CNIL frame for calls are closer to six to eight weeks, plus WER measurement time on your corpus. Duration depends on live vs file and residency. We scope the perimeter up front and give you a firm estimate before we start, including the engine pair.
A transcription connector project?
Let's talk. 30 minutes to scope file or live, residency, and tell you frankly which engine holds on your French audio.
Discuss my transcription project