CIIFragments Studio is CII-accredited: recover up to 20% of your software development spendLearn more

API integration for Transcription

We build your transcription connector

We wire speech-to-text into your business: already recorded calls, field dictation, meetings. Whisper, gpt-transcribe, Deepgram or Voxtral, depending on latency, language and residency. Not a single logo.

  • Senior product team
  • STT in production
  • WER and cost measured on your files
In short

What does a transcription API provide and how do you integrate it into a product?

A transcription API turns an audio recording or phone call into structured text, with speakers identified and keywords extractable. You integrate it into automatic meeting note-taking software, call centres that want to analyse conversations, medical applications that transcribe consultations, or legal platforms that index proceedings. The practical result: spoken information becomes searchable, summaries are generated without human intervention, and data stays in your system without manual transfer.

Use cases

What our clients have transcribed

01

Indexing already recorded support calls

Not the live callbot. Full text 3 parcels, 90 days, restricted access. Support no longer listens for 40 minutes.

02

Field report dictated in the van

The technician speaks, STT first, an LLM next to structure, the ERP receives the file.

03

SRT captions for internal training

Timestamps: still whisper-1 (or an equivalent pipeline). gpt-transcribe is not the default SRT path.

04

Diarisation of a three-voice meeting

gpt-4o-transcribe-diarize, or Deepgram diarize. The CRM updates, it is not the same model as a plain file.

For you

What this changes in your operations

Engineering in service of a measurable result: audio becomes data, typing drops, the engine stays replaceable.

Support finds the sentence, not the tape

The customer said 3 parcels is searchable. Forty minutes of listening become a query, bounded access.

Dictation replaces the form

Field, video, meeting: text lands in the ERP. AI structures next. Transcription first, always.

The engine stays replaceable

One product interface. Switching transcription vendor does not break the business journey.

Compliance is set before recording

Notice to people, legal basis, retention, workplace CNIL rules. The transcript is only enforceable if collection is scoped.

Method

How we ship your transcription connector

01

Scoping

File or live, residency, volume, diarisation. Whisper GPU, Voxtral eu, Deepgram or OpenAI: settled before coding.

02

Measurement

WER on your French files (noise, 8 kHz, names, SIRET). Not an English marketing percentage. You approve.

03

Development

Internal interface, two implementations, 25 MB chunking, hash idempotence, transcript stored on your side.

04

Monitoring

Model, cost, duration, retention. Journal: who transcribed what. Deepgram redaction if needed. No useless PII to the vendor.

The building blocks

What a transcription connector allows

File batch
OpenAI POST /v1/audio/transcriptions, gpt-transcribe recommended. Deepgram POST /v1/listen?model=nova-3. Voxtral Mini Transcribe 2.
Live streaming
Deepgram WebSocket, Voxtral Realtime, OpenAI Realtime. Different latency, price, contract. Do not sell live with a batch API.
Diarisation and timestamps
OpenAI diarisation: another model, diarized_json, auto chunking beyond 30 s. SRT / VTT: still whisper-1. Translation: to English only.
Self-host Whisper
MIT, inference on your side, GPU, VAD. Real sovereignty, no vendor SLA. The quote is not a POST.
Vocabulary

Transcription connector vocabulary

gpt-transcribe
Recommended OpenAI file model. A 2023 article we plug in Whisper is wrong for a new file project, except timestamps, SRT or EN translation.
25 MB
OpenAI file cap. An hour of 16 kHz wav often exceeds it. Chunking or compression first, or Deepgram / self-host Whisper.
Diarisation
Who speaks. At OpenAI, another model (gpt-4o-transcribe-diarize), another payload. Deepgram: diarize option on Nova-3. Two paths, two quotes.
WER
Word error rate. Measured on your corpus (noise, 8 kHz, proper nouns, SIRET). We do not publish a marketing percentage.
Batch vs live
Already recorded file vs WebSocket / Realtime. Different latency, price, contract. Selling live captions with a batch API is a false scoping.
Self-host Whisper
MIT model, GPU on your side, no OpenAI API. Real sovereignty, GPU ops, not a POST. The quote is not the same job as a key.
Good to know

The real constraints of a transcription API

01

25 MB and another model for diarisation

Two official OpenAI traps. Chunking before the call. gpt-4o-transcribe-diarize does not accept prompts like gpt-transcribe. Two API paths.

02

whisper-1 is no longer the file default

gpt-transcribe is. whisper-1 stays for SRT, timestamps, translation to English. A 2023 Whisper API connector is rescope, not a rename of the model ID.

03

Live is not the file

Deepgram WS, Voxtral Realtime, OpenAI Realtime: different latency, different price, different contract. Do not sell streaming with a batch POST, it is not the same product.

04

A call transcript is a processing activity

Information, legal basis, retention, CNIL (workplace recording). No vendor training by default: contracts to reread, not a generic ZDR slogan.

File or live

File transcription or live streaming?

Two STT architectures. The right choice depends on when the audio exists, not on the most cited logo.

CriterionFile batchAlready recordedLive streamingDuring the audio
WhenAfter the call, visio, dictationDuring, captions, agent
OpenAIgpt-transcribe, 25 MBRealtime / gpt-live-transcribe
DeepgramPOST /v1/listen Nova-3WebSocket /v1/listen
MistralVoxtral Mini Transcribe 2Voxtral Realtime (may leave regional)
Whisper GPUNatural (files)VAD + chunks pipeline, another project
SRT / diarisationwhisper-1 / diarize modelDepends on the live vendor
The right caseSupport index, dictation, trainingCaptions, callbot (other page)

The live callbot is the AI voice agent page. Here: STT connector, often file. Two implementations tested behind transcribe(). WER on your corpus, not a bench.

Our expertise

What we measure on a transcription connector

15 d
first file flow in production
2
STT implementations tested
WER
measured on your files, not claimed
4
senior developers on the project

We combine Transcription with

The stack around an STT connector on our projects.

  • OpenAI
  • Deepgram
  • Mistral
  • PostgreSQL
  • Node.js
FAQ

Transcription: your questions

Four steps. Pick file or live, and residency (Vertex/Deepgram EU, Voxtral api.eu, Whisper on-prem). Put an internal transcribe() interface with two implementations tested. Chunking (25 MB OpenAI), retry, hash idempotence, transcript stored on your side. Measure WER on your French audio. The sensitive part is not the POST, it is the architecture choice, diarisation on another model, and the CNIL frame if these are calls.

Two Whispers. whisper-1 at OpenAI: still useful for SRT, timestamps and translation to English, no longer the default for a new file (that is gpt-transcribe). Open-source Whisper: inference on your side, MIT, GPU, no OpenAI API, no vendor SLA. A Whisper API article often mixes the two. We split them at scoping. High KD on that query: the page states the job (STT connector), it does not promise to rank on OpenAI docs.

gpt-transcribe is the file model recommended today (POST /v1/audio/transcriptions). whisper-1 stays for timestamps, SRT, VTT, and /v1/audio/translations (English only). File streaming exists on gpt-* models, not on whisper-1. Diarisation is a third path (gpt-4o-transcribe-diarize). A connector that only exposes a whisper identifier will break on the first need for speakers or captions. We keep the identifier in configuration.

Under conditions. Informing people, legal basis, retention period, recipients. CNIL doctrine on listening and recording calls at the workplace. Separate operations data from tuning data. No useless PII sent; Deepgram redaction if needed. Vendor contracts (training, retention) to reread: we do not promise generic ZDR. The journal (who transcribed what, model, cost) is a deliverable, not an afterthought.

A first file flow (gpt-transcribe or Deepgram) with chunking and internal storage ships in two to three weeks. Two engines, diarisation, SRT, self-host Whisper and a CNIL frame for calls are closer to six to eight weeks, plus WER measurement time on your corpus. Duration depends on live vs file and residency. We scope the perimeter up front and give you a firm estimate before we start, including the engine pair.

A transcription connector project?

Let's talk. 30 minutes to scope file or live, residency, and tell you frankly which engine holds on your French audio.

Discuss my transcription project
Discuss my transcription project