
API integration for HeyGen
We build your HeyGen connector
We wire HeyGen into your product to generate an onboarding, training or follow-up video for each new customer, without a studio.
- Senior product team
- AI video connectors in production
- from scoping to monitoring
What does the HeyGen API provide and when should you generate videos programmatically?
HeyGen generates videos with a speaking avatar from text in over thirty languages. Its API lets your application automatically produce a personalised video: with the recipient's name, a product, a language, without a studio or human intervention. You integrate it in training platforms that generate personalised lessons, sales prospecting tools that create one video per prospect, or e-commerce applications that produce localised product presentations on demand.
What our clients build on the HeyGen API
Onboarding video on account activation
Name, product, language. Triggered by the business event, delivered when ready, stored on your side.
E-learning module regenerated on every script
Three languages, lipsync. A script update no longer waits for a reshoot.
Executive message, consented avatar
Quarterly offer. Written agreement, revocable, consent API call. A stock avatar avoids the topic.
Versioned field tutorial in the store app
Here is how to scan the parcel. Batch of variants (one per branch) on the same template.
What this changes in your production
Engineering in service of a measurable outcome: a variant without a studio, consent in place, a video on your side.
The cost of a variant drops
Onboarding, follow-up, training: name, product, language, without a day of filming. The business triggers, the video arrives.
Export no longer waits for a reshoot
Translation and lipsync. An existing webinar ships to an ES/DE subsidiary without a new set.
Consent comes before the clone
A leader or trainer's face: written consent, scoped, revocable. Same reflex as for a cloned voice.
The file lives in your storage
As soon as the video is ready, we download it to your side. The vendor URL is not treated as eternal hosting.
How we ship your HeyGen connector
Scoping
Avatar API (brand control) or Video Agent (prompt). Consent, languages, volume. The price grid is reread at scoping, we do not invent it.
Prototype
One template, your sentences, your brand names. You approve the render, not us. FR pronunciation lexicon.
Development
Async job, video_id persisted, raw HMAC, Heygen-Event-Id dedup, 2xx < 10 s, storage on your bucket, fallback on fail.
Monitoring
avatar_video.fail, GPU queue, webhook secret (one-shot at creation). Rotation = a window of failed verifications.
What the HeyGen API allows
- Avatar videos and templates
- Brand control: look, voice, script. An Avatar API quote is not a prompt Video Agent quote.
- Video Agent
- Prompt to video. Faster to prototype, less control. Two products, two scopings.
- Translation and lipsync
- 30+ languages. An existing webinar, a subsidiary, without a reshoot. Batch up to 100 variants.
- Signed webhooks
- POST /v3/webhooks/endpoints, one-shot whsec_ secret. HMAC of the raw body only: Event-Id dedup is the replay defence.
HeyGen API vocabulary
- Heygen-Signature
- HMAC-SHA256 hex of the raw body. Heygen-Timestamp to reject beyond about 5 minutes. Timestamp is not in the HMAC: Heygen-Event-Id dedup is mandatory.
- whsec_
- Webhook secret, shown only at creation or rotation. Losing it means rotation and a window of failed verifications (official docs). Off git.
- Video Agent
- Prompt to video. Less brand control than an Avatar API (look, voice, script). Two quotes, not a toggle.
- callback_url
- One-shot alternative at creation (callback_id). Less flexible than a webhook subscription for a product. Useful in a prototype.
- Avatar consent
- POST /v3/avatars/{group_id}/consent. Image rights: written agreement, duration, perimeter, revocable. A stock avatar avoids the legal topic.
- Async GPU
- Tens of seconds to minutes. This is not a phone voice agent. An HTTP timeout on create is normal. The business is the webhook.
The real constraints of the HeyGen API
This is not real-time telephony
A video file delivered later. Real-time voice is the voice agent / ElevenLabs page. Selling live telephony with HeyGen is a bad scoping, not a feature gap.
The webhook secret is one-shot
Creation or rotation only. Losing it forces a rotation and a window of failed signatures. Off git, written procedure, not a secret pasted in a ticket.
HMAC on the body only
Heygen-Event-Id dedup mandatory. Timestamp outside HMAC: a replay with a fresh timestamp passes if you do not dedup. ACK 2xx in 10 s.
Price and URL TTL: at scoping
We do not publish a 2026 grid or a URL lifetime. We download to your bucket. Render residency and DPA: to reread, not asserted here.
HeyGen Avatar API or Video Agent?
Two products in the same v3 API. The right one depends on brand control, not on how fast the first prompt is.
| Criterion | Avatar APIBrand control | Video AgentPrompt |
|---|---|---|
| Control | Look, voice, script | Prompt to video |
| Brand | Tight | Wider, less predictable |
| Face consent | Often the topic (CEO, trainer) | Depends on the avatar chosen |
| Webhook | avatar_video.* | video_agent.* |
| Prototype | More scoping | Faster |
| Batch | Controlled variants, up to 100 | Variants from the prompt |
| The right case | Onboarding, training, offer | Exploration, test volume |
Both stay async. Neither is a callbot. One-shot callback_url vs webhook subscription: the product takes the subscription, the prototype can stay one-shot.
What we measure on a HeyGen integration
The other AI building blocks
HeyGen is combined more often than it is replaced. These options are discussed at scoping.
HeyGenWe build your HeyGen connectorThis page
GeminiText and image still, before moving to the avatar.
MistralThe script and OCR, upstream of the video, EU inference.
ElevenLabsWe build your ElevenLabs connectorTranscriptionWe build your transcription connector
Anthropic (Claude)A model that calls your tools, not one more chat
OpenAIThe model writes into your tools, it does not chat
CursorThe agent reaches your internal tools, not just your codeWe combine HeyGen with
The stack around HeyGen on our projects.
HeyGen: your questions
Four steps. Pick Avatar API or Video Agent. Put consent before any face clone (written agreement, consent call). Async create, persist video_id, HMAC webhooks (raw body, Event-Id, ACK < 10 s), download the file to your bucket. Queue and GPU cap. The sensitive part is not POST /v3/videos, it is the one-shot secret, the dedup, and not treating the HeyGen URL as permanent storage (TTL reread at scoping).
Yes, under conditions. Image rights: written agreement, duration, perimeter, revocable, plus POST /v3/avatars/{group_id}/consent. A stock avatar avoids the topic. Cloning without proof is the same risk as voice cloning: it is scoped before the first generation, not after the campaign. The consent register (who, until when, proof) is a deliverable, not a checkbox in a quote, and it is written before the first render.
No. HeyGen delivers a video file later (seconds to minutes). A voice agent (Twilio, realtime, ElevenLabs) holds a conversation. HeyGen webhooks (success/fail, 10 s, retries 24 h) are an async integration protocol, not a phone line. If the need is to pick up, that is the AI voice agent page. If the need is an onboarding variant, it is here. Mixing the two in one quote is the usual scoping error.
We do not publish a 2026 grid: it is reread on heygen.com/pricing at project scoping. The real cost is GPU compute, variant volume and failures (template fallback). We put a concurrency cap and a budget, a business queue, and we hand you the breakdown before developing. Inventing a price or a URL lifetime here would be a false scoping: both are verified on the current vendor pages, not guessed.
A first useful flow, an onboarding video on a template with webhook and internal storage, ships in two to three weeks. Consented executive avatar, multi-language translation, batch of 100 variants and a consent register are closer to six to eight weeks. Duration depends on the product (Avatar vs Video Agent) and GPU volume. We scope the perimeter up front and give you a firm estimate before we start.
A HeyGen integration project?
Let's talk. 30 minutes to scope avatar or Video Agent, consent, and tell you frankly what async will hold.
Discuss my HeyGen project