
API integration for Anthropic (Claude)
A model that calls your tools, not one more chat
We expose your business tools to the model through MCP, then wire Claude onto long documents and decisions that must be traceable. EU residency goes through a host, not the direct API.
- Senior product team
- MCP servers in production
- residency and budget settled
Why integrate the Anthropic Claude API into a business application?
Anthropic's Claude models stand out for their ability to call your own tools: the model analyses a request, proposes executing an action in your software, and waits for your confirmation before acting. This draws a clear boundary between what the model decides and what your system does. You integrate the Anthropic API when you want an assistant capable of reading a document, qualifying a file, or updating a record without a human transcribing the result. The MCP protocol created by Anthropic also lets you connect the same tool to multiple agents without rebuilding a connector per use case.
Four flows we wire around the model
Long file, sourced answer
Contracts, tenders, technical documentation: the model answers by citing passages, and the stable prefix is cached so re-reading the same file does not cost the same on every question.
Your tools exposed once
An MCP server on top of your ERP, your CRM or your database: declared once, reused by the agent, by your developers' editor and by automations.
An agent that drafts and prepares
Meeting notes, quotes, standard replies: the model produces the draft, writes it into the relevant tool as a proposal, and leaves approval to a human.
Bulk history backfill
Classify or enrich years of archives through the Batch API, at reduced cost, without saturating the real-time flow your users depend on.
What it changes in your product
We do not ship a chat grafted onto the product. We expose your tools cleanly, then let the model use them under control.
Long documents become usable
What nobody read becomes queryable, with sources cited. We measure the share of accepted answers, not the volume processed.
One connector for several uses
MCP avoids rewiring your tools for every new need. The investment happens once, then serves the next ones.
Legal has a clear answer
Inference region, subprocessor, retention: we write the flow diagram, including the fact that the direct API is not enough for residency.
The bill stays under control
Cache on already-read files, bulk processing for history, a cap per journey. Large files stop being the month-end surprise.
What a Claude connector forces as method
Tool surface
Which tools the model may call, with which rights, and which stay out of reach. An MCP server is designed like an internal public API.
Hosting choice
Direct API, Bedrock or Vertex: the residency constraint decides before technical preference. The region is fixed up front.
Cache and budget
Fixed part of the prompt first so it can be cached, bulk work moved to Batch, a cap per journey with an alert.
Writes under control
Every writing tool is idempotent and logged, with an explicit refusal when the model has nothing to conclude from.
What Claude enables
- Explicit tool use
- The model returns a tool call block your server executes or rejects. Responsibility stays with you, which is exactly what an audit wants to see.
- MCP as standard
- The protocol created by Anthropic to expose tools and data. An MCP server written once serves several clients, including your developers' editor.
- Prompt caching
- Reusing a stable prefix sharply cuts the cost of later calls, which changes the economics of large files read again and again.
- Batch processing
- The Batch API handles non-urgent volumes at a reduced rate, useful for history backfills and overnight enrichment.
The Anthropic vocabulary
- Messages API
- The single interface: a list of messages, optionally declared tools, and back either text or a request to call a tool.
- tool_use
- The block through which the model asks for a tool to run. It does not run it: your server decides, and that is where control sits.
- MCP
- Model Context Protocol, an open standard for exposing a tool or data source to a model. Writing an MCP server means investing once for several uses.
- Prompt caching
- Marking the stable part of a prompt so it is not charged at full rate on later calls. Assumes the fixed part comes first.
- Bedrock and Vertex
- The two routes to Claude models from a European region. That is how a residency requirement is met, not through the direct API.
- Batch API
- An asynchronous queue at a reduced rate, suited to volumes that can wait several hours.
What nobody tells you before you sign
The direct API does not solve European residency
Many integrations start on api.anthropic.com and then discover, during legal review or a public tender, that the inference region is not European. Moving to Bedrock or Vertex changes authentication, quotas and sometimes the models available: it is not a last-minute switch.
Reselling a model is not hosting it
A marketplace can offer Claude without inference running on its own infrastructure in the stated region. If residency is contractual, the question is not which platform offers it, but where the compute happens.
An MCP server is an attack surface
Exposing your tools to a model means opening an API. Rights per tool, secrets outside the server, call logging, and above all no destructive tool exposed without a second approval.
Caching is designed, not observed
The saving only exists if the stable part of the prompt really is stable and placed first. A prompt assembled in a varying order cancels the benefit with nothing to signal it.
Direct API or European host
The same model, two access paths. The choice follows the contractual residency constraint, then the tooling you already run.
| Criterion | Direct Anthropic APIThe simplest | Bedrock or Vertex in the EUUnder constraint |
|---|---|---|
| Getting started | One key, a few hours | Cloud account and IAM to scope |
| European residency | Not guaranteed | European region chosen |
| Authentication | x-api-key | IAM or service account |
| New models | Available earliest | Often delayed |
| Billing | Direct with the vendor | Consolidated with your cloud |
| Quotas | The vendor's own | The host's own |
| Suits | Product with no hard constraint | Regulated sector, public tender |
We isolate the call behind your own interface, which lets you start direct and later move to a European host when an enterprise client requires it, without rewriting the business logic.
What we hold to on a Claude project
The other models
The right pick depends on data residency and the business action, not on a ranking.
Anthropic (Claude)A model that calls your tools, not one more chatThis page
GeminiEU residency through Vertex, and the Google ecosystem if you are already there.
MistralThe European option, with inference on api.eu.mistral.ai.
HeyGenWe build your HeyGen connector
ElevenLabsWe build your ElevenLabs connectorTranscriptionWe build your transcription connector
OpenAIThe Responses API and built-in real time, Europe project from creation.
CursorThe agent reaches your internal tools, not just your codeWe combine Claude with
The bricks that make the flow usable
Claude integration: your questions
By starting with the tool surface, not the prompt. We decide which tools the model may call, with which rights, and which stay out of reach. That surface usually takes the form of an MCP server on top of your ERP, your CRM or your database, designed like an internal API with its own rights and logging. Then we wire the Messages API: the model proposes a tool call, your server executes or refuses. Finally we handle operations, meaning the inference region, caching on stable prefixes, the spend cap and the degraded path.
Yes, but not through the direct API, which offers no dedicated European region. The route used is Amazon Bedrock or Google Vertex in a European region, Paris and Frankfurt being the most requested in France. It is not a simple URL change: authentication goes through IAM or a service account, quotas are the host's, and new models sometimes arrive there later. Also beware of platforms reselling Claude without the compute happening on their infrastructure in the stated region: if residency is contractual, the question is where inference runs.
MCP standardises how a tool or data source is exposed to a model. In practice, instead of writing one connector per use, you write an MCP server on top of your system once, and it then serves your product's agent, your developers' editor and your automations. The benefit is a shared investment and a single place to put rights and logging. The trade-off is that an MCP server is an exposed API: it is secured as such, with rights per tool and no destructive action reachable without a second approval.
Development depends mostly on the tool surface to expose and how much control the writes demand: a document reading assistant costs far less than an agent allowed to create records in an ERP. Consumption depends on context size, and that is where large files weigh: caching the stable part of the prompt and moving bulk work to Batch change the bill significantly. We set a cap per journey and its measurement from the first release.
This page covers the Claude connector: the Messages API, tool use, the MCP server, the inference region, the budget. Our agents expertise page covers the upstream framing, meaning which process deserves an agent, how much autonomy to grant it and how to measure the result. If you already know which business action to automate, this page is the right one. If the question is still about scope, start with the expertise page.
Which tools do you want to expose to the model?
Tell us what Claude should be able to read and write. 30 minutes to set the tool surface, the region and the budget.
Discuss my Claude project