Package implementing adapter from DIAL Chat Completions API to Anthropic API
Project description
Python SDK for adapter from DIAL API to Anthropic API
Overview
The package exposes Claude models via two APIs:
| API | Description |
|---|---|
| Chat Completions API | The AI DIAL Chat Completion API adapted to the Anthropic Messages API |
| Anthropic API | The native Anthropic Messages API served in the passthrough mode |
Chat Completions API
The package provides an adapter from the Chat Completions API (ingress) to the Anthropic Messages API (upstream).
Basic request
The standard Chat Completions request fields are supported as follows:
| Field | Support |
|---|---|
messages |
Supported, including the system, developer, tool and function messages. See Multi-modal inputs for the non-text content |
stream, stop, top_p |
Relayed to the Anthropic API as-is |
temperature |
Mapped from the OpenAI [0, 2] range to the Anthropic [0, 1] range |
n |
Supported via parallel requests to the upstream |
max_tokens |
See Maximum completion tokens |
response_format |
See Structured outputs |
tools, tool_choice |
See Function calling |
reasoning_effort |
See Reasoning effort |
seed |
Unsupported by Claude, ignored |
The token usage is reported in the usage object, including completion_tokens_details.reasoning_tokens and the cache counters (see Prompt caching).
Structured outputs
response_format of type json_schema is supported via the Anthropic structured outputs. Claude accepts no other value of additionalProperties but false, so the adapter sets it throughout the schema.
The json_object type is unsupported and ignored.
Maximum completion tokens
Unlike OpenAI models, Claude models require the max_tokens parameter. When the request omits it, the adapter falls back to the default configured by the host application.
We recommend configuring the default on a per-model basis in the DIAL Core config instead, since all the token-related information (like pricing and token limits) is then kept in the same place. The DIAL Core default takes precedence over the adapter one.
DIAL Core configuration
{
"models": {
"${DIAL_DEPLOYMENT_ID}": {
"type": "chat",
"endpoint": "...",
"defaults": {
"max_tokens": 2048
}
}
}
}
Make sure the default doesn't exceed the max output tokens of the model, otherwise the request fails with an error like max_tokens: 10000 > 8192, which is the maximum allowed number of output tokens for claude-....
Function calling
The tools and tool_choice fields are supported, tool_choice including the auto, none, required and named-function modes.
The legacy Functions API (functions and function_call) is supported as well and converted to tools transparently. Claude may generate more than one call per response, while the Functions API allows a single one; the extra calls are discarded in this mode.
Multi-modal inputs
| Content part type | Support |
|---|---|
text |
Supported |
image_url |
The image_url.url field is a file URL |
file |
The file.file_data field is either a data URL or a base64-encoded PDF. The file_id field is unsupported |
input_audio, refusal |
Unsupported |
Request with an image content part
{
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "Describe the image"},
{
"type": "image_url",
"image_url": {"url": "$file_url"}
}
]
}
]
}
Request with a document content part
{
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "Summarize the document"},
{
"type": "file",
"file": {
"filename": "report.pdf",
"file_data": "data:application/pdf;base64,JVBERi0xLjQK..."
}
}
]
}
]
}
Files of any supported type may also be passed as DIAL attachments.
File URL
The $file_url referenced in the examples is one of the three:
| Mode | Example |
|---|---|
| Relative DIAL URL | files/${DIAL_BUCKET}/images/cat.png |
| Public URL | https://example.com/images/cat.png |
| Data URL | data:image/png;base64,iVBORw0KGgo... |
The relative URLs are resolved against the DIAL file storage and downloaded with the caller's API key, which requires the host application to be configured with the storage. Any other URL is downloaded as-is, without the credentials.
Reasoning effort
The reasoning_effort field sets the effort level of the response. It only accepts the OpenAI values, so the Claude-specific xhigh and max levels are reachable via the configuration alone.
DIAL extensions
The features below are the DIAL extensions of the Chat Completions API.
Attachments
The attachments are passed in the custom_content.attachments field of a message. An attachment either points to the file via url — a file URL — or carries it inline in the base64-encoded data field. The type field may be omitted as long as the MIME type is derivable from the URL. The supported types are:
| Type | MIME types |
|---|---|
| Images | image/png, image/jpeg, image/gif, image/webp |
| PDF documents | application/pdf |
| Text documents | text/plain, text/html, text/css, text/javascript, text/x-typescript, text/csv, text/markdown, text/x-python, text/xml, text/rtf, application/json |
The documents are supported only by the models with PDF support.
Request with image attachments
{
"messages": [
{
"role": "user",
"content": "Is there any difference between these images?",
"custom_content": {
"attachments": [
{
"type": "image/png",
"url": "$file_url"
},
{
"type": "image/png",
"data": "iVBORw0KGgo..."
}
]
}
}
]
}
Request with document attachments
{
"messages": [
{
"role": "user",
"content": "Summarize the documents",
"custom_content": {
"attachments": [
{
"type": "application/pdf",
"url": "$file_url"
},
{
"type": "application/pdf",
"data": "JVBERi0xLjQK..."
}
]
}
}
]
}
Setting enable_citations in the configuration makes Claude cite the documents it used; the citations are returned as numbered DIAL attachments.
Configuration
The adapter accepts a per-request configuration object in the custom_fields.configuration field. All its fields are optional; the host application serves its JSON Schema via the DIAL /configuration endpoint.
| Field | Description |
|---|---|
thinking |
Extended thinking configuration |
effort |
Reasoning level of the response |
betas |
List of beta feature flags to enable, e.g. ["token-efficient-tools-2025-02-19"] |
enable_citations |
Enables citations for the document attachments. Defaults to false |
Not every Claude deployment supports every field or beta flag; consult the official documentation before use.
Request with an example configuration
{
"messages": [
{"role": "user", "content": "Hello!"}
],
"custom_fields": {
"configuration": {
"thinking": {"type": "adaptive"},
"effort": "high",
"betas": ["token-efficient-tools-2025-02-19"],
"enable_citations": true
}
}
}
Extended thinking
The thinking object is relayed to the Anthropic API as-is, so any thinking configuration is supported:
| Configuration | Comment |
|---|---|
{"type": "adaptive"} |
The model decides when to think |
{"type": "enabled", "budget_tokens": 1024} |
Thinking with the given limit on reasoning tokens |
{"type": "disabled"} |
Thinking disabled |
The thinking blocks are reported in a dedicated Thinking stage and preserved across conversation turns, so multi-turn tool use works with thinking enabled.
temperature is ignored when thinking is enabled; top_p is ignored as well when thinking is adaptive, since Claude rejects both parameters in these modes.
Reasoning level
The effort field extends the reasoning effort with the Claude-specific levels: low, medium, high, xhigh and max. The value is relayed to the Anthropic API as-is, so the levels added later work without an adapter update.
Setting both effort and reasoning_effort to different values is a validation error.
Web search
Web search gives Claude direct access to real-time web content, allowing it to answer questions with up-to-date information beyond its knowledge cutoff. It is an Anthropic server-side tool: the searches are executed on Anthropic's side, and the final response includes citations for the sources used. See Web search tool in the Anthropic docs.
To enable web search, add a static tool named web_search to the request's tools list. The Anthropic web search tool definition goes into static_function.configuration; the name is defaulted from the static function, so you don't have to repeat it. Being a server-side tool, web search never forces a tool_choice and can be combined with ordinary function tools.
Request with Web search
{
"model": "claude-opus-4-8",
"messages": [
{"role": "user", "content": "What is the weather in NYC?"}
],
"tools": [
{
"type": "static_function",
"static_function": {
"name": "web_search",
"configuration": {
"type": "web_search_20250305"
}
}
}
]
}
The tool definition supports optional fields such as max_uses, allowed_domains, blocked_domains, and user_location. See Tool definition in the Anthropic docs.
Request with configured Web search
{
"model": "claude-opus-4-8",
"messages": [
{"role": "user", "content": "What is the weather in San Francisco?"}
],
"tools": [
{
"type": "static_function",
"static_function": {
"name": "web_search",
"configuration": {
"type": "web_search_20250305",
"max_uses": 5,
"allowed_domains": ["example.com", "trusteddomain.org"],
"user_location": {
"type": "approximate",
"city": "San Francisco",
"region": "California",
"country": "US",
"timezone": "America/Los_Angeles"
}
}
}
}
]
}
Prompt truncation
When max_prompt_tokens is set, the adapter discards the oldest messages until the prompt fits the limit, keeping the system prompt and the last message. The indices of the discarded messages are reported in the statistics.discarded_messages field of the response.
Request with a prompt token limit
{
"max_prompt_tokens": 1024,
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{
"role": "user",
"content": "Summarize the transcript: ${A_TRANSCRIPT_OVER_1024_TOKENS}"
},
{
"role": "assistant",
"content": "The speakers agree to revisit the Q3 plan in October, ..."
},
{"role": "user", "content": "What is the capital of France?"}
]
}
Response with the discarded messages
{
"choices": [
{
"index": 0,
"finish_reason": "stop",
"message": {"role": "assistant", "content": "Paris."}
}
],
"usage": {
"prompt_tokens": 23,
"completion_tokens": 3,
"total_tokens": 26
},
"statistics": {
"discarded_messages": [1, 2]
}
}
The summarization turn alone busts the limit, so both of its messages are discarded and only the system prompt and the last question reach the model.
The token counting is delegated to the Anthropic count tokens endpoint. For the backends that don't implement it, the host application may supply the bundled approximate tokenizer instead, which deliberately overestimates the token count, so that the truncated prompt never overflows the limit.
Prompt caching
Automatic caching
Automatic caching is the simplest way to use prompt caching. A single top-level cache breakpoint instructs Anthropic to automatically apply a cache point to the last cacheable block of the request. This is ideal for multi-turn conversations where the growing message history should be cached automatically. See Automatic caching in the Anthropic docs.
To enable automatic caching, set custom_fields.cache_breakpoint at the top level of the Chat Completion request:
Top-level cache breakpoint
{
"model": "claude-3-5-sonnet-20241022",
"messages": [
{"role": "user", "content": "Hello!"}
],
"custom_fields": {
"cache_breakpoint": {}
}
}
Explicit cache breakpoints
Explicit cache breakpoints give fine-grained control over which parts of the prompt get cached. You can place a cache breakpoint on individual system messages, user/assistant messages, or tool definitions. See Explicit cache breakpoints in the Anthropic docs.
To add a breakpoint, set custom_fields.cache_breakpoint on a message or tool object:
System cache breakpoint
{
"model": "claude-3-5-sonnet-20241022",
"messages": [
{
"role": "system",
"content": "You are a helpful assistant with extensive knowledge.",
"custom_fields": {
"cache_breakpoint": {}
}
},
{"role": "user", "content": "Hello!"}
]
}
Message cache breakpoint
{
"model": "claude-3-5-sonnet-20241022",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{
"role": "user",
"content": "Here is a long document: ...",
"custom_fields": {
"cache_breakpoint": {}
}
},
{"role": "user", "content": "Summarize it."}
]
}
Tools cache breakpoint
{
"model": "claude-3-5-sonnet-20241022",
"messages": [
{"role": "user", "content": "What's the weather?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string"}
},
"required": ["location"]
}
},
"custom_fields": {
"cache_breakpoint": {}
}
}
]
}
TTL support
A cache breakpoint may include an optional ttl field. Supported values are 5m (5 minutes, default) and 1h (one hour). The ttl field is supported on both top-level and explicit breakpoints. See TTL support in the Anthropic docs.
Top-level cache breakpoint with TTL
{
"model": "claude-3-5-sonnet-20241022",
"messages": [
{"role": "user", "content": "Hello!"}
],
"custom_fields": {
"cache_breakpoint": {
"ttl": "1h"
}
}
}
DIAL Core configuration
When a DIAL deployment has multiple upstreams, caching only pays off if the requests sharing a prefix reach the same upstream. Enable the corresponding feature flag in the DIAL Core config to make DIAL Core route them consistently: cacheSupported for explicit breakpoints, autoCachingSupported for automatic caching.
A top-level cache breakpoint may also be preset for all the requests to the deployment via defaults:
DIAL Core configuration
{
"models": {
"${DIAL_DEPLOYMENT_ID}": {
"type": "chat",
"endpoint": "...",
"defaults": {
"custom_fields": {
"cache_breakpoint": {}
}
},
"features": {
"autoCachingSupported": true
},
"upstreams": ["..."]
}
}
}
The cache usage is reported in the usage.prompt_tokens_details object: cached_tokens for the cache hits and cache_write_tokens for the tokens written to the cache.
Anthropic API
The package supports the native Anthropic Messages API in the passthrough mode: the requests are forwarded to the upstream as-is and the upstream errors are relayed to the caller in the native Anthropic error schema.
The exposed API is compatible with the vanilla client from the Anthropic SDK:
from anthropic import Anthropic, AsyncAnthropic
client = Anthropic(api_key="...", base_url="${ADAPTER_ORIGIN}/anthropic")
Usage
Mount the passthrough onto any Starlette/FastAPI host application (e.g. a DIALApp) with mount_anthropic_api. The upstream client is chosen per request by a factory you supply:
from aidial_sdk import DIALApp
from anthropic import AsyncAnthropic
from aidial_adapter_anthropic.passthrough import mount_anthropic_api
app = DIALApp(...)
async def get_client(request):
return AsyncAnthropic(api_key=...)
mount_anthropic_api(app, get_client)
The passthrough is mounted at /anthropic by default; pass path=... to change it. The get_client argument may also be a plain client instance instead of a factory.
Proxied endpoints
The following Anthropic endpoints are forwarded (relative to the mount path):
POST /v1/messages— create a message (streaming and non-streaming)POST /v1/messages/batches— create a message batchPOST /v1/messages/count_tokens— count tokens
Supported backends
The client factory may return any of the Anthropic SDK's async clients: AsyncAnthropic, AsyncAnthropicBedrock, AsyncAnthropicBedrockMantle, AsyncAnthropicVertex, and AsyncAnthropicFoundry.
The Bedrock backends require botocore, which is an optional dependency:
pip install aidial-adapter-anthropic[bedrock]
Endpoints a backend does not implement (e.g. Bedrock has no token-counting or batches route) surface as a 404 error.
Development Environment
This project requires Python ≥3.11 and Poetry ≥2.1.1 for dependency management.
Setup
-
Install Poetry. See the official installation guide.
-
(Optional) Specify custom Python or Poetry executables in
.env.dev. This is useful if multiple versions are installed. By default,pythonandpoetryare used.POETRY_PYTHON=path-to-python-exe POETRY=path-to-poetry-exe
-
Create and activate the virtual environment:
make init_env source .venv/bin/activate
-
Install project dependencies (including linting, formatting, and test tools):
make install
Lint
Run the linting before committing:
make lint
To auto-fix formatting issues run:
make format
Test
Run unit tests locally for available python versions:
make test
Run unit tests for the specific python version:
make test PYTHON=3.13
Clean
To remove the virtual environment and build artifacts run:
make clean
Build
To build the package run:
make build
Publish
To publish the package to PyPI run:
make publish
Git hooks
You may optionally install Git hooks that will automatically run the linting step on Git push. You only need to do it once for the given repository.
make install_git_hooks
[!IMPORTANT] This command doesn't work if you have already installed Git hooks locally or globally.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file aidial_adapter_anthropic-0.17.0.dev1.tar.gz.
File metadata
- Download URL: aidial_adapter_anthropic-0.17.0.dev1.tar.gz
- Upload date:
- Size: 58.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/2.2.1 CPython/3.11.15 Linux/6.17.0-1020-azure
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
657233357a1fc44a72b3f08c02f76c7b8598f02c297818ebd373b710c167c7b2
|
|
| MD5 |
4c9e0063000461abdd4a1710308ded5c
|
|
| BLAKE2b-256 |
b1845e6f0066291bd97856c9d642873ff2b77d633fd29a94e3c19f9bb089deb9
|
File details
Details for the file aidial_adapter_anthropic-0.17.0.dev1-py3-none-any.whl.
File metadata
- Download URL: aidial_adapter_anthropic-0.17.0.dev1-py3-none-any.whl
- Upload date:
- Size: 70.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/2.2.1 CPython/3.11.15 Linux/6.17.0-1020-azure
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cb1fb89b890210c652fa23e5b9b9b6685b0a1fd321cb80c47ed44f82932432de
|
|
| MD5 |
695f77f71a0a8332aae1304df0092085
|
|
| BLAKE2b-256 |
e2e1efffb56baf4c6502e4bc45d9c30a3b8cbe107e702fa5747149f30e2df5e9
|