Skip to main content

Package implementing adapter from DIAL Chat Completions API to Anthropic API

Project description

Python SDK for adapter from DIAL API to Anthropic API

About DIALX

PyPI version Discord


Overview

The package exposes Claude models via two APIs:

API Description
Chat Completions API The AI DIAL Chat Completion API adapted to the Anthropic Messages API
Anthropic API The native Anthropic Messages API served in the passthrough mode

Chat Completions API

The package provides an adapter from the Chat Completions API (ingress) to the Anthropic Messages API (upstream).

Basic request

The standard Chat Completions request fields are supported as follows:

Field Support
messages Supported, including the system, developer, tool and function messages. See Multi-modal inputs for the non-text content
stream, stop, top_p Relayed to the Anthropic API as-is
temperature Mapped from the OpenAI [0, 2] range to the Anthropic [0, 1] range
n Supported via parallel requests to the upstream
max_tokens See Maximum completion tokens
response_format See Structured outputs
tools, tool_choice See Function calling
reasoning_effort See Reasoning effort
seed Unsupported by Claude, ignored

The token usage is reported in the usage object, including completion_tokens_details.reasoning_tokens and the cache counters (see Prompt caching).

Structured outputs

response_format of type json_schema is supported via the Anthropic structured outputs. Claude accepts no other value of additionalProperties but false, so the adapter sets it throughout the schema.

The json_object type is unsupported and ignored.

Maximum completion tokens

Unlike OpenAI models, Claude models require the max_tokens parameter. When the request omits it, the adapter falls back to the default configured by the host application.

We recommend configuring the default on a per-model basis in the DIAL Core config instead, since all the token-related information (like pricing and token limits) is then kept in the same place. The DIAL Core default takes precedence over the adapter one.

DIAL Core configuration
{
  "models": {
    "${DIAL_DEPLOYMENT_ID}": {
      "type": "chat",
      "endpoint": "...",
      "defaults": {
        "max_tokens": 2048
      }
    }
  }
}

Make sure the default doesn't exceed the max output tokens of the model, otherwise the request fails with an error like max_tokens: 10000 > 8192, which is the maximum allowed number of output tokens for claude-....

Function calling

The tools and tool_choice fields are supported, tool_choice including the auto, none, required and named-function modes.

The legacy Functions API (functions and function_call) is supported as well and converted to tools transparently. Claude may generate more than one call per response, while the Functions API allows a single one; the extra calls are discarded in this mode.

Multi-modal inputs

Content part type Support
text Supported
image_url The image_url.url field is a file URL
file The file.file_data field is either a data URL or a base64-encoded PDF. The file_id field is unsupported
input_audio, refusal Unsupported
Request with an image content part
{
  "messages": [
    {
      "role": "user",
      "content": [
        {"type": "text", "text": "Describe the image"},
        {
          "type": "image_url",
          "image_url": {"url": "$file_url"}
        }
      ]
    }
  ]
}
Request with a document content part
{
  "messages": [
    {
      "role": "user",
      "content": [
        {"type": "text", "text": "Summarize the document"},
        {
          "type": "file",
          "file": {
            "filename": "report.pdf",
            "file_data": "data:application/pdf;base64,JVBERi0xLjQK..."
          }
        }
      ]
    }
  ]
}

Files of any supported type may also be passed as DIAL attachments.

File URL

The $file_url referenced in the examples is one of the three:

Mode Example
Relative DIAL URL files/${DIAL_BUCKET}/images/cat.png
Public URL https://example.com/images/cat.png
Data URL data:image/png;base64,iVBORw0KGgo...

The relative URLs are resolved against the DIAL file storage and downloaded with the caller's API key, which requires the host application to be configured with the storage. Any other URL is downloaded as-is, without the credentials.

Reasoning effort

The reasoning_effort field sets the effort level of the response. It only accepts the OpenAI values, so the Claude-specific xhigh and max levels are reachable via the configuration alone.

DIAL extensions

The features below are the DIAL extensions of the Chat Completions API.

Attachments

The attachments are passed in the custom_content.attachments field of a message. An attachment either points to the file via url — a file URL — or carries it inline in the base64-encoded data field. The type field may be omitted as long as the MIME type is derivable from the URL. The supported types are:

Type MIME types
Images image/png, image/jpeg, image/gif, image/webp
PDF documents application/pdf
Text documents text/plain, text/html, text/css, text/javascript, text/x-typescript, text/csv, text/markdown, text/x-python, text/xml, text/rtf, application/json

The documents are supported only by the models with PDF support.

Request with image attachments
{
  "messages": [
    {
      "role": "user",
      "content": "Is there any difference between these images?",
      "custom_content": {
        "attachments": [
          {
            "type": "image/png",
            "url": "$file_url"
          },
          {
            "type": "image/png",
            "data": "iVBORw0KGgo..."
          }
        ]
      }
    }
  ]
}
Request with document attachments
{
  "messages": [
    {
      "role": "user",
      "content": "Summarize the documents",
      "custom_content": {
        "attachments": [
          {
            "type": "application/pdf",
            "url": "$file_url"
          },
          {
            "type": "application/pdf",
            "data": "JVBERi0xLjQK..."
          }
        ]
      }
    }
  ]
}

Setting enable_citations in the configuration makes Claude cite the documents it used; the citations are returned as numbered DIAL attachments.

Configuration

The adapter accepts a per-request configuration object in the custom_fields.configuration field. All its fields are optional; the host application serves its JSON Schema via the DIAL /configuration endpoint.

Field Description
thinking Extended thinking configuration
effort Reasoning level of the response
betas List of beta feature flags to enable, e.g. ["token-efficient-tools-2025-02-19"]
enable_citations Enables citations for the document attachments. Defaults to false

Not every Claude deployment supports every field or beta flag; consult the official documentation before use.

Request with an example configuration
{
  "messages": [
    {"role": "user", "content": "Hello!"}
  ],
  "custom_fields": {
    "configuration": {
      "thinking": {"type": "adaptive"},
      "effort": "high",
      "betas": ["token-efficient-tools-2025-02-19"],
      "enable_citations": true
    }
  }
}
Extended thinking

The thinking object is relayed to the Anthropic API as-is, so any thinking configuration is supported:

Configuration Comment
{"type": "adaptive"} The model decides when to think
{"type": "enabled", "budget_tokens": 1024} Thinking with the given limit on reasoning tokens
{"type": "disabled"} Thinking disabled

The thinking blocks are reported in a dedicated Thinking stage and preserved across conversation turns, so multi-turn tool use works with thinking enabled.

temperature is ignored when thinking is enabled; top_p is ignored as well when thinking is adaptive, since Claude rejects both parameters in these modes.

Reasoning level

The effort field extends the reasoning effort with the Claude-specific levels: low, medium, high, xhigh and max. The value is relayed to the Anthropic API as-is, so the levels added later work without an adapter update.

Setting both effort and reasoning_effort to different values is a validation error.

Web search

Web search gives Claude direct access to real-time web content, allowing it to answer questions with up-to-date information beyond its knowledge cutoff. It is an Anthropic server-side tool: the searches are executed on Anthropic's side, and the final response includes citations for the sources used. See Web search tool in the Anthropic docs.

To enable web search, add a static tool named web_search to the request's tools list. The Anthropic web search tool definition goes into static_function.configuration; the name is defaulted from the static function, so you don't have to repeat it. Being a server-side tool, web search never forces a tool_choice and can be combined with ordinary function tools.

Request with Web search
{
  "model": "claude-opus-4-8",
  "messages": [
    {"role": "user", "content": "What is the weather in NYC?"}
  ],
  "tools": [
    {
      "type": "static_function",
      "static_function": {
        "name": "web_search",
        "configuration": {
          "type": "web_search_20250305"
        }
      }
    }
  ]
}

The tool definition supports optional fields such as max_uses, allowed_domains, blocked_domains, and user_location. See Tool definition in the Anthropic docs.

Request with configured Web search
{
  "model": "claude-opus-4-8",
  "messages": [
    {"role": "user", "content": "What is the weather in San Francisco?"}
  ],
  "tools": [
    {
      "type": "static_function",
      "static_function": {
        "name": "web_search",
        "configuration": {
          "type": "web_search_20250305",
          "max_uses": 5,
          "allowed_domains": ["example.com", "trusteddomain.org"],
          "user_location": {
            "type": "approximate",
            "city": "San Francisco",
            "region": "California",
            "country": "US",
            "timezone": "America/Los_Angeles"
          }
        }
      }
    }
  ]
}

Prompt truncation

When max_prompt_tokens is set, the adapter discards the oldest messages until the prompt fits the limit, keeping the system prompt and the last message. The indices of the discarded messages are reported in the statistics.discarded_messages field of the response.

Request with a prompt token limit
{
  "max_prompt_tokens": 1024,
  "messages": [
    {"role": "system", "content": "You are a helpful assistant."},
    {
      "role": "user",
      "content": "Summarize the transcript: ${A_TRANSCRIPT_OVER_1024_TOKENS}"
    },
    {
      "role": "assistant",
      "content": "The speakers agree to revisit the Q3 plan in October, ..."
    },
    {"role": "user", "content": "What is the capital of France?"}
  ]
}
Response with the discarded messages
{
  "choices": [
    {
      "index": 0,
      "finish_reason": "stop",
      "message": {"role": "assistant", "content": "Paris."}
    }
  ],
  "usage": {
    "prompt_tokens": 23,
    "completion_tokens": 3,
    "total_tokens": 26
  },
  "statistics": {
    "discarded_messages": [1, 2]
  }
}

The summarization turn alone busts the limit, so both of its messages are discarded and only the system prompt and the last question reach the model.

The token counting is delegated to the Anthropic count tokens endpoint. For the backends that don't implement it, the host application may supply the bundled approximate tokenizer instead, which deliberately overestimates the token count, so that the truncated prompt never overflows the limit.

Prompt caching

Automatic caching

Automatic caching is the simplest way to use prompt caching. A single top-level cache breakpoint instructs Anthropic to automatically apply a cache point to the last cacheable block of the request. This is ideal for multi-turn conversations where the growing message history should be cached automatically. See Automatic caching in the Anthropic docs.

To enable automatic caching, set custom_fields.cache_breakpoint at the top level of the Chat Completion request:

Top-level cache breakpoint
{
  "model": "claude-3-5-sonnet-20241022",
  "messages": [
    {"role": "user", "content": "Hello!"}
  ],
  "custom_fields": {
    "cache_breakpoint": {}
  }
}
Explicit cache breakpoints

Explicit cache breakpoints give fine-grained control over which parts of the prompt get cached. You can place a cache breakpoint on individual system messages, user/assistant messages, or tool definitions. See Explicit cache breakpoints in the Anthropic docs.

To add a breakpoint, set custom_fields.cache_breakpoint on a message or tool object:

System cache breakpoint
{
  "model": "claude-3-5-sonnet-20241022",
  "messages": [
    {
      "role": "system",
      "content": "You are a helpful assistant with extensive knowledge.",
      "custom_fields": {
        "cache_breakpoint": {}
      }
    },
    {"role": "user", "content": "Hello!"}
  ]
}
Message cache breakpoint
{
  "model": "claude-3-5-sonnet-20241022",
  "messages": [
    {"role": "system", "content": "You are a helpful assistant."},
    {
      "role": "user",
      "content": "Here is a long document: ...",
      "custom_fields": {
        "cache_breakpoint": {}
      }
    },
    {"role": "user", "content": "Summarize it."}
  ]
}
Tools cache breakpoint
{
  "model": "claude-3-5-sonnet-20241022",
  "messages": [
    {"role": "user", "content": "What's the weather?"}
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the current weather",
        "parameters": {
          "type": "object",
          "properties": {
            "location": {"type": "string"}
          },
          "required": ["location"]
        }
      },
      "custom_fields": {
        "cache_breakpoint": {}
      }
    }
  ]
}
TTL support

A cache breakpoint may include an optional ttl field. Supported values are 5m (5 minutes, default) and 1h (one hour). The ttl field is supported on both top-level and explicit breakpoints. See TTL support in the Anthropic docs.

Top-level cache breakpoint with TTL
{
  "model": "claude-3-5-sonnet-20241022",
  "messages": [
    {"role": "user", "content": "Hello!"}
  ],
  "custom_fields": {
    "cache_breakpoint": {
      "ttl": "1h"
    }
  }
}
DIAL Core configuration

When a DIAL deployment has multiple upstreams, caching only pays off if the requests sharing a prefix reach the same upstream. Enable the corresponding feature flag in the DIAL Core config to make DIAL Core route them consistently: cacheSupported for explicit breakpoints, autoCachingSupported for automatic caching.

A top-level cache breakpoint may also be preset for all the requests to the deployment via defaults:

DIAL Core configuration
{
  "models": {
    "${DIAL_DEPLOYMENT_ID}": {
      "type": "chat",
      "endpoint": "...",
      "defaults": {
        "custom_fields": {
          "cache_breakpoint": {}
        }
      },
      "features": {
        "autoCachingSupported": true
      },
      "upstreams": ["..."]
    }
  }
}

The cache usage is reported in the usage.prompt_tokens_details object: cached_tokens for the cache hits and cache_write_tokens for the tokens written to the cache.


Anthropic API

The package supports the native Anthropic Messages API in the passthrough mode: the requests are forwarded to the upstream as-is and the upstream errors are relayed to the caller in the native Anthropic error schema.

The exposed API is compatible with the vanilla client from the Anthropic SDK:

from anthropic import Anthropic, AsyncAnthropic
client = Anthropic(api_key="...", base_url="${ADAPTER_ORIGIN}/anthropic")

Usage

Mount the passthrough onto any Starlette/FastAPI host application (e.g. a DIALApp) with mount_anthropic_api. The upstream client is chosen per request by a factory you supply:

from aidial_sdk import DIALApp
from anthropic import AsyncAnthropic
from aidial_adapter_anthropic.passthrough import mount_anthropic_api

app = DIALApp(...)

async def get_client(request):
    return AsyncAnthropic(api_key=...)

mount_anthropic_api(app, get_client)

The passthrough is mounted at /anthropic by default; pass path=... to change it. The get_client argument may also be a plain client instance instead of a factory.

Proxied endpoints

The following Anthropic endpoints are forwarded (relative to the mount path):

  • POST /v1/messages — create a message (streaming and non-streaming)
  • POST /v1/messages/batches — create a message batch
  • POST /v1/messages/count_tokens — count tokens

Supported backends

The client factory may return any of the Anthropic SDK's async clients: AsyncAnthropic, AsyncAnthropicBedrock, AsyncAnthropicBedrockMantle, AsyncAnthropicVertex, and AsyncAnthropicFoundry.

The Bedrock backends require botocore, which is an optional dependency:

pip install aidial-adapter-anthropic[bedrock]

Endpoints a backend does not implement (e.g. Bedrock has no token-counting or batches route) surface as a 404 error.


Development Environment

This project requires Python ≥3.11 and Poetry ≥2.1.1 for dependency management.

Setup

  1. Install Poetry. See the official installation guide.

  2. (Optional) Specify custom Python or Poetry executables in .env.dev. This is useful if multiple versions are installed. By default, python and poetry are used.

    POETRY_PYTHON=path-to-python-exe
    POETRY=path-to-poetry-exe
    
  3. Create and activate the virtual environment:

    make init_env
    source .venv/bin/activate
    
  4. Install project dependencies (including linting, formatting, and test tools):

    make install
    

Lint

Run the linting before committing:

make lint

To auto-fix formatting issues run:

make format

Test

Run unit tests locally for available python versions:

make test

Run unit tests for the specific python version:

make test PYTHON=3.13

Clean

To remove the virtual environment and build artifacts run:

make clean

Build

To build the package run:

make build

Publish

To publish the package to PyPI run:

make publish

Git hooks

You may optionally install Git hooks that will automatically run the linting step on Git push. You only need to do it once for the given repository.

make install_git_hooks

[!IMPORTANT] This command doesn't work if you have already installed Git hooks locally or globally.

Project details


Release history Release notifications | RSS feed

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

aidial_adapter_anthropic-0.16.0.dev8.tar.gz (58.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

aidial_adapter_anthropic-0.16.0.dev8-py3-none-any.whl (70.8 kB view details)

Uploaded Python 3

File details

Details for the file aidial_adapter_anthropic-0.16.0.dev8.tar.gz.

File metadata

  • Download URL: aidial_adapter_anthropic-0.16.0.dev8.tar.gz
  • Upload date:
  • Size: 58.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: poetry/2.2.1 CPython/3.11.15 Linux/6.17.0-1020-azure

File hashes

Hashes for aidial_adapter_anthropic-0.16.0.dev8.tar.gz
Algorithm Hash digest
SHA256 eea1ee6bd83fc14896b51fb600d6fe1a281508b30f8648c40ffcad100c48f9ff
MD5 de50e363ccb104aed3e6ae67c99d8f5f
BLAKE2b-256 0c7a8876d54838a1a3d01ea434ab473dae9b39a340b7e21b531fc11d44efac01

See more details on using hashes here.

File details

Details for the file aidial_adapter_anthropic-0.16.0.dev8-py3-none-any.whl.

File metadata

File hashes

Hashes for aidial_adapter_anthropic-0.16.0.dev8-py3-none-any.whl
Algorithm Hash digest
SHA256 6946dc3c5d9d59dc6617d6a829518c2d3751e1a67355bfc5895ca5d0c332b644
MD5 c6917e1670f5d2930f277345094c22fb
BLAKE2b-256 855b588efc831653b23e8724c98a50755216c0b1024af62358f562c646eb4353

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page