Skip to main content

Lightweight, pluggable adapter for multiple LLM APIs (OpenAI, Anthropic, Google)

Project description

LLM API Adapter SDK for Python

PyPI Downloads Tests codecov

Overview

Calling an LLM from Python currently means one of three paths: lock yourself to one provider's SDK, absorb LangChain's framework complexity, or pull in LiteLLM and its 100+ transitive dependencies.

llm-api-adapter is the fourth option: a minimal adapter for OpenAI, Anthropic, and Google. One class, one chat() method, one error hierarchy — regardless of which provider is behind it. Switching providers is changing two arguments.

Why this library?

llm-api-adapter LiteLLM LangChain Provider SDK
Single runtime dependency ✓ (requests) ✗ (100+) ✗ (100+)
Unified reasoning control partial¹ partial¹
Built-in cost accounting via callbacks
Unified error hierarchy partial
Sync streaming
Async API
Number of providers 3 100+ 50+ 1

¹ LiteLLM and LangChain expose reasoning via provider-specific parameters — there is no single unified parameter that works identically across providers.

Small footprint. The only runtime dependency is requests. LiteLLM and LangChain each pull in 100+ transitive packages; this doesn't.

Unified reasoning. One reasoning_level parameter works across OpenAI o-series, Anthropic extended thinking, and Google thinking models — same parameter name, same string levels, no provider-specific kwargs.

Cost accounting, built-in. Every response carries cost_input, cost_output, and cost_total in your chosen currency. No external tooling required.

Predictable errors. One exception hierarchy across all three providers. LLMAPIRateLimitError means the same thing whether you called OpenAI or Anthropic.

Use this when

  • You call LLMs directly — no chains, no agents, no orchestration
  • You want to stay provider-agnostic without adopting a framework
  • You need unified reasoning control or per-request cost tracking

Use something else when

  • LiteLLM — you need 100+ providers, async support, streaming, or a proxy/gateway layer
  • LangChain — you need chains, memory, RAG, or agents
  • Provider SDK directly — you'll never switch and don't need cost tracking

Features

  • Synchronous Streaming: Iterate normalized text deltas with stream_chat() across OpenAI, Anthropic, and Google; receive completed tool calls and the final normalized ChatResponse through optional callbacks.
  • Vision Input: Send images alongside text via ImagePart — URL or raw bytes, all providers handled automatically.
  • Tool / Function Calling: Provider-agnostic tool definitions and normalized tool calls in ChatResponse.tool_calls.
  • Strict JSON Mode: Pass a JSON Schema to chat() and get a parsed object in ChatResponse.parsed_json — provider normalization is handled automatically.
  • Pydantic Integration: Pass a Pydantic model as response_model and get a typed instance back in ChatResponse.parsed_model — no manual schema writing required.
  • Request Timeouts: Per-request timeout control via timeout_s; raises LLMAPITimeoutError on expiry.
  • Flexible Configuration: temperature, max_tokens, top_p, and other parameters passed through to the provider.
  • Pricing Registry: Model prices stored in a bundled JSON registry with per-model input/output rates; overridable per instance.

Installation

To install the SDK, you can use pip:

pip install llm-api-adapter

Note: You will need to obtain API keys from each LLM provider you wish to use (OpenAI, Anthropic, Google). Refer to their respective documentation for instructions on obtaining API keys.

Getting Started

Importing and Setting Up the Adapter

To start using the adapter, you need to import the necessary components:

from llm_api_adapter.models.messages.chat_message import (
    AIMessage, Prompt, UserMessage
)
from llm_api_adapter.universal_adapter import UniversalLLMAPIAdapter

Sending a Simple Request

The SDK supports three types of messages for interacting with the LLM:

  • Prompt: Use Prompt to set the context or initial prompt for the model.
  • UserMessage: Use UserMessage to send messages from the user during a conversation.
  • AIMessage: Use AIMessage to simulate responses from the assistant during a conversation.

Here is an example of how to send a simple request to the adapter:

messages = [
    UserMessage("Hi! Can you explain how artificial intelligence works?")
]

adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5",
    api_key=openai_api_key
)

response = adapter.chat(
    messages=messages,
    max_tokens=max_tokens,
    temperature=temperature,
    top_p=top_p
)
print(response.content)

Parameters

  • max_tokens: The maximum number of tokens to generate in the response. This limits the length of the output.

  • temperature: Controls the randomness of the response. Higher values (e.g., 0.8) make the output more random, while lower values (e.g., 0.2) make it more focused and deterministic. Default value: 1.0 (range: 0 to 2).

  • top_p: Limits the response to a certain cumulative probability. This is used to create more focused and coherent responses by considering only the highest probability options. Default value: 1.0 (range: 0 to 1).

Alternative Message Format

In addition to the built-in message classes, the SDK also supports the standard OpenAI-style message format for quick adoption and compatibility:

messages = [
    {"role": "system", "content": "You are a friendly assistant who answers only yes or no."},
    {"role": "user", "content": "Do you know how AI learns?"},
    {"role": "assistant", "content": "Yes."},
    {"role": "user", "content": "Can you explain it in one sentence?"}
]

response = adapter.chat(messages=messages, max_tokens=50)
print(response.content)

Note
The adapter automatically normalizes message input — you can mix custom message classes and OpenAI-style dicts in one list.

Streaming

stream_chat() is the synchronous streaming counterpart to chat(). It yields normalized visible text deltas as str, independently of the provider's native streaming event format.

from llm_api_adapter import UniversalLLMAPIAdapter

adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5.6-sol",
    api_key="...",
)

chunks = []
for delta in adapter.stream_chat(
    messages=[{"role": "user", "content": "Explain SSE in one sentence."}],
    on_delta=chunks.append,
    on_tool_call=lambda call: print(call.name, call.arguments),
    on_done=lambda response: print(response.usage),
):
    print(delta, end="", flush=True)
  • on_delta(delta) is called for every yielded visible text delta.
  • on_tool_call(tool_call) receives only completed, normalized ToolCall objects after the provider stream finishes.
  • on_done(response) receives the finalized ChatResponse, including usage, pricing, parsed structured output, and any tool calls.

Streaming in v0.6.0 is synchronous and text-first. Async streaming, buffer/backpressure controls, a universal provider-event type, partial tool arguments, and tool-call sequencing are deferred to later releases.

Provider event references: OpenAI Responses streaming, Anthropic streaming, and Google streamGenerateContent.

Handling Errors

Common Errors

The SDK provides a set of standardized errors for easier debugging and integration:

API Errors

  • LLMAPIError: Base class for all API-related errors. This error is also used for any unexpected LLM API errors.

  • LLMAPIAuthorizationError: Raised when authentication or authorization fails.

  • LLMAPIRateLimitError: Raised when rate limits are exceeded.

  • LLMAPITokenLimitError: Raised when token limits are exceeded.

  • LLMAPIClientError: Raised when the client makes an invalid request.

  • LLMAPIServerError: Raised when the server encounters an error.

  • LLMAPITimeoutError: Raised when a request times out.

  • LLMAPIUsageLimitError: Raised when usage limits are exceeded.

  • InvalidToolSchemaError: Raised when a provided tool schema is invalid.

  • InvalidToolArgumentsError: Raised when tool arguments cannot be parsed or validated.

  • ToolChoiceError: Raised when tool_choice is invalid or references an unknown tool.

  • JSONSchemaError: Raised when json_schema or response_model is used incorrectly (combined with each other or with tools), when response_model is not a Pydantic BaseModel subclass, when Pydantic is not installed, or when the model response is not valid JSON.

Config Errors

  • LLMConfigError: Raised when the request configuration is invalid or incompatible.

  • LLMReasoningLevelError: Raised only for Anthropic models when max_tokens is less than reasoning_level.

Provider Error Mapping

Exception OpenAI Anthropic Google
LLMAPIAuthorizationError HTTP 401; InvalidAuthenticationError, AuthenticationError HTTP 401; AuthenticationError, PermissionError HTTP 401/403; PERMISSION_DENIED
LLMAPIRateLimitError HTTP 429; RateLimitError HTTP 429; RateLimitError HTTP 429; RESOURCE_EXHAUSTED
LLMAPITokenLimitError MaxTokensExceededError, TokenLimitError
LLMAPIClientError HTTP 4xx; InvalidRequestError, BadRequestError HTTP 4xx; InvalidRequestError, RequestTooLargeError, NotFoundError HTTP 4xx; INVALID_ARGUMENT, FAILED_PRECONDITION, NOT_FOUND
LLMAPIServerError HTTP 5xx; InternalServerError, ServiceUnavailableError HTTP 5xx; APIError, OverloadedError HTTP 5xx; INTERNAL, UNAVAILABLE
LLMAPITimeoutError requests.Timeout; TimeoutError requests.Timeout requests.Timeout; DEADLINE_EXCEEDED
LLMAPIUsageLimitError UsageLimitError, QuotaExceededError
  • LLMAPITokenLimitError and LLMAPIUsageLimitError are OpenAI-only; equivalent cases for Anthropic and Google fall into LLMAPIClientError (HTTP 4xx).
  • LLMAPIClientError is the default fallback for all unhandled HTTP 4xx responses.
  • LLMAPIServerError is the default fallback for all HTTP 5xx responses.
  • InvalidToolSchemaError, InvalidToolArgumentsError, ToolChoiceError, JSONSchemaError — client-side errors (validated before the request is sent); inherit from LLMAPIClientError.
  • LLMConfigError, LLMReasoningLevelError — configuration errors (parameter validation before the request); LLMReasoningLevelError is only raised for Anthropic models with budget_tokens.

Configuration and Management

Using Different Providers and Models

The SDK allows you to easily switch between LLM providers and specify the model you want to use. Currently supported providers are OpenAI, Anthropic, and Google.

  • OpenAI: You can use models like gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna, gpt-5.5, gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, gpt-5.2, gpt-5.1, gpt-5, gpt-5-mini, gpt-5-nano, gpt-4.1, gpt-4.1-mini, gpt-4.1-nano, gpt-4o, gpt-4o-mini.
  • Anthropic: Available models include claude-fable-5, claude-sonnet-5, claude-opus-4-8, claude-opus-4-7, claude-opus-4-6, claude-sonnet-4-6, claude-opus-4-5, claude-sonnet-4-5, claude-haiku-4-5, claude-opus-4-1.
  • Google: Models such as gemini-3.5-flash, gemini-3.1-pro-preview, gemini-3.1-flash-lite, gemini-3-flash-preview, gemini-2.5-pro, gemini-2.5-flash can be used.

Example:

adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5",
    api_key=openai_api_key
)

To switch to another provider, simply change the organization and model parameters.

Switching Providers

Here is an example of how to switch between different LLM providers using the SDK:

Note: Each instance of UniversalLLMAPIAdapter is tied to a specific provider and model. You cannot change the organization parameter for an existing adapter object. To use a different provider, you must create a new instance.

gpt = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5",
    api_key=openai_api_key
)
gpt_response = gpt.chat(messages=messages)
print(gpt_response.content)

claude = UniversalLLMAPIAdapter(
    organization="anthropic",
    model="claude-sonnet-4-5",
    api_key=anthropic_api_key
)
claude_response = claude.chat(messages=messages)
print(claude_response.content)

google = UniversalLLMAPIAdapter(
    organization="google",
    model="gemini-2.5-flash",
    api_key=google_api_key
)
google_response = google.chat(messages=messages)
print(google_response.content)

Example Use Case

Here is a comprehensive example that showcases all possible message types and interactions:

from llm_api_adapter.models.messages.chat_message import (
    AIMessage, Prompt, UserMessage
)                                               
from llm_api_adapter.universal_adapter import UniversalLLMAPIAdapter

messages = [
    Prompt(
        "You are a friendly assistant who explains complex concepts "
        "in simple terms."
    ),
    UserMessage("Hi! Can you explain how artificial intelligence works?"),
    AIMessage(
        "Sure! Artificial intelligence (AI) is a system that can perform "
        "tasks requiring human-like intelligence, such as recognizing images "
        "or understanding language. It learns by analyzing large amounts of "
        "data, finding patterns, and making predictions."
    ),
    UserMessage("How does AI learn?"),
]

adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5",
    api_key=openai_api_key
)

response = adapter.chat(
    messages=messages,
    max_tokens=256,
    temperature=1.0,
    top_p=1.0
)
print(response.content)

The ChatResponse object returned by chat includes:

  1. model: The model that generated the response.
  2. response_id: Unique identifier for the response.
  3. timestamp: Response generation time.
  4. usage: Object containing input_tokens, output_tokens, and total_tokens.
  5. currency: The currency used for cost calculation.
  6. cost_input: Cost of input tokens.
  7. cost_output: Cost of output tokens.
  8. cost_total: Total combined cost.
  9. content: The generated text response.
  10. finish_reason: Reason why generation stopped (e.g., "stop", "length").
  11. parsed_json: Parsed JSON object when json_schema or response_model was provided, otherwise None.
  12. parsed_model: Typed Pydantic instance when response_model was provided, otherwise None.

Timeout Support

The SDK supports per-request timeouts for all providers.

timeout_s parameter

timeout_s defines the maximum time (in seconds) the SDK will wait for an LLM response.

If the timeout is exceeded, the request is aborted and LLMAPITimeoutError is raised.

Example

chat_params = {
    "messages": messages,
    "timeout_s": 2.5
}
gpt = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5.2",
    api_key=openai_api_key
)
response = gpt.chat(**chat_params)

Notes

  • Timeout is applied uniformly across all providers.
  • The parameter is optional; if omitted, the provider default is used.
  • Timeout affects the full request lifecycle (network + model execution).

Handling timeout errors

Timeouts raise a dedicated exception that can be handled explicitly:

from llm_api_adapter.errors import LLMAPITimeoutError

try:
    response = gpt.chat(**chat_params)
except LLMAPITimeoutError:
    # retry, fallback, or abort
    print("LLM request timed out")

Reasoning Support

This section describes the unified reasoning_level parameter that works the same way for all supported providers and their models.

response = adapter.chat(
    messages=[UserMessage("Solve this step-by-step")],
    reasoning_level=2048,
)

Default behavior

If reasoning_level is not passed, reasoning is:

  • fully disabled where the provider allows it, or
  • reduced to the minimal supported level if it cannot be turned off.

This keeps behavior consistent when switching providers or models.

reasoning_level parameter

reasoning_level is optional and provider‑agnostic.

Supported forms:

  • int — explicit numeric level
  • str — one of: "none", "low", "medium", "high"

Internal mapping:

{
  "none": 0,
  "low": 100,
  "medium": 1000,
  "high": 10000
}

String values are automatically converted to numbers, and numeric values can be normalized back to named levels.

Usage examples

# Named level
response = adapter.chat(
    messages=[UserMessage("Explain this")],
    reasoning_level="medium",
)

# Explicit numeric level
response = adapter.chat(
    messages=[UserMessage("Solve this step-by-step")],
    reasoning_level=2048,
)

# Reasoning disabled (default)
response = adapter.chat(
    messages=[UserMessage("Simple answer, no reasoning")],
)

Provider independence

reasoning_level has the same semantics for all providers:

  • same parameter name
  • same string levels
  • same numeric mapping

This allows switching between OpenAI, Anthropic, and Google without changing reasoning configuration in your code.

Tool / Function Calling

The SDK provides a unified provider‑agnostic tool calling interface.

Tools are defined using ToolSpec, and tool calls are returned in normalized form through ChatResponse.tool_calls.

The adapter does not execute tools. Tool execution must be implemented by the caller.

ToolSpec

from llm_api_adapter.models.tools import ToolSpec

tool = ToolSpec(
    name="get_weather",
    description="Get current weather for a city",
    json_schema={
        "type": "object",
        "properties": {
            "city": {"type": "string"}
        },
        "required": ["city"],
        "additionalProperties": False
    }
)

Tool parameters

chat() supports:

  • tools
  • tool_choice

Tool round‑trip example

import json
from typing import Any, Dict

from llm_api_adapter.models.messages.chat_message import (
    Prompt,
    UserMessage,
    AIMessage,
    ToolMessage,
)
from llm_api_adapter.models.tools import ToolSpec
from llm_api_adapter.universal_adapter import UniversalLLMAPIAdapter


tools = [
    ToolSpec(
        name="get_weather",
        description="Get current weather for a city",
        json_schema={
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
            "additionalProperties": False,
        },
    )
]


def run_tool(name: str, args: Dict[str, Any]) -> Dict[str, Any]:
    if name == "get_weather":
        return {"city": args["city"], "temperature": 22, "unit": "C"}
    raise ValueError(f"Unknown tool: {name}")


adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5.2",
    api_key=openai_api_key,
)

messages = [
    Prompt("If the user asks about weather, call get_weather."),
    UserMessage("What's the weather in Tel Aviv today?")
]

first = adapter.chat(
    messages=messages,
    tools=tools,
    tool_choice="auto",
    max_tokens=1000,
)

if first.tool_calls:
    messages.append(AIMessage(content="", tool_calls=first.tool_calls))

    for tc in first.tool_calls:
        result = run_tool(tc.name, tc.arguments)
        messages.append(
            ToolMessage(
                tool_call_id=tc.call_id,
                content=json.dumps(result)
            )
        )

    final = adapter.chat(
        messages=messages,
        previous_response=first,
        max_tokens=1000
    )
    print(final.content)

previous_response parameter

previous_response accepts the ChatResponse returned by an earlier chat() call.

For OpenAI models that use the Responses API (o-series and newer GPT models), the adapter extracts response_id from the previous response and passes it to the API as previous_response_id. This enables stateful multi-turn conversations where the model retains context server-side, which reduces the tokens you need to send in subsequent turns.

For Anthropic and Google, the parameter is accepted but ignored — context is carried entirely through the messages list regardless.

If you omit previous_response, the call works normally; you just won't get the stateful-session benefit on OpenAI Responses API models.

Structured Output

The SDK supports two ways to get structured output from chat(): a raw json_schema dict, or a Pydantic model via response_model. Both work across all providers without any changes to your code.

Pydantic Integration (response_model)

Pass a Pydantic BaseModel subclass as response_model — the adapter extracts the JSON Schema automatically and returns a typed instance in ChatResponse.parsed_model.

from pydantic import BaseModel
from llm_api_adapter.models.messages.chat_message import Prompt, UserMessage
from llm_api_adapter.universal_adapter import UniversalLLMAPIAdapter

class Person(BaseModel):
    name: str
    age: int

adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5",
    api_key=openai_api_key,
)

response = adapter.chat(
    messages=[
        Prompt("Extract structured data from the user's message."),
        UserMessage("My name is Alice and I'm 30 years old."),
    ],
    response_model=Person,
    max_tokens=200,
)

print(response.parsed_model)  # Person(name='Alice', age=30)
print(response.parsed_json)   # {"name": "Alice", "age": 30}

Note: Pydantic is not a required dependency of this package. Install it separately: pip install pydantic.

response_model cannot be combined with json_schema or tools — passing both raises JSONSchemaError.

Raw JSON Schema (json_schema)

The SDK also supports structured JSON output via a json_schema parameter in chat(). Pass any JSON Schema object — the adapter normalizes it for each provider automatically and returns the parsed result in ChatResponse.parsed_json.

from llm_api_adapter.models.messages.chat_message import Prompt, UserMessage
from llm_api_adapter.universal_adapter import UniversalLLMAPIAdapter

schema = {
    "type": "object",
    "properties": {
        "name": {"type": "string"},
        "age": {"type": "integer"},
    },
    "required": ["name", "age"],
}

adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5",
    api_key=openai_api_key,
)

response = adapter.chat(
    messages=[
        Prompt("Extract structured data from the user's message."),
        UserMessage("My name is Alice and I'm 30 years old."),
    ],
    json_schema=schema,
    max_tokens=200,
)

print(response.content)      # '{"name": "Alice", "age": 30}'
print(response.parsed_json)  # {"name": "Alice", "age": 30}

json_schema parameter

  • Accepts a dict containing a JSON Schema object.
  • Cannot be combined with tools or response_model — passing both raises JSONSchemaError.
  • If the model returns a response that is not valid JSON, JSONSchemaError is raised.

parsed_json field

ChatResponse.parsed_json contains the parsed dict when json_schema or response_model was provided, or None otherwise.

Provider behavior

Provider Implementation
OpenAI (standard) Native response_format.type=json_schema with strict=true
OpenAI (Responses API / o-series) Native text.format.type=json_schema with strict=true
Anthropic Native output_config.format.type=json_schema
Google generationConfig.responseMimeType="application/json" + responseSchema

The adapter automatically handles provider-specific schema constraints, so the same schema works across all providers without changes.

Error handling

from llm_api_adapter.errors import JSONSchemaError

try:
    response = adapter.chat(
        messages=[UserMessage("Give me the data.")],
        json_schema=schema,
        max_tokens=200,
    )
except JSONSchemaError as e:
    print(f"JSON schema error: {e}")

Vision Input

The SDK supports sending images alongside text using ImagePart and the files parameter on UserMessage. Works identically across OpenAI, Anthropic, and Google — wire-format differences are handled automatically.

Import

from llm_api_adapter.models.messages.chat_message import UserMessage
from llm_api_adapter.models.messages.file_parts import ImagePart

Image from URL

msg = UserMessage(
    "What is in this image?",
    files=[ImagePart(url="https://example.com/photo.jpg")]
)
response = adapter.chat(messages=[msg], max_tokens=200)
print(response.content)

Image from bytes

with open("photo.png", "rb") as f:
    image_bytes = f.read()

msg = UserMessage(
    "Describe this image.",
    files=[ImagePart(data=image_bytes, media_type="image/png")]
)
response = adapter.chat(messages=[msg], max_tokens=200)
print(response.content)

Multiple images

msg = UserMessage(
    "Compare these two images.",
    files=[
        ImagePart(url="https://example.com/before.jpg"),
        ImagePart(url="https://example.com/after.jpg"),
    ]
)

Supported formats

ImagePart accepts any image MIME type starting with image/image/jpeg, image/png, image/gif, image/webp, image/heic, image/heif.

When passing a URL, media_type is auto-detected from the file extension. For URLs without an extension, pass media_type explicitly:

ImagePart(url="https://api.example.com/image?id=42", media_type="image/jpeg")

OpenAI-style dict compatibility

The adapter also normalizes OpenAI-style content lists, so existing code that uses dicts works without changes:

messages = [{
    "role": "user",
    "content": [
        {"type": "text", "text": "What is this?"},
        {"type": "image_url", "image_url": {"url": "https://example.com/img.jpg"}},
    ]
}]
response = adapter.chat(messages=messages, max_tokens=200)

Note: ImagePart is supported in v0.5.0; DocumentPart is introduced in v0.5.1. Google already supports audio input, but AudioPart is postponed because Anthropic does not support audio and OpenAI uses a separate audio API (gpt-audio-1.5 with modalities), so there is no common provider-neutral contract yet.

Document Input

The SDK supports PDF documents alongside text using DocumentPart and the files parameter on UserMessage. Provider-specific wire formats are handled automatically.

Import

from llm_api_adapter.models.messages.chat_message import UserMessage
from llm_api_adapter.models.messages.file_parts import DocumentPart

PDF from URL

msg = UserMessage(
    "Summarize this document in one sentence.",
    files=[DocumentPart(url="https://example.com/report.pdf")]
)
response = adapter.chat(messages=[msg], max_tokens=200)
print(response.content)

The PDF MIME type is detected from the .pdf extension. Anthropic, Google, and the OpenAI Responses API accept a document URL. OpenAI Chat Completions models below gpt-5 do not accept a document URL; use bytes instead.

PDF from bytes

with open("report.pdf", "rb") as f:
    pdf_bytes = f.read()

msg = UserMessage(
    "Summarize this document in one sentence.",
    files=[DocumentPart(data=pdf_bytes, media_type="application/pdf")]
)
response = adapter.chat(messages=[msg], max_tokens=200)
print(response.content)

For bytes, the adapter sends the PDF as base64 data in the provider-specific request format. The same DocumentPart works with Anthropic, Google, OpenAI Chat Completions, and the OpenAI Responses API.

File type support

File type Anthropic OpenAI (< gpt-5) OpenAI (gpt-5+) Google
ImagePart (URL)
ImagePart (bytes)
DocumentPart (URL)
DocumentPart (bytes)

Token Usage and Pricing

Token Usage and Pricing Example

google = UniversalLLMAPIAdapter(
    organization="google",
    model="gemini-2.5-flash",
    api_key=google_api_key
)

response = google.chat(**chat_params)

print(response.usage.input_tokens, "tokens", f"({response.cost_input} {response.currency})")
print(response.usage.output_tokens, "tokens", f"({response.cost_output} {response.currency})")
print(response.usage.total_tokens, "tokens", f"({response.cost_total} {response.currency})")

Prices are updated with each release to reflect provider changes. Bundled prices reflect each provider's standard API rates — batch pricing, cached-token discounts, and volume agreements are not accounted for. Use set_in_per_1m / set_out_per_1m to apply your actual rates if needed:

Overriding Pricing or Currency

google = UniversalLLMAPIAdapter(
    organization="google",
    model="gemini-2.5-flash",
    api_key=google_api_key
)

google.pricing.set_in_per_1m(1.5)
google.pricing.set_out_per_1m(3)
google.pricing.set_currency("EUR")

response = google.chat(**chat_params)
print(response.content)
print(response.usage.input_tokens, "tokens", f"({response.cost_input} {response.currency})")
print(response.usage.output_tokens, "tokens", f"({response.cost_output} {response.currency})")
print(response.usage.total_tokens, "tokens", f"({response.cost_total} {response.currency})")

Logging

The library uses Python's standard logging module and does not configure handlers. Loggers are module-based under llm_api_adapter.* (e.g., llm_api_adapter.universal_adapter).

  • Default behavior: No handlers installed, effective level = WARNING.
  • No secrets are logged — API keys and request bodies are excluded. Only event metadata and errors are logged. Adapter and client objects mask the key in __repr__ (e.g. api_key='sk-12345...cdef'), so they are safe to log or print.

Enable logs (console)

import logging

logging.basicConfig(level=logging.INFO)  # or DEBUG
# Optionally limit logging to this library
logging.getLogger("llm_api_adapter").setLevel(logging.DEBUG)

Write logs to a file

import logging

handler = logging.FileHandler("llm_api_adapter.log")
handler.setFormatter(logging.Formatter(
    "%(asctime)s %(levelname)s %(name)s %(message)s"
))
root = logging.getLogger()
root.setLevel(logging.INFO)
root.addHandler(handler)

Per-request correlation (optional)

import logging
logger = logging.getLogger("llm_api_adapter")

req_id = "req-123"
logger = logging.LoggerAdapter(logger, {"request_id": req_id})
logger.info("starting call")

To include request_id in log output, add %(request_id)s to your log formatter.

Reduce noise / silence logs

import logging
logging.getLogger("llm_api_adapter").setLevel(logging.WARNING)   # silence info logs
logging.getLogger("urllib3").setLevel(logging.WARNING)           # if using requests

Env-based log level toggle

# app.py
import logging, os
level = os.getenv("LLM_ADAPTER_LOGLEVEL", "WARNING").upper()
logging.getLogger("llm_api_adapter").setLevel(level)

Tip: For HTTP-level debugging with requests, also set:

import http.client as http_client, logging
http_client.HTTPConnection.debuglevel = 1
logging.getLogger("urllib3").setLevel(logging.DEBUG)

Use this only in development.

Related project

For production applications that need retries, multi-provider failover, circuit breakers, and tool-session recovery, see llm-api-resilience.

Install it with:

python -m pip install llm-api-resilience

llm-api-resilience is built on top of this adapter and keeps its provider-neutral interface unchanged.

Development & Testing

Note
This section is intended for developers working with the source code from GitHub.
It is not relevant for users installing the package from PyPI.

This project uses pytest for testing. Tests are located in the tests/ directory.

Test suites

  • unit: fast, offline, no real provider calls
  • integration: adapter-level integration tests (may use mocked/provider-shaped responses)
  • e2e: real API calls against providers (requires API keys)

Running tests

Run everything:

pytest

Run by marker:

pytest -m unit
pytest -m integration
pytest -m e2e

E2E requirements

E2E tests require provider API keys to be present in environment variables:

  • OPENAI_API_KEY
  • ANTHROPIC_API_KEY
  • GOOGLE_API_KEY

License

This project is licensed under the terms of the MIT License.
See the LICENSE file for details.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llm_api_adapter-0.6.0.tar.gz (49.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

llm_api_adapter-0.6.0-py3-none-any.whl (52.7 kB view details)

Uploaded Python 3

File details

Details for the file llm_api_adapter-0.6.0.tar.gz.

File metadata

  • Download URL: llm_api_adapter-0.6.0.tar.gz
  • Upload date:
  • Size: 49.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for llm_api_adapter-0.6.0.tar.gz
Algorithm Hash digest
SHA256 bc5989f76da233d55019f4ec5254fb6d8a52e4351a27506d26d453d7394ea801
MD5 1ca44a6ca83453cd714e02ef4f705e44
BLAKE2b-256 f02e7c9c12ba4a26870990604bc51ff64fc40a9c7b3ff15ac5438a7723d9e40e

See more details on using hashes here.

File details

Details for the file llm_api_adapter-0.6.0-py3-none-any.whl.

File metadata

File hashes

Hashes for llm_api_adapter-0.6.0-py3-none-any.whl
Algorithm Hash digest
SHA256 c2b95709e473244133daba872732434ea92d9a85f00c204d4c2feba82b0d221d
MD5 a967afefa942ecf5d1ab474e9dc5893d
BLAKE2b-256 f7b4d4fb5a4eafc3be26f3bbdd58753150348ec951abe5293ba92bcdde1b4a16

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page