Unified Python client for OpenAI, Azure OpenAI, Vertex AI, Anthropic, Gemini, DeepSeek, Bedrock, and ChatGPT.

Project description

llmai

llmai is a Python library for working with OpenAI, Azure OpenAI, Vertex AI, Anthropic, Google Gemini, DeepSeek, Bedrock, and ChatGPT through a shared set of message, tool, schema, and response primitives.

Today the repository includes adapters for:

ChatGPT
OpenAI
Azure OpenAI
Vertex AI
DeepSeek
Anthropic
Google Gemini
Amazon Bedrock

Each provider client exposes the same core entrypoint:

generate(..., stream=False)

Why This Exists

Provider SDKs differ in how they represent messages, tool calls, structured output, and streaming events. llmai smooths those differences out so application code can stay closer to one mental model.

Installation

Install the project locally with uv:

uv sync

Or install it in editable mode with pip:

pip install -e .

Quick Start

from llmai import OpenAIClient
from llmai.shared import UserMessage

client = OpenAIClient(api_key="OPENAI_API_KEY")

result = client.generate(
    model="your-openai-model",
    messages=[
        UserMessage(content="Write a two-line poem about clean interfaces."),
    ],
)

print(result.content)
print(result.usage)
print(result.duration_seconds)

For text-only prompts, UserMessage(content="...") is the simplest form. You can also pass explicit content parts like TextContentPart when you need mixed multimodal input or tighter control over message structure.

If you want to swap providers, the overall call shape stays the same. In most cases you only need to change the client class, credentials, and model name.

Azure OpenAI

from llmai import AzureOpenAIClient
from llmai.shared import UserMessage


client = AzureOpenAIClient(
    api_key="AZURE_OPENAI_API_KEY",
    endpoint="https://your-resource.openai.azure.com",
    api_version="2024-10-21",
)

result = client.generate(
    model="your-azure-deployment",
    messages=[
        UserMessage(content="Write a two-line poem about clean interfaces."),
    ],
)

print(result.content)

AzureOpenAIClient uses the official OpenAI SDK's Azure client and supports either API-key auth or Entra token auth. It reads AZURE_OPENAI_API_KEY or AZURE_OPENAI_AD_TOKEN, AZURE_OPENAI_ENDPOINT or AZURE_OPENAI_BASE_URL, AZURE_OPENAI_API_VERSION or OPENAI_API_VERSION, and optional AZURE_OPENAI_DEPLOYMENT by default.

Vertex AI

from llmai import VertexAIClient
from llmai.shared import UserMessage


client = VertexAIClient(
    project="your-gcp-project",
    location="us-central1",
)

result = client.generate(
    model="gemini-2.5-flash",
    messages=[
        UserMessage(content="Write a two-line poem about clean interfaces."),
    ],
)

print(result.content)

VertexAIClient uses the google-genai Vertex AI path internally. It supports ADC or explicit credentials, and also accepts provider-specific VERTEX_PROJECT, VERTEX_LOCATION, and VERTEX_API_KEY envs while still allowing the upstream SDK's standard GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION, and Google API-key env handling.

ChatGPT

from llmai import ChatGPTClient
from llmai.shared import UserMessage


client = ChatGPTClient(access_token="CHATGPT_ACCESS_TOKEN")

result = client.generate(
    model="chatgpt-4o-latest",
    messages=[
        UserMessage(content="Write a two-line poem about clean interfaces."),
    ],
)

print(result.content)

ChatGPTClient targets ChatGPT's Codex backend at https://chatgpt.com/backend-api/codex. It always uses the Responses API internally, and reads CHATGPT_ACCESS_TOKEN or CODEX_ACCESS_TOKEN by default, with optional CHATGPT_ACCOUNT_ID or CODEX_ACCOUNT_ID.

DeepSeek

from llmai import DeepSeekClient
from llmai.shared import JSONSchemaResponse, UserMessage


client = DeepSeekClient()

result = client.generate(
    model="deepseek-chat",
    messages=[
        UserMessage(content="Return a JSON object with one field named answer."),
    ],
    response_format=JSONSchemaResponse(
        name="final_answer",
        json_schema={
            "type": "object",
            "properties": {
                "answer": {"type": "string"},
            },
            "required": ["answer"],
        },
    ),
)

print(result.content)

DeepSeekClient uses the OpenAI SDK against DeepSeek's OpenAI-compatible API and reads DEEPSEEK_API_KEY by default. For structured output, it always uses an internal function-tool schema because DeepSeek does not support response_format={"type":"json_schema"}. During streaming, the internal response tool is surfaced as incremental JSON content chunks, and the stream still ends with parsed JSON on the final completion chunk's content. If you need DeepSeek's server-side strict tool enforcement, point base_url at https://api.deepseek.com/beta.

Amazon Bedrock

from llmai import BedrockClient
from llmai.shared import UserMessage


client = BedrockClient(
    region="us-east-1",
    aws_access_key_id="AWS_ACCESS_KEY_ID",
    aws_secret_access_key="AWS_SECRET_ACCESS_KEY",
)

# Or use Bedrock API-key auth:
# client = BedrockClient(region="us-east-1", api_key="BEDROCK_API_KEY")

result = client.generate(
    model="us.anthropic.claude-3-5-haiku-20241022-v1:0",
    messages=[
        UserMessage(content="Say hello."),
    ],
)

print(result.content)

Structured Output

from pydantic import BaseModel

from llmai import GoogleClient
from llmai.shared import JSONSchemaResponse, UserMessage


class Summary(BaseModel):
    title: str
    bullets: list[str]


client = GoogleClient(api_key="GOOGLE_API_KEY")

result = client.generate(
    model="your-google-model",
    messages=[
        UserMessage(content="Summarize retrieval-augmented generation in simple terms."),
    ],
    response_format=JSONSchemaResponse(json_schema=Summary),
)

print(result.content)

Use JSONSchemaResponse, JSONObjectResponse, or TextResponse to request different response shapes.

Multimodal Content

from llmai import GoogleClient
from llmai.shared import ImageContentPart, TextContentPart, UserMessage


client = GoogleClient(api_key="GOOGLE_API_KEY")

result = client.generate(
    model="your-google-model",
    messages=[
        UserMessage(
            content=[
                TextContentPart(text="Describe this image."),
                ImageContentPart(url="https://example.com/cat.png"),
            ]
        ),
    ],
)

print(result.content)
print(result.thinking)

Use explicit content parts when you need multimodal inputs or want to mix text with images in one message. Normal completion content is surfaced as list[TextContentPart | ImageContentPart] when the provider returns message content, including text-only replies. Reasoning is exposed on ResponseContent.thinking as list[str] when the provider returns one or more thinking blocks, and the same value is also available on the final AssistantMessage.

Tool Calling

from pydantic import BaseModel

from llmai import OpenAIClient
from llmai.shared import Tool, ToolResponseMessage, UserMessage


class WeatherArgs(BaseModel):
    city: str


weather_tool = Tool(
    name="get_weather",
    description="Look up the weather for a city.",
    schema=WeatherArgs,
)

client = OpenAIClient(api_key="OPENAI_API_KEY")

first = client.generate(
    model="your-openai-model",
    messages=[
        UserMessage(content="What is the weather in Kathmandu?"),
    ],
    tools=[weather_tool],
    tool_choice={"tools": ["get_weather"]},
)

for tool_call in first.tool_calls:
    if tool_call.name != "get_weather":
        continue

    follow_up = client.generate(
        model="your-openai-model",
        messages=[
            *first.messages,
            ToolResponseMessage(
                id=tool_call.id,
                content=["It is sunny in Kathmandu."],
            ),
        ],
        tools=[weather_tool],
    )
    print(follow_up.content)

llmai returns tool calls in first.tool_calls and leaves execution to the caller.

Hosted Web Search

llmai also supports a provider-hosted web search tool that is not a function tool:

from llmai import OpenAIClient
from llmai.shared import UserMessage, WebSearchTool

client = OpenAIClient(api_key="OPENAI_API_KEY")

result = client.generate(
    model="your-openai-model",
    messages=[
        UserMessage(content="What was a positive news story from today? Cite sources."),
    ],
    tools=[WebSearchTool()],
    api_type="responses",
)

print(result.content)
print(result.thinking)

You can also target it explicitly in tool_choice:

tool_choice = {
    "mode": "required",
    "tools": ["web_search"],
}

Current llmai behavior:

OpenAI Responses: attaches built-in web_search
Azure OpenAI: follows the same OpenAI adapter surface; service support depends on your Azure API version and deployment
Vertex AI: attaches google_search
ChatGPT/Codex: attaches built-in web_search
Anthropic: attaches Anthropic's hosted web-search tool
Google Gemini: attaches google_search
OpenAI Chat Completions: ignores hosted web_search
DeepSeek: ignores hosted web_search
Amazon Bedrock: ignores hosted web_search

web_search can be mixed with normal function tools in the same request.

Streaming

from llmai import AnthropicClient
from llmai.shared import UserMessage

client = AnthropicClient(api_key="ANTHROPIC_API_KEY")

for chunk in client.generate(
    model="your-anthropic-model",
    messages=[
        UserMessage(content="Explain recursion in one paragraph."),
    ],
    stream=True,
):
    if chunk.type == "content":
        print(chunk.chunk, end="")
    elif chunk.type == "completion":
        print("\nDone:", chunk.usage)

generate(..., stream=True) yields marker chunks with type="event" and event="start" / event="end" around each content, thinking, and tool section. If a provider returns multiple reasoning blocks, each block gets its own thinking start/end pair. The final chunk has type="completion" and includes top-level content, thinking, usage, and accumulated messages.

Package Layout

llmai/openai: OpenAI adapter
llmai/azure: Azure OpenAI adapter
llmai/vertex: Vertex AI adapter
llmai/deepseek: DeepSeek adapter
llmai/anthropic: Anthropic adapter
llmai/google: Google Gemini adapter
llmai/bedrock: Amazon Bedrock adapter
llmai/shared: common message, tool, schema, and response models

Core Types

The shared layer includes the main primitives you will use across providers:

UserMessage, SystemMessage, AssistantMessage
TextContentPart, ImageContentPart
Tool, WebSearchTool, ToolResponseMessage
JSONSchemaResponse, JSONObjectResponse, TextResponse
ResponseContent, ResponseStreamChunk, ResponseStreamContentChunk, ResponseStreamThinkingChunk, ResponseStreamToolChunk, ResponseStreamToolCompleteChunk, ResponseStreamCompletionChunk
ResponseUsage

Project details

Release history Release notifications | RSS feed

0.2.5

May 19, 2026

0.2.4

May 12, 2026

0.2.3

May 11, 2026

0.2.2

Apr 27, 2026

0.2.1

Apr 24, 2026

0.2.0

Apr 24, 2026

0.1.9

Apr 23, 2026

0.1.8

Apr 23, 2026

0.1.7

Apr 23, 2026

0.1.6

Apr 23, 2026

0.1.5

Apr 23, 2026

0.1.4

Apr 23, 2026

0.1.3

Apr 22, 2026

This version

0.1.2

Apr 19, 2026

0.1.1

Apr 19, 2026

0.1.0

Apr 19, 2026

0.0.1

Feb 10, 2026

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llmai-0.1.2.tar.gz (38.7 kB view details)

Uploaded Apr 19, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

llmai-0.1.2-py3-none-any.whl (50.3 kB view details)

Uploaded Apr 19, 2026 Python 3

File details

Details for the file llmai-0.1.2.tar.gz.

File metadata

Download URL: llmai-0.1.2.tar.gz
Upload date: Apr 19, 2026
Size: 38.7 kB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for llmai-0.1.2.tar.gz
Algorithm	Hash digest
SHA256	`a1a25c24fd612cf10be45f8ca038d0aac203a1dc54c4261b34cfcfb6d972c028`
MD5	`d3490a0bb57aaafe17973fbfb98a9fb4`
BLAKE2b-256	`efd5590ee9c1e2600f712f667f9052ae0c10532b946ad916864c7294acca53bb`

See more details on using hashes here.

Provenance

The following attestation bundles were made for llmai-0.1.2.tar.gz:

Publisher: publish.yml on presenton/llmai

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: llmai-0.1.2.tar.gz
- Subject digest: a1a25c24fd612cf10be45f8ca038d0aac203a1dc54c4261b34cfcfb6d972c028
- Sigstore transparency entry: 1340759596
- Sigstore integration time: Apr 19, 2026
Source repository:
- Permalink: presenton/llmai@eea1af5f0ec03750717ca48a91ab09d18410af5c
- Branch / Tag: refs/tags/v0.1.2
- Owner: https://github.com/presenton
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: publish.yml@eea1af5f0ec03750717ca48a91ab09d18410af5c
- Trigger Event: push

File details

Details for the file llmai-0.1.2-py3-none-any.whl.

File metadata

Download URL: llmai-0.1.2-py3-none-any.whl
Upload date: Apr 19, 2026
Size: 50.3 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for llmai-0.1.2-py3-none-any.whl
Algorithm	Hash digest
SHA256	`1d178af0228a42d0830f00010e27e48e50c824760debac7917279552931babc7`
MD5	`fcba31802a4d08fcb54cef41dcd9a225`
BLAKE2b-256	`dcffd063228641aa1d192878a6c4091854c36156103181b18303a6b554032163`

See more details on using hashes here.

Provenance

The following attestation bundles were made for llmai-0.1.2-py3-none-any.whl:

Publisher: publish.yml on presenton/llmai

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: llmai-0.1.2-py3-none-any.whl
- Subject digest: 1d178af0228a42d0830f00010e27e48e50c824760debac7917279552931babc7
- Sigstore transparency entry: 1340759606
- Sigstore integration time: Apr 19, 2026
Source repository:
- Permalink: presenton/llmai@eea1af5f0ec03750717ca48a91ab09d18410af5c
- Branch / Tag: refs/tags/v0.1.2
- Owner: https://github.com/presenton
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: publish.yml@eea1af5f0ec03750717ca48a91ab09d18410af5c
- Trigger Event: push

llmai 0.1.2

Navigation

Verified details

Maintainers

Unverified details

Meta

Project description

llmai

Why This Exists

Installation

Quick Start

Azure OpenAI

Vertex AI

ChatGPT

DeepSeek

Amazon Bedrock

Structured Output

Multimodal Content

Tool Calling

Hosted Web Search

Streaming

Package Layout

Core Types

Project details

Verified details

Maintainers

Unverified details

Meta

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

Provenance

File details

File metadata

File hashes

Provenance