Skip to main content

Aiyallm

Aiyallm is a lightweight Python gateway for OpenAI/Anthropic-style language and vision models.

Aiyallm does not ship a built-in model catalog and never guesses model names. Provider endpoints, model lists, and pricing remain caller-owned. API keys are supplied in one of two ways: a literal api_key value, or an environment variable via api_key_env. If api_key is omitted, Aiyallm automatically reads the local environment variable, and common vendor variables such as DEEPSEEK_API_KEY and DASHSCOPE_API_KEY are inferred from the provider name.

Table of Contents

Features

  • OpenAI and OpenAI-compatible providers
  • Anthropic and Anthropic-compatible providers
  • Chinese OpenAI-compatible providers such as DeepSeek, Qwen/DashScope, Moonshot/Kimi, Zhipu GLM, and SiliconFlow
  • Explicit model selection, provider/model references, and provider route lists
  • Routing policies: fallback, first, round_robin, random, and least_used
  • Vision and multimodal message passthrough with capability filtering
  • Sync and async chat and streaming
  • Detailed token accounting with cache hit/miss fields
  • Optional SQLite or JSONL usage persistence
  • Optional cost calculation from an application-owned price book

Install

pip install -e .

Runtime dependencies are:

  • openai
  • anthropic
  • httpx

Quick Start

There are no default providers. Pass them explicitly.

from aiyallm import Aiyallm

client = Aiyallm(
    providers=[
        {
            "name": "deepseek",
            "account": "billing-deepseek",
            "type": "openai_compatible",
            "base_url": "https://api.deepseek.com",
            # Option 1: set api_key directly.
            "api_key": "your-deepseek-key",
            "models": ["deepseek-flash", "deepseek-v4-pro"],
        },
        {
            "name": "qwen",
            "account": "billing-qwen",
            "type": "openai_compatible",
            "base_url": "https://dashscope.aliyuncs.com/compatible-mode/v1",
            # Option 2: omit api_key and read DASHSCOPE_API_KEY from the local environment.
            "models": ["qwen3.8-flash"],
        },
    ]
)

response = client.chat(
    "用一句话解释什么是路由。",
    model="deepseek/deepseek-flash",
)

print(response.text)
print(response.usage)

If you omit api_key, Aiyallm automatically reads the local environment variable. For example, a deepseek provider reads DEEPSEEK_API_KEY from your shell environment.

models is optional when the provider allows unknown models. Configured providers default to allow_unknown_models=True, but the provider itself must still be declared.

Provider Configuration

Each provider accepts these common fields.

Field Meaning
name Unique provider name used by provider/model routes
account Optional billing or account name recorded in usage
type openai, openai_compatible, anthropic, or anthropic_compatible
base_url Endpoint for compatible services
api_key Explicit API key
api_key_env Optional environment variable name, or list of names, to read the key from
api_key_prompt When True, prompt in the terminal if no key is found; defaults to False
models Optional model names or {"id", "capabilities", ...} mappings
allow_unknown_models Allow model IDs not listed in models; defaults to True
timeout SDK timeout

API keys use one of two paths:

  • Set api_key to a literal key value.
  • Set api_key_env to an environment variable name, or a list of names, to read from.

If api_key is not provided, Aiyallm automatically reads the local environment variable. When api_key_env is omitted, Aiyallm falls back to the common vendor variable for that provider name.

OpenAI-Compatible Examples

providers = [
    {
        "name": "deepseek",
        "type": "openai_compatible",
        "base_url": "https://api.deepseek.com",
        # `api_key_env` is optional; DEEPSEEK_API_KEY is inferred from `name`.
    },
    {
        "name": "qwen",
        "type": "openai_compatible",
        "base_url": "https://dashscope.aliyuncs.com/compatible-mode/v1",
        # DASHSCOPE_API_KEY or QWEN_API_KEY is inferred from `name`.
    },
    {
        "name": "moonshot",
        "type": "openai_compatible",
        "base_url": "https://api.moonshot.cn/v1",
        # MOONSHOT_API_KEY or KIMI_API_KEY is inferred from `name`.
    },
    {
        "name": "glm",
        "type": "openai_compatible",
        "base_url": "https://open.bigmodel.cn/api/paas/v4",
        # ZHIPU_API_KEY / ZHIPUAI_API_KEY / GLM_API_KEY is inferred from `name`.
    },
]

Use the base_url from each provider's current documentation. Aiyallm does not hard-code any endpoint or model name.

Routing

Use provider/model to target a provider.

response = client.chat("Hello", model="qwen/qwen3.8-flash")

Use a route list for fallback.

response = client.chat(
    "Hello",
    model=["deepseek/deepseek-flash", "qwen/qwen3.8-flash"],
    route_policy="fallback",
)

Available policies:

  • fallback - try candidates in order
  • first - use only the first candidate
  • round_robin - rotate through candidates
  • random - choose a random candidate
  • least_used - choose the candidate with the lowest recorded token usage

The default policy is fallback.

Vision and Capability Filtering

Declare capabilities in models and pass require_vision=True when a request needs image input.

from aiyallm import Aiyallm

client = Aiyallm(
    providers=[
        {
            "name": "qwen",
            "type": "openai_compatible",
            "base_url": "https://dashscope.aliyuncs.com/compatible-mode/v1",
            "api_key": "your-key",
            "models": [
                {"id": "qwen3.8-omni-flash", "capabilities": ["text", "vision"]}
            ],
        }
    ]
)

response = client.chat(
    [
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "What is in this image?"},
                {
                    "type": "image_url",
                    "image_url": {"url": "https://example.com/image.jpg"},
                },
            ],
        }
    ],
    model="qwen/qwen3.8-omni-flash",
    require_vision=True,
)

Usage Ledger

Every successful completion is recorded.

response = client.chat("Hello", model="deepseek/deepseek-flash")

print(response.usage.input_tokens)
print(response.usage.output_tokens)
print(response.usage.input_cache_hit_tokens)
print(response.usage.input_cache_miss_tokens)
print(response.usage.output_cache_hit_tokens)
print(response.usage.output_cache_miss_tokens)
print(response.usage.account)
print(response.usage.reasoning_tokens)

print(client.ledger.total())
print(client.ledger.by_model())
print(client.ledger.by_provider())
print(client.ledger.by_account())

client.ledger.export_jsonl("usage.jsonl")

For persistent accounting, pass a SQLite path.

from aiyallm import UsageLedger

ledger = UsageLedger(sqlite_path="data/usage.sqlite3")
client = Aiyallm(providers=providers, ledger=ledger)

Cost Calculation

Token counts are always recorded. Cost is calculated only when the application supplies a price.

from aiyallm import Aiyallm, Price, PriceBook

price_book = PriceBook(
    {
        "deepseek-flash": Price(
            input_per_mtok=0.27,
            output_per_mtok=1.10,
        )
    }
)

client = Aiyallm(providers=providers, price_book=price_book)
response = client.chat("Hello", model="deepseek/deepseek-flash")

print(response.usage.cost_usd)

Prices are US dollars per million tokens.

Streaming and Async

for chunk in client.stream("Hello", model="deepseek/deepseek-flash"):
    print(chunk.content, end="", flush=True)

chat can also choose the output mode directly: once returns one result, while sse returns a standard Server-Sent Events text stream.

for event in client.chat("Hello", model="deepseek/deepseek-flash", output_mode="sse"):
    print(event, end="")
import asyncio


async def main():
    response = await client.achat("Hello", model="deepseek/deepseek-flash")
    print(response.text)

    async for chunk in client.astream("Hello", model="deepseek/deepseek-flash"):
        print(chunk.content, end="", flush=True)

    async for event in client.achat("Hello", model="deepseek/deepseek-flash", output_mode="sse"):
        print(event, end="")


asyncio.run(main())

Extending Providers

For a non-standard provider, subclass aiyallm.BaseProvider and implement complete, acomplete, stream, and astream.

from aiyallm import Aiyallm, BaseProvider


class MyProvider(BaseProvider):
    def complete(self, request):
        ...

Then register the instance directly.

client = Aiyallm(providers=[MyProvider(name="my-provider", models=["my-model"])])

Metadata

Release files for aiyallm 0.1.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for aiyallm 0.1.4
File Size Uploaded
aiyallm-0.1.4.tar.gz 28.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for aiyallm 0.1.4
File Interpreter ABI Platform
aiyallm-0.1.4-py3-none-any.whl Python 3 none any Details

Total release size: 56.4 kB

Release files / aiyallm-0.1.4.tar.gz

Download URL aiyallm-0.1.4.tar.gz
Size 28.0 kB
Tags Source
SHA-256 checksum
How to use checksums
f70d8105c4239601754ad90c41c97a14e27fec7aab27dc2a07dde41bbffcff1e
BLAKE2b-256 checksum
How to use checksums
0ef8f38b1037681ee0da0cad2ed0cd6654f2688dfbe23d3512e91ff1803b6bc6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release files / aiyallm-0.1.4-py3-none-any.whl

Download URL aiyallm-0.1.4-py3-none-any.whl
Size 28.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6ec47b4b5985d02bce05d62945411910736e5ac572c21eba6cea17520a4e0e6b
BLAKE2b-256 checksum
How to use checksums
d7be25575f69462eada672ec45fdff4250838ba83e10a1f8f94779f359d4b6da
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

This release

0.1.4 This release

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page