Skip to main content

Aiyallm

Aiyallm is a lightweight Python gateway for OpenAI/Anthropic-style language and vision models.

Aiyallm does not ship a built-in model catalog and does not read API keys implicitly. Every provider, base_url, API key, and model is declared by the application.

Table of Contents

Features

  • OpenAI and OpenAI-compatible providers
  • Anthropic and Anthropic-compatible providers
  • Chinese OpenAI-compatible providers such as DeepSeek, Qwen/DashScope, Moonshot/Kimi, Zhipu GLM, and SiliconFlow
  • Explicit model selection, provider/model references, and provider route lists
  • Routing policies: fallback, first, round_robin, random, and least_used
  • Vision and multimodal message passthrough with capability filtering
  • Sync and async chat and streaming
  • Detailed token accounting with cache hit/miss fields
  • Optional SQLite or JSONL usage persistence
  • Optional cost calculation from an application-owned price book

Install

pip install -e .

Runtime dependencies are:

  • openai
  • anthropic
  • httpx

Quick Start

There are no default providers. Pass them explicitly.

from aiyallm import Aiyallm

client = Aiyallm(
    providers=[
        {
            "name": "deepseek",
            "account": "billing-deepseek",
            "type": "openai_compatible",
            "base_url": "https://api.deepseek.com",
            "api_key": "your-deepseek-key",
            "models": ["deepseek-chat", "deepseek-reasoner"],
        }
    ]
)

response = client.chat(
    "用一句话解释什么是路由。",
    model="deepseek/deepseek-chat",
)

print(response.text)
print(response.usage)

models is optional when the provider allows unknown models. Configured providers default to allow_unknown_models=True, but the provider itself must still be declared.

Provider Configuration

Each provider accepts these common fields.

Field Meaning
name Unique provider name used by provider/model routes
account Optional billing or account name recorded in usage
type openai, openai_compatible, anthropic, or anthropic_compatible
base_url Endpoint for compatible services
api_key API key, or use api_key_env to read an environment variable
models Optional model names or {"id", "capabilities", ...} mappings
allow_unknown_models Allow model IDs not listed in models; defaults to True
timeout SDK timeout

OpenAI-Compatible Examples

providers = [
    {
        "name": "deepseek",
        "type": "openai_compatible",
        "base_url": "https://api.deepseek.com",
        "api_key_env": "DEEPSEEK_API_KEY",
    },
    {
        "name": "qwen",
        "type": "openai_compatible",
        "base_url": "https://dashscope.aliyuncs.com/compatible-mode/v1",
        "api_key_env": "DASHSCOPE_API_KEY",
    },
    {
        "name": "moonshot",
        "type": "openai_compatible",
        "base_url": "https://api.moonshot.cn/v1",
        "api_key_env": "MOONSHOT_API_KEY",
    },
    {
        "name": "glm",
        "type": "openai_compatible",
        "base_url": "https://open.bigmodel.cn/api/paas/v4",
        "api_key_env": "ZHIPU_API_KEY",
    },
]

Use the base_url from each provider's current documentation. Aiyallm does not hard-code any endpoint or model name.

Routing

Use provider/model to target a provider.

response = client.chat("Hello", model="qwen/qwen-plus")

Use a route list for fallback.

response = client.chat(
    "Hello",
    model=["deepseek/deepseek-chat", "qwen/qwen-plus"],
    route_policy="fallback",
)

Available policies:

  • fallback - try candidates in order
  • first - use only the first candidate
  • round_robin - rotate through candidates
  • random - choose a random candidate
  • least_used - choose the candidate with the lowest recorded token usage

The default policy is fallback.

Vision and Capability Filtering

Declare capabilities in models and pass require_vision=True when a request needs image input.

from aiyallm import Aiyallm

client = Aiyallm(
    providers=[
        {
            "name": "qwen",
            "type": "openai_compatible",
            "base_url": "https://dashscope.aliyuncs.com/compatible-mode/v1",
            "api_key": "your-key",
            "models": [
                {"id": "qwen-vl-plus", "capabilities": ["text", "vision"]}
            ],
        }
    ]
)

response = client.chat(
    [
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "What is in this image?"},
                {
                    "type": "image_url",
                    "image_url": {"url": "https://example.com/image.jpg"},
                },
            ],
        }
    ],
    model="qwen/qwen-vl-plus",
    require_vision=True,
)

Usage Ledger

Every successful completion is recorded.

response = client.chat("Hello", model="deepseek/deepseek-chat")

print(response.usage.input_tokens)
print(response.usage.output_tokens)
print(response.usage.input_cache_hit_tokens)
print(response.usage.input_cache_miss_tokens)
print(response.usage.output_cache_hit_tokens)
print(response.usage.output_cache_miss_tokens)
print(response.usage.account)
print(response.usage.reasoning_tokens)

print(client.ledger.total())
print(client.ledger.by_model())
print(client.ledger.by_provider())
print(client.ledger.by_account())

client.ledger.export_jsonl("usage.jsonl")

For persistent accounting, pass a SQLite path.

from aiyallm import UsageLedger

ledger = UsageLedger(sqlite_path="data/usage.sqlite3")
client = Aiyallm(providers=providers, ledger=ledger)

Cost Calculation

Token counts are always recorded. Cost is calculated only when the application supplies a price.

from aiyallm import Aiyallm, Price, PriceBook

price_book = PriceBook(
    {
        "deepseek-chat": Price(
            input_per_mtok=0.27,
            output_per_mtok=1.10,
        )
    }
)

client = Aiyallm(providers=providers, price_book=price_book)
response = client.chat("Hello", model="deepseek/deepseek-chat")

print(response.usage.cost_usd)

Prices are US dollars per million tokens.

Streaming and Async

for chunk in client.stream("Hello", model="deepseek/deepseek-chat"):
    print(chunk.content, end="", flush=True)
import asyncio


async def main():
    response = await client.achat("Hello", model="deepseek/deepseek-chat")
    print(response.text)

    async for chunk in client.astream("Hello", model="deepseek/deepseek-chat"):
        print(chunk.content, end="", flush=True)


asyncio.run(main())

Extending Providers

For a non-standard provider, subclass aiyallm.BaseProvider and implement complete, acomplete, stream, and astream.

from aiyallm import Aiyallm, BaseProvider


class MyProvider(BaseProvider):
    def complete(self, request):
        ...

Then register the instance directly.

client = Aiyallm(providers=[MyProvider(name="my-provider", models=["my-model"])])

Metadata

Release files for aiyallm 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for aiyallm 0.1.1
File Size Uploaded
aiyallm-0.1.1.tar.gz 23.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for aiyallm 0.1.1
File Interpreter ABI Platform
aiyallm-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 48.8 kB

Release files / aiyallm-0.1.1.tar.gz

Download URL aiyallm-0.1.1.tar.gz
Size 23.8 kB
Tags Source
SHA-256 checksum
How to use checksums
a5839da3549c398de04dece6f39a7677a7b3862bb8d048ca450cf4c3f193cad3
BLAKE2b-256 checksum
How to use checksums
8585a6532697e02f8065826712b3f9a3b11f8ab9cd572f34e168502a328313fa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release files / aiyallm-0.1.1-py3-none-any.whl

Download URL aiyallm-0.1.1-py3-none-any.whl
Size 25.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
98f2874b6fd9e1f5869b9fb496ffdf71fdac6d4d4f3cdb7c392c14b3066fa047
BLAKE2b-256 checksum
How to use checksums
3d6f0a7c266511b8b3c2d861af229cb066ad49cc38bf9210fe7cbb2ea75ef1bc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page