Skip to main content

Aiya

Aiya is a lightweight Python gateway for OpenAI/Anthropic-style language and vision models.

Aiya does not ship a built-in model catalog and does not read API keys implicitly. Every provider, base_url, API key, and model is declared by the application.

Table of Contents

Features

  • OpenAI and OpenAI-compatible providers
  • Anthropic and Anthropic-compatible providers
  • Chinese OpenAI-compatible providers such as DeepSeek, Qwen/DashScope, Moonshot/Kimi, Zhipu GLM, and SiliconFlow
  • Explicit model selection, provider/model references, and provider route lists
  • Routing policies: fallback, first, round_robin, random, and least_used
  • Vision and multimodal message passthrough with capability filtering
  • Sync and async chat and streaming
  • Detailed token accounting with cache hit/miss fields
  • Optional SQLite or JSONL usage persistence
  • Optional cost calculation from an application-owned price book

Install

pip install -e .

Runtime dependencies are:

  • openai
  • anthropic
  • httpx

Quick Start

There are no default providers. Pass them explicitly.

from aiya import Aiya

client = Aiya(
    providers=[
        {
            "name": "deepseek",
            "account": "billing-deepseek",
            "type": "openai_compatible",
            "base_url": "https://api.deepseek.com",
            "api_key": "your-deepseek-key",
            "models": ["deepseek-chat", "deepseek-reasoner"],
        }
    ]
)

response = client.chat(
    "用一句话解释什么是路由。",
    model="deepseek/deepseek-chat",
)

print(response.text)
print(response.usage)

models is optional when the provider allows unknown models. Configured providers default to allow_unknown_models=True, but the provider itself must still be declared.

Provider Configuration

Each provider accepts these common fields.

Field Meaning
name Unique provider name used by provider/model routes
account Optional billing or account name recorded in usage
type openai, openai_compatible, anthropic, or anthropic_compatible
base_url Endpoint for compatible services
api_key API key, or use api_key_env to read an environment variable
models Optional model names or {"id", "capabilities", ...} mappings
allow_unknown_models Allow model IDs not listed in models; defaults to True
timeout SDK timeout

OpenAI-Compatible Examples

providers = [
    {
        "name": "deepseek",
        "type": "openai_compatible",
        "base_url": "https://api.deepseek.com",
        "api_key_env": "DEEPSEEK_API_KEY",
    },
    {
        "name": "qwen",
        "type": "openai_compatible",
        "base_url": "https://dashscope.aliyuncs.com/compatible-mode/v1",
        "api_key_env": "DASHSCOPE_API_KEY",
    },
    {
        "name": "moonshot",
        "type": "openai_compatible",
        "base_url": "https://api.moonshot.cn/v1",
        "api_key_env": "MOONSHOT_API_KEY",
    },
    {
        "name": "glm",
        "type": "openai_compatible",
        "base_url": "https://open.bigmodel.cn/api/paas/v4",
        "api_key_env": "ZHIPU_API_KEY",
    },
]

Use the base_url from each provider's current documentation. Aiya does not hard-code any endpoint or model name.

Routing

Use provider/model to target a provider.

response = client.chat("Hello", model="qwen/qwen-plus")

Use a route list for fallback.

response = client.chat(
    "Hello",
    model=["deepseek/deepseek-chat", "qwen/qwen-plus"],
    route_policy="fallback",
)

Available policies:

  • fallback - try candidates in order
  • first - use only the first candidate
  • round_robin - rotate through candidates
  • random - choose a random candidate
  • least_used - choose the candidate with the lowest recorded token usage

The default policy is fallback.

Vision and Capability Filtering

Declare capabilities in models and pass require_vision=True when a request needs image input.

from aiya import Aiya

client = Aiya(
    providers=[
        {
            "name": "qwen",
            "type": "openai_compatible",
            "base_url": "https://dashscope.aliyuncs.com/compatible-mode/v1",
            "api_key": "your-key",
            "models": [
                {"id": "qwen-vl-plus", "capabilities": ["text", "vision"]}
            ],
        }
    ]
)

response = client.chat(
    [
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "What is in this image?"},
                {
                    "type": "image_url",
                    "image_url": {"url": "https://example.com/image.jpg"},
                },
            ],
        }
    ],
    model="qwen/qwen-vl-plus",
    require_vision=True,
)

Usage Ledger

Every successful completion is recorded.

response = client.chat("Hello", model="deepseek/deepseek-chat")

print(response.usage.input_tokens)
print(response.usage.output_tokens)
print(response.usage.input_cache_hit_tokens)
print(response.usage.input_cache_miss_tokens)
print(response.usage.output_cache_hit_tokens)
print(response.usage.output_cache_miss_tokens)
print(response.usage.account)
print(response.usage.reasoning_tokens)

print(client.ledger.total())
print(client.ledger.by_model())
print(client.ledger.by_provider())
print(client.ledger.by_account())

client.ledger.export_jsonl("usage.jsonl")

For persistent accounting, pass a SQLite path.

from aiya import UsageLedger

ledger = UsageLedger(sqlite_path="data/usage.sqlite3")
client = Aiya(providers=providers, ledger=ledger)

Cost Calculation

Token counts are always recorded. Cost is calculated only when the application supplies a price.

from aiya import Aiya, Price, PriceBook

price_book = PriceBook(
    {
        "deepseek-chat": Price(
            input_per_mtok=0.27,
            output_per_mtok=1.10,
        )
    }
)

client = Aiya(providers=providers, price_book=price_book)
response = client.chat("Hello", model="deepseek/deepseek-chat")

print(response.usage.cost_usd)

Prices are US dollars per million tokens.

Streaming and Async

for chunk in client.stream("Hello", model="deepseek/deepseek-chat"):
    print(chunk.content, end="", flush=True)
import asyncio


async def main():
    response = await client.achat("Hello", model="deepseek/deepseek-chat")
    print(response.text)

    async for chunk in client.astream("Hello", model="deepseek/deepseek-chat"):
        print(chunk.content, end="", flush=True)


asyncio.run(main())

Extending Providers

For a non-standard provider, subclass aiya.BaseProvider and implement complete, acomplete, stream, and astream.

from aiya import Aiya, BaseProvider


class MyProvider(BaseProvider):
    def complete(self, request):
        ...

Then register the instance directly.

client = Aiya(providers=[MyProvider(name="my-provider", models=["my-model"])])

Metadata

Release files for aiyallm 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for aiyallm 0.1.0
File Size Uploaded
aiyallm-0.1.0.tar.gz 24.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for aiyallm 0.1.0
File Interpreter ABI Platform
aiyallm-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 48.9 kB

Release files / aiyallm-0.1.0.tar.gz

Download URL aiyallm-0.1.0.tar.gz
Size 24.0 kB
Tags Source
SHA-256 checksum
How to use checksums
237785a7b4b51046f9a4ca431a2120b8c632d4864bea834c02b76d712fba95a0
BLAKE2b-256 checksum
How to use checksums
0a13e36fcb103501a78a3f64869d97ef58c9df1a85254900815d73619bef8753
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release files / aiyallm-0.1.0-py3-none-any.whl

Download URL aiyallm-0.1.0-py3-none-any.whl
Size 24.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9ce28b136a8f09b1732d8da6533631156b547b759503c126300f76ca05ef20b5
BLAKE2b-256 checksum
How to use checksums
f513b4f5fc031a641448dbd496ceda622bc39724136dfb1a0c7372b68dda32da
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page