Skip to main content

vv-llm

中文文档

Universal LLM interface layer for Python. One API, 17 backends, sync & async.

pip install vv-llm

Supported Backends

OpenAI | Anthropic | DeepSeek | Gemini | Qwen | Groq | Mistral | Moonshot | MiniMax | Yi | ZhiPuAI | Baichuan | StepFun | xAI | Xiaomi | Ernie | Local

Also supports Azure OpenAI, Vertex AI, and AWS Bedrock deployments.

Quick Start

Configure

from vv_llm.settings import settings

settings.load({
    "endpoints": [
        {
            "id": "openai-default",
            "api_base": "https://api.openai.com/v1",
            "api_key": "sk-...",
        }
    ],
    "backends": {
        "openai": {
            "models": {
                "gpt-4o": {
                    "id": "gpt-4o",
                    "endpoints": ["openai-default"],
                }
            }
        }
    }
})

Typed sync (canonical request)

from vv_llm.chat_clients import create_chat_client, BackendType
from vv_llm import ChatRequest, ChatRequestOptions, ThinkingPreference

client = create_chat_client(BackendType.OpenAI, model="gpt-4o")
resp = client.create(
    ChatRequest(
        model="gpt-4o",
        messages=[{"role": "user", "content": "Explain RAG in one sentence"}],
        options=ChatRequestOptions(
            thinking=ThinkingPreference.default(),
            max_tokens=512,
        ),
    )
)
print(resp.content)

ChatRequest is the normalized runtime request. At a contract boundary, ChatRequest.from_contract(...) decodes canonical JSON (where model is required and options.stream is nested), while to_contract() omits runtime transport controls such as headers and query parameters.

Use ThinkingPreference.default() to preserve the provider default, enabled() or enabled(budget_tokens=...) to opt in, and disabled() to opt out explicitly.

Keyword API

create_completion(...) accepts keyword arguments. Pass thinking explicitly when a provider supports Anthropic-style thinking control; omit it to use the provider default:

resp = client.create_completion(
    messages=[{"role": "user", "content": "Answer directly"}],
    thinking={"type": "disabled"},
)

Middleware, Retry, And Metadata

Wrap a client with MiddlewareChatClient for middleware hooks, classified retry, and execution metadata:

from vv_llm import ChatMiddlewareV1, ChatRequest, MiddlewareChatClient, RetryPolicy

class TraceMiddleware(ChatMiddlewareV1):
    def on_request(self, context, request):
        context.attributes["trace_id"] = "request-42"
        return request

runtime = MiddlewareChatClient(
    client,
    [TraceMiddleware()],
    retry_policy=RetryPolicy(max_attempts=3, total_timeout=20),
)
result = runtime.create_with_metadata(
    ChatRequest(messages=[{"role": "user", "content": "Answer directly"}])
)

print(result.response.content)
print(result.metadata.provider, result.metadata.attempts, result.metadata.latency_ms)

ErrorKind distinguishes authentication, rate limiting, network, timeout, invalid request, context length, content policy, missing model, provider internal, serialization, and configuration failures. The default retry policy retries only transient kinds and respects retry-after-ms plus numeric or HTTP-date Retry-After, exponential backoff, jitter, and an optional total deadline.

Explicit Registry And Fallback

Fallback is opt-in and ordered. Every registration declares model capabilities, so an incompatible route is skipped without sending a request:

from vv_llm import FallbackChatClient, FallbackRoute, ProviderRegistry

registry = ProviderRegistry()
registry.register(
    "primary",
    lambda: primary_client,
    capabilities=primary_client.capabilities,
)
registry.register(
    "secondary",
    lambda: secondary_client,
    capabilities=secondary_client.capabilities,
)
runtime = FallbackChatClient(
    registry,
    [
        FallbackRoute("primary", "primary-model"),
        FallbackRoute("secondary", "secondary-model"),
    ],
)

Authentication and invalid-request errors do not fall back by default. Streaming may switch routes only while establishing the stream or before its first visible chunk; after output begins, later errors are returned without replay.

Streaming

from vv_llm import ChatRequest

for chunk in client.create(ChatRequest(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Write a haiku"}],
    stream=True,
)):
    if chunk.content:
        print(chunk.content, end="")

Async

import asyncio
from vv_llm.chat_clients import create_async_chat_client, BackendType
from vv_llm import ChatRequest

async def main():
    client = create_async_chat_client(BackendType.OpenAI, model="gpt-4o")
    resp = await client.create(ChatRequest(
        model="gpt-4o",
        messages=[{"role": "user", "content": "hello"}],
    ))
    print(resp.content)

asyncio.run(main())

HTTP transport clients

The http_client argument accepts httpx2.Client for sync calls and httpx2.AsyncClient for async calls. This lets applications provide a custom transport (for example, an offline MockTransport) while keeping OpenAI 3.x and Anthropic 1.x clients on the same HTTPX2 runtime:

import httpx2
from vv_llm.chat_clients import BackendType, create_chat_client

transport = httpx2.MockTransport(
    lambda request: httpx2.Response(200, json={"choices": []}, request=request)
)
http_client = httpx2.Client(transport=transport)
client = create_chat_client(BackendType.OpenAI, model="gpt-4o", http_client=http_client)

Use an endpoint proxy setting when vv-llm should construct the transport client itself. Legacy httpx.Client and httpx.AsyncClient instances are not accepted; use the matching HTTPX2 client type instead.

Embedding & Rerank

from vv_llm.settings import settings

settings.load({
    "endpoints": [
        {
            "id": "siliconflow",
            "api_base": "https://api.siliconflow.cn/v1",
            "api_key": "sk-...",
        }
    ],
    "backends": {},
    "embedding_backends": {
        "siliconflow": {
            "models": {
                "BAAI/bge-large-zh-v1.5": {
                    "id": "BAAI/bge-large-zh-v1.5",
                    "endpoints": ["siliconflow"],
                    "protocol": "openai_embeddings",
                }
            }
        }
    },
    "rerank_backends": {
        "siliconflow": {
            "models": {
                "BAAI/bge-reranker-v2-m3": {
                    "id": "BAAI/bge-reranker-v2-m3",
                    "endpoints": ["siliconflow"],
                    "protocol": "custom_json_http",
                    "request_mapping": {
                        "method": "POST",
                        "path": "/rerank",
                        "body_template": {
                            "model": "${model_id}",
                            "query": "${query}",
                            "documents": "${documents}",
                        },
                    },
                    "response_mapping": {
                        "results_path": "$.results[*]",
                        "field_map": {
                            "index": "$.index",
                            "relevance_score": "$.relevance_score",
                        },
                    },
                }
            }
        }
    },
})
from vv_llm.embedding_clients import create_embedding_client
from vv_llm.rerank_clients import create_rerank_client

embedding_client = create_embedding_client("siliconflow", model="BAAI/bge-large-zh-v1.5")
embedding_resp = embedding_client.create_embeddings(input="hello world")
print(len(embedding_resp.data[0].embedding))

rerank_client = create_rerank_client("siliconflow", model="BAAI/bge-reranker-v2-m3")
rerank_resp = rerank_client.rerank(
    query="Apple",
    documents=["apple", "banana", "fruit", "vegetable"],
)
print(rerank_resp.results[0].index, rerank_resp.results[0].relevance_score)
import asyncio
from vv_llm.embedding_clients import create_async_embedding_client
from vv_llm.rerank_clients import create_async_rerank_client

async def main():
    embedding_client = create_async_embedding_client("siliconflow", model="BAAI/bge-large-zh-v1.5")
    rerank_client = create_async_rerank_client("siliconflow", model="BAAI/bge-reranker-v2-m3")

    emb = await embedding_client.create_embeddings(input=["a", "b"])
    rr = await rerank_client.rerank(query="Apple", documents=["apple", "banana"])
    print(len(emb.data), len(rr.results))

asyncio.run(main())

Reasoning effort

Model capabilities expose reasoning_efforts: omitted/null means unknown, [] means unsupported, and a list declares effective choices. Unspecified effort uses the provider default; none is an explicit model-dependent value.

Select a different model through the request model or an endpoint binding's model_id. A conflicting extra_body.model is rejected even with passthrough.

Use client.create(request, capability_policy=CapabilityPolicy.STRICT) or client.create_completion(..., capability_policy=CapabilityPolicy.STRICT) for validation before sending. The default is WARN; PASSTHROUGH skips model support checks. Conflicting controls still fail. Responses maps effort to reasoning.effort and Anthropic maps it to output_config.effort.

Endpoint binding capabilities partially override model metadata; lists replace inherited lists. Registry model_capabilities supplies per-model fallback metadata. Each route preserves the requested effort and skips incompatible models.

Existing keyword and typed calls remain valid. Leaving effort unspecified still uses the provider default. Calls that previously sent an unrecognized value can now emit a warning under the default policy; opting into strict validation rejects both unsupported values and unknown model support before a provider request. This applies to sync/async and completion/streaming paths.

After configuring a DeepSeek Flash endpoint binding:

from vv_llm import CapabilityPolicy, ChatRequest, ChatRequestOptions
from vv_llm.chat_clients import BackendType, create_chat_client

client = create_chat_client(BackendType.DeepSeek, model="deepseek-flash")
print(client.capabilities.reasoning_efforts)
print(client.capabilities.reasoning_effort_aliases)

response = client.create(
    ChatRequest(
        model=client.model,
        messages=[{"role": "user", "content": "Compute 37 * 19."}],
        options=ChatRequestOptions(reasoning_effort="xhigh", max_tokens=256),
    ),
    capability_policy=CapabilityPolicy.STRICT,
)
# The provider receives xhigh unchanged; its documented effective target is high.

The keyword API uses the same policy and model-specific validation:

response = client.create_completion(
    messages=[{"role": "user", "content": "Compute 37 * 19."}],
    reasoning_effort="high",
    capability_policy=CapabilityPolicy.STRICT,
    max_tokens=256,
)

Use effective choices for a model selector; compatibility aliases are additional accepted inputs, not separate intensities. An included none is an explicit off control. Effort omission does not enable or disable thinking: use ThinkingPreference for that separate model capability. GLM-5.3/FLASH reject xhigh in strict mode and require thinking; GLM-5.2 permits explicit disabled thinking. In the recorded GLM-5.2 live checks, none/minimal still returned reasoning content, while explicit disabled thinking did not.

Offline capabilities/validation/fallback and configured effort/thinking/streaming examples are described in the examples guide.

reasoning_effort_aliases maps documented compatibility inputs to effective choices. Aliases are accepted only when their target remains in reasoning_efforts; requests retain the original input. Both lists and alias maps on bindings replace inherited fields. DeepSeek exposes low/high/max, plus none for off, with minimal → low, medium/xhigh → high and ultra → max. Aliases are not extra selectable intensities.

Features

  • Unified interface — canonical ChatRequest execution across all providers, with create_completion / create_stream retained for compatibility
  • Embedding & rerank — unified sync/async retrieval clients with normalized outputs
  • Type-safe factory — create_chat_client(BackendType.X) returns the correct client type
  • Multi-endpoint — select enabled endpoints by ascending priority, preserving configuration order within each tier
  • Tool calling — normalized tool/function calling across providers
  • Multimodal — text + image inputs where supported
  • Thinking/reasoning — access chain-of-thought from Claude, DeepSeek Reasoner, etc.
  • Token counting — per-model tokenizers (tiktoken, deepseek-tokenizer, qwen-tokenizer)
  • Rate limiting — RPM/TPM controls with memory, Redis, or DiskCache backends
  • Context length control — automatic message truncation to fit model limits
  • Prompt caching — Anthropic prompt caching support
  • Retry with backoff — configurable retry logic for transient failures
  • Versioned middleware — stable v1 request, response, and error hooks outside provider adapters
  • Classified errors — provider-neutral error kinds with retryability and request context
  • Explicit fallback — registered, ordered, capability-aware routes with no hidden provider switching
  • Scripted testing — deterministic completion/error/stream scripts for conformance tests

Model endpoint bindings accept an optional priority integer of at least 1 (default: 1). Explicit endpoint_id selection takes precedence. from vv_llm.settings import order_endpoints exposes the same stable ordering: order_endpoints(endpoints, preferred_endpoint_id=None) returns a new list. A preferred endpoint moves ahead only within its priority tier.

The package includes vv-llm-contract 1.2.0. Read contract metadata, the model catalog, and integrity status through vv_llm.contract:

from vv_llm.contract import contract_info, load_catalog, verify_contract

info = contract_info()
assert info.contract_version == "1.2.0"
assert verify_contract().ok
catalog = load_catalog()

Maintainers can validate the packaged copy with pdm run contract-check and update it from a verified release directory with pdm run contract-sync --source PATH.

Python Capability Matrix

Surface Python support Boundary
Middleware MiddlewareChatClient and AsyncMiddlewareChatClient; v1 request/response/error hooks and metadata Opt-in wrapper around a chat client
Fallback FallbackChatClient and AsyncFallbackChatClient; ordered, capability-aware routes Stream fallback is limited to setup/before the first visible chunk
Retry RetryPolicy plus sync/async executors; classified transient errors, Retry-After, backoff, jitter, deadline Authentication and invalid-request errors are not retried by default
Deterministic testing Scripted clients, vendored protocol fixtures, and unit tests No network access; live checks require explicit opt-in
Chat providers Anthropic native adapter; 15 OpenAI-compatible adapters; Local adapter Sync/async and streaming are normalized; tools, structured output, multimodal input, and thinking remain model/provider dependent
Embedding Sync/async configured clients openai_embeddings, SiliconFlow, Cohere, Voyage, and custom JSON HTTP protocols
Rerank Sync/async configured clients OpenAI-compatible, Cohere, Jina, Voyage, SiliconFlow, and custom JSON HTTP protocols

Chat Provider Matrix

Adapter Providers Transport and common behavior
Native Anthropic Native sync/async chat, streaming, tools, vision, thinking, and prompt-cache handling
OpenAI-compatible OpenAI, DeepSeek, Gemini, Groq, MiniMax, Mistral, Moonshot, Qwen, Yi, ZhiPuAI, Baichuan, StepFun, xAI, Xiaomi, Ernie Shared sync/async request and stream normalization; actual tools, structured output, multimodal, and reasoning support follows the vendored model catalog and provider endpoint
Local Local Same configured adapter shape for sync/async and streaming; endpoint behavior is deployment-specific

Examples

Runnable examples are in examples/: basic_chat.py, streaming.py, tools.py, multimodal.py, and contract_json.py cover the main typed request paths. async_streaming.py, typed_thinking.py, middleware_metadata.py, and registry_fallback.py cover focused extensions; the last one is deterministic and offline. reasoning_capabilities.py is also offline; reasoning_effort.py sends one configured request with optional streaming.

Cache Usage Semantics

OpenAI-compatible chat completions report cache reads through usage.prompt_tokens_details.cached_tokens. usage.prompt_tokens remains the total input token count, so consumers can calculate uncached input as prompt_tokens - cached_tokens. This path intentionally does not populate Anthropic's cache_read_input_tokens field because Anthropic defines its base input_tokens as uncached input.

For generic OpenAI-compatible backends, omitted cache-read fields remain unknown, while an explicit cached_tokens: 0 is preserved as an observed zero. Moonshot may omit both top-level cached_tokens and prompt_tokens_details on a cold request; only in that fully omitted case does vv-llm project prompt_tokens_details.cached_tokens = 0 from the provider contract. Explicit null or invalid cache values remain unknown.

Utilities

from vv_llm.chat_clients import format_messages, get_token_counts, get_message_token_counts
Function Description
format_messages Normalize multimodal/tool messages across formats
get_token_counts Count tokens for a text string
get_message_token_counts Count tokens for a message list

Optional Dependencies

pip install 'vv-llm[redis]'      # Redis rate limiting
pip install 'vv-llm[diskcache]'  # DiskCache rate limiting
pip install 'vv-llm[server]'     # FastAPI token server
pip install 'vv-llm[vertex]'     # Google Vertex AI
pip install 'vv-llm[bedrock]'    # AWS Bedrock

Project Structure

src/vv_llm/
  _contract/      # Versioned schemas, fixtures, catalog, and consumer lock
  chat_clients/    # Per-backend clients + factory
  embedding_clients/  # Embedding clients + factory
  rerank_clients/     # Rerank clients + factory
  retrieval_clients/  # Shared retrieval client internals
  settings/        # Configuration management
  types/           # Type definitions & enums
  utilities/       # Rate limiting, retry, media processing, token counting
  server/          # Optional token counting server

tests/unit/        # Unit tests
tests/live/        # Live integration tests (requires real API keys)

User, Maintainer, And Release Workflows

Users

Install the package and configure endpoints through the public Settings API. Contract artifacts are read from the installed package; no repository checkout or local path is required at runtime.

Maintainers

pdm install -d          # Install dev dependencies
pdm run contract-check  # Validate only the vendored contract lock
pdm run contract-sync --source PATH  # Refresh from an explicit source tree
# Or: VV_LLM_CONTRACT_SOURCE=PATH pdm run contract-sync
# Compare an explicit source tree with the vendor:
python scripts/sync_contract.py --check --source PATH
pdm run lint            # Ruff linter
pdm run format-check    # Ruff format check
pdm run type-check      # Ty type checker
pdm run test            # Unit tests

For an intentional live smoke check, provide private settings through the existing tests/dev_settings.py mechanism and opt in explicitly:

VV_LLM_RUN_LIVE_TESTS=1 python tests/live/run_live_tests.py test_deepseek_contract_smoke.py

The smoke output contains only provider/model, response shape, usage counters, and exit status; it does not print credentials or response content.

Reasoning effort live checks

For an opt-in transport smoke using an explicit private settings file:

python tests/live/reasoning_effort_smoke.py --settings /secure/path/llm_settings.json \
  --backend deepseek --aliases --invalid-probe --limit 18

The smoke tests each declared value on credentialed model/transport routes, up to 80 requests by default. --aliases includes compatibility inputs; --invalid-probe tests an invalid value; --model backend:model selects a model; --report saves sanitized observations. --include-catalog opts into default-endpoint bindings absent from the local model list. Acceptance alone does not prove intensity behavior. The SDK timeout is configured per request, not a total wall-clock limit. See the recorded live results.

Release publishers

pdm build
python scripts/smoke_wheel.py

The release CI performs contract-check, unit tests, linting, package build, and isolated wheel smoke before publication. Live API checks are intentionally not part of release CI.

License

MIT

Metadata

Release files for vv-llm 0.7.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vv-llm 0.7.3
File Size Uploaded
vv_llm-0.7.3.tar.gz 117.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for vv-llm 0.7.3
File Interpreter ABI Platform
vv_llm-0.7.3-py3-none-any.whl Python 3 none any Details

Total release size: 262.6 kB

Release files / vv_llm-0.7.3.tar.gz

Download URL vv_llm-0.7.3.tar.gz
Size 117.9 kB
Tags Source
SHA-256 checksum
How to use checksums
5115ea1d9318120f94e24e898fec8c41e7738d54673355dce93e4a2cf803ef4c
BLAKE2b-256 checksum
How to use checksums
07cf765961a9f5c4c499672d2a239dd2a8ceb12d89b9a9a2b68b0b5a4f90b8e0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / vv_llm-0.7.3-py3-none-any.whl

Download URL vv_llm-0.7.3-py3-none-any.whl
Size 144.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
692b4222adfe76eb030b72b1ba451879af0f3837204f332e00c8716d0ac8c712
BLAKE2b-256 checksum
How to use checksums
02f339d97efebf74f4bf38e2ab5ad60ef3fd381670c805b57017a26a20c189d7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

0.7.5

2 release files

0.7.4

2 release files

This release

0.7.3 This release

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.7

2 release files

0.6.6

2 release files

0.6.5

2 release files

0.6.4

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.99

2 release files

0.3.98

2 release files

0.3.97

2 release files

0.3.96

2 release files

0.3.94

2 release files

0.3.93

2 release files

0.3.92

2 release files

0.3.91

2 release files

0.3.86

2 release files

0.3.85

2 release files

0.3.84

2 release files

0.3.83

2 release files

0.3.82

2 release files

0.3.81

2 release files

0.3.80

2 release files

0.3.79

2 release files

0.3.78

2 release files

0.3.77

2 release files

0.3.75

2 release files

0.3.74

2 release files

0.3.73

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page