Skip to main content

edgequake-litellm

PyPI Python Python CI License

edgequake-litellm is a LiteLLM-compatible Python package backed by the Rust edgequake-llm core. The intent is simple: keep the LiteLLM call shape, replace the Python network path with a native implementation, and preserve operational features such as streaming, tool calling, embeddings, and provider routing.

# Before
import litellm

# After
import edgequake_litellm as litellm

Install

pip install edgequake-litellm

Supported wheel targets:

Platform Architectures
Linux (glibc) x86_64, aarch64
Linux (musl) x86_64, aarch64
macOS x86_64, arm64
Windows x86_64

The package uses abi3-py39, so one wheel per platform covers Python 3.9+.

Scope note: this package covers the LiteLLM-compatible chat and embedding API surface. The Rust crate also ships image-generation providers, but those APIs are not exposed through edgequake-litellm yet.

Quick Start

import asyncio
import edgequake_litellm as litellm

messages = [{"role": "user", "content": "Explain Rust ownership in one sentence."}]

# Sync
resp = litellm.completion("openai/gpt-5.6-terra", messages, max_tokens=128)
print(resp.choices[0].message.content)

# Async
async def main() -> None:
    resp = await litellm.acompletion("anthropic/claude-sonnet-5-5", messages)
    print(resp.content)

    stream = await litellm.acompletion("openai/gpt-5.6-terra", messages, stream=True)
    async for chunk in stream:
        print(chunk.choices[0].delta.content or "", end="", flush=True)

asyncio.run(main())

Embeddings:

import edgequake_litellm as litellm

result = litellm.embedding(
    "openai/text-embedding-3-small",
    ["hello world", "rust is fast"],
)

print(result.data[0].embedding[:3])
print(len(result[0]))

Model Discovery

Programmatic model listing, capability filtering, and name/fuzzy search — backed by the Rust discovery engine:

import edgequake_litellm as litellm

# List providers (unified catalog — includes cohere, nvidia, etc.)
print(litellm.list_providers())

# Filter by capabilities (live discovery)
models = litellm.discovery.find_models(
    requires_vision=True,
    min_context_length=100_000,
    max_output_tokens=32_768,
)

# Offline capability search (no API keys)
static = litellm.discovery.find_static_models(requires_thinking=True)

# Search by name / fuzzy with input & output length bounds
hits = litellm.discovery.search_static_models_by_name(
    "claude sonnet",
    fuzzy=True,
    min_context_length=200_000,
    min_output_tokens=16_384,
)
for hit in hits:
    print(f"{hit.model.provider}/{hit.model.id} score={hit.score:.2f} ({hit.match_kind})")

# Exact lookup by ID or display name
model = litellm.discovery.lookup_model_by_name("openai", "GPT-4.1")

See docs/discovery.md for the full Rust + Python API reference.

Provider Routing

Pass provider/model as the model argument:

Provider Example
OpenAI openai/gpt-5.6-terra
Azure OpenAI azure/my-gpt56-deployment
Anthropic anthropic/claude-sonnet-5-5
Gemini gemini/gemini-3.8-flash
Vertex AI vertexai/gemini-3.8-flash
xAI xai/grok-4.7
OpenRouter openrouter/meta-llama/llama-3.1-70b-instruct
NVIDIA NIM nvidia/meta/llama-3.1-8b-instruct
Mistral mistral/mistral-medium-3-5
AWS Bedrock bedrock/amazon.nova-lite-v1:0
HuggingFace huggingface/meta-llama/Meta-Llama-3.1-8B-Instruct
OpenAI Compatible openai-compatible/deepseek-chat
Ollama ollama/llama3.2
LM Studio lmstudio/local-model
VSCode Copilot vscode-copilot/auto
Mock mock/test-model

Embedding-only backend:

Provider Example
Jina jina/jina-embeddings-v3

Supported Features

Provider Chat Stream Tools Embeddings Notes
OpenAI Yes Yes Yes Yes includes max_completion_tokens handling
Azure OpenAI Yes Yes Yes Yes deployment-based routing
Anthropic Yes Yes Yes No Claude extended thinking surfaced in response metadata
Gemini Yes Yes Yes Yes Google AI Studio
Vertex AI Yes Yes Yes Yes GCP auth / ADC
xAI Yes Yes Yes No Grok
OpenRouter Yes Yes Yes No gateway models
NVIDIA NIM Yes Yes Yes No OpenAI-compatible hosted NIM
Mistral Yes Yes Yes Yes native embeddings
AWS Bedrock Yes Yes Yes Yes backed by the Rust Bedrock feature
HuggingFace Yes Yes Limited No Inference API
OpenAI Compatible Yes Yes Yes Yes Groq, Together, DeepSeek, custom gateways
Ollama Yes Yes Yes Yes local runtime
LM Studio Yes Yes Yes Yes local OpenAI-compatible server
VSCode Copilot Yes Yes Yes Yes direct auth by default, proxy optional
Jina No No No Yes embeddings only
Mock Yes No Yes Yes unit tests / local development

Application Attribution

Propagate caller identity to upstream providers (OpenAI OpenAI-Project, OpenRouter referer/title, Ollama X-Client-Request-Id, etc.) and OTEL spans.

import edgequake_litellm as eq
from edgequake_litellm import ApplicationContext, get_provider_attribution

# Per-call kwargs
eq.completion(
    "openrouter/anthropic/claude-sonnet-5-5",
    [{"role": "user", "content": "hi"}],
    application_id="my-backend",
    application_name="My Service",
    application_url="https://app.example.com",
    request_id="req-123",
)

# Reusable context
ctx = ApplicationContext(application_id="my-backend", request_id="req-456")
eq.completion("mock/test-model", [{"role": "user", "content": "hi"}], application_context=ctx)

# Catalog: full | passthrough | observability_only | none
assert get_provider_attribution("openai") == "full"
assert get_provider_attribution("ollama") == "passthrough"

Ingress from a web framework: ApplicationContext.from_headers(request.headers).

Defaults from env: EDGEQUAKE_APP_ID, EDGEQUAKE_APP_NAME, EDGEQUAKE_APP_URL, EDGEQUAKE_TENANT_ID.

See migration guide and observability.

Environment Setup

Provider Required environment
OpenAI OPENAI_API_KEY
Azure OpenAI AZURE_OPENAI_ENDPOINT, AZURE_OPENAI_API_KEY, AZURE_OPENAI_DEPLOYMENT_NAME
Anthropic ANTHROPIC_API_KEY
Gemini GEMINI_API_KEY or GOOGLE_API_KEY
Vertex AI GOOGLE_CLOUD_PROJECT and ADC or GOOGLE_ACCESS_TOKEN
xAI XAI_API_KEY
OpenRouter OPENROUTER_API_KEY
NVIDIA NIM NVIDIA_API_KEY
Mistral MISTRAL_API_KEY
AWS Bedrock standard AWS credential chain plus AWS_REGION
HuggingFace HF_TOKEN or HUGGINGFACE_TOKEN
OpenAI Compatible OPENAI_COMPATIBLE_BASE_URL, optional OPENAI_COMPATIBLE_API_KEY
Ollama optional OLLAMA_HOST
LM Studio optional LMSTUDIO_HOST
VSCode Copilot optional VSCODE_COPILOT_PROXY_URL; otherwise reuse the official VS Code Copilot auth cache
Jina JINA_API_KEY

Module defaults:

import edgequake_litellm as litellm

litellm.set_default_provider("anthropic")
litellm.set_default_model("claude-sonnet-5-5")

Environment defaults:

  • LITELLM_EDGE_PROVIDER
  • LITELLM_EDGE_MODEL
  • LITELLM_EDGE_TIMEOUT
  • LITELLM_EDGE_MAX_RETRIES
  • LITELLM_EDGE_VERBOSE

LiteLLM Compatibility

Implemented:

  • completion()
  • acompletion()
  • embedding()
  • aembedding()
  • stream=True on acompletion()
  • stream() async generator
  • response.choices[0].message.content
  • response.to_dict()
  • AuthenticationError, RateLimitError, NotFoundError, Timeout
  • list_providers()
  • detect_provider()
  • discovery.discover_all() / adiscover_all()
  • discovery.find_models() / afind_models()
  • discovery.find_static_models()
  • discovery.search_models() / search_static_models_by_name() / asearch_models()
  • discovery.lookup_model_by_name()
  • discovery.get_model_info("provider/model")

Behavior notes:

  • synchronous streaming is intentionally not supported; use acompletion(..., stream=True) or stream()
  • unsupported or extra keyword arguments are dropped for LiteLLM parity
  • per-call api_key, api_base, and timeout parameters are accepted at the Python layer but not yet wired into the Rust core for every provider

Provider Examples

OpenAI-compatible custom gateway:

export OPENAI_COMPATIBLE_BASE_URL=https://api.groq.com/openai/v1
export OPENAI_COMPATIBLE_API_KEY=...
import edgequake_litellm as litellm

resp = litellm.completion(
    "openai-compatible/llama-3.3-70b-versatile",
    [{"role": "user", "content": "Write a one-line changelog summary."}],
)
print(resp.content)

Vertex AI:

export GOOGLE_CLOUD_PROJECT=my-project
gcloud auth application-default login
resp = litellm.completion(
    "vertexai/gemini-3.8-flash",
    [{"role": "user", "content": "Summarise this design review."}],
)

Jina embeddings:

import edgequake_litellm as litellm

vectors = litellm.embedding(
    "jina/jina-embeddings-v3",
    ["retrieval query", "retrieval document"],
)
print(len(vectors[0]))

Development

git clone https://github.com/raphaelmansuy/edgequake-llm.git
cd edgequake-llm/edgequake-litellm

python -m venv .venv
source .venv/bin/activate

pip install "maturin>=1.7" "pytest>=9.0.3" "pytest-asyncio>=0.24" "ruff>=0.3" "mypy>=1.8"
pip install . -v

pytest -q -k "not e2e"
ruff check python/
mypy python/edgequake_litellm --ignore-missing-imports

Release

Release tags are separate from the Rust crate:

  • Rust crate: vX.Y.Z
  • Python package: py-vX.Y.Z

Publish flow for edgequake-litellm:

  1. bump edgequake-litellm/Cargo.toml
  2. bump edgequake-litellm/pyproject.toml
  3. update CHANGELOG.md
  4. push the release-prep commit
  5. wait for python-ci.yml to go green
  6. push py-vX.Y.Z

python-publish.yml builds the sdist and wheels, smoke-tests the native wheels, publishes to PyPI, and can attach built artifacts to the GitHub Release.

Changelog

See CHANGELOG.md for the current release line and published history.

License

Apache-2.0. See ../LICENSE-APACHE.

Metadata

Release files for edgequake-litellm 0.10.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for edgequake-litellm 0.10.4
File Size Uploaded
edgequake_litellm-0.10.4.tar.gz 1.3 MB Details

Built distributions (wheels)

Table of built distributions (wheels) for edgequake-litellm 0.10.4
File
edgequake_litellm-0.10.4-cp39-abi3-win_amd64.whl CPython 3.9 abi3 Windows x86-64 Details
edgequake_litellm-0.10.4-cp39-abi3-musllinux_1_2_x86_64.whl CPython 3.9 abi3 Linux musl 1.2+ x86-64 Details
edgequake_litellm-0.10.4-cp39-abi3-musllinux_1_2_aarch64.whl CPython 3.9 abi3 Linux musl 1.2+ ARM64 Details
edgequake_litellm-0.10.4-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.9 abi3 Linux glibc 2.17+ x86-64 Details
edgequake_litellm-0.10.4-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl CPython 3.9 abi3 Linux glibc 2.17+ ARM64 Details
edgequake_litellm-0.10.4-cp39-abi3-macosx_11_0_arm64.whl CPython 3.9 abi3 macOS 11.0+ ARM64 Details
edgequake_litellm-0.10.4-cp39-abi3-macosx_10_12_x86_64.whl CPython 3.9 abi3 macOS 10.12+ x86-64 Details

Total release size: 64.3 MB

Release files / edgequake_litellm-0.10.4.tar.gz

Download URL edgequake_litellm-0.10.4.tar.gz
Size 1.3 MB
Tags Source
SHA-256 checksum
How to use checksums
5b311abdac74b61c19ddf7f243be2efcf4342ed03a3bbcda28d110c735d54988
BLAKE2b-256 checksum
How to use checksums
1537d42b263840ac7a19f4cd1b0a644a255b01f5e926136dd79e1ee93f08a988
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / edgequake_litellm-0.10.4-cp39-abi3-win_amd64.whl

Download URL edgequake_litellm-0.10.4-cp39-abi3-win_amd64.whl
Size 7.6 MB
Tags CPython 3.9 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
664eb68ca6bc76525afa1aec1a248c86ed6227ed2f7cf728e16bece2276ecd39
BLAKE2b-256 checksum
How to use checksums
ee796472a783d636c10887b6ccbbce8f5ff68dd776f4c4569f7dd2427bfbee84
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / edgequake_litellm-0.10.4-cp39-abi3-musllinux_1_2_x86_64.whl

Download URL edgequake_litellm-0.10.4-cp39-abi3-musllinux_1_2_x86_64.whl
Size 9.7 MB
Tags CPython 3.9 Linux musl 1.2+ x86-64 abi3
SHA-256 checksum
How to use checksums
22a14b797fe174bb173876d4673152975045000427d89b8184468e7ec642a351
BLAKE2b-256 checksum
How to use checksums
269ee4499e15e08154bca9f4af9ebc573f3895294dc35a9b7dda2b95f345dd09
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / edgequake_litellm-0.10.4-cp39-abi3-musllinux_1_2_aarch64.whl

Download URL edgequake_litellm-0.10.4-cp39-abi3-musllinux_1_2_aarch64.whl
Size 9.6 MB
Tags CPython 3.9 Linux musl 1.2+ ARM64 abi3
SHA-256 checksum
How to use checksums
0944104c239b052ffbe68d6174249805423d73fb3a9da22f24bd9106a2fcc075
BLAKE2b-256 checksum
How to use checksums
3419ba84e21949175fcff11b391449783191c788abf0be1637d3bef80f8f1d06
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / edgequake_litellm-0.10.4-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL edgequake_litellm-0.10.4-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 9.4 MB
Tags CPython 3.9 Linux glibc 2.17+ x86-64 abi3
SHA-256 checksum
How to use checksums
2c454c29cba69eeeb53257131f26f6e1a5924032b192af4be255310b4f47c913
BLAKE2b-256 checksum
How to use checksums
268436f3de5c2ba737a376ebe8bf93db197d1fa2a92272a815be3b73efc8c9e8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / edgequake_litellm-0.10.4-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl

Download URL edgequake_litellm-0.10.4-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Size 9.4 MB
Tags CPython 3.9 Linux glibc 2.17+ ARM64 abi3
SHA-256 checksum
How to use checksums
ef4235fa444fb96df945cbe08e68f07d96dd8e6a2499bcde431013e2597d02bb
BLAKE2b-256 checksum
How to use checksums
064dcb09a1398efca317a9fa6db3eda4c63f23a2690ef1c8d077470a38287a46
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / edgequake_litellm-0.10.4-cp39-abi3-macosx_11_0_arm64.whl

Download URL edgequake_litellm-0.10.4-cp39-abi3-macosx_11_0_arm64.whl
Size 8.5 MB
Tags CPython 3.9 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
b9a0882fc6846d6900e1636009385c07b88c74bee55f3e9fea902c60af415eef
BLAKE2b-256 checksum
How to use checksums
c618080d95ed291e3c3fa1e8b1b7d88bc1efaa04ae6cc51adff7bdcaf26db17f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / edgequake_litellm-0.10.4-cp39-abi3-macosx_10_12_x86_64.whl

Download URL edgequake_litellm-0.10.4-cp39-abi3-macosx_10_12_x86_64.whl
Size 8.7 MB
Tags CPython 3.9 abi3 macOS 10.12+ x86-64
SHA-256 checksum
How to use checksums
5dc281098101b1e4c60e95a41055dee302902eea883383f382df1c48fe91df05
BLAKE2b-256 checksum
How to use checksums
cf94dd4a364d4c5dcc3732b507e4a0187cbd5e4aa12fad957464c404ab1e4190
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.10.4 This release

8 release files

0.9.0

8 release files

0.6.12

8 release files

0.6.7

8 release files

0.4.0

8 release files

0.3.0

7 release files

0.2.0

7 release files

0.1.4

8 release files

0.1.3

8 release files

0.1.2

8 release files

0.1.1

8 release files

0.1.0

8 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page