edgequake-litellm
edgequake-litellm is a LiteLLM-compatible Python package backed by the Rust edgequake-llm core. The intent is simple: keep the LiteLLM call shape, replace the Python network path with a native implementation, and preserve operational features such as streaming, tool calling, embeddings, and provider routing.
# Before
import litellm
# After
import edgequake_litellm as litellm
Install
pip install edgequake-litellm
Supported wheel targets:
| Platform | Architectures |
|---|---|
| Linux (glibc) | x86_64, aarch64 |
| Linux (musl) | x86_64, aarch64 |
| macOS | x86_64, arm64 |
| Windows | x86_64 |
The package uses abi3-py39, so one wheel per platform covers Python 3.9+.
Scope note: this package covers the LiteLLM-compatible chat and embedding API
surface. The Rust crate also ships image-generation providers, but those APIs
are not exposed through edgequake-litellm yet.
Quick Start
import asyncio
import edgequake_litellm as litellm
messages = [{"role": "user", "content": "Explain Rust ownership in one sentence."}]
# Sync
resp = litellm.completion("openai/gpt-5.6-terra", messages, max_tokens=128)
print(resp.choices[0].message.content)
# Async
async def main() -> None:
resp = await litellm.acompletion("anthropic/claude-sonnet-5-5", messages)
print(resp.content)
stream = await litellm.acompletion("openai/gpt-5.6-terra", messages, stream=True)
async for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)
asyncio.run(main())
Embeddings:
import edgequake_litellm as litellm
result = litellm.embedding(
"openai/text-embedding-3-small",
["hello world", "rust is fast"],
)
print(result.data[0].embedding[:3])
print(len(result[0]))
Model Discovery
Programmatic model listing, capability filtering, and name/fuzzy search — backed by the Rust discovery engine:
import edgequake_litellm as litellm
# List providers (unified catalog — includes cohere, nvidia, etc.)
print(litellm.list_providers())
# Filter by capabilities (live discovery)
models = litellm.discovery.find_models(
requires_vision=True,
min_context_length=100_000,
max_output_tokens=32_768,
)
# Offline capability search (no API keys)
static = litellm.discovery.find_static_models(requires_thinking=True)
# Search by name / fuzzy with input & output length bounds
hits = litellm.discovery.search_static_models_by_name(
"claude sonnet",
fuzzy=True,
min_context_length=200_000,
min_output_tokens=16_384,
)
for hit in hits:
print(f"{hit.model.provider}/{hit.model.id} score={hit.score:.2f} ({hit.match_kind})")
# Exact lookup by ID or display name
model = litellm.discovery.lookup_model_by_name("openai", "GPT-4.1")
See docs/discovery.md for the full Rust + Python API reference.
Provider Routing
Pass provider/model as the model argument:
| Provider | Example |
|---|---|
| OpenAI | openai/gpt-5.6-terra |
| Azure OpenAI | azure/my-gpt56-deployment |
| Anthropic | anthropic/claude-sonnet-5-5 |
| Gemini | gemini/gemini-3.8-flash |
| Vertex AI | vertexai/gemini-3.8-flash |
| xAI | xai/grok-4.7 |
| OpenRouter | openrouter/meta-llama/llama-3.1-70b-instruct |
| NVIDIA NIM | nvidia/meta/llama-3.1-8b-instruct |
| Mistral | mistral/mistral-medium-3-5 |
| AWS Bedrock | bedrock/amazon.nova-lite-v1:0 |
| HuggingFace | huggingface/meta-llama/Meta-Llama-3.1-8B-Instruct |
| OpenAI Compatible | openai-compatible/deepseek-chat |
| Ollama | ollama/llama3.2 |
| LM Studio | lmstudio/local-model |
| VSCode Copilot | vscode-copilot/auto |
| Mock | mock/test-model |
Embedding-only backend:
| Provider | Example |
|---|---|
| Jina | jina/jina-embeddings-v3 |
Supported Features
| Provider | Chat | Stream | Tools | Embeddings | Notes |
|---|---|---|---|---|---|
| OpenAI | Yes | Yes | Yes | Yes | includes max_completion_tokens handling |
| Azure OpenAI | Yes | Yes | Yes | Yes | deployment-based routing |
| Anthropic | Yes | Yes | Yes | No | Claude extended thinking surfaced in response metadata |
| Gemini | Yes | Yes | Yes | Yes | Google AI Studio |
| Vertex AI | Yes | Yes | Yes | Yes | GCP auth / ADC |
| xAI | Yes | Yes | Yes | No | Grok |
| OpenRouter | Yes | Yes | Yes | No | gateway models |
| NVIDIA NIM | Yes | Yes | Yes | No | OpenAI-compatible hosted NIM |
| Mistral | Yes | Yes | Yes | Yes | native embeddings |
| AWS Bedrock | Yes | Yes | Yes | Yes | backed by the Rust Bedrock feature |
| HuggingFace | Yes | Yes | Limited | No | Inference API |
| OpenAI Compatible | Yes | Yes | Yes | Yes | Groq, Together, DeepSeek, custom gateways |
| Ollama | Yes | Yes | Yes | Yes | local runtime |
| LM Studio | Yes | Yes | Yes | Yes | local OpenAI-compatible server |
| VSCode Copilot | Yes | Yes | Yes | Yes | direct auth by default, proxy optional |
| Jina | No | No | No | Yes | embeddings only |
| Mock | Yes | No | Yes | Yes | unit tests / local development |
Application Attribution
Propagate caller identity to upstream providers (OpenAI OpenAI-Project, OpenRouter referer/title, Ollama X-Client-Request-Id, etc.) and OTEL spans.
import edgequake_litellm as eq
from edgequake_litellm import ApplicationContext, get_provider_attribution
# Per-call kwargs
eq.completion(
"openrouter/anthropic/claude-sonnet-5-5",
[{"role": "user", "content": "hi"}],
application_id="my-backend",
application_name="My Service",
application_url="https://app.example.com",
request_id="req-123",
)
# Reusable context
ctx = ApplicationContext(application_id="my-backend", request_id="req-456")
eq.completion("mock/test-model", [{"role": "user", "content": "hi"}], application_context=ctx)
# Catalog: full | passthrough | observability_only | none
assert get_provider_attribution("openai") == "full"
assert get_provider_attribution("ollama") == "passthrough"
Ingress from a web framework: ApplicationContext.from_headers(request.headers).
Defaults from env: EDGEQUAKE_APP_ID, EDGEQUAKE_APP_NAME, EDGEQUAKE_APP_URL, EDGEQUAKE_TENANT_ID.
See migration guide and observability.
Environment Setup
| Provider | Required environment |
|---|---|
| OpenAI | OPENAI_API_KEY |
| Azure OpenAI | AZURE_OPENAI_ENDPOINT, AZURE_OPENAI_API_KEY, AZURE_OPENAI_DEPLOYMENT_NAME |
| Anthropic | ANTHROPIC_API_KEY |
| Gemini | GEMINI_API_KEY or GOOGLE_API_KEY |
| Vertex AI | GOOGLE_CLOUD_PROJECT and ADC or GOOGLE_ACCESS_TOKEN |
| xAI | XAI_API_KEY |
| OpenRouter | OPENROUTER_API_KEY |
| NVIDIA NIM | NVIDIA_API_KEY |
| Mistral | MISTRAL_API_KEY |
| AWS Bedrock | standard AWS credential chain plus AWS_REGION |
| HuggingFace | HF_TOKEN or HUGGINGFACE_TOKEN |
| OpenAI Compatible | OPENAI_COMPATIBLE_BASE_URL, optional OPENAI_COMPATIBLE_API_KEY |
| Ollama | optional OLLAMA_HOST |
| LM Studio | optional LMSTUDIO_HOST |
| VSCode Copilot | optional VSCODE_COPILOT_PROXY_URL; otherwise reuse the official VS Code Copilot auth cache |
| Jina | JINA_API_KEY |
Module defaults:
import edgequake_litellm as litellm
litellm.set_default_provider("anthropic")
litellm.set_default_model("claude-sonnet-5-5")
Environment defaults:
LITELLM_EDGE_PROVIDERLITELLM_EDGE_MODELLITELLM_EDGE_TIMEOUTLITELLM_EDGE_MAX_RETRIESLITELLM_EDGE_VERBOSE
LiteLLM Compatibility
Implemented:
completion()acompletion()embedding()aembedding()stream=Trueonacompletion()stream()async generatorresponse.choices[0].message.contentresponse.to_dict()AuthenticationError,RateLimitError,NotFoundError,Timeoutlist_providers()detect_provider()discovery.discover_all()/adiscover_all()discovery.find_models()/afind_models()discovery.find_static_models()discovery.search_models()/search_static_models_by_name()/asearch_models()discovery.lookup_model_by_name()discovery.get_model_info("provider/model")
Behavior notes:
- synchronous streaming is intentionally not supported; use
acompletion(..., stream=True)orstream() - unsupported or extra keyword arguments are dropped for LiteLLM parity
- per-call
api_key,api_base, andtimeoutparameters are accepted at the Python layer but not yet wired into the Rust core for every provider
Provider Examples
OpenAI-compatible custom gateway:
export OPENAI_COMPATIBLE_BASE_URL=https://api.groq.com/openai/v1
export OPENAI_COMPATIBLE_API_KEY=...
import edgequake_litellm as litellm
resp = litellm.completion(
"openai-compatible/llama-3.3-70b-versatile",
[{"role": "user", "content": "Write a one-line changelog summary."}],
)
print(resp.content)
Vertex AI:
export GOOGLE_CLOUD_PROJECT=my-project
gcloud auth application-default login
resp = litellm.completion(
"vertexai/gemini-3.8-flash",
[{"role": "user", "content": "Summarise this design review."}],
)
Jina embeddings:
import edgequake_litellm as litellm
vectors = litellm.embedding(
"jina/jina-embeddings-v3",
["retrieval query", "retrieval document"],
)
print(len(vectors[0]))
Development
git clone https://github.com/raphaelmansuy/edgequake-llm.git
cd edgequake-llm/edgequake-litellm
python -m venv .venv
source .venv/bin/activate
pip install "maturin>=1.7" "pytest>=9.0.3" "pytest-asyncio>=0.24" "ruff>=0.3" "mypy>=1.8"
pip install . -v
pytest -q -k "not e2e"
ruff check python/
mypy python/edgequake_litellm --ignore-missing-imports
Release
Release tags are separate from the Rust crate:
- Rust crate:
vX.Y.Z - Python package:
py-vX.Y.Z
Publish flow for edgequake-litellm:
- bump
edgequake-litellm/Cargo.toml - bump
edgequake-litellm/pyproject.toml - update
CHANGELOG.md - push the release-prep commit
- wait for
python-ci.ymlto go green - push
py-vX.Y.Z
python-publish.yml builds the sdist and wheels, smoke-tests the native wheels, publishes to PyPI, and can attach built artifacts to the GitHub Release.
Changelog
See CHANGELOG.md for the current release line and published history.
License
Apache-2.0. See ../LICENSE-APACHE.
Metadata
Release files for edgequake-litellm 0.10.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| edgequake_litellm-0.10.4.tar.gz | 1.3 MB | Details |
Built distributions (wheels)
| File | Reset | |||
|---|---|---|---|---|
| edgequake_litellm-0.10.4-cp39-abi3-win_amd64.whl | CPython 3.9 | abi3 | Windows x86-64 | Details |
| edgequake_litellm-0.10.4-cp39-abi3-musllinux_1_2_x86_64.whl | CPython 3.9 | abi3 | Linux musl 1.2+ x86-64 | Details |
| edgequake_litellm-0.10.4-cp39-abi3-musllinux_1_2_aarch64.whl | CPython 3.9 | abi3 | Linux musl 1.2+ ARM64 | Details |
| edgequake_litellm-0.10.4-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl | CPython 3.9 | abi3 | Linux glibc 2.17+ x86-64 | Details |
| edgequake_litellm-0.10.4-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl | CPython 3.9 | abi3 | Linux glibc 2.17+ ARM64 | Details |
| edgequake_litellm-0.10.4-cp39-abi3-macosx_11_0_arm64.whl | CPython 3.9 | abi3 | macOS 11.0+ ARM64 | Details |
| edgequake_litellm-0.10.4-cp39-abi3-macosx_10_12_x86_64.whl | CPython 3.9 | abi3 | macOS 10.12+ x86-64 | Details |
Total release size: 64.3 MB
Release files / edgequake_litellm-0.10.4.tar.gz
| Download URL | edgequake_litellm-0.10.4.tar.gz |
|---|---|
| Size | 1.3 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5b311abdac74b61c19ddf7f243be2efcf4342ed03a3bbcda28d110c735d54988
|
|
BLAKE2b-256 checksum How to use checksums |
1537d42b263840ac7a19f4cd1b0a644a255b01f5e926136dd79e1ee93f08a988
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / edgequake_litellm-0.10.4-cp39-abi3-win_amd64.whl
| Download URL | edgequake_litellm-0.10.4-cp39-abi3-win_amd64.whl |
|---|---|
| Size | 7.6 MB |
| Tags | CPython 3.9 Windows x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
664eb68ca6bc76525afa1aec1a248c86ed6227ed2f7cf728e16bece2276ecd39
|
|
BLAKE2b-256 checksum How to use checksums |
ee796472a783d636c10887b6ccbbce8f5ff68dd776f4c4569f7dd2427bfbee84
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / edgequake_litellm-0.10.4-cp39-abi3-musllinux_1_2_x86_64.whl
| Download URL | edgequake_litellm-0.10.4-cp39-abi3-musllinux_1_2_x86_64.whl |
|---|---|
| Size | 9.7 MB |
| Tags | CPython 3.9 Linux musl 1.2+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
22a14b797fe174bb173876d4673152975045000427d89b8184468e7ec642a351
|
|
BLAKE2b-256 checksum How to use checksums |
269ee4499e15e08154bca9f4af9ebc573f3895294dc35a9b7dda2b95f345dd09
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / edgequake_litellm-0.10.4-cp39-abi3-musllinux_1_2_aarch64.whl
| Download URL | edgequake_litellm-0.10.4-cp39-abi3-musllinux_1_2_aarch64.whl |
|---|---|
| Size | 9.6 MB |
| Tags | CPython 3.9 Linux musl 1.2+ ARM64 abi3 |
|
SHA-256 checksum How to use checksums |
0944104c239b052ffbe68d6174249805423d73fb3a9da22f24bd9106a2fcc075
|
|
BLAKE2b-256 checksum How to use checksums |
3419ba84e21949175fcff11b391449783191c788abf0be1637d3bef80f8f1d06
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / edgequake_litellm-0.10.4-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
| Download URL | edgequake_litellm-0.10.4-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl |
|---|---|
| Size | 9.4 MB |
| Tags | CPython 3.9 Linux glibc 2.17+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
2c454c29cba69eeeb53257131f26f6e1a5924032b192af4be255310b4f47c913
|
|
BLAKE2b-256 checksum How to use checksums |
268436f3de5c2ba737a376ebe8bf93db197d1fa2a92272a815be3b73efc8c9e8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / edgequake_litellm-0.10.4-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
| Download URL | edgequake_litellm-0.10.4-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl |
|---|---|
| Size | 9.4 MB |
| Tags | CPython 3.9 Linux glibc 2.17+ ARM64 abi3 |
|
SHA-256 checksum How to use checksums |
ef4235fa444fb96df945cbe08e68f07d96dd8e6a2499bcde431013e2597d02bb
|
|
BLAKE2b-256 checksum How to use checksums |
064dcb09a1398efca317a9fa6db3eda4c63f23a2690ef1c8d077470a38287a46
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / edgequake_litellm-0.10.4-cp39-abi3-macosx_11_0_arm64.whl
| Download URL | edgequake_litellm-0.10.4-cp39-abi3-macosx_11_0_arm64.whl |
|---|---|
| Size | 8.5 MB |
| Tags | CPython 3.9 abi3 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
b9a0882fc6846d6900e1636009385c07b88c74bee55f3e9fea902c60af415eef
|
|
BLAKE2b-256 checksum How to use checksums |
c618080d95ed291e3c3fa1e8b1b7d88bc1efaa04ae6cc51adff7bdcaf26db17f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / edgequake_litellm-0.10.4-cp39-abi3-macosx_10_12_x86_64.whl
| Download URL | edgequake_litellm-0.10.4-cp39-abi3-macosx_10_12_x86_64.whl |
|---|---|
| Size | 8.7 MB |
| Tags | CPython 3.9 abi3 macOS 10.12+ x86-64 |
|
SHA-256 checksum How to use checksums |
5dc281098101b1e4c60e95a41055dee302902eea883383f382df1c48fe91df05
|
|
BLAKE2b-256 checksum How to use checksums |
cf94dd4a364d4c5dcc3732b507e4a0187cbd5e4aa12fad957464c404ab1e4190
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|