Resilient cascading LLM inference across multiple providers with failover, circuit breaking, and retry cooldowns
Project description
llm-pycascade
Resilient cascading LLM inference with automatic failover, circuit breaking, and retry cooldowns.
A faithful Python port of the llm-cascade Rust crate.
Features
- Cascade inference — define ordered fallback chains across multiple LLM providers
- Automatic failover — if a provider fails, the next in the cascade is tried immediately
- Circuit breaking / cooldowns — providers that fail repeatedly are put on exponential backoff
- Retry-After support — respects HTTP 429
Retry-Afterheaders - Failed-prompt persistence — conversations that exhaust all providers are saved as timestamped JSON
- Multi-provider — OpenAI, Anthropic, Google Gemini, and Ollama out of the box
- OpenAI-compatible — works with vLLM, LiteLLM, Together AI, and any OpenAI-compatible endpoint
- Tool/function calling — first-class support across all providers
- Async — built on
httpxandaiosqlitefor non-blocking I/O - Keyring integration — optional
keyringpackage for secure API key storage
Architecture
┌──────────────┐ ┌──────────────────────────────────────────┐
│ User Code │────▶│ run_cascade() │
└──────────────┘ │ │
│ ┌─────────┐ ┌─────────┐ ┌─────────┐ │
│ │ Entry 1 │─▶│ Entry 2 │─▶│ Entry 3 │ │
│ │OpenAI │ │Anthropic│ │ Gemini │ │
│ │gpt-4o │ │claude.. │ │gemini.. │ │
│ └─────────┘ └─────────┘ └─────────┘ │
│ ▼ ▼ ▼ │
│ ┌─────────────────────────────────────┐ │
│ │ Cooldown Tracker (SQLite) │ │
│ │ ┌───────────────────────────────┐ │ │
│ │ │ is_on_cooldown() / set_cooldown│ │ │
│ │ └───────────────────────────────┘ │ │
│ └─────────────────────────────────────┘ │
│ ▼ ▼ ▼ │
│ ┌─────────────────────────────────────┐ │
│ │ Attempt Log (SQLite) │ │
│ │ ┌───────────────────────────────┐ │ │
│ │ │ log_attempt(status, latency, │ │ │
│ │ │ tokens) │ │ │
│ │ └───────────────────────────────┘ │ │
│ └─────────────────────────────────────┘ │
│ ▼ (all failed) │
│ ┌─────────────────────────────────────┐ │
│ │ save_failed_conversation() → JSON │ │
│ └─────────────────────────────────────┘ │
└──────────────────────────────────────────┘
Installation
pip
pip install llm-pycascade
uv
uv add llm-pycascade
With optional keyring support
pip install llm-pycascade[keyring]
Configuration
Create a TOML config file (defaults to ~/.config/llm-pycascade/config.toml):
# config.example.toml — full example configuration
[providers.openai]
type = "openai"
api_key_env = "OPENAI_API_KEY"
[providers.anthropic]
type = "anthropic"
api_key_env = "ANTHROPIC_API_KEY"
[providers.gemini]
type = "gemini"
api_key_env = "GEMINI_API_KEY"
[providers.ollama]
type = "ollama"
# No API key needed for local Ollama
[cascades.primary]
entries = [
{ provider = "openai", model = "gpt-4o" },
{ provider = "anthropic", model = "claude-sonnet-4-20250514" },
{ provider = "gemini", model = "gemini-2.0-flash" },
{ provider = "ollama", model = "llama3.1" },
]
[cascades.fast]
entries = [
{ provider = "anthropic", model = "claude-haiku-4-20250507" },
{ provider = "openai", model = "gpt-4o-mini" },
{ provider = "ollama", model = "mistral" },
]
[database]
path = "~/.local/share/llm-pycascade/db.sqlite"
[failure_persistence]
dir = "~/.local/share/llm-pycascade/failed_prompts"
Usage
Basic cascade
import asyncio
from llm_pycascade import (
AppConfig,
Conversation,
init_db,
load_config,
run_cascade,
)
async def main():
config = load_config()
db_path = config.database.path.replace("~", "/home/user")
conn = await init_db(db_path)
conversation = Conversation.single_user_prompt("Explain quantum computing in one sentence.")
try:
response = await run_cascade("primary", conversation, config, conn)
print(response.text_only())
finally:
await conn.close()
asyncio.run(main())
With tools
from llm_pycascade import ContentBlock, Message, ToolDefinition
tools = [
ToolDefinition(
name="get_weather",
description="Get the current weather for a city",
parameters={
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name"},
},
"required": ["city"],
},
)
]
conversation = Conversation.with_tools(
messages=[Message.user("What's the weather in Tokyo?")],
tools=tools,
)
response = await run_cascade("primary", conversation, config, conn)
for block in response.content:
if block.type == ContentBlockType.TEXT:
print(f"Text: {block.text}")
elif block.type == ContentBlockType.TOOL_CALL:
print(f"Tool call: {block.name}({block.arguments})")
Environment variable override
export LLM_PYCASCADE_CONFIG=/path/to/custom/config.toml
Key Types
| Type | Module | Description |
|---|---|---|
run_cascade() |
cascade |
Main entry point — runs a named cascade |
AppConfig |
config |
Top-level configuration loaded from TOML |
ProviderConfig |
config |
Per-provider settings (type, API key, base URL) |
CascadeConfig |
config |
Ordered list of provider/model entries |
CascadeEntry |
config |
Single provider/model pair in a cascade |
Conversation |
models |
Multi-turn conversation with optional tools |
Message |
models |
Single message (role + content) |
MessageRole |
models |
Enum: system, user, assistant, tool |
LlmResponse |
models |
Provider response with content blocks and usage |
ContentBlock |
models |
Discriminated union: text or tool_call |
ToolDefinition |
models |
Tool/function schema definition |
ProviderError |
error |
Provider-level failure (HTTP, parse, missing key) |
CascadeError |
error |
All-providers-exhausted failure |
Cooldown & Backoff
| Failure # | Cooldown Duration | Capped At |
|---|---|---|
| 1st | 60s | — |
| 2nd | 120s | — |
| 3rd | 240s | — |
| 4th | 480s | — |
| 5th | 960s | — |
| 6th+ | 1920s+ | 3600s (1h) |
Special cases:
- HTTP 429 with
Retry-After: uses the header value directly instead of exponential backoff - Cooldown check: the entry is silently skipped if still on cooldown
- Cooldown expiry: cooldowns are checked against real-time; expired entries are automatically available
License
MIT — see LICENSE.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file llm_pycascade-0.1.0.tar.gz.
File metadata
- Download URL: llm_pycascade-0.1.0.tar.gz
- Upload date:
- Size: 81.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
743093412538ff0a0f8b7f37ef41ac260fcaafeb05a4c2d3970fbee581570cac
|
|
| MD5 |
d8f150a94ba08a1ed8eb1beaf58eb39e
|
|
| BLAKE2b-256 |
38e33f672e9fc5a29121542c05bf22f98ec088ba658a383d7e65cf324bad6744
|
Provenance
The following attestation bundles were made for llm_pycascade-0.1.0.tar.gz:
Publisher:
python-publish.yml on paluigi/llm-pycascade
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
llm_pycascade-0.1.0.tar.gz -
Subject digest:
743093412538ff0a0f8b7f37ef41ac260fcaafeb05a4c2d3970fbee581570cac - Sigstore transparency entry: 2170082738
- Sigstore integration time:
-
Permalink:
paluigi/llm-pycascade@892564225783beae708435bd736a4e8779ecb647 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/paluigi
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@892564225783beae708435bd736a4e8779ecb647 -
Trigger Event:
release
-
Statement type:
File details
Details for the file llm_pycascade-0.1.0-py3-none-any.whl.
File metadata
- Download URL: llm_pycascade-0.1.0-py3-none-any.whl
- Upload date:
- Size: 28.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c2beaa247da172b43bbdd49987f60146eb95214398d3819fe3925e6d36f6859f
|
|
| MD5 |
93cc1198ca1a443f022a8fc2e171d117
|
|
| BLAKE2b-256 |
f9adce94e31b049ca4d883590aa91c6502a6645fd38c0c3c481721a973f80ad1
|
Provenance
The following attestation bundles were made for llm_pycascade-0.1.0-py3-none-any.whl:
Publisher:
python-publish.yml on paluigi/llm-pycascade
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
llm_pycascade-0.1.0-py3-none-any.whl -
Subject digest:
c2beaa247da172b43bbdd49987f60146eb95214398d3819fe3925e6d36f6859f - Sigstore transparency entry: 2170082830
- Sigstore integration time:
-
Permalink:
paluigi/llm-pycascade@892564225783beae708435bd736a4e8779ecb647 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/paluigi
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@892564225783beae708435bd736a4e8779ecb647 -
Trigger Event:
release
-
Statement type: