claude-spillway
A quota-aware failover proxy for Claude Code.
claude-spillway sits between Claude Code and Anthropic's API. It forwards
requests to Anthropic as normal, but watches the rate-limit headers on every
response. When your quota (the rolling 5-hour / 7-day usage window on a
Claude Pro/Max/Team subscription, or the classic request/token limits on an
API key) runs low, it automatically "spills over" POST /v1/messages traffic
to Ollama Cloud instead — and switches back to
Anthropic once quota recovers.
It was built to solve a very specific problem: you're paying for both a Claude subscription and Ollama Cloud, and you'd rather burn through your (often cheaper, more relaxed) Ollama quota automatically instead of getting hard-blocked by Anthropic rate limits mid-session.
How it works
Claude Code --ANTHROPIC_BASE_URL--> claude-spillway --> Anthropic API
|
'--(quota low)--> Ollama Cloud
(Anthropic-
compatible
/v1/messages)
- Every response from Anthropic carries rate-limit headers
(
anthropic-ratelimit-unified-5h-utilization,anthropic-ratelimit-unified-7d-utilizationfor subscription plans, oranthropic-ratelimit-{requests,tokens}-{limit,remaining}for API-key billing). claude-spillway parses these on every call — no extra API calls needed to know your usage. - When the worst remaining ratio across all known signals drops below
fallback_threshold_pct(default 10%), subsequentPOST /v1/messagescalls are routed to Ollama Cloud instead, with the model name rewritten per yourmodel_mappingconfig. - Requests to Anthropic don't need credentials configured in
claude-spillway: whatever
Authorization/x-api-keyheader Claude Code sends is forwarded as-is. Only Ollama Cloud needs an API key in the config (Ollama's Anthropic-compatible endpoint only acceptsAuthorization: Bearer, notx-api-key— see ollama/ollama#16922). - A background probe reads your quota from the OAuth usage endpoint — the one
Claude Code itself reads for
/usage— which consumes no quota and also reports when each window resets. Once the remaining ratio climbs back aboverecovery_threshold_pct(default 20%, intentionally higher than the fallback threshold to avoid flapping), traffic switches back to Anthropic. That endpoint is OAuth-only and is not part of the published API, so when it is unavailable the proxy falls back to the rate-limit headers of relayed traffic — and, only while in fallback where no traffic is flowing, to a minimal/v1/messagesrequest that does cost a little quota.
Why only /v1/messages fails over
Only the actual inference call (POST /v1/messages) is subject to
failover. Auxiliary endpoints such as /v1/messages/count_tokens or model
listing always go to Anthropic, never to Ollama. This is deliberate: Ollama
Cloud's Anthropic-compatibility shim is known to hang and even restart the
server when it receives requests to endpoints it doesn't support
(ollama/ollama#13949),
which would be far worse than just missing a token count.
A note on Ollama Cloud quota visibility
Ollama Cloud still publishes no documented quota API
(ollama/ollama#15663,
#16448), and its inference
responses carry no rate-limit headers at all. But its own dashboard reads an
undocumented GET /api/usage, which takes the same API key as inference and is
not itself counted as a request — so claude-spillway polls that for the session
and weekly utilization plus a per-model request count.
Two caveats. It is undocumented, so a shape change degrades to "no data" rather than an error. And unlike Anthropic, Ollama reports no reset times, so nothing in this tool can tell you when an Ollama window turns over.
About Ollama's reset times (estimated)
No API reports them (see .github/TODO.md for the full
investigation), so the monitor estimates instead: utilization cannot fall until
a window resets, so the first observed rise marks the start of a fresh window,
and "that rise + the window's maximum length" (5h for session, 7d for weekly)
is the latest the reset could have been. The estimate is an upper bound, not
the truth, so the TUI prefixes it with a ~.
Estimates are display-only. Routing decisions never read them -
burn_rate_balance treats a window with an unknown reset as freshly started,
which needs no estimate at all.
The status endpoint and TUI still also report self-tracked counters (requests relayed, failures, last status code) covering only the traffic that passed through this proxy.
Choosing a backend
With both sides' quota visible, routing.policy decides between them while
neither is critical:
anthropic_first(default) — stay on Anthropic until it runs low, then fail over. The original behaviour.weekly_balance— while both short windows are comfortable (balance_session_floor_pct), prefer whichever side has more of its weekly window left, switching only once the other side is ahead bybalance_margin_pctso two near-equal backends don't oscillate.burn_rate_balance— compare, per window, how much is left relative to how long the window still has to run (remaining / (time to reset / window length)), take the tightest window per backend, and prefer the side whose tightest reading is better. A reset time we don't know (Ollama never reports one) is assumed to be a window that just started — the most generous reading — so the unknown never manufactures urgency.anthropic_priority_weight(default 1.1) leans the comparison toward the Claude subscription you are already paying for; at parity it wins, and it needs to be ~9% worse before traffic moves to Ollama. Use it as a safety valve when one side is burning quota faster than its reset will allow.
Two guards apply under every policy:
- Ollama exhaustion. Never fail over into an Ollama account that is itself
nearly out (
ollama_min_remaining_pct) — that trades one dead end for another. - Reverse failover. Ollama Cloud can get slow or fail outright when a model
is busy. After
ollama_failure_thresholdconsecutive failures, traffic goes back to Anthropic provided its 5-hour window still hasreverse_failover_min_5h_pctleft to serve with, and Ollama is left alone forreverse_failover_cooldown_seconds.
A hard guard outranks both: if Anthropic drops below
quota.fallback_threshold_pct, traffic fails over regardless of policy.
Installation
Requires Python 3.11+ and uv.
uv tool install claude-spillway
Or run it without installing:
uvx --from claude-spillway claude-spillway serve
To work on it instead, clone the repository:
git clone https://github.com/akivajp/claude-spillway.git
cd claude-spillway
uv sync
Quick start
-
Put the example config in the location
servereads by default, and set your Ollama Cloud API key (get one at https://ollama.com/settings/keys):mkdir -p ~/.config/claude-spillway curl -o ~/.config/claude-spillway/config.yaml \ https://raw.githubusercontent.com/akivajp/claude-spillway/main/config.example.yaml export OLLAMA_API_KEY=your-ollama-cloud-api-key
-
Start the proxy:
claude-spillway serve -
Point Claude Code at it and launch as usual:
export ANTHROPIC_BASE_URL=http://127.0.0.1:8787 claude
-
(Optional) In another terminal, watch quota status live:
claude-spillway monitorclaude-spillway never holds Anthropic credentials of its own — it borrows the one Claude Code sends. So
monitorshows a "waiting" message until at least one real request has passed through. After that it keeps polling on its own and stays live even while you are idle, because reading the usage endpoint costs no quota. -
(Optional) Or watch the same thing in a browser, at http://127.0.0.1:8787/_spillway/. Both
serveandmonitorprint the address, so there is nothing to memorise.The page is served by the proxy itself, so it needs no separate process and loads nothing from the network. It shows both backends side by side with their remaining quota, reset countdowns and the thresholds it switches on, and refreshes on its own. It turns red the moment the proxy stops answering, which is worth knowing quickly: while
ANTHROPIC_BASE_URLpoints here, Claude Code cannot reach Anthropic either.
You can try the whole flow without any real credentials using the bundled
fake upstream servers — see
scripts/manual_smoketest/.
To keep it running in the background instead of starting it by hand, install it as a systemd user service — one script does the whole thing:
./scripts/install-service.sh
See docs/service.md for what it sets up, and how to upgrade or remove it.
Running on Windows with WSL2? See docs/wsl-windows.md for wiring up both the Windows-side and the WSL-side Claude Code extension.
Configuration
See config.example.yaml for the full reference.
When -c is omitted, the config file is looked up in this order:
- the path in the
CLAUDE_SPILLWAY_CONFIGenvironment variable ~/.config/claude-spillway/config.yaml(honours$XDG_CONFIG_HOME; on Windows,%APPDATA%\claude-spillway\config.yaml)
If neither exists, claude-spillway starts on its built-in defaults. Key fields:
| Field | Default | Description |
|---|---|---|
listen.host / listen.port |
127.0.0.1 / 8787 |
Where claude-spillway listens |
anthropic.base_url |
https://api.anthropic.com |
Anthropic API endpoint |
ollama.base_url |
https://ollama.com |
Ollama Cloud endpoint |
ollama.api_key |
— | Ollama Cloud API key (supports ${ENV_VAR}) |
quota.fallback_threshold_pct |
10.0 |
Remaining % below which we fail over |
quota.recovery_threshold_pct |
20.0 |
Remaining % above which we switch back |
quota.probe_interval_seconds |
60.0 |
How often the background probe refreshes the quota reading |
quota.use_usage_endpoint |
true |
Read quota from the OAuth usage endpoint (consumes no quota) |
routing.policy |
anthropic_first |
anthropic_first, weekly_balance or burn_rate_balance |
routing.anthropic_priority_weight |
1.1 |
burn_rate_balance: how much to favour Anthropic |
routing.ollama_min_remaining_pct |
5.0 |
Never fail over once Ollama is this low |
routing.ollama_failure_threshold |
5 |
Consecutive Ollama failures that send traffic back |
model_mapping.rules / model_mapping.default |
— | Anthropic model name -> Ollama model name |
CLI flags on claude-spillway serve (--host, --port,
--fallback-threshold-pct, --recovery-threshold-pct, --log-level)
override the config file. Run claude-spillway --help /
claude-spillway serve --help / claude-spillway monitor --help for
details.
Interface language
CLI help text and the monitor TUI are shown in English by default, and in
Japanese when the environment locale asks for it (LC_ALL, LC_MESSAGES,
LANG or LANGUAGE starting with ja). Set CLAUDE_SPILLWAY_LANG to
override the detection explicitly:
CLAUDE_SPILLWAY_LANG=en claude-spillway monitor # force English
CLAUDE_SPILLWAY_LANG=ja claude-spillway monitor # force Japanese
Log records emitted through logging are always in English, so they stay
easy to grep and to search for online.
Development
uv run pytest # unit + integration tests (no network access needed)
uv run ruff check . # lint
Tests exercise the full request/response path (including the mode switch
and hysteresis logic) against httpx.MockTransport, so they run offline and
don't touch real Anthropic/Ollama quota.
Status
Early-stage, built for personal use and shared in case it's useful to others. Not affiliated with Anthropic or Ollama.
Support
If claude-spillway saves you from hitting a rate limit mid-session, you can buy me a coffee.
License
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file claude_spillway-0.5.0.tar.gz.
File metadata
- Download URL: claude_spillway-0.5.0.tar.gz
- Upload date:
- Size: 57.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2b824a4b175aff0197a79dbba885e7ee7ee3dda76fad5673c50fa3c151f4fd8a
|
|
| MD5 |
a9f2331124fa4229b1107681cb8a8823
|
|
| BLAKE2b-256 |
62b51493555fb95a85c5b68d187dab32c5ad29454abf44d96e4c978b352f3624
|
File details
Details for the file claude_spillway-0.5.0-py3-none-any.whl.
File metadata
- Download URL: claude_spillway-0.5.0-py3-none-any.whl
- Upload date:
- Size: 67.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
96cd789be7ef63d0b499969362b4149896b78c9aaccb63e28b75147a472ed18c
|
|
| MD5 |
05e28bc1eac8c65dd7016160524995bd
|
|
| BLAKE2b-256 |
7979b0235d00013b5cc7ab2352244407a8ace0aa4ef9952d32fd7c502ec353f6
|