Skip to main content

claude-spillway

PyPI Python License: MIT

日本語版 README はこちら

A quota-aware failover proxy for Claude Code.

claude-spillway sits between Claude Code and Anthropic's API. It forwards requests to Anthropic as normal, but watches the rate-limit headers on every response. When your quota (the rolling 5-hour / 7-day usage window on a Claude Pro/Max/Team subscription, or the classic request/token limits on an API key) runs low, it automatically "spills over" POST /v1/messages traffic to Ollama Cloud instead — and switches back to Anthropic once quota recovers.

It was built to solve a very specific problem: you're paying for both a Claude subscription and Ollama Cloud, and you'd rather burn through your (often cheaper, more relaxed) Ollama quota automatically instead of getting hard-blocked by Anthropic rate limits mid-session.

How it works

Claude Code --ANTHROPIC_BASE_URL--> claude-spillway --> Anthropic API
                                          |
                                          '--(quota low)--> Ollama Cloud
                                                             (Anthropic-
                                                              compatible
                                                              /v1/messages)
  • Every response from Anthropic carries rate-limit headers (anthropic-ratelimit-unified-5h-utilization, anthropic-ratelimit-unified-7d-utilization for subscription plans, or anthropic-ratelimit-{requests,tokens}-{limit,remaining} for API-key billing). claude-spillway parses these on every call — no extra API calls needed to know your usage.
  • When the worst remaining ratio across all known signals drops below fallback_threshold_pct (default 10%), subsequent POST /v1/messages calls are routed to Ollama Cloud instead, with the model name rewritten per your model_mapping config.
  • Requests to Anthropic don't need credentials configured in claude-spillway: whatever Authorization/x-api-key header Claude Code sends is forwarded as-is. Only Ollama Cloud needs an API key in the config (Ollama's Anthropic-compatible endpoint only accepts Authorization: Bearer, not x-api-key — see ollama/ollama#16922).
  • A background probe reads your quota from the OAuth usage endpoint — the one Claude Code itself reads for /usage — which consumes no quota and also reports when each window resets. Once the remaining ratio climbs back above recovery_threshold_pct (default 20%, intentionally higher than the fallback threshold to avoid flapping), traffic switches back to Anthropic. That endpoint is OAuth-only and is not part of the published API, so when it is unavailable the proxy falls back to the rate-limit headers of relayed traffic — and, only while in fallback where no traffic is flowing, to a minimal /v1/messages request that does cost a little quota.

Why only /v1/messages fails over

Only the actual inference call (POST /v1/messages) is subject to failover. Auxiliary endpoints such as /v1/messages/count_tokens or model listing always go to Anthropic, never to Ollama. This is deliberate: Ollama Cloud's Anthropic-compatibility shim is known to hang and even restart the server when it receives requests to endpoints it doesn't support (ollama/ollama#13949), which would be far worse than just missing a token count.

A note on Ollama Cloud quota visibility

Ollama Cloud still publishes no documented quota API (ollama/ollama#15663, #16448), and its inference responses carry no rate-limit headers at all. But its own dashboard reads an undocumented GET /api/usage, which takes the same API key as inference and is not itself counted as a request — so claude-spillway polls that for the session and weekly utilization plus a per-model request count.

Two caveats. It is undocumented, so a shape change degrades to "no data" rather than an error. And unlike Anthropic, Ollama reports no reset times, so nothing in this tool can tell you when an Ollama window turns over.

About Ollama's reset times (estimated)

No API reports them (see .github/TODO.md for the full investigation), so the monitor estimates instead: utilization cannot fall until a window resets, so the first observed rise marks the start of a fresh window, and "that rise + the window's maximum length" (5h for session, 7d for weekly) is the latest the reset could have been. The estimate is an upper bound, not the truth, so the TUI prefixes it with a ~.

Estimates are display-only. Routing decisions never read them - burn_rate_balance treats a window with an unknown reset as freshly started, which needs no estimate at all.

The status endpoint and TUI still also report self-tracked counters (requests relayed, failures, last status code) covering only the traffic that passed through this proxy.

Choosing a backend

With both sides' quota visible, routing.policy decides between them while neither is critical:

  • anthropic_first (default) — stay on Anthropic until it runs low, then fail over. The original behaviour.
  • weekly_balance — while both short windows are comfortable (balance_session_floor_pct), prefer whichever side has more of its weekly window left, switching only once the other side is ahead by balance_margin_pct so two near-equal backends don't oscillate.
  • burn_rate_balance — compare, per window, how much is left relative to how long the window still has to run (remaining / (time to reset / window length)), take the tightest window per backend, and prefer the side whose tightest reading is better. A reset time we don't know (Ollama never reports one) is assumed to be a window that just started — the most generous reading — so the unknown never manufactures urgency. anthropic_priority_weight (default 1.1) leans the comparison toward the Claude subscription you are already paying for; at parity it wins, and it needs to be ~9% worse before traffic moves to Ollama. Use it as a safety valve when one side is burning quota faster than its reset will allow.

Two guards apply under every policy:

  • Ollama exhaustion. Never fail over into an Ollama account that is itself nearly out (ollama_min_remaining_pct) — that trades one dead end for another.
  • Reverse failover. Ollama Cloud can get slow or fail outright when a model is busy. After ollama_failure_threshold consecutive failures, traffic goes back to Anthropic provided its 5-hour window still has reverse_failover_min_5h_pct left to serve with, and Ollama is left alone for reverse_failover_cooldown_seconds.

A hard guard outranks both: if Anthropic drops below quota.fallback_threshold_pct, traffic fails over regardless of policy.

Installation

Requires Python 3.11+ and uv.

uv tool install claude-spillway

Or run it without installing:

uvx --from claude-spillway claude-spillway serve

To work on it instead, clone the repository:

git clone https://github.com/akivajp/claude-spillway.git
cd claude-spillway
uv sync

Quick start

  1. Put the example config in the location serve reads by default, and set your Ollama Cloud API key (get one at https://ollama.com/settings/keys):

    mkdir -p ~/.config/claude-spillway
    curl -o ~/.config/claude-spillway/config.yaml \
      https://raw.githubusercontent.com/akivajp/claude-spillway/main/config.example.yaml
    export OLLAMA_API_KEY=your-ollama-cloud-api-key
    
  2. Start the proxy:

    claude-spillway serve
    
  3. Point Claude Code at it and launch as usual:

    export ANTHROPIC_BASE_URL=http://127.0.0.1:8787
    claude
    
  4. (Optional) In another terminal, watch quota status live:

    claude-spillway monitor
    

    claude-spillway never holds Anthropic credentials of its own — it borrows the one Claude Code sends. So monitor shows a "waiting" message until at least one real request has passed through. After that it keeps polling on its own and stays live even while you are idle, because reading the usage endpoint costs no quota.

You can try the whole flow without any real credentials using the bundled fake upstream servers — see scripts/manual_smoketest/.

Running on Windows with WSL2? See docs/wsl-windows.md for keeping the proxy resident as a systemd user service and wiring up both the Windows and the WSL Claude Code extension.

Configuration

See config.example.yaml for the full reference.

When -c is omitted, the config file is looked up in this order:

  1. the path in the CLAUDE_SPILLWAY_CONFIG environment variable
  2. ~/.config/claude-spillway/config.yaml (honours $XDG_CONFIG_HOME; on Windows, %APPDATA%\claude-spillway\config.yaml)

If neither exists, claude-spillway starts on its built-in defaults. Key fields:

Field Default Description
listen.host / listen.port 127.0.0.1 / 8787 Where claude-spillway listens
anthropic.base_url https://api.anthropic.com Anthropic API endpoint
ollama.base_url https://ollama.com Ollama Cloud endpoint
ollama.api_key Ollama Cloud API key (supports ${ENV_VAR})
quota.fallback_threshold_pct 10.0 Remaining % below which we fail over
quota.recovery_threshold_pct 20.0 Remaining % above which we switch back
quota.probe_interval_seconds 60.0 How often the background probe refreshes the quota reading
quota.use_usage_endpoint true Read quota from the OAuth usage endpoint (consumes no quota)
routing.policy anthropic_first anthropic_first, weekly_balance or burn_rate_balance
routing.anthropic_priority_weight 1.1 burn_rate_balance: how much to favour Anthropic
routing.ollama_min_remaining_pct 5.0 Never fail over once Ollama is this low
routing.ollama_failure_threshold 5 Consecutive Ollama failures that send traffic back
model_mapping.rules / model_mapping.default Anthropic model name -> Ollama model name

CLI flags on claude-spillway serve (--host, --port, --fallback-threshold-pct, --recovery-threshold-pct, --log-level) override the config file. Run claude-spillway --help / claude-spillway serve --help / claude-spillway monitor --help for details.

Interface language

CLI help text and the monitor TUI are shown in English by default, and in Japanese when the environment locale asks for it (LC_ALL, LC_MESSAGES, LANG or LANGUAGE starting with ja). Set CLAUDE_SPILLWAY_LANG to override the detection explicitly:

CLAUDE_SPILLWAY_LANG=en claude-spillway monitor  # force English
CLAUDE_SPILLWAY_LANG=ja claude-spillway monitor  # force Japanese

Log records emitted through logging are always in English, so they stay easy to grep and to search for online.

Development

uv run pytest      # unit + integration tests (no network access needed)
uv run ruff check . # lint

Tests exercise the full request/response path (including the mode switch and hysteresis logic) against httpx.MockTransport, so they run offline and don't touch real Anthropic/Ollama quota.

Status

Early-stage, built for personal use and shared in case it's useful to others. Not affiliated with Anthropic or Ollama.

Support

If claude-spillway saves you from hitting a rate limit mid-session, you can buy me a coffee.

Buy Me a Coffee

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

claude_spillway-0.4.1.tar.gz (45.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

claude_spillway-0.4.1-py3-none-any.whl (54.0 kB view details)

Uploaded Python 3

File details

Details for the file claude_spillway-0.4.1.tar.gz.

File metadata

  • Download URL: claude_spillway-0.4.1.tar.gz
  • Upload date:
  • Size: 45.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for claude_spillway-0.4.1.tar.gz
Algorithm Hash digest
SHA256 4705676c522988399a922fd2603ecac1bab94ffeb9bdaffaa94fc61ebd0b606c
MD5 88ec1e15943d4d0e9a74042d035446f1
BLAKE2b-256 341f34e5f2ec3eb2c5a499f57e1b0ad854e064991c609b337bca48fd741a460f

See more details on using hashes here.

File details

Details for the file claude_spillway-0.4.1-py3-none-any.whl.

File metadata

  • Download URL: claude_spillway-0.4.1-py3-none-any.whl
  • Upload date:
  • Size: 54.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for claude_spillway-0.4.1-py3-none-any.whl
Algorithm Hash digest
SHA256 55ed5ca89cec7e21f900c75ae3bfdba7a912cb087faa8a0ec334ef4d77e916ad
MD5 ec9abe1e7c12c2800a0d46b2533c788d
BLAKE2b-256 b36fb4f8f5c71047bd5017c9f0cfb211d8c7927743991be91f738a6e550fe7dd

See more details on using hashes here.

Release history Release notifications | RSS feed

0.6.2

2 files

0.6.1

2 files

0.6.0

2 files

0.5.0

2 files

This release

0.4.1 This release

2 files

0.4.0

2 files

0.3.1

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page