Skip to main content

freewhirr

Spend nothing until you choose to.

freewhirr is a free-first router for OpenRouter. It walks an ordered list of $0 models, fails over when one of them rate-limits, errors, times out, or blows the context window, and only then — if you opt in — touches a paid model under budgets you set yourself.

Named after Whir of Invention: find the right artifact, stay inside the mana value.

Use it as:

  • a Python library (sync and async)
  • an OpenAI-compatible proxy so apps in any language can point base_url at it
  • a small CLI
pip install freewhirr

Why

OpenRouter publishes a rotating catalog of free models. They are generous, and they are also flaky: 429s, provider outages, tiny context windows. Paying humans usually hard-code one cheap paid model and leak money on work a free model would have done.

freewhirr flips that default.

Policy What happens when free models are exhausted
stop Raise a clear error. Never spend money.
paid Escalate to an ordered paid list, only inside the guardrails you configured. Once a budget is hit, behave like stop.

30-second quickstart

python -m pip install freewhirr
export OPENROUTER_API_KEY=sk-or-v1-...
from freewhirr import Freewhirr

with Freewhirr() as client:
    result = client.complete(
        messages=[{"role": "user", "content": "Explain vector clocks in one paragraph."}]
    )
    print(result.model)
    print(result.content)

Same thing from the CLI:

freewhirr chat "Explain vector clocks in one paragraph."

The API key is read from OPENROUTER_API_KEY. There is no other vendor lock-in: settings also come from a YAML or TOML file, environment variables, or constructor kwargs.

Exhaustion policies

stop — never spend money

# examples/stop.yaml
exhaustion_policy: stop
include_discovered_free: true
from freewhirr import Freewhirr, FreeModelsExhausted, load_config

client = Freewhirr(load_config("examples/stop.yaml"))
try:
    client.complete(messages=[{"role": "user", "content": "hi"}])
except FreeModelsExhausted as exc:
    print("free tier exhausted; nothing was billed")
    print(exc.attempts)

paid — escalate with guardrails

# examples/paid.yaml
exhaustion_policy: paid
include_discovered_free: true
paid_models:
  - openai/gpt-4o-mini
max_paid_per_request: 0.05
max_paid_per_day: 1.00
max_paid_per_month: 10.00
max_paid_share: 0.1          # at most 10% of successful requests
paid_only_high_complexity: true

Paid fallback is skipped — and the call fails like stop — when any of these are true:

  • the daily or monthly spend ledger would go over the cap
  • the model's estimated worst-case cost exceeds max_paid_per_request
  • another paid success would push the paid share above max_paid_share
  • paid_only_high_complexity is on and the request was not tagged high
result = client.complete(
    messages=[{"role": "user", "content": "Draft the migration RFC."}],
    complexity="high",
)

Ledgers live on disk (SQLite by default, or a JSON file). They are local to the machine that ran the request. Nothing is phoned home.

OpenAI- and Anthropic-compatible proxy

Any language that can speak the OpenAI HTTP API — or the Anthropic Messages API — can sit behind freewhirr.

freewhirr --config examples/stop.yaml serve --port 8787
# examples/proxy_openai.py
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8787/v1", api_key="not-used")
response = client.chat.completions.create(
    model="freewhirr",
    messages=[{"role": "user", "content": "Give me a one-line haiku about buses."}],
)
print(response.choices[0].message.content)

The proxy:

  • accepts POST /v1/chat/completions (including SSE streaming)
  • accepts POST /v1/responses (OpenAI Responses API, including Responses SSE) so current Codex CLI (wire_api = "responses") works
  • accepts POST /v1/messages in Anthropic's format (tool_use / tool_result, system prompts, stop reasons) and translates to OpenAI-style upstream
  • streams Anthropic SSE (message_start, content_block_*, message_delta, message_stop)
  • estimates POST /v1/messages/count_tokens from the prompt (heuristic, not a billed tokenizer)
  • lists the current free catalog at GET /v1/models
  • honors X-Freewhirr-Complexity: high for the paid-complexity gate
  • ignores the client's model field unless you set honor_requested_model: true (so a hardcoded gpt-4o-mini in an app cannot accidentally skip the free list)

Claude Code's ANTHROPIC_BASE_URL is the gateway root without /v1 (http://127.0.0.1:8787). OpenAI-compatible clients include /v1.

Use with…

Agent harnesses that speak OpenAI Chat Completions or Anthropic Messages can point at the proxy. Most of them require tool-calling-capable models; freewhirr skips catalog entries whose OpenRouter supported_parameters omit tools / tool_choice / functions (see require_tool_support below).

Start the proxy first:

export OPENROUTER_API_KEY=sk-or-v1-...
freewhirr --config examples/stop.yaml serve --port 8787

Claude Code

Set the gateway root (no /v1) and a credential. Claude Code then calls /v1/messages and optional /v1/messages/count_tokens.

export ANTHROPIC_BASE_URL=http://127.0.0.1:8787
export ANTHROPIC_API_KEY=not-used-by-proxy
claude

Or persist in ~/.claude/settings.json:

{
  "env": {
    "ANTHROPIC_BASE_URL": "http://127.0.0.1:8787",
    "ANTHROPIC_API_KEY": "not-used-by-proxy"
  }
}

Use ANTHROPIC_AUTH_TOKEN instead if the client should send a bearer token. Needs models that can call tools. Settings-file env wins over the shell.

Source: Connect Claude Code to an LLM gateway, gateway protocol.

OpenCode

Current docs use opencode.json / opencode.jsonc with a custom OpenAI-compatible provider. baseURL includes /v1.

{
  "$schema": "https://opencode.ai/config.json",
  "model": "freewhirr/auto",
  "providers": {
    "freewhirr": {
      "package": "@opencode/ai/providers/openai-compatible",
      "settings": { "baseURL": "http://127.0.0.1:8787/v1" },
      "models": {
        "auto": { "modelID": "freewhirr" }
      }
    }
  }
}

To send the built-in Anthropic provider through the Messages API instead, override providers.anthropic.settings.baseURL to http://127.0.0.1:8787/v1 (OpenCode appends /messages). Older docs still show a provider.<id>.options.baseURL shape with npm: "@ai-sdk/openai-compatible". Agent use needs tools.

Source: OpenCode providers (v2), classic providers.

Codex (OpenAI Codex CLI)

Current Codex speaks only the Responses API. wire_api accepts only responses (it is also the default). Put the provider in user-level ~/.codex/config.toml — project .codex/config.toml cannot set model_provider / model_providers.

model = "freewhirr"
model_provider = "freewhirr"

[model_providers.freewhirr]
name = "freewhirr"
base_url = "http://127.0.0.1:8787/v1"
env_key = "OPENROUTER_API_KEY"
wire_api = "responses"

base_url is the OpenAI root including /v1; Codex posts to {base_url}/responses. The proxy key is unused for auth — env_key just satisfies Codex. Do not set wire_api = "chat" (Codex will refuse to start). Do not name the provider openai, ollama, or lmstudio (reserved). Needs a tool-calling-capable free model for agent use.

Source: Codex config reference, advanced configuration.

Cursor

Cursor Settings → Models: enable OpenAI API Key, turn on Override OpenAI Base URL, and add a custom model id such as freewhirr.

OpenAI API Key:          not-used-by-proxy
Override OpenAI Base URL: https://<public-https-host>/v1

Caveat: Cursor's backend, not the editor, calls the base URL. localhost and LAN addresses resolve on Cursor's servers and will not reach your machine. Expose the proxy with a public HTTPS tunnel (ngrok, Cloudflare Tunnel, …). Agent mode needs tool-calling models. There is no first-party docs.cursor.com BYOK page; this matches Cursor staff on the official forum.

Source: Cursor forum — local LLM.

Pi

Pi (badlogic/pi-mono) loads custom endpoints from ~/.pi/agent/models.json. Use openai-completions for this proxy.

{
  "providers": {
    "freewhirr": {
      "baseUrl": "http://127.0.0.1:8787/v1",
      "api": "openai-completions",
      "apiKey": "not-used-by-proxy",
      "models": [{ "id": "freewhirr" }]
    }
  }
}

$NAME / ${NAME} interpolation works for apiKey. Opening /model reloads the file. Prefer a tool-capable model for the coding agent.

Source: Pi — Choose a Model.

Goose

Goose's OpenAI provider is the documented path for OpenAI-compatible proxies. OPENAI_HOST is the root with no path; OPENAI_BASE_PATH is the chat completions suffix. Goose requires tool calling for anything beyond plain chat (disable all extensions if you must use a no-tools model).

export OPENAI_API_KEY=not-used-by-proxy
export OPENAI_HOST=http://127.0.0.1:8787
export OPENAI_BASE_PATH=v1/chat/completions
goose session

Or goose configure and pick OpenAI. A 404 usually means OPENAI_BASE_PATH is wrong. Keys placed only in config.yaml are ignored.

Source: Goose providers.

Berd

TODO. Berd is Block's desktop UI over a Goose ACP sidecar. The public README covers setup, bundling, and enterprise distribution seams. It does not document a Berd-specific custom OpenAI or Anthropic base URL, env var, or settings field. Do not invent one. If you control the bundled Goose sidecar, configure that sidecar using the Goose section above.

Searched: block/berd README, repository docs tree, and public web results for "Berd custom OpenAI base URL" — no official end-user provider page.

Omnigent

Run omni setup and add a Gateway (base URL + key), or write ~/.omnigent/config.yaml. Use the OpenAI-compatible /v1 URL for OpenAI-style harnesses.

# ~/.omnigent/config.yaml
providers:
  freewhirr:
    kind: gateway
    default: true
    openai:
      base_url: http://127.0.0.1:8787/v1
      api_key: not-used-by-proxy
      wire_api: chat
      models:
        default: freewhirr

Omnigent can also launch Claude Code / Codex against a gateway. Claude Code wants the Anthropic root without /v1; Codex should use wire_api: responses against http://127.0.0.1:8787/v1. Agent harnesses need tools.

Source: Omnigent models, harness configuration.

Hermes

Hermes Agent (Nous Research) talks to any OpenAI-compatible /v1/chat/completions endpoint. Interactive: hermes model → "Custom endpoint". Or edit ~/.hermes/config.yaml (model.base_url, not the legacy model.api_base):

# ~/.hermes/config.yaml
model:
  default: freewhirr
  provider: custom
  base_url: http://127.0.0.1:8787/v1
  api_key: not-used-by-proxy

OPENAI_BASE_URL only overrides the openai-api provider, not custom. Tool-calling models recommended.

Source: Hermes providers, environment variables.

Cost-savings stats

Every request is logged: models tried, outcome, latency, tokens, and actual cost from OpenRouter's usage.cost when present (otherwise an estimate from the catalog). Compare that to an all-paid baseline (baseline_paid_model, default openai/gpt-4o-mini):

$ freewhirr stats
freewhirr stats
  Requests:          142
  Free successes:    131 (94.2%)
  Paid successes:    8 (5.8%)
  Failures:          3
  Spent:             $0.0410
  All-paid baseline: $1.8800
  Estimated saved:   $1.8390

Estimated saved is baseline − spent. It is an accounting aid, not an invoice.

Privacy

Most free OpenRouter endpoints may train on prompts. If that is unacceptable, set:

deny_data_collection: true

or FREEWHIRR_DENY_DATA_COLLECTION=true.

freewhirr then sends OpenRouter's provider preference {"data_collection": "deny"} on every completion. OpenRouter will only route to providers that do not collect prompts for training.

Stealth/cloaked models and any catalog entry whose description says prompts are logged or used for training are also removed from the rotation when this flag is on. freewhirr watch and freewhirr models --stealth still list them, marked excluded.

This will shrink — sometimes to zero — the set of free endpoints. OpenRouter will return an error that no endpoint matches your data policy. That is the tradeoff working as designed. stop will then fail closed; paid may still escalate to a paid model that honors the same preference.

Stealth and new free models

OpenRouter sometimes drops stealth (also called cloaked) models: anonymous preview endpoints, usually $0, so a lab can gather feedback. They are free because prompts and outputs are logged and may be used for training. See OpenRouter's Stealth Program EULA.

Do not send secrets, source you cannot leak, or personal data through a stealth model. promote_stealth is off by default for that reason.

Live catalog snapshot (2026-10-10, GET https://openrouter.ai/api/v1/models): 458 models, 19 free (total_count matches). Listings expose id, name, canonical_slug, description, context_length, created, pricing, supported_parameters, architecture, plus top_provider, hugging_face_id, knowledge_cutoff, expiration_date, reasoning, and default_parameters. No current listing used the words stealth, cloaked, or logging. The openrouter/ namespace was routers (auto, auto-beta, free, fusion, pareto-code, bodybuilder), not stealth drops. Two named free models now advertise a 256000 context (cohere/north-mini-code:free, nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free), so that length is only a weak signal. Historical stealth models such as openrouter/horizon-alpha and openrouter/horizon-beta used the openrouter/ namespace, a codename, 256000 context, and copy that called the model cloaked and warned that prompts are logged. Heuristics target that shape so the next drop is caught.

freewhirr models --new        # $0 models first seen since the last snapshot
freewhirr models --stealth    # scored as stealth/cloaked
freewhirr watch --once        # one scan (also the default poller)
freewhirr watch --interval 60

watch persists first-seen timestamps in the local ledger, diffs against the last catalog, and optionally POSTs JSON to FREEWHIRR_ALERT_WEBHOOK.

promote_stealth: true          # try detected stealth models first
alert_webhook: https://example.invalid/hook
watch_interval: 300

When deny_data_collection is on, stealth models and any model whose description mentions logging or training stay visible to watch / --stealth but are excluded from chat, the proxy, and freewhirr models (the rotation list).

See OpenRouter's provider routing docs for data_collection and the related zdr (zero data retention) flag.

How routing works

  1. Load free models from your list, OpenRouter's /models catalog (pricing 0), or both. The catalog is cached in memory for model_cache_ttl seconds (default one hour).
  2. Try each free model in order. Retry the same model with jittered exponential backoff on 429s, 5xx, and timeouts. Fail over immediately on context-length errors, most 4xx, and optional quality-check failures. When the request includes tools / functions and require_tool_support is on, skip models whose OpenRouter metadata omits tool calling. Unknown catalog entries are still tried.
  3. If a quality check is set (JSON Schema, minimum length, tool-call validation, or any callable) and it rejects the response, treat that like a failover and try the next model. Malformed tool calls (invalid JSON arguments or an unknown tool name) fail the same way when validate_tool_calls is on. Quality checks are skipped for streaming.
  4. When the free list is exhausted, apply the exhaustion policy.
  5. Streaming failovers only happen before the first token. After tokens have been sent to the caller, a mid-stream error is surfaced rather than silently switching models.

Configuration

Lookup order: defaults < YAML/TOML file < environment variables < constructor kwargs.

Files, first match wins:

  • FREEWHIRR_CONFIG
  • ./freewhirr.yaml, ./freewhirr.yml, ./freewhirr.toml
  • $XDG_CONFIG_HOME/freewhirr/ (same filenames)
Setting Env var Default
API key OPENROUTER_API_KEY (required)
Base URL OPENROUTER_BASE_URL https://openrouter.ai/api/v1
Policy FREEWHIRR_EXHAUSTION_POLICY stop
Free model pin list FREEWHIRR_FREE_MODELS (discover)
Paid model list FREEWHIRR_PAID_MODELS []
Merge discovered free models FREEWHIRR_INCLUDE_DISCOVERED true
Max $ / request FREEWHIRR_MAX_PAID_PER_REQUEST unset
Max $ / day FREEWHIRR_MAX_PAID_PER_DAY unset
Max $ / month FREEWHIRR_MAX_PAID_PER_MONTH unset
Max paid share (0–1) FREEWHIRR_MAX_PAID_SHARE unset
Paid only if complexity=high FREEWHIRR_PAID_ONLY_HIGH_COMPLEXITY false
data_collection: deny FREEWHIRR_DENY_DATA_COLLECTION false
Catalog TTL seconds FREEWHIRR_MODEL_CACHE_TTL 3600
HTTP timeout FREEWHIRR_TIMEOUT 60
Extra retries / model FREEWHIRR_MAX_RETRIES 2
Ledger path FREEWHIRR_STATE_PATH ~/.freewhirr/state.sqlite
Ledger backend FREEWHIRR_STATE_BACKEND sqlite (json also works)
Savings baseline model FREEWHIRR_BASELINE_MODEL openai/gpt-4o-mini
Skip models that cannot call tools FREEWHIRR_REQUIRE_TOOL_SUPPORT true
Fail over on bad tool calls FREEWHIRR_VALIDATE_TOOL_CALLS true
Promote stealth models to the front FREEWHIRR_PROMOTE_STEALTH false
Watch poll interval (seconds) FREEWHIRR_WATCH_INTERVAL 300
Watch / stealth webhook FREEWHIRR_ALERT_WEBHOOK unset
freewhirr config          # resolved settings, key redacted
freewhirr models          # free list the next call would walk

Library surface

from freewhirr import (
    Freewhirr,
    FreewhirrConfig,
    ExhaustionPolicy,
    MinLengthCheck,
    JsonSchemaCheck,
    ToolCallCheck,
)

client = Freewhirr(
    FreewhirrConfig(
        api_key="...",
        exhaustion_policy=ExhaustionPolicy.STOP,
    )
)

result = client.complete(messages=[...], quality=MinLengthCheck(40))
async_result = await client.acomplete(messages=[...])

for chunk in client.stream(messages=[...]):
    print(chunk["choices"][0]["delta"].get("content") or "", end="")

# OpenAI-shaped namespace
client.chat.completions.create(messages=[...])

ChatResult exposes .content, .model, .cost_usd, .usage, .models_tried, .paid, and .to_openai().

FAQ

Does this make live OpenRouter calls in CI? No. Tests mock HTTP with respx. You can run pytest without a key.

Will deny_data_collection break free models? Often, yes. Most free endpoints collect prompts. The flag is an explicit privacy/availability tradeoff, not a silent default.

Can I pin a model from the OpenAI SDK? By default the proxy ignores model so existing clients cannot skip the free list. Set honor_requested_model: true if you want a requested free model tried first.

How is "money saved" computed? Each success stores the actual (or estimated) cost and a baseline cost for the same token counts on baseline_paid_model. freewhirr stats subtracts those. If you never configured a paid baseline in the catalog, a GPT-4o-mini price is used so the number is still in the right order of magnitude.

What happens if the catalog request fails? If you supplied free_models, those are still used. If you rely on discovery alone, the call fails with a clear error — it does not fall through to paid unless you listed paid models and the policy is paid.

Does streaming support failover? Yes, until the first token is forwarded. After that, switching models would corrupt the stream, so the error is returned.

Where should I store the API key? Environment variable or a gitignored local config file. Never commit it.

Development

See CONTRIBUTING.md. Releases use PyPI Trusted Publishing; see RELEASING.md. Short version:

python -m pip install -e ".[dev]"
ruff check src tests examples
pytest

License

MIT. See LICENSE.

Metadata

Release files for freewhirr 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for freewhirr 0.1.0
File Size Uploaded
freewhirr-0.1.0.tar.gz 65.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for freewhirr 0.1.0
File Interpreter ABI Platform
freewhirr-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 117.2 kB

Release files / freewhirr-0.1.0.tar.gz

Download URL freewhirr-0.1.0.tar.gz
Size 65.2 kB
Tags Source
SHA-256 checksum
How to use checksums
c2d06be9080c8c4fe99d63ef1d3867d670cd4495f46e5b49432a89c1d5c247f5
BLAKE2b-256 checksum
How to use checksums
0d48f6c2f47aebdfc73bd8850cb26934f32f37632d47e84aced1397d400ea19d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release files / freewhirr-0.1.0-py3-none-any.whl

Download URL freewhirr-0.1.0-py3-none-any.whl
Size 52.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3c541ae90cae7ebc631b9c30ca5d5b9e6bae9e17ebaa7b29afc2977479b21994
BLAKE2b-256 checksum
How to use checksums
8561a7d8fc7a3ec6c311a4fd6e856ea8e7a2bd70ca9ed6682a7bb4a2389a1378
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page