Skip to main content

Frugon

Your LLM bill is leaking — see exactly where, on your machine.

Free, local, open-source LLM cost analyzer. Point Frugon at your LLM call logs and see — on your machine — how much you'd save by switching or routing models.

PyPI License: MIT CI Python Platforms

Your data never leaves your machine. Your keys go straight to your own providers. Nothing reaches us.

Frugon analyzing a log file and recommending a routing split

Install & run

# one-shot (no install)
uvx frugon analyze ./logs.jsonl

# permanent install
uv tool install frugon          # or: pipx install frugon / pip install frugon
frugon analyze ./logs.jsonl

# for --measure (optional): samples real prompts through your own provider keys
uv tool install 'frugon[measure]'   # or: pip install 'frugon[measure]'
frugon analyze ./logs.jsonl --measure

No logs yet? See Getting your logs below, or run frugon analyze --demo to see it work on a bundled sample.

Getting your logs

frugon reads JSONL files in the OpenAI request/response format. There are two ways to produce them.

Option A — frugon capture (proxy shim)

frugon capture is a local HTTP proxy that sits between your app and your provider. Every call is forwarded unchanged to your real provider and saved as one JSONL line.

# Start the shim (default port 8787, output file capture.jsonl)
frugon capture --out ./logs.jsonl

# Then point your app's base URL at the shim instead of api.openai.com:
OPENAI_BASE_URL=http://127.0.0.1:8787 your-app           # bash / zsh
$env:OPENAI_BASE_URL="http://127.0.0.1:8787"; your-app   # PowerShell (Windows)
# or in code: client = OpenAI(base_url="http://127.0.0.1:8787/v1")

Options: --port, --out, --upstream (override the forwarding target), --verbose (print one line per captured call to verify it's recording), --proxy (opt in to route upstream calls through a proxy — by default frugon ignores any ambient HTTP_PROXY / HTTPS_PROXY, so your API key never passes through a third-party proxy). The shim adds no latency overhead on localhost and makes no calls to any frugon endpoint.

Option B — write JSONL directly

If you already capture logs (e.g. via middleware or a provider SDK callback), write one JSON object per line with this shape:

{
  "model": "gpt-4-turbo",
  "request": {
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user",   "content": "Summarise this document: ..."}
    ]
  },
  "response": {
    "choices": [{"message": {"content": "Here is the summary: ..."}}]
  },
  "usage": {
    "prompt_tokens": 312,
    "completion_tokens": 84
  },
  "timestamp": "2024-11-01T14:22:01Z"
}

usage.prompt_tokens / usage.completion_tokens — preferred when present; frugon falls back to its own tokenizer when absent. timestamp is optional but enables frugon to project costs over a real observed span. model is required; everything else degrades gracefully.

5-minute path from install to first analysis

uv tool install frugon          # or: pipx install frugon / pip install frugon
frugon capture --out ./logs.jsonl &   # start the proxy in the background
# ... run your app, make some LLM calls ...
frugon analyze ./logs.jsonl     # see the cost breakdown and routing recommendation

What it does

  • Cost analysis — fully local, no LLM calls, no network. Tokenizers + pricing + arithmetic on your machine.
  • Quality visibility (--measure, optional) — samples your traffic through candidate models using your own API keys, sent directly to your own providers. Never to us. --measure needs pip install 'frugon[measure]' and a provider API key (OPENAI_API_KEY, etc.); calls go to your own provider, never to us. On --demo, sampling is pinned to a single OpenAI model so the try-out needs only OPENAI_API_KEY; on your own logs, --measure samples the actual recommendation.
  • Routing recommendation — "move these X% of calls to a cheaper model and save ~$Y/mo; keep the hard Z% where they are." Comes with an explicit quality caveat so you know what you're trading. Run frugon models to see the model names available for --candidates (optionally frugon models gpt-4o to filter by substring).
  • Share the result — add --report savings.html (or .md) to write a clean, shareable report you can drop into a PR, a Slack thread, or a budget review.
  • Fast on real logs — everything runs locally and is comfortable well past 100k records. The bundled ~56,100-call demo (frugon analyze --demo) prices in a few seconds. Very large logs (>200k records) may take a little longer; Frugon shows a live progress bar and a one-line heads-up so you can see it working. There's no hard limit.

Example output

$ frugon analyze --demo --candidates claude-sonnet-4-5,gpt-4.1,claude-haiku-4-5,gemini-2.5-flash,deepseek-v4-flash

┌─ frugon · cost analysis ────────────────────────────────────────────────────┐
│                                                                             │
│   Analyzed      56,100 calls  ·  baseline gpt-5.5 (your current model)      │
│   Current spend $549.46 / mo                                                │
│                                                                             │
│     Route  36,100 easy calls (64.4%)  →  deepseek-v4-flash   within         │
│   tolerance                                                                 │
│     Keep   10,000 hard calls (17.8%)  →  gpt-5.5                            │
│     Keep   10,000 already on deepseek-v4-flash (17.8%)   already optimal    │
│   — no action                                                               │
│                                                                             │
│   New spend     $343.91 / mo                                                │
│                                                                             │
│   SAVING        $205.55 / mo    ·    37.4% lower                            │
│                                                                             │
└─────────────────────────────────────────────────────────────────────────────┘
                                                                               
  Candidates considered                                                        
  claude-sonnet-4-5  $452.23 / mo  17.7% lower  Strong   considered            
  gpt-4.1            $405.89 / mo  26.1% lower  Capable  considered            
  claude-haiku-4-5   $377.82 / mo  31.2% lower  Capable  considered            
  gemini-2.5-flash   $356.35 / mo  35.1% lower  Strong   considered            
  deepseek-v4-flash  $343.91 / mo  37.4% lower  Strong   recommended           
  Each candidate is shown under the same quality-preserving split (easy calls  
  to the candidate, hard calls kept on baseline); the biggest saving is the    
  headline recommendation, and when savings tie at the precision shown the    
  higher quality tier wins. Run --measure --judge to score each candidate's    
  quality.                                                                     

  Accounting   36,100 routed + 10,000 kept (gpt-5.5) + 10,000 already on 
               cheaper deepseek-v4-flash  =  56,100 analyzed
  Upper bound  a full swap to deepseek-v4-flash saves ~98.1% — run with 
               --verbose for detail
  Quality tier gpt-5.5: Elite  →  deepseek-v4-flash: Strong   (LMArena)
  Prices       synced 2026-07-02
  Quality      synced 2026-07-02

⚠ Quality is not verified — 'within tolerance' is an offline estimate;
  run --measure to confirm it on your real outputs before you switch.

  Your data never leaves your machine. Your keys go to your own providers.
→ Route every call automatically and hold the savings:  https://frugon.rodiun.io

Recommendations use a curated set of current top models across providers, drawn
from OpenRouter usage rankings. Prices synced 2026-07-02 from the LiteLLM 
registry. Run `frugon update` for the full live roster.
This is bundled sample data — run `frugon analyze <your-logs>` for a 
recommendation on your own logs.

Your numbers depend on your logs and your locally synced pricing/quality data. Run frugon analyze --demo --candidates claude-sonnet-4-5,gpt-4.1,claude-haiku-4-5,gemini-2.5-flash,deepseek-v4-flash to see the same output on your machine.

Quality tiers for reasoning models reflect the model at its default/typical reasoning effort — effort changes how many tokens a call spends thinking, not its per-token rate, so it never affects the price shown above.

How it's different

A provider's billing dashboard tells you what you already spent, and a raw token counter prices a single call — Frugon prices your real logs against every model, locally, and tells you which calls to move and which to keep.

Realistic savings

Based on RouteLLM's published research (LMSYS):

Traffic mix Typical saving
General mixed workload 30 – 50%
Easy / repetitive (high MT-Bench similarity) up to ~85%
Hard reasoning / MMLU-heavy ~30%

Your actual number comes from your logs. Frugon never inflates — it shows what the math says for your data.

Limitations

Frugon volunteers its edges. A few things are worth stating plainly before you rely on a number.

analyze is an offline estimate. The cost analysis reads your logs, prices them, and applies an easy/hard split heuristic without making a single model call. It never sees a candidate model's actual output, so "within tolerance" is a projection, not a measurement. The quality tiers come from LMArena and the savings bands come from RouteLLM research; both are population priors drawn from public leaderboards and studies. They describe how models tend to compare across many users, not a guarantee for your specific prompts. To turn the estimate into a measurement on your own traffic, run --measure, and --measure --judge to score it.

Sampling can miss the tail. --measure and --judge grade a sample of your prompts, not every one. A rare or unusual case may never appear in the sample, and an average across the sample can hide a model that degrades badly on a small slice of hard inputs. Raise --samples to widen coverage, and re-run on a fresh sample to see whether the verdict holds.

A tie can hide a shared failure. The judge is pairwise. It sees your current model's answer and the candidate's answer, anonymised as A and B in a randomised order, and it defaults to a tie unless one answer is clearly and materially better. That design removes label and position bias, but it also means two answers that fail the same way score as a tie: the judge is asked which answer is better, not whether either is correct on its own. A tie tells you the candidate is no worse than your current model, not that either one is right. When you need the ground truth on a prompt, read the sampled outputs yourself; --measure prints them side by side.

One draw does not measure variance. Each sampled prompt is answered once, and Frugon sends no temperature setting, so your provider's default applies. A single answer cannot tell you how much a model's output varies from run to run, and a one-draw verdict can mislead on a model that is nondeterministic. Re-running --measure on a fresh sample is the manual way to check whether a verdict is stable.

Per-call analysis cannot see second-order cost. Frugon prices the calls in your logs. It cannot see the cost a weak answer creates downstream: a follow-up call it triggers, a retry, or a human stepping in to review it. Retry attribution and human-review time are not in the logs, so a model that looks cheaper per call is not automatically cheaper once those effects are counted. Sampling real outputs with --judge is the closest built-in check on whether a cheaper model actually holds quality on your prompts.

On the judge, and on local models. When the judge is one of the models it is scoring, its verdict is partly a self-assessment; Frugon detects this and prints a caution naming the affected models, and you can pass --judge-model to use an independent judge. A model served on your own machine works here too: it can be sampled as a candidate or named as the judge, keyless and local. A local model absent from the pricing table has its measurement cost flagged as unpriced rather than guessed.

Your number is whatever your own logs and your own evaluation say.

Is this you?

  • Agent builders — your GPT-4o agents are expensive; most easy hops don't need them.
  • AI dev teams — monthly LLM bill is real; routing pays for itself in days.
  • RAG & support — retrieval + rerank is cheap; the final answer call doesn't have to be Opus.
  • Data-ETL pipelines — batch extraction is 100% repeatable; mini models handle it fine.
  • Indie hackers — every dollar saved is a dollar of runway.

Keep the savings

This is a one-time snapshot. Want it to keep routing automatically and hold the savings? → frugon.rodiun.io

Star the repo if this saved you money.

Contributing

Bug reports and pull requests are welcome — see CONTRIBUTING.md. Frugon is deliberately small: six commands (analyze, capture, models, update, pricing, quality), three capabilities (cost analysis, quality visibility, routing recommendation). Gateways, live routing proxies, web UIs, and multi-tenant accounts are out of scope by design.


Built by Rodiun. MIT licensed.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

frugon-0.2.5.tar.gz (1.3 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

frugon-0.2.5-py3-none-any.whl (514.0 kB view details)

Uploaded Python 3

File details

Details for the file frugon-0.2.5.tar.gz.

File metadata

  • Download URL: frugon-0.2.5.tar.gz
  • Upload date:
  • Size: 1.3 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for frugon-0.2.5.tar.gz
Algorithm Hash digest
SHA256 f321781f1dabe3c889ff046c91c54fd77f350632abd194072498377a2a8f9bbb
MD5 61d4e38afa07ba424eee5859f642b6ae
BLAKE2b-256 0bfd40b944c4f8d4d53c57e3f6713619c60186834c36f97e8f9dc0315ada8896

See more details on using hashes here.

Provenance

The following attestation bundles were made for frugon-0.2.5.tar.gz:

Publisher: release.yml on Rodiun/frugon

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file frugon-0.2.5-py3-none-any.whl.

File metadata

  • Download URL: frugon-0.2.5-py3-none-any.whl
  • Upload date:
  • Size: 514.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for frugon-0.2.5-py3-none-any.whl
Algorithm Hash digest
SHA256 21f8a699be751b3c913c35442df176fe8c8620594209d535d3eed62faa6811df
MD5 07595d935d602dc48c1fedf795fb1c26
BLAKE2b-256 c937688a37ad66e530c9aae5d9304354a63631ea430ea7772030101ab3abc949

See more details on using hashes here.

Provenance

The following attestation bundles were made for frugon-0.2.5-py3-none-any.whl:

Publisher: release.yml on Rodiun/frugon

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page