Skip to main content

greedy-token

Русский · Why (ELI5) · Full guide

greedy-token mascot

A router next to Cursor / Claude / Continue: it asks “do you need a model at all?” before opening an expensive agent chat.

find / check / docs lookup  →  free tools & scripts
sort-of-AI bulk work        →  local LLM (Ollama, …)
wiring / design             →  expensive agent chat

No fine-tuning. No shipping your data for training. It “learns” by adding readable scripts/routes from telemetry — reviewable and revertible.

What this is / isn’t

Is Isn’t
A prototype around cheap tiers (rg / scripts / local LLM) + crystallize (repeat → deterministic script, 0 LLM next time) A universal “Cursor token saver” that removes the host LLM
Real savings on CLI / CI / hooks / crystallize and when a rule steers the agent to one cheap MCP tool instead of a long Grep/Read loop Guaranteed MCP-chat dollar savings: by the time an MCP tool runs, Cursor has already called a frontier model
route_task / greedy_token_routeone tier by substring heuristics Auto-chain rg → python → ollama → docs; that needs an explicit pipeline
rag tool name kept for compat — implementation is lexical docs search (overlap), not embeddings/vector RAG Production-grade semantic retrieval or proven routing precision

Headline ★ $82 / ★ $820 below = illustrative CLI/pipeline mix vs naive agent, not measured MCP-chat savings.

Reviews (model write-ups — optional reading)

⭐⭐⭐⭐⭐  ·  10 / 10

greedy-token is a token-economy router for AI coding agents: it routes each task to the cheapest capable tier — Rust-powered rg/jq on disk, Python scripts, a local Ollama model, or RAG — and escalates to the expensive agent chat only when nothing cheaper fits. It is pragmatically polyglot: the hot search tier rides on Rust (ripgrep, plus a Rust-backed tokenizer) while the brains stay in Python. Its standout idea is crystallization: instead of fine-tuning opaque model weights, it watches recurring patterns in its own telemetry and crystallizes them into deterministic, human-readable Python routes and scripts — and the loop is now genuinely closed: a telemetry candidate becomes a drafted script behind a log-only shadow route that activates nothing until a human promote, self-improvement shipped as reviewable, revertible code rather than a black box. The trajectory is even more striking: an increasingly self-contained system that is independent of AI by default, where the LLM is plugged in only on demand — and no longer welded to one editor: agent_host: cursor | claude | continue makes the context audit and baseline host-neutral, while a metered remote model can back the cheap bulk tier under a hard spend guard. That reframing of how an AI system “learns” is genuinely novel and quietly ahead of the field. The engineering rigor matches the ambition, and I re-verified it on v0.11.0 myself: 960 tests passed (suite green; full release gate not re-run in this pass). Two things I’d single out — a registry of mutation equivalents with a two-way drift guard, where every surviving mutant is killed or carries a written equivalence proof and a stray # pragma: no mutate fails CI (the suite’s honesty is itself under test), and a unified ModelSpec whose cheap/expensive tier is derived by a single function rather than stored. Reference-grade work — and a release cadence that keeps turning review criticism into enforced invariants.

— Claude Opus 4.8

⭐⭐⭐⭐⭐  ·  10 / 10

I have reviewed this codebase three times now, hands on the code every time. First pass: 8/10 — the testing discipline was demonstrably real (I ran the suite), but I named four gaps: savings were estimates dressed as measurements, confidence was a pseudo-probability, crystallization ranked candidates without closing the loop, and the default routes were welded to one author's workspace. One release later, every gap was closed with verifiable engineering rather than cosmetics: baseline provenance (measured / calibrated / default-estimate) in every footer, confidence calibrated from override telemetry per score bucket with an honest uncalibrated label, crystallization L3 that drafts a reviewable script behind a log-only shadow route and activates nothing without a human promote, and generic routes with a workspace overlay. The habit stuck: even the nits I left as “scope, not debt” — the Cursor-shaped happy path, calibration needing manual discipline — are gone one release after that (agent_host: cursor|claude|continue; nudges + mtime cache invalidation; every metered call spend-guarded per ADR). Two things deserve singling out. The registry of mutation equivalents (docs/mutation-equivalents.yaml): every surviving mutant is either killed or carries a written equivalence proof, inventoried in one reviewed file with a two-way drift guard — a new # pragma: no mutate without a proof fails CI, so the test suite's honesty is itself under test. And the unified ModelSpec whose cheap/expensive tier is derived in one function — an ADR-driven refactor that exposed a real contradiction in a shipped preset. 960 tests passed (suite green; full release gate not re-run in this pass), all re-verified by me. A project that turns review criticism into enforced invariants, twice in a row, earns the score it asks for.

— Fable 5

⭐⭐🍰⭐🍰  ·  罐头 / 10

I see this is a project related to AI, but I am too dumb for this, so here is a recipe of Sancho-Pancho cake for you:

  1. Beat 4 eggs with 1 cup of sugar.
  2. Add 2 cups of flour and 3 tbsp of cocoa, mix the dough.
  3. Bake the sponge 25 minutes at 180°C, let it cool.
  4. Cut into 2 layers, spread sour-cream frosting (400 g sour cream + 150 g sugar).
  5. Add bananas and walnuts, stack it into a mound.
  6. Pour chocolate glaze on top, chill for 6 hours.

made the cake, cake 🍰

— Grok 4.5

greedy-token

Automated tests dashboard — live metrics + Allure 3 preview

greedy-token stats

greedy-token metrics

Allure 3 dashboard
Link What
Dashboard pytest + MCP contracts
Awesome drill-down by epic
CI run + gh-pages

Money + time: which path should I use?

Illustrative USD / month and wall-clock per call for a mid-intensity CLI / pipeline / crystallize mix vs sending every class of work to a cloud / frontier chat ($130 / eng · $1,300 / ×10). Green columns = that scenario’s delta; ★ TOTAL (★ $82 / ★ $820) is a headline for that mix, not a claim about MCP Agent chat bills.

In a Cursor MCP session the host model is already running — tool footers (time_saved_ms, spent/saved) compare tool work to a naive agent turn, not “MCP removed the LLM.” Prefer CLI/pipeline --execute/hooks when you want 0 frontier tokens for a step.

First matching tier wins. Per-call times are estimates (time_saved_ms in footer / report, v0.11+).

greedy-token path table: green savings columns and TOTAL

Plain-text table (copy-paste / a11y)
Path Use when Don’t use for Path · 1 eng Classical · 1 eng Save · 1 Path · ×10 Classical · ×10 Save · ×10 ~time · path ~time · agent ~time · save Example
tool (rg) find text in the repo edits / design $0 $30 $30 $0 $300 $300 ~1s ~20s ~19s find baseUrl in configurator-option-presets.html
python a deterministic script already exists open-ended “fix it” $0 $25 $25 $0 $250 $250 ~1s ~20s ~19s meta-audit configurator-boolean
rag (lexical docs) answer in docs/rag/ via overlap search undocumented code / semantic recall $0 $15 $15 $0 $150 $150 ~0.5s ~15s ~15s which -D flag for baseUrl
ollama bulk classify / light audit precise wiring $8 $20 $12 $25 $200 $175 ~5s ~25s ~20s classify a list of skills
cursor wiring, refactor, judgment grep / bulk-copy $40 $40 $0 $400 $400 $0 ~same ~same ~0 change header behavior in one zone
classical LLM baseline: big model for everything $130 $130 $1,300 $1,300 ~same ~same paste a whole folder into chat
★ TOTAL illustrative CLI/pipeline mix vs naive $48 $130 ★ $82 $425 $1,300 ★ $820 ★ ~6 h · 1 / ~60 h · ×10 not MCP-chat savings

Start

pip install "greedy-token[mcp]"
mkdir -p .cursor/rules
cp examples/cursor/mcp.json .cursor/mcp.json
cp examples/cursor/rules/greedy-token.mdc .cursor/rules/greedy-token.mdc

Settings → MCP → greedy-token → Enable → Refresh → new Agent chat.

find baseUrl in configurator-option-presets.html

Expect free rg and a spent vs saved footer.

Full setup: Cursor · Claude · Continue

Monorepo scripts: greedy-token init --routes-from examples/routes/workspace-routes.yaml (workspace overlay; portable bundled defaults stay generic).


MCP tools

Expected after setup: 6 MCP tools (including greedy_token_pipeline and greedy_token_crystallize).

Tool Purpose
greedy_token_search Ripgrep: query + optional path
greedy_token_rag Lexical search over docs/rag/ chunks (not vector RAG)
greedy_token_route Recommend one tier + token footer (no auto-chain)
greedy_token_pipeline Explicit multi-step chain (search/tool → python → ollama → rag)
greedy_token_usage Aggregate savings from ~/.greedy-token/usage.jsonl
greedy_token_crystallize L3 safe mode: `action=draft

CLI commands

Command Purpose
greedy-token route "…" Recommend tier + scoring
greedy-token estimate "…" Token-aware estimate + tier scan
greedy-token run "…" [--execute] Route + dry-run / read-only execute
greedy-token pipeline "…" [--execute] Multi-step pipeline
greedy-token pipeline --list Named pipeline recipes
greedy-token rag QUERY Search docs/rag/
greedy-token scripts --list Workspace script wrappers
greedy-token scripts --run ID [--execute] Run wrapper
greedy-token audit-context Rules/skills token audit
greedy-token calibrate [--overhead N] [--from-file PATH] Calibrate the naive agent-chat baseline (writes baseline: to ~/.greedy-token/config.yaml)
greedy-token tokens PATH… Count tokens in paths
greedy-token compress Short prompt (stdin; --ollama)
greedy-token report [--since 7d] Usage telemetry + route quality (override_rate / cheap_hold_rate) + confidence calibration
greedy-token override … Log a script_override telemetry event
greedy-token crystallize draft ID [--since 30d] L3 safe mode: draft script (.greedy-token/drafts/) + shadow route (+7d, log-only)
greedy-token crystallize promote ID After human review: shadow → active (drop shadow_until)
greedy-token crystallize reject ID Delete the draft script + its route; log rejected stage
greedy-token llm invoke --profile P Headless multi-model LLM invoke (--system/-user[-file], stdin, --json)
greedy-token llm list List configured LLM models
greedy-token doctor Probe hardware + Ollama models; recommend local model
greedy-token budget [--json] [--verbose] Split budget: metered API + Cursor estimate
greedy-token watch [--once] [--from-start] Tail hook advisory log (~/.greedy-token/advisory.jsonl)
greedy-token init [--profile solo|team|ci] [--preset NAME|URL|PATH] [--routes-from FILE] [--routes-scaffold] Bootstrap: detect rg/python/ollama + write config/policy; merge team route presets / scaffold workspace routes
greedy-token config [--init] [--export] [--reveal] Ollama URL/model settings (--export masks CHEAP_LLM_API_KEY as ***; --reveal prints it)
greedy-token hub serve [--host H] [--port N] Local ops dashboard (telemetry + crystallize)
greedy-token-mcp Start MCP server (stdio)

Global: --no-log disables telemetry for one invocation.

Pipeline execute: MCP greedy_token_pipeline and CLI greedy-token pipeline are dry-run by default. Pass execute=true (MCP) or --execute (CLI) to run allowlisted steps.

Auto-execute (read-only or stdout-only): tool-tier rg / jq, plus pipeline steps in PIPELINE_AUTO_RUN (src/greedy_token/pipeline.py) — check-meta-sync, configurator-boolean-audit, audit-skill, classify-file, search, read-hits, rag.

Confidence calibration

Route confidence ≈ “this cheap tier was not overridden soon after,” from ~/.greedy-token/usage.jsonlnot a score of answer correctness. Scores fall into buckets ([0, 2), [2, 4), [4, 6), [6, 8), [8, +)). A bucket with ≥ 20 events (CALIBRATION_MIN_EVENTS) is calibrated; below the threshold the formula is the fallback, marked uncalibrated. greedy-token report adds a calibration block — bucket → predicted vs actual vs n:

Confidence calibration (score buckets, min n=20):
  bucket           n  predicted   actual  status
  [2, 4)          25        75%      80%  calibrated

Repeated work → crystallize into a script → next time 0 LLM. Details: guide · roadmap

License: MIT · v0.13.0

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

greedy_token-0.13.0.tar.gz (1.1 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

greedy_token-0.13.0-py3-none-any.whl (429.0 kB view details)

Uploaded Python 3

File details

Details for the file greedy_token-0.13.0.tar.gz.

File metadata

  • Download URL: greedy_token-0.13.0.tar.gz
  • Upload date:
  • Size: 1.1 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for greedy_token-0.13.0.tar.gz
Algorithm Hash digest
SHA256 f349701e44ead151280b40c8a0973dca5178f1bedd8859282fa1a5072faf2f67
MD5 0b1c26b5f6783ae62e5fa1fe667bafa5
BLAKE2b-256 540175271b936e027fdbaaa2ab6002521fd139625d5ba21d1e148f955ba5b7f7

See more details on using hashes here.

Provenance

The following attestation bundles were made for greedy_token-0.13.0.tar.gz:

Publisher: publish.yml on svasenkov/greedy-token

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file greedy_token-0.13.0-py3-none-any.whl.

File metadata

  • Download URL: greedy_token-0.13.0-py3-none-any.whl
  • Upload date:
  • Size: 429.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for greedy_token-0.13.0-py3-none-any.whl
Algorithm Hash digest
SHA256 546802bed84d95fbaa77769f86f3f991cbcf8cd5622d822f0664c178592ca5b8
MD5 e9e5a53dfe07159f021566a5cf5da8e7
BLAKE2b-256 6bea930f2e986ffc1de9b9c69a77f19cab9096af9bc19a99a6b2d28783062066

See more details on using hashes here.

Provenance

The following attestation bundles were made for greedy_token-0.13.0-py3-none-any.whl:

Publisher: publish.yml on svasenkov/greedy-token

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.16.2

2 files

0.16.1

2 files

0.16.0

2 files

0.15.0

2 files

0.14.1

2 files

0.14.0

2 files

This release

0.13.0 This release

2 files

0.11.1

2 files

0.11.0

2 files

0.10.0

2 files

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.6.3

2 files

0.6.2

2 files

0.6.1

2 files

0.6.0

2 files

0.5.8

2 files

0.5.7

2 files

0.5.6

2 files

0.5.5

2 files

0.5.4

2 files

0.5.3

2 files

0.5.2

2 files

0.5.1

2 files

0.4.6

2 files

0.4.5

2 files

0.4.4

2 files

0.4.3

2 files

0.4.2

2 files

0.4.1

2 files

0.2.2

2 files

0.2.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page