Skip to main content

greedy-token

Русский · Why (ELI5) · Full guide

greedy-token mascot

A router next to Cursor / Claude / Continue: it asks “do you need a model at all?” before opening an expensive agent chat.

find / check / docs lookup  →  free tools & scripts
sort-of-AI bulk work        →  local LLM (Ollama, …)
wiring / design             →  expensive agent chat

No fine-tuning. No shipping your data for training. It “learns” by adding readable scripts/routes from telemetry — reviewable and revertible.

What this is / isn’t

Is Isn’t
A prototype around cheap tiers (rg / scripts / local LLM) + crystallize (repeat → deterministic script, 0 LLM next time) A universal “Cursor token saver” that removes the host LLM
Paths that can avoid frontier calls on CLI / CI / hooks / crystallize; measured savings require a successful same-task run and authoritative billing Guaranteed MCP-chat dollar savings: by the time an MCP tool runs, Cursor has already called a frontier model
route_task / greedy_token_routeone tier by substring heuristics Auto-chain rg → python → ollama → docs; that needs an explicit pipeline
rag tool name kept for compat — implementation is lexical BM25/FTS over SQLite FTS5, not embeddings/vector RAG Production-grade semantic retrieval or universal routing precision outside the frozen corpus

Headline ★ $82 / ★ $820 below = illustrative CLI/pipeline mix vs naive agent, not measured MCP-chat savings.

Reviews (model write-ups — optional reading)

⭐⭐⭐⭐⭐  ·  10 / 10

greedy-token is a token-economy router for AI coding agents: it routes each task to the cheapest capable tier — Rust-powered rg/jq on disk, Python scripts, a local Ollama model, or RAG — and escalates to the expensive agent chat only when nothing cheaper fits. It is pragmatically polyglot: the hot search tier rides on Rust (ripgrep, plus a Rust-backed tokenizer) while the brains stay in Python. Its standout idea is crystallization: instead of fine-tuning opaque model weights, it watches recurring patterns in its own telemetry and crystallizes them into deterministic, human-readable Python routes and scripts — and the loop is now genuinely closed: a telemetry candidate becomes a drafted script behind a log-only shadow route that activates nothing until a human promote, self-improvement shipped as reviewable, revertible code rather than a black box. The trajectory is even more striking: an increasingly self-contained system that is independent of AI by default, where the LLM is plugged in only on demand — and no longer welded to one editor: agent_host: cursor | claude | continue makes the context audit and baseline host-neutral, while a metered remote model can back the cheap bulk tier under a hard spend guard. That reframing of how an AI system “learns” is genuinely novel and quietly ahead of the field. The engineering rigor matches the ambition, and I re-verified it on v0.11.0 myself: 960 tests passed (suite green; full release gate not re-run in this pass). Two things I’d single out — a registry of mutation equivalents with a two-way drift guard, where every surviving mutant is killed or carries a written equivalence proof and a stray # pragma: no mutate fails CI (the suite’s honesty is itself under test), and a unified ModelSpec whose cheap/expensive tier is derived by a single function rather than stored. Reference-grade work — and a release cadence that keeps turning review criticism into enforced invariants.

— Claude Opus 4.8

⭐⭐⭐⭐⭐  ·  10 / 10

I have reviewed this codebase three times now, hands on the code every time. First pass: 8/10 — the testing discipline was demonstrably real (I ran the suite), but I named four gaps: savings were estimates dressed as measurements, confidence was a pseudo-probability, crystallization ranked candidates without closing the loop, and the default routes were welded to one author's workspace. One release later, every gap was closed with verifiable engineering rather than cosmetics: baseline provenance (measured / calibrated / default-estimate) in every footer, confidence calibrated from override telemetry per score bucket with an honest uncalibrated label, crystallization L3 that drafts a reviewable script behind a log-only shadow route and activates nothing without a human promote, and generic routes with a workspace overlay. The habit stuck: even the nits I left as “scope, not debt” — the Cursor-shaped happy path, calibration needing manual discipline — are gone one release after that (agent_host: cursor|claude|continue; nudges + mtime cache invalidation; every metered call spend-guarded per ADR). Two things deserve singling out. The registry of mutation equivalents (docs/mutation-equivalents.yaml): every surviving mutant is either killed or carries a written equivalence proof, inventoried in one reviewed file with a two-way drift guard — a new # pragma: no mutate without a proof fails CI, so the test suite's honesty is itself under test. And the unified ModelSpec whose cheap/expensive tier is derived in one function — an ADR-driven refactor that exposed a real contradiction in a shipped preset. 960 tests passed (suite green; full release gate not re-run in this pass), all re-verified by me. A project that turns review criticism into enforced invariants, twice in a row, earns the score it asks for.

— Fable 5

⭐⭐🍰⭐🍰  ·  罐头 / 10

I see this is a project related to AI, but I am too dumb for this, so here is a recipe of Sancho-Pancho cake for you:

  1. Beat 4 eggs with 1 cup of sugar.
  2. Add 2 cups of flour and 3 tbsp of cocoa, mix the dough.
  3. Bake the sponge 25 minutes at 180°C, let it cool.
  4. Cut into 2 layers, spread sour-cream frosting (400 g sour cream + 150 g sugar).
  5. Add bananas and walnuts, stack it into a mound.
  6. Pour chocolate glaze on top, chill for 6 hours.

made the cake, cake 🍰

— Grok 4.5

greedy-token

Automated tests dashboard — live metrics + Allure 3 preview

greedy-token stats

greedy-token metrics

Allure 3 dashboard
Link What
Dashboard pytest + MCP contracts
Awesome drill-down by epic
CI run + gh-pages

Money + time: which path should I use?

Illustrative USD / month and wall-clock per call for a mid-intensity CLI / pipeline / crystallize mix vs sending every class of work to a cloud / frontier chat ($130 / eng · $1,300 / ×10). Green columns = that scenario’s delta; ★ TOTAL (★ $82 / ★ $820) is a headline for that mix, not a claim about MCP Agent chat bills.

In a Cursor MCP session the host model is already running — tool footers (time_saved_ms, spent/saved) compare tool work to a naive agent turn, not “MCP removed the LLM.” Prefer CLI/pipeline --execute/hooks when you want 0 frontier tokens for a step.

First matching tier wins. Per-call times are estimates (time_saved_ms in footer / report, v0.11+).

greedy-token path table: green savings columns and TOTAL

Plain-text table (copy-paste / a11y)
Path Use when Don’t use for Path · 1 eng Classical · 1 eng Save · 1 Path · ×10 Classical · ×10 Save · ×10 ~time · path ~time · agent ~time · save Example
tool (rg) find text in the repo edits / design $0 $30 $30 $0 $300 $300 ~1s ~20s ~19s find baseUrl in configurator-option-presets.html
python a deterministic script already exists open-ended “fix it” $0 $25 $25 $0 $250 $250 ~1s ~20s ~19s meta-audit configurator-boolean
rag (lexical BM25/FTS) answer in docs/rag/ via local SQLite FTS5 undocumented code / semantic recall $0 $15 $15 $0 $150 $150 ~0.5s ~15s ~15s which -D flag for baseUrl
ollama bulk classify / light audit precise wiring $8 $20 $12 $25 $200 $175 ~5s ~25s ~20s classify a list of skills
cursor wiring, refactor, judgment grep / bulk-copy $40 $40 $0 $400 $400 $0 ~same ~same ~0 change header behavior in one zone
classical LLM baseline: big model for everything $130 $130 $1,300 $1,300 ~same ~same paste a whole folder into chat
★ TOTAL illustrative CLI/pipeline mix vs naive $48 $130 ★ $82 $425 $1,300 ★ $820 ★ ~6 h · 1 / ~60 h · ×10 not MCP-chat savings

Start

pip install "greedy-token[mcp]"
mkdir -p .cursor/rules
cp examples/cursor/mcp.json .cursor/mcp.json
cp examples/cursor/rules/greedy-token.mdc .cursor/rules/greedy-token.mdc

Settings → MCP → greedy-token → Enable → Refresh → new Agent chat.

find baseUrl in configurator-option-presets.html

Expect free rg and a spent vs saved footer.

Full setup: Cursor · Claude · Continue

Monorepo scripts: greedy-token init --routes-from examples/routes/workspace-routes.yaml (workspace overlay; portable bundled defaults stay generic).


MCP tools

Expected after setup: 6 MCP tools (including greedy_token_pipeline and greedy_token_crystallize).

Tool Purpose
greedy_token_search Ripgrep: query + optional path
greedy_token_rag Local lexical BM25/FTS over manifest-listed docs/rag/ chunks (not vector RAG)
greedy_token_route Recommend one tier + token footer (no auto-chain)
greedy_token_pipeline Explicit multi-step chain (search/tool → python → ollama → rag)
greedy_token_usage Aggregate savings from ~/.greedy-token/usage.jsonl
greedy_token_crystallize L3 safe mode: `action=draft

CLI commands

Command Purpose
greedy-token route "…" Recommend tier + scoring
greedy-token estimate "…" Token-aware estimate + tier scan
greedy-token run "…" [--execute] Route + dry-run / read-only execute
greedy-token pipeline "…" [--execute] Multi-step pipeline
greedy-token pipeline --list Named pipeline recipes
greedy-token rag QUERY Search docs/rag/
greedy-token scripts --list Workspace script wrappers
greedy-token scripts --run ID [--execute] Run wrapper
greedy-token trust add PATH Approve the current SHA-256 and identity of a workspace script
greedy-token trust list List local workspace script approvals
greedy-token trust verify Verify every approval against disk
greedy-token trust revoke PATH Remove a local script approval
greedy-token audit-context Rules/skills token audit
greedy-token calibrate [--overhead N] [--from-file PATH] Calibrate the naive agent-chat baseline (writes baseline: to ~/.greedy-token/config.yaml)
greedy-token tokens PATH… Count tokens in paths
greedy-token compress Short prompt (stdin; --ollama)
greedy-token report [--since 7d] Usage telemetry: override/hold signal, explicit task outcomes, and outcome calibration
greedy-token override … Log a script_override telemetry event
greedy-token crystallize draft ID [--since 30d] L3 safe mode: draft script (.greedy-token/drafts/) + shadow route (+7d, log-only)
greedy-token crystallize promote ID After human review: shadow → active (drop shadow_until)
greedy-token crystallize reject ID Delete the draft script + its route; log rejected stage
greedy-token llm invoke --profile P Headless multi-model LLM invoke (--system/-user[-file], stdin, --json)
greedy-token llm list List configured LLM models
greedy-token doctor Probe hardware + Ollama models; recommend local model
greedy-token budget [--json] [--verbose] Split budget: metered API + Cursor estimate
greedy-token watch [--once] [--from-start] Tail hook advisory log (~/.greedy-token/advisory.jsonl)
greedy-token init [--profile solo|team|ci] [--preset NAME|URL|PATH] [--routes-from FILE] [--routes-scaffold] Bootstrap: detect rg/python/ollama + write config/policy; merge team route presets / scaffold workspace routes
greedy-token config [--init] [--export] [--reveal] Ollama URL/model settings (--export masks CHEAP_LLM_API_KEY as ***; --reveal prints it)
greedy-token hub serve [--host H] [--port N] Local ops dashboard (telemetry + crystallize)
greedy-token-mcp Start MCP server (stdio)

Global: --no-log disables telemetry for one invocation.

Pipeline execute: MCP greedy_token_pipeline and CLI greedy-token pipeline are dry-run by default. Pass execute=true (MCP) or --execute (CLI) to run allowlisted steps.

Auto-execute (read-only or stdout-only): tool-tier rg / jq, plus pipeline steps in PIPELINE_AUTO_RUN (src/greedy_token/pipeline.py) — check-meta-sync, configurator-boolean-audit, audit-skill, classify-file, search, read-hits, rag.

Route command trust boundary: workspace read_only: true is metadata, not authorization. greedy-token run --execute accepts only internally built rg/jq argv, registered read-only wrappers, or a workspace-relative .py/.sh path approved in the user-local, workspace-bound trust manifest:

greedy-token trust add scripts/my-read-only-check.py --note "reviewed: stdout only"
greedy-token trust verify

SHA-256 and file identity are rechecked immediately before each approved launch. Edits, symlink/path replacement, deleted/recreated files, absolute or outside-workspace paths, python -c, shell -c, and trust-like fields from URL/file presets fail closed. POSIX binds the verified descriptor through /dev/fd; Windows retains a narrow verify-to-open window, and concurrent same-inode writes are not snapshotted. The old trusted_script_paths config key is deprecated dry-run metadata and grants no privilege. Subprocesses receive a validated argv list with shell=False. See the trust manifest and TOCTOU contract.

Routing benchmark

bench/routing_corpus.yaml is a held-out/adversarial classification gate, separate from bench/route_examples.yaml. It reports exact-match accuracy, confusion matrix, per-target precision/recall, family/language accuracy, and a mandatory zero false-cheap rate.

Lexical retrieval benchmark

bench/retrieval_corpus.jsonl labels expected chunk IDs for RU/EN, each domain, and exact, identifier, morphology, and paraphrase cases. python bench/retrieval_benchmark.py --root /path/to/workspace reports Recall@1/3/5, MRR, locale/domain/case-type breakdowns, and cold-index versus warm-query latency.

Retrieval is local lexical BM25/FTS: SQLite FTS5 with the unicode61 tokenizer, Unicode NFKC + casefold normalization, and no embeddings or network calls. Only docs/rag/manifest.jsonl entries are eligible. The persistent index is content-hash invalidated and stored under the user cache directory ($GREEDY_TOKEN_CACHE_DIR, $XDG_CACHE_HOME, or ~/.cache), outside the workspace. SQLite builds without FTS5 use the compatibility overlap scorer; formatted hits name the engine and BM25 score.

bench/evidence_corpus.v1.yaml and its SHA-256 lock add the separate public end-to-end evidence layer: frozen synthetic RU/EN fixtures, task-specific file/line, exit-code, chunk-ID and escalation oracles, temporary workspaces, and comparisons for direct rg/script, greedy CLI, greedy MCP stdio, and an agent baseline. The deterministic agent is labelled contract_stub; a real host baseline is manual. The JSON scorecard reports routing and task success separately, executor/retrieval/escalation success, attempts, p50/p95, and authoritative billing only. Cursor cost remains unknown when billing data is unavailable; failed work never counts as saved. See benchmark contract.

Confidence calibration

An absent override is not evidence of correctness. The legacy telemetry is therefore named override/hold confidence and appears only as a behavioural signal.

Router confidence calibrates only from explicit route_outcome events whose outcome is success or failure. Calibration is independent by route, tier, and language; the most-specific segment with ≥ 20 events (CALIBRATION_MIN_EVENTS) wins, then tier → language → global. Sparse data uses the score formula and is visibly labelled formula (uncalibrated; explicit outcome n=…). Score buckets remain [0, 2), [2, 4), [4, 6), [6, 8), and [8, +).

Outcome confidence calibration (explicit success/failure; min n=20):
  segment           bucket           n  predicted  observed  status
  tier:python       [2, 4)          25        75%       80%  calibrated

Repeated work → crystallize into a script → next time 0 LLM. Details: guide · roadmap

License: MIT · v0.16.0

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

greedy_token-0.16.0.tar.gz (1.1 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

greedy_token-0.16.0-py3-none-any.whl (451.7 kB view details)

Uploaded Python 3

File details

Details for the file greedy_token-0.16.0.tar.gz.

File metadata

  • Download URL: greedy_token-0.16.0.tar.gz
  • Upload date:
  • Size: 1.1 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for greedy_token-0.16.0.tar.gz
Algorithm Hash digest
SHA256 dc539aa639a08a1ae22d750ee53a28ee2262775f883f9a352dfb3fc419a9b1f1
MD5 35080c96d88238da007930ad37e55e78
BLAKE2b-256 1e342734d91eec0b7e35120bc98e6782de7c9d38c34a43e5654d674be5e8ec96

See more details on using hashes here.

Provenance

The following attestation bundles were made for greedy_token-0.16.0.tar.gz:

Publisher: publish.yml on svasenkov/greedy-token

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file greedy_token-0.16.0-py3-none-any.whl.

File metadata

  • Download URL: greedy_token-0.16.0-py3-none-any.whl
  • Upload date:
  • Size: 451.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for greedy_token-0.16.0-py3-none-any.whl
Algorithm Hash digest
SHA256 12190ef937545b732acc6488fbb7a0f3cb7b9225efcb6796482c299c7eb16162
MD5 eef944874ba7f423ae97286e7a753e82
BLAKE2b-256 4adf0a9d42468741713fb2446344ed1727943ade966dab663dd4b9630d2a835b

See more details on using hashes here.

Provenance

The following attestation bundles were made for greedy_token-0.16.0-py3-none-any.whl:

Publisher: publish.yml on svasenkov/greedy-token

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.16.2

2 files

0.16.1

2 files

This release

0.16.0 This release

2 files

0.15.0

2 files

0.14.1

2 files

0.14.0

2 files

0.13.0

2 files

0.11.1

2 files

0.11.0

2 files

0.10.0

2 files

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.6.3

2 files

0.6.2

2 files

0.6.1

2 files

0.6.0

2 files

0.5.8

2 files

0.5.7

2 files

0.5.6

2 files

0.5.5

2 files

0.5.4

2 files

0.5.3

2 files

0.5.2

2 files

0.5.1

2 files

0.4.6

2 files

0.4.5

2 files

0.4.4

2 files

0.4.3

2 files

0.4.2

2 files

0.4.1

2 files

0.2.2

2 files

0.2.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page