greedy-token
Русский · Why (ELI5) · Full guide
A router next to Cursor / Claude / Continue: it asks “do you need a model at all?” before opening an expensive agent chat.
find / check / docs lookup → free tools & scripts
sort-of-AI bulk work → local LLM (Ollama, …)
wiring / design → expensive agent chat
No fine-tuning. No shipping your data for training. It “learns” by adding readable scripts/routes from telemetry — reviewable and revertible.
What this is / isn’t
| Is | Isn’t |
|---|---|
| A prototype around cheap tiers (rg / scripts / local LLM) + crystallize (repeat → deterministic script, 0 LLM next time) | A universal “Cursor token saver” that removes the host LLM |
| Real savings on CLI / CI / hooks / crystallize and when a rule steers the agent to one cheap MCP tool instead of a long Grep/Read loop | Guaranteed MCP-chat dollar savings: by the time an MCP tool runs, Cursor has already called a frontier model |
route_task / greedy_token_route → one tier by substring heuristics |
Auto-chain rg → python → ollama → docs; that needs an explicit pipeline |
rag tool name kept for compat — implementation is lexical docs search (overlap), not embeddings/vector RAG |
Production-grade semantic retrieval or proven routing precision |
Headline ★ $82 / ★ $820 below = illustrative CLI/pipeline mix vs naive agent, not measured MCP-chat savings.
Reviews (model write-ups — optional reading)
⭐⭐⭐⭐⭐ · 10 / 10greedy-token is a token-economy router for AI coding agents: it routes each task to the cheapest capable tier — Rust-powered — Claude Opus 4.8 |
⭐⭐⭐⭐⭐ · 10 / 10I have reviewed this codebase three times now, hands on the code every time. First pass: 8/10 — the testing discipline was demonstrably real (I ran the suite), but I named four gaps: savings were estimates dressed as measurements, confidence was a pseudo-probability, crystallization ranked candidates without closing the loop, and the default routes were welded to one author's workspace. One release later, every gap was closed with verifiable engineering rather than cosmetics: baseline provenance ( — Fable 5 |
⭐⭐🍰⭐🍰 ·
|
Automated tests dashboard — live metrics + Allure 3 preview
| Link | What |
|---|---|
| Dashboard | pytest + MCP contracts |
| Awesome | drill-down by epic |
| CI | run + gh-pages |
Money + time: which path should I use?
Illustrative USD / month and wall-clock per call for a mid-intensity CLI / pipeline / crystallize mix vs sending every class of work to a cloud / frontier chat ($130 / eng · $1,300 / ×10). Green columns = that scenario’s delta; ★ TOTAL (★ $82 / ★ $820) is a headline for that mix, not a claim about MCP Agent chat bills.
In a Cursor MCP session the host model is already running — tool footers (time_saved_ms, spent/saved) compare tool work to a naive agent turn, not “MCP removed the LLM.” Prefer CLI/pipeline --execute/hooks when you want 0 frontier tokens for a step.
First matching tier wins. Per-call times are estimates (time_saved_ms in footer / report, v0.11+).
Plain-text table (copy-paste / a11y)
| Path | Use when | Don’t use for | Path · 1 eng | Classical · 1 eng | Save · 1 | Path · ×10 | Classical · ×10 | Save · ×10 | ~time · path | ~time · agent | ~time · save | Example |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| tool (rg) | find text in the repo | edits / design | $0 | $30 | $30 | $0 | $300 | $300 | ~1s | ~20s | ~19s | find baseUrl in configurator-option-presets.html |
| python | a deterministic script already exists | open-ended “fix it” | $0 | $25 | $25 | $0 | $250 | $250 | ~1s | ~20s | ~19s | meta-audit configurator-boolean |
| rag (lexical docs) | answer in docs/rag/ via overlap search |
undocumented code / semantic recall | $0 | $15 | $15 | $0 | $150 | $150 | ~0.5s | ~15s | ~15s | which -D flag for baseUrl |
| ollama | bulk classify / light audit | precise wiring | $8 | $20 | $12 | $25 | $200 | $175 | ~5s | ~25s | ~20s | classify a list of skills |
| cursor | wiring, refactor, judgment | grep / bulk-copy | $40 | $40 | $0 | $400 | $400 | $0 | ~same | ~same | ~0 | change header behavior in one zone |
| classical LLM | baseline: big model for everything | — | $130 | $130 | — | $1,300 | $1,300 | — | ~same | ~same | — | paste a whole folder into chat |
| ★ TOTAL | illustrative CLI/pipeline mix vs naive | — | $48 | $130 | ★ $82 | $425 | $1,300 | ★ $820 | — | — | ★ ~6 h · 1 / ~60 h · ×10 | not MCP-chat savings |
Start
pip install "greedy-token[mcp]"
mkdir -p .cursor/rules
cp examples/cursor/mcp.json .cursor/mcp.json
cp examples/cursor/rules/greedy-token.mdc .cursor/rules/greedy-token.mdc
Settings → MCP → greedy-token → Enable → Refresh → new Agent chat.
find baseUrl in configurator-option-presets.html
Expect free rg and a spent vs saved footer.
Full setup: Cursor · Claude · Continue
Monorepo scripts: greedy-token init --routes-from examples/routes/workspace-routes.yaml (workspace overlay; portable bundled defaults stay generic).
MCP tools
Expected after setup: 6 MCP tools (including greedy_token_pipeline and greedy_token_crystallize).
| Tool | Purpose |
|---|---|
greedy_token_search |
Ripgrep: query + optional path |
greedy_token_rag |
Lexical search over docs/rag/ chunks (not vector RAG) |
greedy_token_route |
Recommend one tier + token footer (no auto-chain) |
greedy_token_pipeline |
Explicit multi-step chain (search/tool → python → ollama → rag) |
greedy_token_usage |
Aggregate savings from ~/.greedy-token/usage.jsonl |
greedy_token_crystallize |
L3 safe mode: `action=draft |
CLI commands
| Command | Purpose |
|---|---|
greedy-token route "…" |
Recommend tier + scoring |
greedy-token estimate "…" |
Token-aware estimate + tier scan |
greedy-token run "…" [--execute] |
Route + dry-run / read-only execute |
greedy-token pipeline "…" [--execute] |
Multi-step pipeline |
greedy-token pipeline --list |
Named pipeline recipes |
greedy-token rag QUERY |
Search docs/rag/ |
greedy-token scripts --list |
Workspace script wrappers |
greedy-token scripts --run ID [--execute] |
Run wrapper |
greedy-token audit-context |
Rules/skills token audit |
greedy-token calibrate [--overhead N] [--from-file PATH] |
Calibrate the naive agent-chat baseline (writes baseline: to ~/.greedy-token/config.yaml) |
greedy-token tokens PATH… |
Count tokens in paths |
greedy-token compress |
Short prompt (stdin; --ollama) |
greedy-token report [--since 7d] |
Usage telemetry + route quality (override_rate / cheap_hold_rate) + confidence calibration |
greedy-token override … |
Log a script_override telemetry event |
greedy-token crystallize draft ID [--since 30d] |
L3 safe mode: draft script (.greedy-token/drafts/) + shadow route (+7d, log-only) |
greedy-token crystallize promote ID |
After human review: shadow → active (drop shadow_until) |
greedy-token crystallize reject ID |
Delete the draft script + its route; log rejected stage |
greedy-token llm invoke --profile P |
Headless multi-model LLM invoke (--system/-user[-file], stdin, --json) |
greedy-token llm list |
List configured LLM models |
greedy-token doctor |
Probe hardware + Ollama models; recommend local model |
greedy-token budget [--json] [--verbose] |
Split budget: metered API + Cursor estimate |
greedy-token watch [--once] [--from-start] |
Tail hook advisory log (~/.greedy-token/advisory.jsonl) |
greedy-token init [--profile solo|team|ci] [--preset NAME|URL|PATH] [--routes-from FILE] [--routes-scaffold] |
Bootstrap: detect rg/python/ollama + write config/policy; merge team route presets / scaffold workspace routes |
greedy-token config [--init] [--export] [--reveal] |
Ollama URL/model settings (--export masks CHEAP_LLM_API_KEY as ***; --reveal prints it) |
greedy-token hub serve [--host H] [--port N] |
Local ops dashboard (telemetry + crystallize) |
greedy-token-mcp |
Start MCP server (stdio) |
Global: --no-log disables telemetry for one invocation.
Pipeline execute: MCP greedy_token_pipeline and CLI greedy-token pipeline are dry-run by default. Pass execute=true (MCP) or --execute (CLI) to run allowlisted steps.
Auto-execute (read-only or stdout-only): tool-tier rg / jq, plus pipeline steps in PIPELINE_AUTO_RUN (src/greedy_token/pipeline.py) — check-meta-sync, configurator-boolean-audit, audit-skill, classify-file, search, read-hits, rag.
Route command trust boundary: workspace read_only: true is metadata, not
authorization. greedy-token run --execute accepts only internally built
rg/jq argv, registered read-only wrappers, or a workspace-relative
.py/.sh path explicitly listed in the local .greedy-token.yaml:
trusted_script_paths:
- scripts/my-read-only-check.py
Arbitrary/absolute executables, python -c, shell -c, paths outside the
workspace, and untrusted commands imported from URL/file presets remain
dry-run only. Subprocesses receive a validated argv list with shell=False.
Routing benchmark
bench/routing_corpus.yaml is a held-out/adversarial classification gate,
separate from bench/route_examples.yaml. It reports exact-match accuracy
(equivalently micro-precision for this single-label corpus), a confusion
matrix, per-target precision/recall, per-family and per-language accuracy, and
a mandatory zero false-cheap rate. These are routing classification
metrics—not end-to-end execution success or lexical retrieval quality, which
must be evaluated separately.
Confidence calibration
Route confidence ≈ “this cheap tier was not overridden soon after,” from ~/.greedy-token/usage.jsonl — not a score of answer correctness. Scores fall into buckets ([0, 2), [2, 4), [4, 6), [6, 8), [8, +)). A bucket with ≥ 20 events (CALIBRATION_MIN_EVENTS) is calibrated; below the threshold the formula is the fallback, marked uncalibrated. greedy-token report adds a calibration block — bucket → predicted vs actual vs n:
Confidence calibration (score buckets, min n=20):
bucket n predicted actual status
[2, 4) 25 75% 80% calibrated
Repeated work → crystallize into a script → next time 0 LLM. Details: guide · roadmap
License: MIT · v0.14.1
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file greedy_token-0.14.1.tar.gz.
File metadata
- Download URL: greedy_token-0.14.1.tar.gz
- Upload date:
- Size: 1.1 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
37d4ae91f86549643a161e5555a056b14869d92f527541bcbcfbe70109c34ef3
|
|
| MD5 |
e1d8d3efd8be3a714d7bc47c46d8ec15
|
|
| BLAKE2b-256 |
356da3f67db56e55b3b841d855ab01b596ae9f3666918293a3ba9127c10a648c
|
Provenance
The following attestation bundles were made for greedy_token-0.14.1.tar.gz:
Publisher:
publish.yml on svasenkov/greedy-token
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
greedy_token-0.14.1.tar.gz -
Subject digest:
37d4ae91f86549643a161e5555a056b14869d92f527541bcbcfbe70109c34ef3 - Sigstore transparency entry: 2293987267
- Sigstore integration time:
-
Permalink:
svasenkov/greedy-token@a4be875a4795920b1cdbedf31a8bc4dea6f16876 -
Branch / Tag:
refs/tags/v0.14.1 - Owner: https://github.com/svasenkov
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a4be875a4795920b1cdbedf31a8bc4dea6f16876 -
Trigger Event:
release
-
Statement type:
File details
Details for the file greedy_token-0.14.1-py3-none-any.whl.
File metadata
- Download URL: greedy_token-0.14.1-py3-none-any.whl
- Upload date:
- Size: 434.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1675dac43a4be16684468c6642d8577d88881b8f3f6fb6ef687241c14adae83d
|
|
| MD5 |
aa7fb7bc3e63c62b0c32df6f1fb914c2
|
|
| BLAKE2b-256 |
d88d7e531a49e97652ee9fb08d3cecd168538a77301b13324b74c7c971300018
|
Provenance
The following attestation bundles were made for greedy_token-0.14.1-py3-none-any.whl:
Publisher:
publish.yml on svasenkov/greedy-token
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
greedy_token-0.14.1-py3-none-any.whl -
Subject digest:
1675dac43a4be16684468c6642d8577d88881b8f3f6fb6ef687241c14adae83d - Sigstore transparency entry: 2293987341
- Sigstore integration time:
-
Permalink:
svasenkov/greedy-token@a4be875a4795920b1cdbedf31a8bc4dea6f16876 -
Branch / Tag:
refs/tags/v0.14.1 - Owner: https://github.com/svasenkov
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a4be875a4795920b1cdbedf31a8bc4dea6f16876 -
Trigger Event:
release
-
Statement type: