greedy-token
Русский · Why (ELI5) · Full guide
A router next to Cursor / Claude / Continue: it asks “do you need a model at all?” before opening an expensive agent chat.
find / check / docs lookup → free tools & scripts
sort-of-AI bulk work → local LLM (Ollama, …)
wiring / design → expensive agent chat
No fine-tuning. No shipping your data for training. It “learns” by adding readable scripts/routes from telemetry — reviewable and revertible.
What this is / isn’t
| Is | Isn’t |
|---|---|
| A prototype around cheap tiers (rg / scripts / local LLM) + crystallize (repeat → deterministic script, 0 LLM next time) | A universal “Cursor token saver” that removes the host LLM |
| Paths that can avoid frontier calls on CLI / CI / hooks / crystallize; measured savings require a successful same-task run and authoritative billing | Guaranteed MCP-chat dollar savings: by the time an MCP tool runs, Cursor has already called a frontier model |
route_task / greedy_token_route → one tier by substring heuristics |
Auto-chain rg → python → ollama → docs; that needs an explicit pipeline |
rag tool name kept for compat — implementation is lexical BM25/FTS over SQLite FTS5, not embeddings/vector RAG |
Production-grade semantic retrieval or universal routing precision outside the frozen corpus |
Headline ★ $82 / ★ $820 below = illustrative CLI/pipeline mix vs naive agent, not measured MCP-chat savings.
Reviews (model write-ups — optional reading)
⭐⭐⭐⭐⭐ · 10 / 10greedy-token is a token-economy router for AI coding agents: it routes each task to the cheapest capable tier — Rust-powered — Claude Opus 4.8 |
⭐⭐⭐⭐⭐ · 10 / 10I have reviewed this codebase three times now, hands on the code every time. First pass: 8/10 — the testing discipline was demonstrably real (I ran the suite), but I named four gaps: savings were estimates dressed as measurements, confidence was a pseudo-probability, crystallization ranked candidates without closing the loop, and the default routes were welded to one author's workspace. One release later, every gap was closed with verifiable engineering rather than cosmetics: baseline provenance ( — Fable 5 |
⭐⭐🍰⭐🍰 ·
|
Automated tests dashboard — live metrics + Allure 3 preview
| Link | What |
|---|---|
| Dashboard | pytest + MCP contracts |
| Awesome | drill-down by epic |
| CI | run + gh-pages |
Money + time: which path should I use?
Illustrative USD / month and wall-clock per call for a mid-intensity CLI / pipeline / crystallize mix vs sending every class of work to a cloud / frontier chat ($130 / eng · $1,300 / ×10). Green columns = that scenario’s delta; ★ TOTAL (★ $82 / ★ $820) is a headline for that mix, not a claim about MCP Agent chat bills.
In a Cursor MCP session the host model is already running — tool footers (time_saved_ms, spent/saved) compare tool work to a naive agent turn, not “MCP removed the LLM.” Prefer CLI/pipeline --execute/hooks when you want 0 frontier tokens for a step.
First matching tier wins. Per-call times are estimates (time_saved_ms in footer / report, v0.11+).
Plain-text table (copy-paste / a11y)
| Path | Use when | Don’t use for | Path · 1 eng | Classical · 1 eng | Save · 1 | Path · ×10 | Classical · ×10 | Save · ×10 | ~time · path | ~time · agent | ~time · save | Example |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| tool (rg) | find text in the repo | edits / design | $0 | $30 | $30 | $0 | $300 | $300 | ~1s | ~20s | ~19s | find baseUrl in configurator-option-presets.html |
| python | a deterministic script already exists | open-ended “fix it” | $0 | $25 | $25 | $0 | $250 | $250 | ~1s | ~20s | ~19s | meta-audit configurator-boolean |
| rag (lexical BM25/FTS) | answer in docs/rag/ via local SQLite FTS5 |
undocumented code / semantic recall | $0 | $15 | $15 | $0 | $150 | $150 | ~0.5s | ~15s | ~15s | which -D flag for baseUrl |
| ollama | bulk classify / light audit | precise wiring | $8 | $20 | $12 | $25 | $200 | $175 | ~5s | ~25s | ~20s | classify a list of skills |
| cursor | wiring, refactor, judgment | grep / bulk-copy | $40 | $40 | $0 | $400 | $400 | $0 | ~same | ~same | ~0 | change header behavior in one zone |
| classical LLM | baseline: big model for everything | — | $130 | $130 | — | $1,300 | $1,300 | — | ~same | ~same | — | paste a whole folder into chat |
| ★ TOTAL | illustrative CLI/pipeline mix vs naive | — | $48 | $130 | ★ $82 | $425 | $1,300 | ★ $820 | — | — | ★ ~6 h · 1 / ~60 h · ×10 | not MCP-chat savings |
Start
pip install "greedy-token[mcp]"
mkdir -p .cursor/rules
cp examples/cursor/mcp.json .cursor/mcp.json
cp examples/cursor/rules/greedy-token.mdc .cursor/rules/greedy-token.mdc
Settings → MCP → greedy-token → Enable → Refresh → new Agent chat.
find baseUrl in configurator-option-presets.html
Expect free rg and a spent vs saved footer.
Full setup: Cursor · Claude · Continue
Monorepo scripts: greedy-token init --routes-from examples/routes/workspace-routes.yaml (workspace overlay; portable bundled defaults stay generic).
MCP tools
Expected after setup: 6 MCP tools (including greedy_token_pipeline and greedy_token_crystallize).
| Tool | Purpose |
|---|---|
greedy_token_search |
Ripgrep: query + optional path |
greedy_token_rag |
Local lexical BM25/FTS over manifest-listed docs/rag/ chunks (not vector RAG) |
greedy_token_route |
Recommend one tier + token footer (no auto-chain) |
greedy_token_pipeline |
Explicit multi-step chain (search/tool → python → ollama → rag) |
greedy_token_usage |
Aggregate savings from ~/.greedy-token/usage.jsonl |
greedy_token_crystallize |
L3 safe mode: `action=draft |
CLI commands
| Command | Purpose |
|---|---|
greedy-token route "…" |
Recommend tier + scoring |
greedy-token estimate "…" |
Token-aware estimate + tier scan |
greedy-token run "…" [--execute] |
Route + dry-run / read-only execute |
greedy-token pipeline "…" [--execute] |
Multi-step pipeline |
greedy-token pipeline --list |
Named pipeline recipes |
greedy-token rag QUERY |
Search docs/rag/ |
greedy-token scripts --list |
Workspace script wrappers |
greedy-token scripts --run ID [--execute] |
Run wrapper |
greedy-token trust add PATH |
Approve the current SHA-256 and identity of a workspace script |
greedy-token trust list |
List local workspace script approvals |
greedy-token trust verify |
Verify every approval against disk |
greedy-token trust revoke PATH |
Remove a local script approval |
greedy-token audit-context |
Rules/skills token audit |
greedy-token calibrate [--overhead N] [--from-file PATH] |
Calibrate the naive agent-chat baseline (writes baseline: to ~/.greedy-token/config.yaml) |
greedy-token tokens PATH… |
Count tokens in paths |
greedy-token compress |
Short prompt (stdin; --ollama) |
greedy-token report [--since 7d] |
Usage telemetry: override/hold signal, explicit task outcomes, and outcome calibration |
greedy-token override … |
Log a script_override telemetry event |
greedy-token crystallize draft ID [--since 30d] |
L3 safe mode: draft script (.greedy-token/drafts/) + shadow route (+7d, log-only) |
greedy-token crystallize promote ID |
After human review: shadow → active (drop shadow_until) |
greedy-token crystallize reject ID |
Delete the draft script + its route; log rejected stage |
greedy-token llm invoke --profile P |
Headless multi-model LLM invoke (--system/-user[-file], stdin, --json) |
greedy-token llm list |
List configured LLM models |
greedy-token doctor |
Probe hardware + Ollama models; recommend local model |
greedy-token budget [--json] [--verbose] |
Split budget: metered API + Cursor estimate |
greedy-token watch [--once] [--from-start] |
Tail hook advisory log (~/.greedy-token/advisory.jsonl) |
greedy-token init [--profile solo|team|ci] [--preset NAME|URL|PATH] [--routes-from FILE] [--routes-scaffold] |
Bootstrap: detect rg/python/ollama + write config/policy; merge team route presets / scaffold workspace routes |
greedy-token config [--init] [--export] [--reveal] |
Ollama URL/model settings (--export masks CHEAP_LLM_API_KEY as ***; --reveal prints it) |
greedy-token hub serve [--host H] [--port N] |
Local ops dashboard (telemetry + crystallize) |
greedy-token-mcp |
Start MCP server (stdio) |
Global: --no-log disables telemetry for one invocation.
Pipeline execute: MCP greedy_token_pipeline and CLI greedy-token pipeline are dry-run by default. Pass execute=true (MCP) or --execute (CLI) to run allowlisted steps.
Auto-execute (read-only or stdout-only): tool-tier rg / jq, plus pipeline steps in PIPELINE_AUTO_RUN (src/greedy_token/pipeline.py) — check-meta-sync, configurator-boolean-audit, audit-skill, classify-file, search, read-hits, rag.
Route command trust boundary: workspace read_only: true is metadata, not
authorization. greedy-token run --execute accepts only internally built
rg/jq argv, registered read-only wrappers, or a workspace-relative
.py/.sh path approved in the user-local, workspace-bound trust manifest:
greedy-token trust add scripts/my-read-only-check.py --note "reviewed: stdout only"
greedy-token trust verify
SHA-256 and file identity are rechecked immediately before each approved
launch. Edits, symlink/path replacement, deleted/recreated files, absolute or
outside-workspace paths, python -c, shell -c, and trust-like fields from
URL/file presets fail closed. POSIX binds the verified descriptor through
/dev/fd; Windows retains a narrow verify-to-open window, and concurrent
same-inode writes are not snapshotted. The old trusted_script_paths config key
is deprecated dry-run metadata and grants no privilege. Subprocesses receive a
validated argv list with shell=False. See the
trust manifest and TOCTOU contract.
Routing benchmark
bench/routing_corpus.yaml is a held-out/adversarial classification gate,
separate from bench/route_examples.yaml. It reports exact-match accuracy,
confusion matrix, per-target precision/recall, family/language accuracy, and a
mandatory zero false-cheap rate.
Lexical retrieval benchmark
bench/retrieval_corpus.jsonl labels expected chunk IDs for RU/EN, each domain,
and exact, identifier, morphology, and paraphrase cases.
python bench/retrieval_benchmark.py --root /path/to/workspace reports
Recall@1/3/5, MRR, locale/domain/case-type breakdowns, and cold-index versus
warm-query latency.
Retrieval is local lexical BM25/FTS: SQLite FTS5 with the unicode61
tokenizer, Unicode NFKC + casefold normalization, and no embeddings or network
calls. Only docs/rag/manifest.jsonl entries are eligible. The persistent index
is content-hash invalidated and stored under the user cache directory
($GREEDY_TOKEN_CACHE_DIR, $XDG_CACHE_HOME, or ~/.cache), outside the
workspace. SQLite builds without FTS5 use the compatibility overlap scorer;
formatted hits name the engine and BM25 score.
bench/evidence_corpus.v1.yaml and its SHA-256 lock add the separate public
end-to-end evidence layer: frozen synthetic RU/EN fixtures, task-specific
file/line, exit-code, chunk-ID and escalation oracles, temporary workspaces,
and comparisons for direct rg/script, greedy CLI, greedy MCP stdio, and an
agent baseline. The deterministic agent is labelled contract_stub; a real
host baseline is manual. The JSON scorecard reports routing and task success
separately, executor/retrieval/escalation success, attempts, p50/p95, and
authoritative billing only. Cursor cost remains unknown when billing data is
unavailable; failed work never counts as saved. See benchmark contract.
Confidence calibration
An absent override is not evidence of correctness. The legacy telemetry is therefore named override/hold confidence and appears only as a behavioural signal.
Router confidence calibrates only from explicit route_outcome events whose
outcome is success or failure. Calibration is independent by route, tier,
and language; the most-specific segment with ≥ 20 events
(CALIBRATION_MIN_EVENTS) wins, then tier → language → global. Sparse data
uses the score formula and is visibly labelled formula (uncalibrated; explicit outcome n=…). Score buckets remain [0, 2), [2, 4), [4, 6),
[6, 8), and [8, +).
Outcome confidence calibration (explicit success/failure; min n=20):
segment bucket n predicted observed status
tier:python [2, 4) 25 75% 80% calibrated
Repeated work → crystallize into a script → next time 0 LLM. Details: guide · roadmap
License: MIT · v0.16.1
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file greedy_token-0.16.1.tar.gz.
File metadata
- Download URL: greedy_token-0.16.1.tar.gz
- Upload date:
- Size: 1.2 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c62f6e070a1023ff67a694f3adc14f3c8009d88c606a4b48f48833edf8dd288c
|
|
| MD5 |
19aff0a8cbc27ee32348b45de2561a1f
|
|
| BLAKE2b-256 |
9fb0032db937073b1a52159ffd9d751bb95c33952dd9e60163fcd0ded052b7bf
|
Provenance
The following attestation bundles were made for greedy_token-0.16.1.tar.gz:
Publisher:
publish.yml on svasenkov/greedy-token
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
greedy_token-0.16.1.tar.gz -
Subject digest:
c62f6e070a1023ff67a694f3adc14f3c8009d88c606a4b48f48833edf8dd288c - Sigstore transparency entry: 2581374507
- Sigstore integration time:
-
Permalink:
svasenkov/greedy-token@912ff0306a94eb1abbc98d5c22f40495e0788983 -
Branch / Tag:
refs/tags/v0.16.1 - Owner: https://github.com/svasenkov
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@912ff0306a94eb1abbc98d5c22f40495e0788983 -
Trigger Event:
release
-
Statement type:
File details
Details for the file greedy_token-0.16.1-py3-none-any.whl.
File metadata
- Download URL: greedy_token-0.16.1-py3-none-any.whl
- Upload date:
- Size: 456.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
91b348070b1d4f161404f738dde0d9106568c2dfded7c196e1847e53ded2deb1
|
|
| MD5 |
4af4b1ff11a91ca8d1dcf988750d2bb6
|
|
| BLAKE2b-256 |
6cb12873156510e66eb3683271bbf710b700a98ee3267f5a60f8fce858810171
|
Provenance
The following attestation bundles were made for greedy_token-0.16.1-py3-none-any.whl:
Publisher:
publish.yml on svasenkov/greedy-token
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
greedy_token-0.16.1-py3-none-any.whl -
Subject digest:
91b348070b1d4f161404f738dde0d9106568c2dfded7c196e1847e53ded2deb1 - Sigstore transparency entry: 2581374511
- Sigstore integration time:
-
Permalink:
svasenkov/greedy-token@912ff0306a94eb1abbc98d5c22f40495e0788983 -
Branch / Tag:
refs/tags/v0.16.1 - Owner: https://github.com/svasenkov
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@912ff0306a94eb1abbc98d5c22f40495e0788983 -
Trigger Event:
release
-
Statement type: