greedy-token
Русская версия: README-RU.md
You work in Cursor — greedy-token sits next to the agent (CLI + MCP) so everyday tasks don’t always open a full agent chat.
It routes each task to the cheapest matching tier (tool → python → ollama → rag → cursor; first pattern match wins). Pipeline chains multiple tiers in one call. Escalation to Cursor agent chat only when no cheaper route matches. Each response includes a Token economy footer vs a naive full-context chat.
In Cursor: your task → greedy-token (MCP/CLI)
↓
route (one tier per task):
tool → python → ollama → rag → cursor
first matching route in routes.yaml wins; ollama tier skipped if server down
↓
pipeline (optional, multi-step):
e.g. check-meta-sync then audit-skill …
composes tool / python / ollama / rag steps — not a separate tier
↓
escalation: Cursor agent chat when no cheaper route matches
What it does
| Layer | When | LLM cost |
|---|---|---|
| tool (rg) | find / grep / search | ~0 |
| python | scripts, meta-sync, gen-env | ~0 |
| ollama | bulk classify, skill audit | local only |
| rag | lookup in docs/rag/ |
small read |
| cursor | wiring, refactor, architecture | full agent chat |
Tier order: TIER_ORDER in router.py / routes.yaml — first pattern match wins within the walk tool → python → ollama → rag → cursor. Not every tier runs on every task. Ollama tier is skipped when the server is down.
Scope & roadmap
Today the happy path is Cursor + Ollama + monorepo. CLI and MCP are IDE-agnostic; paid APIs and alternate local runtimes are not wired yet.
Full matrix (✅ / ❌ / 🔜) + acceptance criteria + GitHub issues: docs/ROADMAP.md · docs/ROADMAP-RU.md
| Area | ✅ today | 🔜 v0.5+ |
|---|---|---|
| Executors | tool, python, ollama, rag |
cloud_llm, openai_compat local |
| Agent host | Cursor MCP + token baseline | Claude Desktop, Continue |
| Config | OLLAMA_URL / OLLAMA_MODEL |
local_llm.provider, cloud_llm.provider |
Install
Python 3.12+ (CI and PyPI builds use 3.12).
pip install greedy-token
# with Cursor MCP server:
pip install "greedy-token[mcp]"
# editable (monorepo):
cd projects/greedy-token-home/dev && ./scripts/install.sh
export GREEDY_TOKEN_ROOT=/path/to/workspace # optional; auto-detect when markers exist
Cursor integration (recommended)
Full guide (any workspace / PyPI): docs/cursor-setup.md · docs/cursor-setup-RU.md
Starter kit in this repo (copy into your project):
| Template | Copy to |
|---|---|
examples/cursor/mcp.json |
.cursor/mcp.json |
examples/cursor/rules/token-economy.mdc |
.cursor/rules/token-economy.mdc |
pip install "greedy-token[mcp]"
mkdir -p .cursor/rules
# from a greedy-token clone, or paste from the docs:
cp examples/cursor/mcp.json .cursor/mcp.json
cp examples/cursor/rules/token-economy.mdc .cursor/rules/token-economy.mdc
Then: Settings → MCP → greedy-token → Enable → Refresh → new Agent chat.
Expected: 5 MCP tools (including greedy_token_pipeline).
MCP tools
| Tool | Purpose |
|---|---|
greedy_token_search |
Ripgrep: query + optional path |
greedy_token_rag |
Search docs/rag/ chunks |
greedy_token_route |
Recommend tier + token footer |
greedy_token_pipeline |
Multi-step chain (python → ollama → rag) |
greedy_token_usage |
Aggregate savings from ~/.greedy-token/usage.jsonl |
Every tool response ends with a Token economy block — show it to the user.
Pipeline (multi-step)
pipeline: meta-audit configurator-boolean
or:
pipeline: check-meta-sync then audit-skill configurator-boolean
Named recipes (pipeline --list):
| Recipe | Steps |
|---|---|
meta-audit |
python → ollama |
meta-rag |
python → rag |
search-rag |
rg → rag |
Footer includes per-step savings table:
Per-step savings (if each step were a separate naive Cursor chat):
# step executor ms spent baseline saved billing
1 check-meta-sync python 83 0 9,487 9,487 local script
2 audit-skill ollama 2698 2,507 9,499 6,992 local Ollama
Saved by executor (sum of per-step savings):
python (script) steps=1 spent ~0 saved ~9,487
ollama (local LLM) steps=1 spent ~2,507 saved ~6,992
CLI commands
| Command | Purpose |
|---|---|
greedy-token route "…" |
Recommend tier + scoring |
greedy-token estimate "…" |
Token-aware estimate + tier scan |
greedy-token run "…" [--execute] |
Route + dry-run / read-only execute |
greedy-token pipeline "…" [--execute] |
Multi-step pipeline |
greedy-token pipeline --list |
Named pipeline recipes |
greedy-token rag QUERY |
Search docs/rag/ |
greedy-token scripts --list |
Workspace script wrappers |
greedy-token scripts --run ID [--execute] |
Run wrapper |
greedy-token audit-context |
Rules/skills token audit |
greedy-token tokens PATH… |
Count tokens in paths |
greedy-token compress |
Short prompt (stdin; --ollama) |
greedy-token report [--since 7d] |
Usage telemetry aggregate |
greedy-token config [--init] [--export] |
Ollama URL/model settings |
greedy-token-mcp |
Start MCP server (stdio) |
Global: --no-log disables telemetry for one invocation.
Pipeline execute: MCP
greedy_token_pipelineand CLIgreedy-token pipelineare dry-run by default. Passexecute=true(MCP) or--execute(CLI) to run allowlisted steps.
Testing
Requires Python 3.12+ (same as CI). GitHub Actions runs pytest + Allure 3 (quality gate, GitHub Pages report; optional TestOps upload when repo vars are set). Line and branch coverage on src/greedy_token/ must stay at 100% (branch = true, fail_under = 100).
CI ethalon: .github/_ethalon/ (action pins in gha-actions.yaml) → runnable .github/workflows/. Same pattern as monorepo tests-java/.github/_ethalon/. Sync: ./scripts/sync-github-workflows.sh; CI runs ./scripts/check-github-workflows-sync.sh before pytest.
cd projects/greedy-token-home/dev && ./scripts/install.sh
source .venv/bin/activate
cd ../greedy-token
python -m coverage run -m pytest tests/ -v --alluredir=build/allure-results
python -m coverage report --include='src/greedy_token/*'
npx --yes allure@3.13.0 quality-gate build/allure-results --config allurerc.mjs
npx --yes allure@3.13.0 generate build/allure-results --config allurerc.mjs -o build/allure-report
Coverage: branch = true and fail_under = 100 on src/greedy_token/ (see [tool.coverage.run] / [tool.coverage.report] in pyproject.toml). CI runs coverage run + coverage report on every push/PR.
Pyramid slices: layer per module in tests/pyramid_layers.py → Allure label layer + pytest marker (-m unit|component|integration|e2e). CI matrix job pyramid runs each slice separately.
Optional integration tests (real monorepo files) run when the checkout includes stacks/java-spring/; set GREEDY_TOKEN_ROOT to override the workspace root.
TestOps: project 5276 on allure.autotests.cloud. CI uploads when repo secret ALLURE_TOKEN is set (ALLURE_PROJECT_ID defaults to 5276, override via repo variable). Pyramid layers (unit / component / integration) are set via Allure label layer in tests/pyramid_layers.py — same keys as Java @Layer and TestOps mappings. Human-readable names use @allure.title / @allure.feature / @allure.story / @allure.epic on each test, and @allure.parent_suite / @allure.suite on each module (pytestmark) for TestOps folder names — JUnit @DisplayName / @Feature equivalent.
Examples
# Search (0 LLM tokens)
greedy-token run "find baseUrl in configurator-option-presets.html" --execute
# RAG lookup
greedy-token rag "baseUrl -D flag"
# Ollama tier
greedy-token route "audit skill configurator-boolean"
# Pipeline dry-run
greedy-token pipeline "pipeline: meta-audit configurator-boolean"
# Pipeline execute (python + ollama)
greedy-token pipeline "check-meta-sync then audit-skill configurator-boolean" --execute
# Savings report
greedy-token report --since 7d
Token economy footer
Single-tool responses include:
- This call — executor, spent, billing (local vs cloud)
- Cursor baseline — rules + task + overhead
- Saved vs naive Cursor chat
Pipeline adds per-step baseline / spent / saved and saved by executor.
Note: MCP executor steps are local/cheap. Agent chat wrapper (rules + your message + reply) still uses Cursor tokens.
Usage telemetry
Log file: ~/.greedy-token/usage.jsonl (disable: GREEDY_TOKEN_LOG=0).
Each event: tier, est_tokens, cursor_baseline, cursor_saved, duration_ms.
Pipeline logs one event per step. When the log exceeds GREEDY_TOKEN_LOG_MAX_BYTES (default 5 MiB), it rotates to usage.jsonl.1, .2, …; report reads the active log and archives.
Environment
| Var | Default |
|---|---|
GREEDY_TOKEN_ROOT |
auto-detect or required |
OLLAMA_URL |
from config or http://localhost:11434 |
OLLAMA_MODEL |
from config or qwen2.5-coder:7b-instruct-q4_K_M |
GREEDY_TOKEN_LOG |
~/.greedy-token/usage.jsonl |
GREEDY_TOKEN_LOG_MAX_BYTES |
5242880 (5 MiB) |
GREEDY_TOKEN_LOG_MAX_FILES |
5 rotated archives |
Ollama config
Priority (low → high): defaults → ~/.greedy-token/config.yaml → $GREEDY_TOKEN_ROOT/.greedy-token.yaml → OLLAMA_* env.
greedy-token config init
greedy-token config init --model llama3.2 --url http://192.168.1.10:11434
greedy-token config
eval "$(greedy-token config --export)"
Routing config
| File | Purpose |
|---|---|
src/greedy_token/config/routes.yaml |
Task routing patterns |
src/greedy_token/config/pipelines.yaml |
Named pipeline recipes |
--execute safety
Auto-execute (read-only or stdout-only): rg, jq, check-meta-sync, pipeline steps in allowlist.
Rsync / migrate / batch-inventory — dry-run only unless run manually.
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file greedy_token-0.4.6.tar.gz.
File metadata
- Download URL: greedy_token-0.4.6.tar.gz
- Upload date:
- Size: 796.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1c71393b18f3e596b22555140841fafb3b2f9d45e246363314c89626ec6311c3
|
|
| MD5 |
27ae7b338c3b80fa0b778bc0b03f99ae
|
|
| BLAKE2b-256 |
b909fda0f56b5d810497a5b8953935fdfbfd14acff57693a9eaae120a5001bba
|
Provenance
The following attestation bundles were made for greedy_token-0.4.6.tar.gz:
Publisher:
publish.yml on svasenkov/greedy-token
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
greedy_token-0.4.6.tar.gz -
Subject digest:
1c71393b18f3e596b22555140841fafb3b2f9d45e246363314c89626ec6311c3 - Sigstore transparency entry: 2129774845
- Sigstore integration time:
-
Permalink:
svasenkov/greedy-token@537f96e71f47d0d7fb9a6e26f3742208fc5a66d8 -
Branch / Tag:
refs/tags/v0.4.6 - Owner: https://github.com/svasenkov
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@537f96e71f47d0d7fb9a6e26f3742208fc5a66d8 -
Trigger Event:
release
-
Statement type:
File details
Details for the file greedy_token-0.4.6-py3-none-any.whl.
File metadata
- Download URL: greedy_token-0.4.6-py3-none-any.whl
- Upload date:
- Size: 327.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1d49b32595214da61a4d28e0ed8b5a20fe863d475144edcdcefecf437b00b6b5
|
|
| MD5 |
fe6145fab6d3b1c093b09193eedb349f
|
|
| BLAKE2b-256 |
572812d7feddae775f71eb2e9639e1858ad3733c9ccc35503ff2e743258fa0dd
|
Provenance
The following attestation bundles were made for greedy_token-0.4.6-py3-none-any.whl:
Publisher:
publish.yml on svasenkov/greedy-token
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
greedy_token-0.4.6-py3-none-any.whl -
Subject digest:
1d49b32595214da61a4d28e0ed8b5a20fe863d475144edcdcefecf437b00b6b5 - Sigstore transparency entry: 2129774949
- Sigstore integration time:
-
Permalink:
svasenkov/greedy-token@537f96e71f47d0d7fb9a6e26f3742208fc5a66d8 -
Branch / Tag:
refs/tags/v0.4.6 - Owner: https://github.com/svasenkov
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@537f96e71f47d0d7fb9a6e26f3742208fc5a66d8 -
Trigger Event:
release
-
Statement type: