Skip to main content

greedy-token

Русская версия: README-RU.md

greedy-token mascot

You work in Cursor — greedy-token sits next to the agent (CLI + MCP) so everyday tasks don’t always open a full agent chat.

It routes each task to the cheapest matching tier (toolpythonollamaragcursor; first pattern match wins). Pipeline chains multiple tiers in one call. Escalation to Cursor agent chat only when no cheaper route matches. Each response includes a Token economy footer vs a naive full-context chat.

In Cursor:  your task  →  greedy-token (MCP/CLI)
                 ↓
     route (one tier per task):
       tool → python → ollama → rag → cursor
       first matching route in routes.yaml wins; ollama tier skipped if server down
                 ↓
     pipeline (optional, multi-step):
       e.g. check-meta-sync then audit-skill …
       composes tool / python / ollama / rag steps — not a separate tier
                 ↓
     escalation: Cursor agent chat when no cheaper route matches

What it does

Layer When LLM cost
tool (rg) find / grep / search ~0
python scripts, meta-sync, gen-env ~0
ollama bulk classify, skill audit local only
rag lookup in docs/rag/ small read
cursor wiring, refactor, architecture full agent chat

Tier order: TIER_ORDER in router.py / routes.yaml — first pattern match wins within the walk tool → python → ollama → rag → cursor. Not every tier runs on every task. Ollama tier is skipped when the server is down.

Scope & roadmap

Today the happy path is Cursor + Ollama + monorepo. CLI and MCP are IDE-agnostic; paid APIs and alternate local runtimes are not wired yet.

Full matrix (✅ / ❌ / 🔜) + acceptance criteria + GitHub issues: docs/ROADMAP.md · docs/ROADMAP-RU.md

Area ✅ today 🔜 v0.5+
Executors tool, python, ollama, rag cloud_llm, openai_compat local
Agent host Cursor MCP + token baseline Claude Desktop, Continue
Config OLLAMA_URL / OLLAMA_MODEL local_llm.provider, cloud_llm.provider

Install

Python 3.12+ (CI and PyPI builds use 3.12).

pip install greedy-token
# with Cursor MCP server:
pip install "greedy-token[mcp]"
# editable (monorepo):
cd projects/greedy-token-home/dev && ./scripts/install.sh
export GREEDY_TOKEN_ROOT=/path/to/workspace   # optional; auto-detect when markers exist

Cursor integration (recommended)

Full guide (any workspace / PyPI): docs/cursor-setup.md · docs/cursor-setup-RU.md

Starter kit in this repo (copy into your project):

Template Copy to
examples/cursor/mcp.json .cursor/mcp.json
examples/cursor/rules/token-economy.mdc .cursor/rules/token-economy.mdc
pip install "greedy-token[mcp]"
mkdir -p .cursor/rules
# from a greedy-token clone, or paste from the docs:
cp examples/cursor/mcp.json .cursor/mcp.json
cp examples/cursor/rules/token-economy.mdc .cursor/rules/token-economy.mdc

Then: Settings → MCP → greedy-token → Enable → Refreshnew Agent chat.

Expected: 5 MCP tools (including greedy_token_pipeline).

MCP tools

Tool Purpose
greedy_token_search Ripgrep: query + optional path
greedy_token_rag Search docs/rag/ chunks
greedy_token_route Recommend tier + token footer
greedy_token_pipeline Multi-step chain (python → ollama → rag)
greedy_token_usage Aggregate savings from ~/.greedy-token/usage.jsonl

Every tool response ends with a Token economy block — show it to the user.

Pipeline (multi-step)

pipeline: meta-audit configurator-boolean

or:

pipeline: check-meta-sync then audit-skill configurator-boolean

Named recipes (pipeline --list):

Recipe Steps
meta-audit python → ollama
meta-rag python → rag
search-rag rg → rag

Footer includes per-step savings table:

Per-step savings (if each step were a separate naive Cursor chat):
   #  step                   executor     ms   spent  baseline     saved  billing
   1  check-meta-sync        python       83       0     9,487     9,487  local script
   2  audit-skill            ollama     2698   2,507     9,499     6,992  local Ollama

Saved by executor (sum of per-step savings):
  python (script)              steps=1  spent ~0      saved ~9,487
  ollama (local LLM)           steps=1  spent ~2,507  saved ~6,992

CLI commands

Command Purpose
greedy-token route "…" Recommend tier + scoring
greedy-token estimate "…" Token-aware estimate + tier scan
greedy-token run "…" [--execute] Route + dry-run / read-only execute
greedy-token pipeline "…" [--execute] Multi-step pipeline
greedy-token pipeline --list Named pipeline recipes
greedy-token rag QUERY Search docs/rag/
greedy-token scripts --list Workspace script wrappers
greedy-token scripts --run ID [--execute] Run wrapper
greedy-token audit-context Rules/skills token audit
greedy-token tokens PATH… Count tokens in paths
greedy-token compress Short prompt (stdin; --ollama)
greedy-token report [--since 7d] Usage telemetry aggregate
greedy-token config [--init] [--export] Ollama URL/model settings
greedy-token-mcp Start MCP server (stdio)

Global: --no-log disables telemetry for one invocation.

Pipeline execute: MCP greedy_token_pipeline and CLI greedy-token pipeline are dry-run by default. Pass execute=true (MCP) or --execute (CLI) to run allowlisted steps.

Testing

Requires Python 3.12+ (same as CI). GitHub Actions runs pytest + Allure 3 (quality gate, GitHub Pages report; optional TestOps upload when repo vars are set). Line and branch coverage on src/greedy_token/ must stay at 100% (branch = true, fail_under = 100).

CI ethalon: .github/_ethalon/ (action pins in gha-actions.yaml) → runnable .github/workflows/. Same pattern as monorepo tests-java/.github/_ethalon/. Sync: ./scripts/sync-github-workflows.sh; CI runs ./scripts/check-github-workflows-sync.sh before pytest.

cd projects/greedy-token-home/dev && ./scripts/install.sh
source .venv/bin/activate
cd ../greedy-token
python -m coverage run -m pytest tests/ -v --alluredir=build/allure-results
python -m coverage report --include='src/greedy_token/*'
npx --yes allure@3.13.0 quality-gate build/allure-results --config allurerc.mjs
npx --yes allure@3.13.0 generate build/allure-results --config allurerc.mjs -o build/allure-report

Coverage: branch = true and fail_under = 100 on src/greedy_token/ (see [tool.coverage.run] / [tool.coverage.report] in pyproject.toml). CI runs coverage run + coverage report on every push/PR.

Pyramid slices: layer per module in tests/pyramid_layers.py → Allure label layer + pytest marker (-m unit|component|integration|e2e). CI matrix job pyramid runs each slice separately.

Optional integration tests (real monorepo files) run when the checkout includes stacks/java-spring/; set GREEDY_TOKEN_ROOT to override the workspace root.

TestOps: project 5276 on allure.autotests.cloud. CI uploads when repo secret ALLURE_TOKEN is set (ALLURE_PROJECT_ID defaults to 5276, override via repo variable). Pyramid layers (unit / component / integration) are set via Allure label layer in tests/pyramid_layers.py — same keys as Java @Layer and TestOps mappings. Human-readable names use @allure.title / @allure.feature / @allure.story / @allure.epic on each test, and @allure.parent_suite / @allure.suite on each module (pytestmark) for TestOps folder names — JUnit @DisplayName / @Feature equivalent.

Examples

# Search (0 LLM tokens)
greedy-token run "find baseUrl in configurator-option-presets.html" --execute

# RAG lookup
greedy-token rag "baseUrl -D flag"

# Ollama tier
greedy-token route "audit skill configurator-boolean"

# Pipeline dry-run
greedy-token pipeline "pipeline: meta-audit configurator-boolean"

# Pipeline execute (python + ollama)
greedy-token pipeline "check-meta-sync then audit-skill configurator-boolean" --execute

# Savings report
greedy-token report --since 7d

Token economy footer

Single-tool responses include:

  • This call — executor, spent, billing (local vs cloud)
  • Cursor baseline — rules + task + overhead
  • Saved vs naive Cursor chat

Pipeline adds per-step baseline / spent / saved and saved by executor.

Note: MCP executor steps are local/cheap. Agent chat wrapper (rules + your message + reply) still uses Cursor tokens.

Usage telemetry

Log file: ~/.greedy-token/usage.jsonl (disable: GREEDY_TOKEN_LOG=0).

Each event: tier, est_tokens, cursor_baseline, cursor_saved, duration_ms.

Pipeline logs one event per step. When the log exceeds GREEDY_TOKEN_LOG_MAX_BYTES (default 5 MiB), it rotates to usage.jsonl.1, .2, …; report reads the active log and archives.

Environment

Var Default
GREEDY_TOKEN_ROOT auto-detect or required
OLLAMA_URL from config or http://localhost:11434
OLLAMA_MODEL from config or qwen2.5-coder:7b-instruct-q4_K_M
GREEDY_TOKEN_LOG ~/.greedy-token/usage.jsonl
GREEDY_TOKEN_LOG_MAX_BYTES 5242880 (5 MiB)
GREEDY_TOKEN_LOG_MAX_FILES 5 rotated archives

Ollama config

Priority (low → high): defaults → ~/.greedy-token/config.yaml$GREEDY_TOKEN_ROOT/.greedy-token.yamlOLLAMA_* env.

greedy-token config init
greedy-token config init --model llama3.2 --url http://192.168.1.10:11434
greedy-token config
eval "$(greedy-token config --export)"

Routing config

File Purpose
src/greedy_token/config/routes.yaml Task routing patterns
src/greedy_token/config/pipelines.yaml Named pipeline recipes

--execute safety

Auto-execute (read-only or stdout-only): rg, jq, check-meta-sync, pipeline steps in allowlist.

Rsync / migrate / batch-inventory — dry-run only unless run manually.

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

greedy_token-0.4.6.tar.gz (796.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

greedy_token-0.4.6-py3-none-any.whl (327.6 kB view details)

Uploaded Python 3

File details

Details for the file greedy_token-0.4.6.tar.gz.

File metadata

  • Download URL: greedy_token-0.4.6.tar.gz
  • Upload date:
  • Size: 796.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for greedy_token-0.4.6.tar.gz
Algorithm Hash digest
SHA256 1c71393b18f3e596b22555140841fafb3b2f9d45e246363314c89626ec6311c3
MD5 27ae7b338c3b80fa0b778bc0b03f99ae
BLAKE2b-256 b909fda0f56b5d810497a5b8953935fdfbfd14acff57693a9eaae120a5001bba

See more details on using hashes here.

Provenance

The following attestation bundles were made for greedy_token-0.4.6.tar.gz:

Publisher: publish.yml on svasenkov/greedy-token

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file greedy_token-0.4.6-py3-none-any.whl.

File metadata

  • Download URL: greedy_token-0.4.6-py3-none-any.whl
  • Upload date:
  • Size: 327.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for greedy_token-0.4.6-py3-none-any.whl
Algorithm Hash digest
SHA256 1d49b32595214da61a4d28e0ed8b5a20fe863d475144edcdcefecf437b00b6b5
MD5 fe6145fab6d3b1c093b09193eedb349f
BLAKE2b-256 572812d7feddae775f71eb2e9639e1858ad3733c9ccc35503ff2e743258fa0dd

See more details on using hashes here.

Provenance

The following attestation bundles were made for greedy_token-0.4.6-py3-none-any.whl:

Publisher: publish.yml on svasenkov/greedy-token

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.16.2

2 files

0.16.1

2 files

0.16.0

2 files

0.15.0

2 files

0.14.1

2 files

0.14.0

2 files

0.13.0

2 files

0.11.1

2 files

0.11.0

2 files

0.10.0

2 files

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.6.3

2 files

0.6.2

2 files

0.6.1

2 files

0.6.0

2 files

0.5.8

2 files

0.5.7

2 files

0.5.6

2 files

0.5.5

2 files

0.5.4

2 files

0.5.3

2 files

0.5.2

2 files

0.5.1

2 files

This release

0.4.6 This release

2 files

0.4.5

2 files

0.4.4

2 files

0.4.3

2 files

0.4.2

2 files

0.4.1

2 files

0.2.2

2 files

0.2.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page