tokenecon — Token-Economics for Agentic AI
A dependency-free Python implementation of the token-economics cost model and tiered model routing design from the paper Token-Economics for Agentic AI: A Cost Model and Tiered Routing Reference (K. Pandey, Oct 2026).
The problem it solves: a single model call has a predictable price; a looping agent does not. Every tool call, observation, retry, and retrieved document adds another round-trip through a priced model, and teams usually discover the total from the cloud bill instead of the architecture. This package gives you two things:
- A cost model — decompose per-task spend into prompt, completion, tool-call, and retrieval tokens, priced per model tier, so you can forecast and budget before deployment instead of reconciling after it.
- A tiered router — classify each request by difficulty, send it to the cheapest tier capable of handling it, escalate on low confidence, and enforce per-task budgets that degrade gracefully instead of failing.
Install
pip install tokenecon
No dependencies. Python 3.9+.
From source: git clone https://github.com/karmendra8386/agentic-ai-token-economics.git && cd agentic-ai-token-economics && pip install .
Quickstart
from tokenecon import TieredRouter, DEFAULT_TIERS, MockModel
router = TieredRouter(tiers=list(DEFAULT_TIERS), budget=0.05) # 5¢ per task
receipt = router.run("Summarize our Q3 cloud spend in two sentences.", MockModel())
print(receipt.pretty())
# request : Summarize our Q3 cloud spend in two sentences.
# difficulty : 0.213
# outcome : ok
# step 1: tier=small in=210 out=120 cost=$0.000138 conf=0.70
# total_cost : $0.000138
Every run produces a receipt: tier used, tokens, cost, and confidence per step. Receipts are the unit of cost observability — accumulate them and they become the dataset your routing policy improves from.
CLI
# Run 5 sample tasks through the router and compare against the all-large baseline
tokenecon demo
# Same, under a 1¢ per-task budget (watch guardrails kick in)
tokenecon demo --budget 0.01
# Forecast spend for a list of requests (one per line, no model calls)
tokenecon estimate requests.txt
# Bring your own tiers (JSON: {"tiers": [{"name": ..., "input_per_mtok": ...,
# "output_per_mtok": ..., "capability": ...}]})
tokenecon demo --tiers my-tiers.json
How the router works
- Classify — each request gets a difficulty score (0–1) from surface signals: length, analytical keywords ("compare", "design", "prove"), question marks, multi-step phrasing. Simple and replaceable; graduate to a learned classifier with labeled traffic.
- Select — cheapest tier whose capability clears difficulty + safety margin. Falls back to the strongest tier; the router never refuses work.
- Escalate — every response carries a confidence signal; below the quality threshold, the request cascades to the next tier (try cheap, escalate on doubt).
- Budget — two guardrails, both degrade instead of failing:
- Before spending: if the cheapest first step exceeds the budget, serve a cached answer for $0.
- Before escalating: if the next tier up would break the budget, keep the current answer and mark the run degraded.
The tiering arithmetic: with planning fraction α ≈ 0.1 and price ratio r ≈ 30
between tiers, tiered cost is roughly 13% of the all-strong baseline —
tokenecon.tiering_ratio(0.1, 30).
Honest limitations
- Pricing is illustrative. Default tiers are placeholders. Substitute your provider's actual price list before making decisions.
- Token counting is heuristic (~4 chars/token for English). Production use needs the provider's tokenizer or usage API.
- Confidence comes from the mock model in demos. Production needs a real signal: log-probabilities, a verifier model, or sampling agreement.
- Single-request scope. Multi-turn compounding (tool-call overhead, retrieval accumulation) is modeled in the cost equations but not executed by the router.
- See the paper (§6) for the full limitations discussion.
Cite
If you use this in your work, please cite the paper:
K. Pandey, "Token-Economics for Agentic AI: A Cost Model and Tiered Routing Reference," Zenodo, DOI 10.5281/zenodo.23195511, Oct. 2026.
A CITATION.cff is included for automated citation tooling.
License
MIT — see LICENSE.
Metadata
Release files for tokenecon 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tokenecon-0.1.0.tar.gz | 13.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tokenecon-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 26.3 kB
Release files / tokenecon-0.1.0.tar.gz
| Download URL | tokenecon-0.1.0.tar.gz |
|---|---|
| Size | 13.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cbcb9cbea4a0e0188daa2e26ae408deb691689527cc582f7720864b32cbf4c21
|
|
BLAKE2b-256 checksum How to use checksums |
4cf47f39007b6f41de1d417be5ffe8e96362b586caebc5ccce89bc2c00f0d92f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency logRelease files / tokenecon-0.1.0-py3-none-any.whl
| Download URL | tokenecon-0.1.0-py3-none-any.whl |
|---|---|
| Size | 13.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
148fe310ceaeef0011cdff0814f477b65a32ff31359bcbc7f5016ef122979032
|
|
BLAKE2b-256 checksum How to use checksums |
2a775f9028601793e9a27c6640f98dc5c16fa865c2d6f005b99c7af1e013497d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency log