Skip to main content

tokenecon — Token-Economics for Agentic AI

A dependency-free Python implementation of the token-economics cost model and tiered model routing design from the paper Token-Economics for Agentic AI: A Cost Model and Tiered Routing Reference (K. Pandey, Oct 2026).

The problem it solves: a single model call has a predictable price; a looping agent does not. Every tool call, observation, retry, and retrieved document adds another round-trip through a priced model, and teams usually discover the total from the cloud bill instead of the architecture. This package gives you two things:

  1. A cost model — decompose per-task spend into prompt, completion, tool-call, and retrieval tokens, priced per model tier, so you can forecast and budget before deployment instead of reconciling after it.
  2. A tiered router — classify each request by difficulty, send it to the cheapest tier capable of handling it, escalate on low confidence, and enforce per-task budgets that degrade gracefully instead of failing.

Install

pip install tokenecon

No dependencies. Python 3.9+.

From source: git clone https://github.com/karmendra8386/agentic-ai-token-economics.git && cd agentic-ai-token-economics && pip install .

Quickstart

from tokenecon import TieredRouter, DEFAULT_TIERS, MockModel

router = TieredRouter(tiers=list(DEFAULT_TIERS), budget=0.05)  # 5¢ per task
receipt = router.run("Summarize our Q3 cloud spend in two sentences.", MockModel())
print(receipt.pretty())
# request    : Summarize our Q3 cloud spend in two sentences.
# difficulty : 0.213
# outcome    : ok
#   step 1: tier=small in=210 out=120 cost=$0.000138 conf=0.70
# total_cost : $0.000138

Every run produces a receipt: tier used, tokens, cost, and confidence per step. Receipts are the unit of cost observability — accumulate them and they become the dataset your routing policy improves from.

CLI

# Run 5 sample tasks through the router and compare against the all-large baseline
tokenecon demo

# Same, under a 1¢ per-task budget (watch guardrails kick in)
tokenecon demo --budget 0.01

# Forecast spend for a list of requests (one per line, no model calls)
tokenecon estimate requests.txt

# Bring your own tiers (JSON: {"tiers": [{"name": ..., "input_per_mtok": ...,
#   "output_per_mtok": ..., "capability": ...}]})
tokenecon demo --tiers my-tiers.json

How the router works

  1. Classify — each request gets a difficulty score (0–1) from surface signals: length, analytical keywords ("compare", "design", "prove"), question marks, multi-step phrasing. Simple and replaceable; graduate to a learned classifier with labeled traffic.
  2. Select — cheapest tier whose capability clears difficulty + safety margin. Falls back to the strongest tier; the router never refuses work.
  3. Escalate — every response carries a confidence signal; below the quality threshold, the request cascades to the next tier (try cheap, escalate on doubt).
  4. Budget — two guardrails, both degrade instead of failing:
    • Before spending: if the cheapest first step exceeds the budget, serve a cached answer for $0.
    • Before escalating: if the next tier up would break the budget, keep the current answer and mark the run degraded.

The tiering arithmetic: with planning fraction α ≈ 0.1 and price ratio r ≈ 30 between tiers, tiered cost is roughly 13% of the all-strong baseline — tokenecon.tiering_ratio(0.1, 30).

Honest limitations

  • Pricing is illustrative. Default tiers are placeholders. Substitute your provider's actual price list before making decisions.
  • Token counting is heuristic (~4 chars/token for English). Production use needs the provider's tokenizer or usage API.
  • Confidence comes from the mock model in demos. Production needs a real signal: log-probabilities, a verifier model, or sampling agreement.
  • Single-request scope. Multi-turn compounding (tool-call overhead, retrieval accumulation) is modeled in the cost equations but not executed by the router.
  • See the paper (§6) for the full limitations discussion.

Cite

If you use this in your work, please cite the paper:

K. Pandey, "Token-Economics for Agentic AI: A Cost Model and Tiered Routing Reference," Zenodo, DOI 10.5281/zenodo.23195511, Oct. 2026.

A CITATION.cff is included for automated citation tooling.

License

MIT — see LICENSE.

Metadata

Release files for tokenecon 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokenecon 0.1.0
File Size Uploaded
tokenecon-0.1.0.tar.gz 13.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tokenecon 0.1.0
File Interpreter ABI Platform
tokenecon-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 26.3 kB

Release files / tokenecon-0.1.0.tar.gz

Download URL tokenecon-0.1.0.tar.gz
Size 13.3 kB
Tags Source
SHA-256 checksum
How to use checksums
cbcb9cbea4a0e0188daa2e26ae408deb691689527cc582f7720864b32cbf4c21
BLAKE2b-256 checksum
How to use checksums
4cf47f39007b6f41de1d417be5ffe8e96362b586caebc5ccce89bc2c00f0d92f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / tokenecon-0.1.0-py3-none-any.whl

Download URL tokenecon-0.1.0-py3-none-any.whl
Size 13.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
148fe310ceaeef0011cdff0814f477b65a32ff31359bcbc7f5016ef122979032
BLAKE2b-256 checksum
How to use checksums
2a775f9028601793e9a27c6640f98dc5c16fa865c2d6f005b99c7af1e013497d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.0

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page