Reduce AI context loss by 2x. Graph-backed checkpoint and resume for any LLM session.

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

Shweta-Mishra-ai

These details have not been verified by PyPI

Project description

TokenMizer

Keep your AI context alive across sessions.

Graph-backed memory · session checkpointing · intelligent compression
Drop-in proxy for Claude, GPT, Gemini, Grok, DeepSeek, Ollama — any LLM.

The Problem

Every AI session has a context limit. When you hit it:

The model forgets every decision, rationale, and context built over hours
You waste 10–30 minutes re-explaining the project every new session
Large files (CSV, PDF, Excel) eat your entire token budget instantly

How TokenMizer Solves It

TokenMizer is a local proxy between your app and any LLM. Every request goes through a pipeline that builds a live knowledge graph, compresses inputs, caches responses, and auto-checkpoints before context runs out.

Your App  →  TokenMizer (:8000)  →  Claude / GPT / Gemini / any LLM
                    │
          ┌─────────┴──────────────┐
          │   6-Layer Pipeline     │
          │   L0  File Intel       │  CSV/PDF/Excel → schema + sample
          │   L1  Compression      │  15–40% input reduction
          │   L2  Output Trim      │  5–15% output reduction
          │   L3  Semantic Cache   │  100% on repeated queries
          │   L4  Graph Memory     │  session continuity
          │   L5  Prompt Cache     │  90% on repeated system prompts
          └────────────────────────┘

Architecture

Decision Memory — 4-State Model

Status	Meaning	In Resume
🟢 `ACTIVE`	Current — in effect	✅ Always
🟡 `SUPERSEDED`	Replaced by newer decision	⚠️ 7 days
🔴 `INVALIDATED`	Explicitly wrong/cancelled	⚠️ Always (warning)
⬜ `ARCHIVED`	Old but valid, not relevant	❌ Never

History is never deleted. "Why did we switch from React to Next.js?" — always answerable.

Quick Start

1. Install

# Recommended
pip install "tokenmizer[anthropic,cache]"

# All providers
pip install "tokenmizer[anthropic,openai,gemini,cohere,cache]"

# No key? Use Ollama (free, local)
brew install ollama && ollama pull llama3
pip install tokenmizer

2. Set your API key

export TOKENMIZER_ANTHROPIC_API_KEY=sk-ant-...
# or: TOKENMIZER_OPENAI_API_KEY, TOKENMIZER_GEMINI_API_KEY, etc.

3. Start

tokenmizer serve
# → Proxy:     http://localhost:8000/v1/chat/completions
# → Dashboard: http://localhost:8000
# → API docs:  http://localhost:8000/docs

4. Use — change one line

from openai import OpenAI

client = OpenAI(
    api_key="your-key",
    base_url="http://localhost:8000/v1",  # ← only this changes
)

response = client.chat.completions.create(
    model="claude-sonnet-4-6",
    messages=[{"role": "user", "content": "Let's build an auth service"}],
    extra_body={"session_id": "my-project"},  # enables graph memory
)

⚠️ Streaming is not supported yet. Set stream: false (or leave it unset — that's the default) in every request. Requests with stream: true return HTTP 501. True SSE streaming is planned for v0.3.

Cursor: Settings → Models → your model entry → disable "Stream responses" Continue.dev: in config.json, set "streamResponses": false for the TokenMizer provider entry

Claude Code Integration

Option A — Plugin (recommended)

# Add TokenMizer as a plugin marketplace
/plugin marketplace add Shweta-Mishra-ai/tokenmizer

# Install
/plugin install tokenmizer@Shweta-Mishra-ai/tokenmizer

Then use skills directly:

/tokenmizer:checkpoint my-project      → save session to graph memory
/tokenmizer:resume my-project          → load previous session (300 tokens)
/tokenmizer:resume my-project full     → full 600-token context
/tokenmizer:analyze /data/sales.csv    → analyze file (99% token savings)
/tokenmizer:stats                      → token savings report

Option B — MCP server

Add to ~/.claude/settings.json:

{
  "mcpServers": {
    "tokenmizer": {
      "command": "python3",
      "args": ["-m", "tokenmizer.mcp.server"],
      "env": { "TOKENMIZER_URL": "http://localhost:8000" }
    }
  }
}

Other Tools

Cursor / Continue.dev / any OpenAI-compatible tool:

API Base URL:  http://localhost:8000/v1

Session Resume

tokenmizer checkpoint my-project
tokenmizer resume my-project

Goal: Build FastAPI auth service with JWT + PostgreSQL
Done: Project setup | User model | Login endpoint | Fix 422 | 18 tests passing
In progress: Refresh token rotation
Decided: PostgreSQL (concurrent writes) | bcrypt | Redis for refresh tokens
Changed: ~~React~~ → Next.js (better SEO)
Files: api/auth.py, api/models.py, config.py
Continue: Implement token refresh endpoint

247 tokens replaces 25,000+ tokens of conversation history.

File Intelligence

from tokenmizer.filters.file_intelligence import FileIntelligence

fi = FileIntelligence()
result = fi.process(open("sales.csv","rb").read(), "sales.csv",
                    token_budget=500, query="which regions underperforming")
# 412,000 tokens → 447 tokens  (99.9% saved)

File	Savings
CSV (50k rows)	99.9%
PDF (200 pages)	98.8%
Excel (10 sheets)	99.7%
JSON (1k items)	95%

Works Alongside Caveman & CodeBurn

TokenMizer complements — does not replace — these tools:

Tool	What it does
Caveman	Output tokens shorter (~65%)
CodeBurn	Input context trimming
TokenMizer	Graph memory + resume + file intelligence + cache

Tip: If using Caveman, set terse_output: enabled: false in tokenmizer.yaml to avoid conflicting system prompts.

Supported Providers

Model strings pass through unchanged — the newest models work out of the box: claude-fable-5, claude-opus-4-8, claude-sonnet-5, claude-haiku-4-5, GPT-4o/o-series, Gemini 1.5/2.0, and any Ollama/OpenRouter model.

Provider	Env var
Anthropic (Claude)	`TOKENMIZER_ANTHROPIC_API_KEY`
OpenAI	`TOKENMIZER_OPENAI_API_KEY`
Google Gemini	`TOKENMIZER_GEMINI_API_KEY`
DeepSeek	`TOKENMIZER_DEEPSEEK_API_KEY`
Mistral	`TOKENMIZER_MISTRAL_API_KEY`
Grok (xAI)	`TOKENMIZER_GROK_API_KEY`
Cohere	`TOKENMIZER_COHERE_API_KEY`
OpenRouter	`TOKENMIZER_OPENROUTER_API_KEY`
Ollama	No key — free, local

Configuration

# tokenmizer.yaml
provider: anthropic
default_model: claude-sonnet-4-6

graph_checkpoint:
  enabled: true
  trigger_at_percent: 0.85
  use_llm_extraction: false     # true = 80%+ recall, needs key (~$0.001/turn)

compression:
  enabled: true

cache:
  enabled: true
  max_size: 10000

state_backend: memory           # memory | redis (production)

All settings via env vars: TOKENMIZER_PROVIDER, TOKENMIZER_API_KEY, etc.

Docker

# Quick start
docker-compose up tokenmizer

# With Redis (production)
ANTHROPIC_API_KEY=sk-ant-... docker-compose up

# With proxy auth
TOKENMIZER_API_KEY=strong-key docker-compose up

API Reference

Endpoint	Method	Description
`/v1/chat/completions`	POST	OpenAI-compatible proxy
`/api/resume/{id}`	GET	Get resume context
`/api/checkpoint`	POST	Manual checkpoint
`/api/decision/invalidate`	POST	Mark decision as invalid
`/api/graph/{id}`	GET	Session graph stats
`/api/stats`	GET	Token savings analytics
`/health`	GET	Health check
`/docs`	GET	Swagger UI

Security

API key auth — TOKENMIZER_API_KEY (constant-time comparison)
Secret/PII redaction applied once at ingestion, before graph storage, checkpoint storage, AND every LLM call (main chat and the background extraction model — these are separate, the redaction gap between them was a real bug, now fixed)
Session-isolated cache (sensitive data never shared across sessions)
Basic prompt-injection keyword filter — catches copy-pasted jailbreak templates only; not a security boundary against a motivated adversary. See SECURITY.md for exactly what it does and doesn't catch.
CORS restricted to configured origins by default

Benchmarks

python benchmarks/checkpoint_accuracy/runner_v2.py
pytest tests/ -v

Benchmark v2 — Graph vs plain Summary (3 sessions, heuristic-only, measured 2026-07-02 on v0.2.4):

Method	Task Recall	Decision Recall	File Recall	Info Preserved
TokenMizer Graph	76%	85%	100%	87%
Plain Summary baseline	76%	70%	92%	79%
Δ advantage	0%	+15%	+8%	+8%

Avg resume size: 254 tokens vs ~1,500+ tokens of raw history. (n=3 synthetic sessions — small sample; treat as directional, reproduce with the command above.)

Enable use_llm_extraction: true for hybrid extraction (LLM + heuristic merge).

On LLM/hybrid recall numbers — read this before trusting any percentage here: earlier versions of this README quoted "90-100% hybrid recall" sourced from runner_v3.py's MockLLMProvider. That mock sampled its fake output directly from the same ground-truth dict used to score recall — circular by construction, guaranteed to look good regardless of what the real extraction logic did. It measured nothing about actual LLM extraction quality. That number has been removed rather than replaced with a better-sounding one we can't back up.

What runner_v3.py now actually does:

Default mode verifies HybridExtractor.merge()'s logic contract against fixtures with deliberately known overlap (corroborated / LLM-only / heuristic-only items) — confirms merge never drops an item either source found, and applies confidence tiers (0.95 corroborated, 0.80 LLM-only, 0.65 heuristic-only) correctly. This is a real, non-circular check, but it's a logic-contract test, not a recall measurement.
--live mode calls a real configured provider (ANTHROPIC_API_KEY or OPENAI_API_KEY) and scores its actual output against ground truth. This is the only path that produces a number meaningful enough to put in a table. Run it yourself — we're not publishing a live-mode number here because n=3 sessions is too small a sample to generalize, and publishing one without a large, ongoing benchmark would just be swapping one unsubstantiated number for another.

Heuristic-only numbers above (76-100%) ARE real, deterministic, reproducible measurements — runner_v2.py runs actual heuristic extraction against actual ground truth with no LLM and no mocking involved, which is why those numbers are presented with confidence and the LLM ones currently are not.

Why TokenMizer and not X?

Engineers ask this every time. Honest answers:

Why not just use Git history? Git stores what changed, not why you decided to change it. You can't ask Git "what did we decide about auth?" or "why did we switch from MySQL to PostgreSQL?" TokenMizer stores decisions with trigger, reason, and evidence — not diffs.

Why not RAG (retrieval-augmented generation)? RAG retrieves relevant chunks — it doesn't model decision state. If you switched from bcrypt to Argon2 mid-session, RAG might retrieve both and confuse the model about which is current. TokenMizer tracks decision supersession explicitly: old decision is marked SUPERSEDED, new decision is ACTIVE. Resume context only includes current state.

Why not a plain summary at the start of each session? Summaries lose structure. You can't query "all superseded decisions" or "what triggered the auth change" from a blob of text. Our benchmark shows graph memory preserves +5% more information than a summary baseline — and unlike summaries, the graph is queryable, editable, and grows incrementally without re-summarizing everything each turn.

Why not Mem0 or Zep? Mem0 and Zep store facts ("user prefers Python"). TokenMizer stores decisions with rationale — the full causal chain: what was decided, what replaced it, why, what evidence triggered the change, and how confidence shifted. If you need "remember my name across sessions," use Mem0. If you need "remember that we switched from PostgreSQL to SQLite because of cost, and here's the evidence," use TokenMizer.

Why not just a longer context window? Longer context = higher cost + slower inference + model attention dilution on long histories. TokenMizer compresses a 50-turn session into ~246 tokens of structured context — not by summarizing, but by extracting what actually matters: goals, active decisions, current tasks, recent errors.

CLI

tokenmizer serve [--port 8000]
tokenmizer checkpoint <session-id>
tokenmizer resume <session-id> [--level standard|full|critical]
tokenmizer stats

Note on file analysis: /tokenmizer:analyze (used from inside Claude Code, see Claude Code Integration above) is real and works — it's a plugin skill (.claude-plugin/skills/analyze/) that calls FileIntelligence directly via an inline Python snippet, independent of the CLI/API layer. What does not exist is a bare tokenmizer analyze <file> terminal command or a /api/analyze HTTP endpoint — useful if you want file analysis from a plain shell or a non-Claude-Code tool (Cursor, a script, curl, etc.) rather than inside Claude Code specifically. Found during a documentation accuracy pass: an earlier version of this README listed tokenmizer analyze <file> in this CLI section as if it were a cli.py command — it never was. Removed from here rather than left in place pointing at something that would fail. Tracked as a real, wanted gap — contributions adding a /api/analyze endpoint + thin CLI wrapper (following the existing pattern in cli.py) are welcome.

Contributing

See CONTRIBUTING.md. Graph extraction contributions are the highest priority.

git clone https://github.com/Shweta-Mishra-ai/tokenmizer
pip install -e ".[dev]"
pytest tests/ -v && ruff check tokenmizer/

License

MIT © Shweta Mishra

_{Built for developers who spend too much time re-explaining their projects to AI.}

Project details

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

Shweta-Mishra-ai

These details have not been verified by PyPI

Release history Release notifications | RSS feed

0.3.1

Jul 3, 2026

0.3.0

Jul 3, 2026

0.2.6

Jul 2, 2026

This version

0.2.5

Jul 2, 2026

0.2.4

Jul 2, 2026

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

tokenmizer-0.2.5.tar.gz (192.7 kB view details)

Uploaded Jul 2, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

tokenmizer-0.2.5-py3-none-any.whl (128.9 kB view details)

Uploaded Jul 2, 2026 Python 3

File details

Details for the file tokenmizer-0.2.5.tar.gz.

File metadata

Download URL: tokenmizer-0.2.5.tar.gz
Upload date: Jul 2, 2026
Size: 192.7 kB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for tokenmizer-0.2.5.tar.gz
Algorithm	Hash digest
SHA256	`33d259832a1d3654a21de96dabb58d1c337a28d187eb09459a69f34d650767a7`
MD5	`cc4e017045fa1377d6d8849f1655d8dc`
BLAKE2b-256	`8dd35e5cf7948442f5d52101f981dea7bd6cab68eba18d20b03bf2f5e3e6a3c6`

See more details on using hashes here.

Provenance

The following attestation bundles were made for tokenmizer-0.2.5.tar.gz:

Publisher: release.yml on Shweta-Mishra-ai/tokenmizer

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: tokenmizer-0.2.5.tar.gz
- Subject digest: 33d259832a1d3654a21de96dabb58d1c337a28d187eb09459a69f34d650767a7
- Sigstore transparency entry: 2047895561
- Sigstore integration time: Jul 2, 2026
Source repository:
- Permalink: Shweta-Mishra-ai/tokenmizer@aaaa90262619bfa1cefb01420f8dbe58bed29fe4
- Branch / Tag: refs/tags/v0.2.5
- Owner: https://github.com/Shweta-Mishra-ai
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: release.yml@aaaa90262619bfa1cefb01420f8dbe58bed29fe4
- Trigger Event: release

File details

Details for the file tokenmizer-0.2.5-py3-none-any.whl.

File metadata

Download URL: tokenmizer-0.2.5-py3-none-any.whl
Upload date: Jul 2, 2026
Size: 128.9 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for tokenmizer-0.2.5-py3-none-any.whl
Algorithm	Hash digest
SHA256	`99dfd72be255990a4e5459e628508fcc74881f9455ba7764b0d0d7359b2fe9e0`
MD5	`afb609a153289509a9495777aab7b579`
BLAKE2b-256	`8a890447967e17c3629ab006aa2c5fa603dacee6898293ef3b00321314320f43`

See more details on using hashes here.

Provenance

The following attestation bundles were made for tokenmizer-0.2.5-py3-none-any.whl:

Publisher: release.yml on Shweta-Mishra-ai/tokenmizer

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: tokenmizer-0.2.5-py3-none-any.whl
- Subject digest: 99dfd72be255990a4e5459e628508fcc74881f9455ba7764b0d0d7359b2fe9e0
- Sigstore transparency entry: 2047895578
- Sigstore integration time: Jul 2, 2026
Source repository:
- Permalink: Shweta-Mishra-ai/tokenmizer@aaaa90262619bfa1cefb01420f8dbe58bed29fe4
- Branch / Tag: refs/tags/v0.2.5
- Owner: https://github.com/Shweta-Mishra-ai
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: release.yml@aaaa90262619bfa1cefb01420f8dbe58bed29fe4
- Trigger Event: release

tokenmizer 0.2.5

Navigation

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Meta

Classifiers

Project description

TokenMizer

The Problem

How TokenMizer Solves It

Architecture

Decision Memory — 4-State Model

Quick Start

1. Install

2. Set your API key

3. Start

4. Use — change one line

Claude Code Integration

Option A — Plugin (recommended)

Option B — MCP server

Other Tools

Session Resume

File Intelligence

Works Alongside Caveman & CodeBurn

Supported Providers

Configuration

Docker

API Reference

Security

Benchmarks

Why TokenMizer and not X?

CLI

Contributing

License

Project details

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

Provenance

File details

File metadata

File hashes

Provenance