Skip to main content

TOKENFOLD

Send less noise. Fit more context. Pay for fewer input tokens.

Up to 92% fewer input tokens · lossless by default · exact receipts

CLI · Python · TypeScript · Rust · proxy · MCP · local-first · provider-neutral

CI Coverage GitHub Release PyPI npm Rust License

Measured results · Platform · Quick start · Lossy pruning · Select model · Reproduce


Measured results

Compression

Search / RAG results Repetitive JSON API responses Tool schemas
91.7% fewer tokens 67.6% fewer tokens 63.9% fewer tokens 45.6% fewer tokens
84,032 → 6,951 50-record payload 30-record payload 1.8 MB OpenAI fixture
opt-in, recoverable lossless lossless lossless

All four use exact o200k_base counts in balanced preset. The 91.7% showcase keeps the planted 503 result inline and makes the other 98 long rows retrievable by hash. Run it with python examples/lossy_pruning.py. The versioned inputs and source provenance for these headline figures live in tests/fixtures/readme_metrics.json.

Fine-tuned relevance ranking

Tokenfold Select is an optional external query-aware model for choosing what fills a tight context window after structural compression ends. It is distributed and evaluated separately; Tokenfold Core does not invoke it.

Token budget kept Tokenfold Select task success BM25 reference Lift
50% 86.3% 79.4% +8.7%
25% 70.3% 60.7% +15.8%
10% 39.9% 37.7% +5.8%

At a 25% budget, the fine-tuned ranker keeps the answer 70% of the time and beats the strongest free heuristic by 15.8%. Critical-content survival is 100% at every measured budget through allocator force-keep. Results are three-seed repeated subsampling over roughly 73,000 training and 24,000 held-out fixtures per run; the model card publishes the training recipe and full baseline set. Here, task success means the literal gold answer survived selection under the stated budget. These are source-reported external results, pinned by model revision in the metric manifest; Core CI does not reproduce them.

One platform, two engines

Capability What you get
Tokenfold Core Fast deterministic compression for JSON, schemas, provider requests, logs, diffs, and command output
Recoverable pruning Drop low-signal JSON rows only after storing them locally; fetch any omission with tokenfold retrieve
Tokenfold Select LoRA-fine-tuned, query-aware span ranking with up to 15.8% lift over BM25 at the same budget
Budget control Conservative, balanced, and aggressive presets; nine task scopes; --target-tokens with honest best-effort reporting
Exact receipts Before/after token counts, applied transforms, warnings, provenance, and final status on every call
Realized savings gain, stats, and session report measured savings; learn proposes policy improvements without silently applying them
Every integration CLI, Python, TypeScript, Rust, HTTP proxy, and MCP share the same Core engine and report shape
Local-first safety Secret redaction, protected-content gates, reversible structural transforms, and no hosted data processor

Lossless or recoverable lossy

Lossless — default Recoverable lossy — opt-in
What it does Minifies, folds repeated keys into columns, and stores repeated values once Adds deterministic ranking and retrieval markers for selected array rows
Best on Uniform records, schemas, logs, diffs Search results, mixed event feeds, agent traces, long arrays
Measured here 45.6–67.6% fewer tokens 51.7–91.7% fewer tokens
Recovery All data remains in the payload Every emitted marker resolves through the local retrieval store

Quick start

Install the surface that fits your stack:

pip install tokenfold       # Python 3.9+
npm install tokenfold       # Node.js 22+
cargo add tokenfold-core    # Rust library
cargo install tokenfold-cli # Rust CLI

Or download a checksummed CLI build for Linux, macOS, or Windows from GitHub Releases and verify it with the adjacent .sha256 file.

Preview first, then compress:

tokenfold inspect payload.json --format json
tokenfold compress payload.json --format json --output payload.compact.json

Python uses the same Core engine and typed receipt:

import json
from pathlib import Path

from tokenfold import InputFormat, Preset, compress

result = compress(
    Path("request.json").read_bytes(),
    format=InputFormat.OPENAI_JSON,
    preset=Preset.BALANCED,
)
compressed_request = json.loads(result.payload)
print(f"saved {result.report.saved_tokens} tokens ({result.saved_pct():.1f}%)")
Surface Best for Start here
CLI Files and stdin tokenfold compress, inspect, decode, retrieve
Python Applications and evaluation pipelines pip install tokenfold
TypeScript Node.js applications and automation npm install tokenfold
Rust Native embedding cargo add tokenfold-core
HTTP proxy Transparent provider-shaped traffic Build tokenfold-proxy
MCP Agents and editors tokenfold mcp serve

tokenfold init --agent claude-code merges a project-scoped .mcp.json entry without replacing other servers; tokenfold doctor --agent claude-code verifies it. Codex can use the same stdio server with codex mcp add tokenfold -- tokenfold mcp serve. See docs/configuration.md for tested MCP JSON/TOML and every environment override. Trusted filters for Git, build, and test output are available through tokenfold filters list.

Runnable examples, one per surface, all under examples/:

python examples/quickstart.py      # Python: compress messages, request bodies, JSON data
node examples/quickstart.mjs       # TypeScript/Node: compress, inspect, read the receipt
cargo run -p tokenfold-core --example quickstart   # Rust: the embedded core API
python examples/lossy_pruning.py   # CLI: opt-in recoverable pruning, end to end

examples/quickstart.ipynb is the guided tour of the whole Python surface: everything quickstart.py shows, plus previews, budgets and presets, provider payloads, recoverable pruning with retrieval, and where Tokenfold Select fits.

Recoverable lossy pruning

Lossless folding has a ceiling on heterogeneous arrays. --prune ranks rows, keeps the strongest signals, and replaces selected rows with compact {"$tf_ref": {...}} handles. A row leaves the payload only after the local store accepts it.

tokenfold compress examples/incident_feed.json --format json \
  --prune --keep-ratio 0.35 --output feed.compact.json

On the bundled 40-event feed:

Mode Exact tokens Reduction Events kept Incident kept
Lossless 6,144 14.9% 40 ✅
--keep-ratio 0.35 3,485 51.7% 13 ✅
--keep-ratio 0.05 2,294 68.2% 1 ✅

The planted 503 with success: false and retries: 7 survives every shown setting because typed failure signals outrank position and length — at --keep-ratio 0.05 it is the only row kept. Long rows amortize marker overhead further: the 100-result showcase reaches 91.7%.

Fetch a dropped row:

tokenfold retrieve cb13cc59cca0c218c579cd1d4b3cbab58d6dea265eb995cc9c00faf0cd0a6856
# {"seq":1,"ts":"2026-08-15T00:01:11Z","subsystem":"index-writer",...}

What the flags mean:

  • --keep-ratio is an aggression hint over eligible array items. Lower keeps fewer rows; it is not a whole-document guarantee.
  • --target-tokens is the whole-document goal. Tokenfold stops when it reaches the target losslessly and reports best_effort when the safe transform set cannot reach it (unreachable when protected content alone exceeds the target).
  • --preserve <path> protects a named array; nested paths conservatively protect their nearest eligible ancestor.
  • Generic JSON only: lossy pruning does not run on OpenAI or Anthropic message payloads.
  • Storage is fail-closed: refused rows stay inline, and detected secret-shaped bytes are never persisted.

Preview the projected savings with no store writes:

tokenfold inspect examples/incident_feed.json --format json \
  --prune --keep-ratio 0.35

The same flags, and the same fail-closed contract, from Python and TypeScript:

from tokenfold import InputFormat, PruningPolicy, compress, retrieve

pruning = PruningPolicy(keep_ratio=0.35)
result = compress(feed_bytes, format=InputFormat.JSON, pruning=pruning)
original = retrieve(marker["$tf_ref"])   # any dropped row, verbatim
import { compress, retrieve } from "tokenfold";

const { payload, report } = await compress(feed, {
  format: "json",
  pruning: { keepRatio: 0.35 },
});
const original = await retrieve(hash); // any dropped row, verbatim
Current Phase 1 constraints

Treat $tf_ref as reserved and do not enable lossy pruning on documents that already contain retrieval markers. Filesystem entries are published under a cross-process lock, and a refused batch leaves every candidate inline.

Preview is a projection rather than a filesystem transaction, so a real run may keep more rows if storage becomes unavailable.

Tokenfold Select

When structure ends, rank what matters.

Tokenfold Select is an Apache-2.0 LoRA adapter on ibm-granite/granite-embedding-reranker-english-r2. It scores candidate spans against a query; your allocator applies the token budget and force-keeps required content. It is a separately distributed companion: Core remains model-free and deterministic and neither loads nor invokes Select.

Tokenfold Core Tokenfold Select
Best at Structural compression Query-conditioned span ranking
Runtime Static Rust binary Granite reranker + LoRA adapter
Output Compressed payload + exact receipt Ranking logits
Python: load the model and score spans
from pathlib import Path

import torch
from huggingface_hub import snapshot_download
from peft import PeftModel
from transformers import AutoModelForSequenceClassification, AutoTokenizer

base_id = "ibm-granite/granite-embedding-reranker-english-r2"
repo_dir = Path(snapshot_download("snchimata/tokenfold-select"))
adapter_dir = repo_dir / "adapter"
tokenizer = AutoTokenizer.from_pretrained(adapter_dir)
base = AutoModelForSequenceClassification.from_pretrained(
    base_id,
    dtype=torch.float32,
)
model = PeftModel.from_pretrained(base, adapter_dir).eval()

def score(query: str, spans: list[str]) -> list[float]:
    if not spans:
        return []
    encoded = tokenizer(
        [query] * len(spans),
        spans,
        padding=True,
        truncation=True,
        max_length=8192,
        return_tensors="pt",
    )
    with torch.no_grad():
        output = model(
            input_ids=encoded["input_ids"],
            attention_mask=encoded["attention_mask"],
        )
    return output.logits.view(-1).float().tolist()

See the Tokenfold Select model card for setup, evaluation, training data, and limitations.

Safety and auditability

  • Never larger: Core keeps a transform only when exact recounting shows a reduction; a lossy branch must also beat the lossless result.
  • Reversible structure: JSON folds must pass an exact round trip.
  • Protected content: provider system messages and latest-user content are held behind format-aware safety gates.
  • Clear provenance: exact tokenizer counts, heuristics, and extrapolations are labeled separately.
  • Actionable receipts: every result lists savings, transforms, warnings, retrieval state, and final status.
  • Local control: detected secrets are redacted before reports or storage; policy learning changes configuration only with --apply.

The offline fidelity gate uses deterministic lexical-overlap, critical-token, and containment proxies. Those checks catch regressions but do not establish semantic equivalence or downstream task success. Runtime quality fields are absent unless a versioned evaluator has supplied data; applications should run their own representative task evaluation before enabling lossy pruning.

Reproduce the results

# 91.7% maximum-compression showcase, mixed-feed curve, retrieval, and preserve
python examples/lossy_pruning.py

# Lossless transform benchmarks
cargo bench -p tokenfold-core

# Small exact-token CLI example
cargo run --release --locked -p tokenfold-cli -- \
  inspect examples/api_response.json --format json

The bundled 30-record API response reports 3,812 → 1,376 tokens, a 63.9% lossless reduction. Benchmark sources and thresholds live in CHANGELOG.md and crates/tokenfold-core/benches/THRESHOLDS.toml.

Contributing

Issues and pull requests are welcome. Run the relevant checks before opening a PR:

cargo fmt --all --check
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspace --locked
python eval/run_fidelity.py --gate --profile smoke-first-consumer
cd packages/tokenfold && npm ci && npm test

License

Apache-2.0


Start with one representative payload, inspect the receipt, and see how many tokens your application can stop sending today.

pip install tokenfold

If Tokenfold earns a place in your stack, a ⭐ on GitHub helps the next team find it.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

tokenfold-0.5.0.tar.gz (154.7 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

tokenfold-0.5.0-cp39-abi3-win_amd64.whl (3.1 MB view details)

Uploaded CPython 3.9+Windows x86-64

tokenfold-0.5.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (3.4 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ x86-64

tokenfold-0.5.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (3.3 MB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ ARM64

tokenfold-0.5.0-cp39-abi3-macosx_11_0_arm64.whl (3.2 MB view details)

Uploaded CPython 3.9+macOS 11.0+ ARM64

tokenfold-0.5.0-cp39-abi3-macosx_10_12_x86_64.whl (3.2 MB view details)

Uploaded CPython 3.9+macOS 10.12+ x86-64

File details

Details for the file tokenfold-0.5.0.tar.gz.

File metadata

  • Download URL: tokenfold-0.5.0.tar.gz
  • Upload date:
  • Size: 154.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for tokenfold-0.5.0.tar.gz
Algorithm Hash digest
SHA256 d6e0f997dbd0ecde219533bcfd3345d17ad35a8e51fda7e7b5d8e748f68f0061
MD5 2891e80b99ad2f7d2e3815298f123877
BLAKE2b-256 e66e4633918e3dab7313b9a6cec54a85c0d3a78460564309e6f02d5cb9a778cf

See more details on using hashes here.

File details

Details for the file tokenfold-0.5.0-cp39-abi3-win_amd64.whl.

File metadata

  • Download URL: tokenfold-0.5.0-cp39-abi3-win_amd64.whl
  • Upload date:
  • Size: 3.1 MB
  • Tags: CPython 3.9+, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for tokenfold-0.5.0-cp39-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 e5d3196a104d626b631c25aa7a60cd92eb31a2413803ecc808b550ea39b54f3e
MD5 719af2de2d30e676a73c5caba3c232de
BLAKE2b-256 1218ab02fcbb1b1387d14c326b8b9a8b4a92b4d74a765244b1c8fccd2a7b75ad

See more details on using hashes here.

File details

Details for the file tokenfold-0.5.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for tokenfold-0.5.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 3e5d331dd6d953fb6f26262bee1618f01497a9369f76f8c80e098f13766a2dcb
MD5 019c1a12aab24b22a59cac94b877bf30
BLAKE2b-256 6a7ee0fe19efabd7cf400df4ee1f9a11ac38f702e523aae54a7155dd4905c9c8

See more details on using hashes here.

File details

Details for the file tokenfold-0.5.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for tokenfold-0.5.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 ec9ccc4fa67dd0da0345582e8340f860fac2999ede37dfadc19e27fa3a166574
MD5 f12ecdaf86d67ed7648c17832dc6e052
BLAKE2b-256 a3b1de6afad330c3d4759c3c138ce7c0781dffbf199b09aafcba08e9ab9f0b0c

See more details on using hashes here.

File details

Details for the file tokenfold-0.5.0-cp39-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for tokenfold-0.5.0-cp39-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 95c8e9b6bb29626d310b4948bec9875199ddeb365572f40f118cf96ae9d1a10b
MD5 670d7424346a425157fd197190b238c9
BLAKE2b-256 f266662991b02ed9e201abbf47af8673656515765bc6880e60a72d9e916c384c

See more details on using hashes here.

File details

Details for the file tokenfold-0.5.0-cp39-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for tokenfold-0.5.0-cp39-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 22573074a41b1fe645f62558ec14d448419428fd73f5288c824ad193eaa1444c
MD5 40aed483ad9e0628cb142f43ec8dba26
BLAKE2b-256 ab61b90a00dc3783297f51ec8f55621870471f423cadfeedc997e01b07f9b8c1

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.5.0 This release

6 files

0.4.1

6 files

0.4.0

6 files

0.3.4

6 files

0.3.3

6 files

0.3.2

6 files

0.3.1

6 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page