Skip to main content

tokentoken

License: MIT Python 3.9+

Compile human prompts into dense LLM machine language — save up to 70% on token costs.

tokentoken is a Python SDK and CLI that compresses verbose text into a minimal, LLM-native representation that any capable model can interpret without a codebook — and recovers it on demand.

Based on the research paper Large Language Models Do Not Always Need Readable Language (arXiv:2606.19857).

Features

  • 14 compression modes — every prompt variant from the paper (default + bt_p1–bt_p13)
  • 4 providers — OpenAI, Google Gemini, Anthropic, and local Ollama
  • Cross-model readable — text compressed by one LLM can be read by another
  • Built-in analytics — token savings, retention ratio, readability and efficiency metrics
  • Agentic workflows — agent memory compression, multi-agent messaging, context-window extension
  • Clean CLI — inline text or file input, one-line errors, meaningful exit codes

Installation

Requires Python 3.9+.

git clone https://github.com/thesumitbanik/tokentoken.git
cd tokentoken
pip install .

For development:

pip install -e ".[dev]"

Configuration

API keys are read from environment variables — never hardcode them:

export GEMINI_API_KEY="..."      # default provider
export OPENAI_API_KEY="..."
export ANTHROPIC_API_KEY="..."
# Ollama needs no key (local, http://localhost:11434)
Provider --provider Environment variable Example --model
Google Gemini gemini (default) GEMINI_API_KEY gemini-3.5-flash (default)
OpenAI openai OPENAI_API_KEY gpt-4
Anthropic anthropic ANTHROPIC_API_KEY claude-3-5-sonnet
Ollama ollama — llama3

OpenAI-compatible endpoints

Point the openai provider at any OpenAI-compatible server — vLLM, LM Studio, LocalAI, OpenRouter, or Ollama's /v1 API:

export OPENAI_API_KEY="not-needed"   # placeholder for keyless local servers
tokentoken compress --text "..." --provider openai --model my-model \
  --base-url http://localhost:1234/v1
tt = TokenToken(provider="openai", model="my-model", base_url="http://localhost:1234/v1")

No flag needed if you prefer environment configuration: the OpenAI client reads OPENAI_BASE_URL automatically.

Quick Start

Python SDK

from tokentoken import TokenToken, CompressionMode

# Uses GEMINI_API_KEY from the environment (or pass api_key="...")
tt = TokenToken(provider="gemini", model="gemini-3.5-flash")

result = tt.compress("Your verbose text here...", mode=CompressionMode.DEFAULT)

print(f"Original:   {result.original_tokens} tokens")
print(f"Compressed: {result.compressed_tokens} tokens")
print(f"Reduction:  {result.savings_pct}%")

# Any capable LLM can interpret it — no codebook required
answer = tt.decompress(result.dense_text, question="What is the main topic?")

CLI

# Compress a file
tokentoken compress input.txt --out dense.txt

# Or inline text — no file needed
tokentoken compress --text "Your verbose text here..."

# Read it back
tokentoken decompress dense.txt --question "What is the main topic?"

# Explore
tokentoken --help
tokentoken list-modes

CLI Reference

Run any command with --help for full options.

Command Description
compress Compress a file or inline --text; writes --out (default dense.txt)
decompress Interpret compressed text; optional --question to answer from it
analyze Measure compression potential; add --compressed to compare against an output
list-modes List all 14 compression modes with paper references
estimate Count tokens in a file
agent-memory Compress agent conversation histories from a directory
multi-agent Compress an inter-agent message between --sender and --receiver
cross-model Verify compressed text remains readable by a different model
extend-context Compress a long document down to a --target-tokens budget

Common options: --provider (default gemini), --model (default gemini-3.5-flash), --out/-o, --mode (for compress), and --base-url (OpenAI-compatible endpoints).

Failures print a single-line error and exit with code 1 — never a traceback.

Python API

TokenToken(provider, model, api_key=None, host=None, base_url=None)

Method Returns Description
compress(text, mode=CompressionMode.DEFAULT) CompressionResult Compress text into dense form
decompress(text, question=None) str Interpret compressed text / answer a question
compress_for_agent_memory(memories, session_id=None) AgentMemoryResult Compress conversation histories
compress_for_multi_agent(message, sender, receiver) MultiAgentMessage Compress inter-agent messages
check_cross_model_compatibility(text, reader_provider, reader_model) CrossModelCompatibility Test readability across models
extend_context_window(long_text, target_tokens=200000) str Compress long documents to a token budget
get_compression_analytics(original, compressed) dict Detailed compression metrics

Utility functions are exported at package level: count_tokens, check_breakeven_threshold, chunk_text, calculate_compression_metrics, estimate_readability, calculate_token_efficiency, validate_compression_result.

CompressionResult

Field Type Description
dense_text str The compressed output
original_tokens int Input token count
compressed_tokens int Output token count
savings_pct float Token reduction in percent
retention_ratio float Output size ÷ input size
readability_metrics dict Dale-Chall estimate, difficult-word ratio, word count
efficiency_metrics dict Tokens per word, chars per token, compression potential
validation dict Quality checks + chain-of-thought tax warning

Compression Modes

Select with --mode bt_p7 (CLI) or mode=CompressionMode.BT_P7 (SDK).

Mode Paper ref Description
default C.1 Default compression prompt
bt_p1 C.2.1 Adaptive Symbolic Collapse
bt_p2 C.2.2 Refined Zero-Overhead Compression
bt_p3 C.2.3 Minimal Lossless Objective
bt_p4 C.2.4 Structured Omnilingual Mapping
bt_p5 C.2.5 Canonical Omnilingual-Symbolic
bt_p6 C.2.6 Structured Mapping Control
bt_p7 C.2.7 Canonical Compression Objective
bt_p8 C.2.8 Fixed Symbolic Mapping Rules
bt_p9 C.2.9 Structured Semantic Mapping
bt_p10 C.2.10 LLM-Native Compressor
bt_p11 C.2.11 Compact Symbolic Mapping
bt_p12 C.2.12 Free-Emergence Attention Checklist
bt_p13 C.2.13 ASCII Anchor Skeleton

How It Works

---
config:
  look: handDrawn
  theme: neutral
---
flowchart LR
    A["Verbose human text"] --> B["Compress: paper prompt + your LLM"]
    B --> C["Dense text (~30% of tokens)"]
    C --> D["Decompress: any capable LLM"]
    D --> E["Answers / full meaning"]

Compression swaps verbose natural language for high-density multilingual and symbolic forms. No decoder, codebook, or fine-tuning is involved — interpretation is a capability the reading LLM already has.

Use Cases

  • Document QA - compress long documents while preserving semantic fidelity for question answering
  • Agent memory - compress conversation histories to reduce storage while maintaining reliable recall
  • Multi-agent communication - cut context overhead between collaborating agents
  • Context window extension - handle documents that exceed model limits by compressing chunks

Contributing

git clone https://github.com/sumitbanik/tokentoken.git
cd tokentoken
pip install -e ".[dev]"
pytest

Issues and pull requests are welcome.

License

MIT — see LICENSE.

Metadata

Release files for tokentoken 1.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokentoken 1.0.1
File Size Uploaded
tokentoken-1.0.1.tar.gz 31.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tokentoken 1.0.1
File Interpreter ABI Platform
tokentoken-1.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 55.4 kB

Release files / tokentoken-1.0.1.tar.gz

Download URL tokentoken-1.0.1.tar.gz
Size 31.7 kB
Tags Source
SHA-256 checksum
How to use checksums
779a61951ea6b20b1725f4b4300b560bb0c7b1bb8687bfd54d3aff1444ef0e28
BLAKE2b-256 checksum
How to use checksums
9f5dd1bbfe22a91918d5b7f3ef848d09911820043f47f497338f24657efe42be
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.11

Release files / tokentoken-1.0.1-py3-none-any.whl

Download URL tokentoken-1.0.1-py3-none-any.whl
Size 23.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
21d78b06ce4df08a6f0b7916d210b8f0b97a5f48b1ca0f00f2de1d813e7a6eea
BLAKE2b-256 checksum
How to use checksums
26f4e2e7e67d97727f02db317c97d4c2e900bb0967916da2e9d95687ae57cd4f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.11

Release history Release notifications | RSS feed

1.0.2

2 release files

This release

1.0.1 This release

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page