Skip to main content

HyperToken

Atom-level token optimization for LLM context windows.

PyPI Python License

pip install hypertok
from hypertok import HyperToken

Zero required dependencies. Pure Python. Works offline.

The project is HyperToken; the package is hypertok. The shorter name on PyPI belongs to an unrelated project that also installs a top-level hypertoken/ module, so sharing it would let the two overwrite each other.


The problem

Every context-window compressor works on windows: slide a fixed span across the document, score it, keep the best ones. That is convenient and it throws away the thing that makes a document a document. A 512-token window cuts a function in half, separates a table row from its caption, and drops the heading that told you what the paragraph was about.

HyperToken works one level down, on atoms — the smallest structural element the source actually has:

Source An atom is
PDF / plain text one heading, paragraph, table row, footnote, citation
Markdown one heading, list item, fenced block, table row, quote
Source code one function or method, with its scope recorded
JSON / API payload one leaf value, at its full path
HTML one block-level element

…and one level down again, on units — the sentences inside an atom. The optimizer selects atoms, then trims units inside the atoms it keeps. Nothing is cut mid-structure, and every emitted token is traceable to where it came from.

Results

python benchmarks/bench.py — 16 queries over 3 documents, measuring the fraction of gold answer spans that survive compression:

Budget truncate (head) truncate (tail) fixed chunks + BM25 HyperToken
80 tokens 12.5% 25.0% 50.0% 93.8%
150 tokens 37.5% 37.5% 81.2% 93.8%
300 tokens 68.8% 75.0% 87.5% 100.0%
600 tokens 100.0% 100.0% 87.5% 100.0%

At 300 tokens HyperToken retains every answer while emitting fewer tokens (242 avg) than the chunk baseline (259 avg) that retains 87.5%. The budget is a hard ceiling: across all presets and budgets, zero runs overshot.

Quick start

from hypertok import HyperToken

ht = HyperToken(preset="rag")

result = ht.optimize(
    open("contract.txt").read(),
    budget=2000,
    query="termination notice period",
)

print(result.text)     # the optimized context, ≤ 2000 tokens
print(result.stats)    # 18420 -> 1987 tokens (16433 saved, 89.2% reduction)

One-liner:

from hypertok import optimize

text = optimize(document, budget=1000, query="what changed in v4.2?").text

No budget, no dropping — just squeeze out the waste:

ht.compress(scraped_html).text     # every atom survives, each one smaller

Several documents into one shared budget:

ht.optimize_many([doc_a, doc_b, doc_c], budget=4000, query=q, weights=[3, 1, 1])

Command line

hypertok optimize report.pdf --budget 2000 --query "revenue by region" --stats
hypertok compress notes.md --preset aggressive -o small.md
hypertok atoms contract.txt -n 20        # see how a document decomposes
hypertok count report.pdf --tokenizer tiktoken
hypertok presets
hypertok · preset=rag · tokenizer=heuristic
  3,309 → 194 tokens   █·······················  94.1% saved
  budget 200 · within budget · headroom 6
  atoms: 8 kept, 1 compressed, 81 dropped
  by structure label:
    body                 2,562 →    111  █··············· (32 atoms)
    citation               332 →      0  ················ (15 atoms)
    heading                180 →     36  ███············· (27 atoms)

How it works

Five stages. Each one is a replaceable module.

 source ──▶ atomize ──▶ graph ──▶ score ──▶ select ──▶ compress ──▶ render
             │            │         │          │           │
     structure-aware   reading   relevance   budgeted    sentence
        atoms with      order    density     knapsack    trimming
        type labels   headings   structure   + support   inside kept
                     references  redundancy   pull-in       atoms

1 · Atomize. Structure-aware decomposition. For markup the structure is read off the syntax; for PDF-extracted text it is inferred with heuristics that recover headings, captions, footnotes, citations, table rows and list items — the label vocabulary from the SAGER corpus study. Hard-wrapped lines are rejoined and words broken across a line break are healed, which is free tokens before anything else runs.

2 · Graph. Atoms are linked by reading-order adjacency, heading containment, and cross-references ("see Table III" → the Table III caption).

3 · Score. Four signals per atom: BM25 query relevance with conservative stemming, information density, a structural prior per label, and a SimHash near-duplicate penalty. Every component is exposed for audit.

4 · Select. A value-density greedy pass with three corrections that matter more than exact knapsack packing:

  • Support pull-in — admitting an atom admits its heading ancestors and referenced captions, so kept content stays legible.
  • Redundancy rejection — an atom restating admitted content is skipped.
  • Marginal trimming — the first atom that does not fit is compressed into the remaining space rather than dropped.

The plan is then rendered and re-measured; if the real string overshoots, the lowest-value content is dropped and it is rendered again. budget is a guarantee, not a hint.

5 · Compress. Lossless normalization always (Unicode folding, whitespace, decoration). Then, by level, stock-phrase simplification, repeated-phrase abbreviation, sentence trimming, and function-word pruning. Numbers, currency, dates, percentages, URLs, emails, identifiers and code spans are protected at every level, as are negations and modals (not, must, shall, except) — deleting a "not" from a contract clause is not a token saving.

Presets

Preset What it does
safe Lossless only. Never rewords, never trims a sentence. For load-bearing text.
balanced Default. Normalization, phrase simplification, sentence trimming.
rag Relevance-heavy, strong cross-chunk dedup, page annotations for citations.
summary No query attached; density and structure carry selection.
code Never reformats code. Drops whole definitions, not lines.
aggressive Maximum reduction that still reads as English.
extreme Function-word pruning. Telegraphic; models read it, people mostly cannot.

Every field is overridable:

from hypertok import HyperToken, ScoreWeights, SelectorConfig

ht = HyperToken(
    preset="rag",
    weights=ScoreWeights(relevance=0.8, informativeness=0.1, coverage=0.1),
    selector=SelectorConfig(support=False, redundancy_threshold=0.75),
)

The audit trail

Every run explains itself. This is the part that makes atom-level optimization usable in production rather than just smaller.

result = ht.optimize(doc, budget=1000, query="refund policy")

for record in result.dropped[:3]:
    print(record.type, record.tokens_before, record.reason)
# citation  41  not selected
# footnote  22  redundant with 8f2a1c04 (0.94)
# body     118  no budget remaining

result.manifest()   # full JSON: per-atom decision, cost, score, reason
from hypertok.report import format_report, summarize
print(format_report(result))
summarize(result)["by_type"]   # tokens kept vs dropped, per structure label

Tokenizers

The default is a dependency-free estimator calibrated against cl100k_base and biased slightly high, so a plan that fits the estimate fits the real tokenizer. For exact counts:

pip install 'hypertok[tiktoken]'   # or [hf], [pdf], [html], [all]
HyperToken(tokenizer="tiktoken")            # cl100k_base
HyperToken(tokenizer="gpt-4o")              # model name → its encoding
HyperToken(tokenizer="hf:bert-base-uncased")
HyperToken(tokenizer=my_tokenizer)          # anything with .count(str) -> int

Budgets that account for the whole prompt

A context window is shared. Budgeting only the evidence is why prompts overflow.

from hypertok import Budget

budget = Budget(
    total=8000,
    reserve_output=1500,    # room for the model to answer
    reserve_system=400,
    reserve_history=1200,
)
ht.optimize(document, budget, query)   # gets 4900 tokens

Working with the pieces

The stages are usable on their own.

from hypertok import HyperToken, AtomType

ht = HyperToken()
doc = ht.atomize(open("paper.pdf"))          # or a path, string, or list of chunks

for atom in doc.of_type(AtomType.TABLE_ROW):
    print(atom.page, atom.text)

scored = ht.score(doc, query="revenue")
print(scored[0].explain())
# body#12 value=0.8421 (coverage=1.000, density=0.774, redundancy=0.000, ...)
from hypertok import AtomGraph
graph = AtomGraph(doc)
graph.support(atom.id)      # what this atom needs to be legible
graph.descendants(head.id)  # everything under a heading

Optional extras

Extra Adds
hypertok[tiktoken] Exact OpenAI-family token counts
hypertok[hf] Hugging Face tokenizer counts
hypertok[pdf] PDF ingestion via pypdf
hypertok[html] Robust HTML parsing via BeautifulSoup
hypertok[all] All of the above

PDF extraction is fault-tolerant by design: pages that fail are recorded in Document.meta["failed_pages"] and the batch continues, the same posture that let the SAGER corpus study finish 1,059 of 1,076 PDFs without aborting.

Background

HyperToken's data model comes from SAGER — Structure-Aware Graph Evidence Retrieval for Document Intelligence (Meet Jethwa), which processed 1,076 PDFs into 492,530 evidence atoms and showed that an evidence-graph substrate beats flat chunk retrieval on messy real-world corpora (MRR@10 0.167 → 0.812).

That work applied the substrate to retrieval. HyperToken applies it to budgeting — the problem you hit next, once retrieval works and the results no longer fit.

Development

git clone https://github.com/meet2147/hypertoken
cd hypertoken
pip install -e '.[dev,all]'

pytest                          # test suite
ruff check src tests            # lint
python benchmarks/bench.py      # reproduce the results table

Point the corpus harness at a real directory to check ingestion robustness, structure recovery, throughput and budget compliance across many documents:

python benchmarks/corpus_run.py ~/corpus --budgets 500 2000 --json report.json

It reports failures with reasons, flags documents whose atomization looks wrong, and exits non-zero on any budget violation — so it works as a CI gate over a fixture corpus.

License

Apache-2.0. See LICENSE.

Metadata

Release files for hypertok 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for hypertok 0.1.0
File Size Uploaded
hypertok-0.1.0.tar.gz 99.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for hypertok 0.1.0
File Interpreter ABI Platform
hypertok-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 193.2 kB

Release files / hypertok-0.1.0.tar.gz

Download URL hypertok-0.1.0.tar.gz
Size 99.3 kB
Tags Source
SHA-256 checksum
How to use checksums
654d162fffc494053e0582d26b74fad1348296b0d99e054988301f6534499446
BLAKE2b-256 checksum
How to use checksums
7ded9f649733eb4a08daf0382c2e61b3a7dd0cd19be1e77b3e12322a8fe1b8bc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.8

Release files / hypertok-0.1.0-py3-none-any.whl

Download URL hypertok-0.1.0-py3-none-any.whl
Size 93.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
41370037b0e05ffd78964b4e1e6424dd23cb89e84dedcf0fc42797d7baf2c6f1
BLAKE2b-256 checksum
How to use checksums
2dfa6e27b4d3ccf514e99ad26eb9a1f0646002d6ff2dc10e4807013b37303dae
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.8

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page