Skip to main content

ctok

ctok reconstructs Claude token counts offline, with no API call or network access. It is unofficial and is not affiliated with Anthropic.

The reconstruction targets counts. Claude does not expose token boundaries, so tokenize() returns one valid minimum-cost tiling, not a claim about Anthropic's exact segmentation. The research behind the model is described in On the biology of Claude's tokenizer.

Count a string

ctok is on PyPI, so uvx fetches and runs it with no install. With no --version it reports both generations, since the same string does not cost the same on each.

uvx ctok "hello, world"
  v3.0
  stream: '⟨bow⟩hello⟨eow⟩,⟨eow⟩⟨bow⟩world⟨eow⟩'
  tokens: ['⟨bow⟩hello⟨eow⟩', ',⟨eow⟩', '⟨bow⟩world⟨eow⟩']

  content 3 + frame 7 = 10

  v4.8
  stream: '⟨bow⟩hello⟨eow⟩,⟨eow⟩⟨bow⟩world⟨eow⟩'
  tokens: ['⟨bow⟩h', 'ello⟨eow⟩', ',⟨eow⟩', '⟨bow⟩world⟨eow⟩']

  content 4 + frame 6 = 10

--version 4.7 picks one family instead.

As a library

from ctok import token_count, tokenize

token_count("hello, world")           # 10, using the v3 family
token_count("hello, world", "4.7")    # 15
token_count("hello, world", "4.8")    # 10

tokens = tokenize("NASA likes tokenizers")
assert len(tokens) == token_count("NASA likes tokenizers")

Supported families

version is a string such as "4.7". Components are compared as integers, so "4.10" sorts after "4.9". Python reads the float literal 4.10 as 4.1, so passing a non-string version raises TypeError.

requested version family model generation
"3.0" <= version < "4.7" v3, the default Claude 3 through Opus 4.6
"4.7" <= version < "4.8" v4.7 Opus 4.7
version >= "4.8" v4.8 Opus 4.8, Sonnet 5, and Fable 5

v4.8+ uses the v4.7 vocabulary with a six-token frame. Opus 5 alone makes trailing ASCII whitespace free, which ctok intentionally does not model.

How it works

For one user message, ctok:

  1. normalizes the text, including NFC and family-specific quote folding;
  2. rewrites it into a stream with word, case, and byte markers. For BMP combining marks, Unicode Alphabetic decides the split: Alphabetic marks stay with the word and non-Alphabetic marks, other than variation selectors, stand outside it;
  3. finds a minimum-cost tiling over the measured vocabulary and UTF-8 byte fallback;
  4. adds the measured message frame.

token_count(text) is len(tokenize(text)). The output notation makes internal structure visible:

notation meaning
⟨bow⟩, ⟨eow⟩ word boundaries
⟨shift⟩, ⟨caps⟩ case rewrites
⟨0xNN⟩ a byte-fallback token
⟨pad⟩ part of the single-message frame

Measured accuracy

These results compare ctok with recorded count_tokens responses:

corpus role v3 exact v4.7 exact
Goldfish, 350 languages and 962,054 rows mining 962,053 962,054
MultiPL-E, 22 programming languages held out 22 22
Rosetta Code, 1,741 documents mining 1,741 1,741
Rosetta Code, separate 250 documents held out 250 250
UDHR, 501 languages mining (in-sample since 2026-08-12) 501 501

v4.8+ is omitted because the table gates physical vocabularies. It reuses v4.7's vocabulary.

The stored measurement sets contain no under-counts: 0 of 2,276,929 v3 texts and 0 of 2,328,425 v4.7 texts. This does not guarantee the result for arbitrary input. Goldfish, the main Rosetta sample, and UDHR informed piece selection. MultiPL-E and the separate Rosetta sample did not.

Run the public gates with:

uv run pytest
uv run python tests/gates.py

Vocabulary evidence

The two vocabulary files contain 48,792 v3 pieces and 15,283 v4.7 pieces. Every entry has a fixed membership witness or is one of the structural marker atoms checked by the test suite.

from ctok import pieces, witness

len(pieces("4.7"))                       # 15283
witness("⟨bow⟩the⟨eow⟩", "4.7")
# {'probe': 'the', 'raw': 12, 'kind': 'raw'}

A witness says that one marked piece costs one token in a calibrated probe. It does not prove the encoder rewrite or resolve ties between equal-cost tilings. tests/test_witness.py checks every published witness and requires complete witnessed-or-special coverage.

License

MIT

Metadata

Release files for ctok 1.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ctok 1.3.0
File Size Uploaded
ctok-1.3.0.tar.gz 3.1 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for ctok 1.3.0
File Interpreter ABI Platform
ctok-1.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 3.5 MB

Release files / ctok-1.3.0.tar.gz

Download URL ctok-1.3.0.tar.gz
Size 3.1 MB
Tags Source
SHA-256 checksum
How to use checksums
f41812a6a67aea1422ff742c1b20f2a61f5ea733c74d755b94db88be5c9a930f
BLAKE2b-256 checksum
How to use checksums
5561f1e14874aec35db437c98af271316ffb47035e7d018ed65a6ccd0480fb57
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 1, 2026.

Transparency log

Release files / ctok-1.3.0-py3-none-any.whl

Download URL ctok-1.3.0-py3-none-any.whl
Size 461.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
84442682eeb2013ecb86861f9477b4ad3e19cc174099df98e9b5016450013217
BLAKE2b-256 checksum
How to use checksums
07be85fdcb8b6c901f0dfaf7da4bac125820cd501f73e69f7550131bb685a143
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 1, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.3.0 This release

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page