Skip to main content

llm-ids

Encode data as token-efficient, LLM-friendly strings.

Hex, base64, and UUID strings are the usual way to represent IDs and secrets, but generic tokenizers compress them poorly, costing many tokens per ID: 3.3-9.7 bits/token depending on tokenizer and the ID charset. llm-ids uses a bijective numeration over an alphabet of whole vocabulary tokens instead: 11.6-13.7 bits per token.

Install

pip install llm-ids

Usage

import hashlib, string
from llm_ids import Alphabet, Codec

int_codec = Codec(Alphabet.O200K_V1, charset=string.digits)
token_id = int_codec.to_token_id(str(42))                              # 'ACLE'
assert int_codec.from_token_id(token_id) == "42"

alnum_codec = Codec(Alphabet.O200K_V1, charset=string.ascii_uppercase + string.digits)
token_id = alnum_codec.to_token_id("A1B2C3")                           # 'Stocks.bunifu'
assert alnum_codec.from_token_id(token_id) == "A1B2C3"

hex_codec = Codec(Alphabet.O200K_V1, charset=string.digits + "abcdef")
digest_hex = hashlib.sha256(b"hello").hexdigest()
token_id = hex_codec.to_token_id(digest_hex)                           # 'Tournament_you...'
assert hex_codec.from_token_id(token_id) == digest_hex

Available alphabets (see src/llm_ids/alphabets.py):

Alphabet member use for bits/token
SHARED_V1 OpenAI (o200k/cl100k), Llama 3.x, and Qwen3 at once >=13.25
O200K_V1 GPT-4o, GPT-4.1, GPT-5 family >=13.36
CL100K_V1 GPT-4, GPT-3.5-turbo >=13.67
MISTRAL_TEKKEN_V1 Mistral NeMo, Small/Large 2.x+ >=12.14
DEEPSEEK_V1 DeepSeek V3, R1 >=11.58

Use a solo alphabet when the target tokenizer is known. Use SHARED_V1 when it isn't, or spans several of those families, since it costs almost nothing (>=13.25 vs. >=13.36-13.67 bits/token). Mistral and DeepSeek get their own alphabets since their vocabularies barely overlap with the others.

NOTE: Gemma is not supported. SentencePiece has no fixed pretoken boundary, so token counts are not provably fixed the way they are for BPE tokenizers.

How it works

  • llm_ids.id_codec: integer <-> token_id string, using the token-boundary trick above.
  • llm_ids.bijective: generic bijective numeration; preserves leading "zero" symbols that plain positional notation would drop.
  • llm_ids.char_token_id: strings over a caller-declared charset, built on bijective + id_codec.
  • llm_ids.codec.Codec: the alphabet-and-charset-bound entry point above; use this.

Development

This project uses uv.

uv sync --group dev
uv run pytest

To regenerate an alphabet's data file from its source tokenizer vocabulary (requires the tiktoken/tokenizers/huggingface_hub dev dependencies):

uv run python scripts/build_alphabets.py

Commits follow Conventional Commits (fix:, feat:, feat!:/BREAKING CHANGE:), checked on every PR. A maintainer triggers a release manually from the Actions tab (Release workflow), which runs Commitizen to bump the version, update CHANGELOG.md, tag, and create a GitHub Release -- publishing to PyPI happens automatically from there.

Metadata

Release files for llm-ids 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-ids 0.3.0
File Size Uploaded
llm_ids-0.3.0.tar.gz 532.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-ids 0.3.0
File Interpreter ABI Platform
llm_ids-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.1 MB

Release files / llm_ids-0.3.0.tar.gz

Download URL llm_ids-0.3.0.tar.gz
Size 532.4 kB
Tags Source
SHA-256 checksum
How to use checksums
b563372bcecacaab3df10a9e8a9d391efa7db9b852b0876b3b9acfa9f00407a3
BLAKE2b-256 checksum
How to use checksums
88619ab8e1363c0dc9da5313e7ee6933d3e63ec921146c368e319e0cda50c823
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release files / llm_ids-0.3.0-py3-none-any.whl

Download URL llm_ids-0.3.0-py3-none-any.whl
Size 534.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
65d78c3a91ee97fce9922811b1e841eef78b9103aae726b14c8c820db617de42
BLAKE2b-256 checksum
How to use checksums
53ee45cd04e44c13c7ded1037ea8a51f1faccca2e840e1ed9d0636f34d0f72d4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page