Skip to main content

llmbelt 🧰

A tiny, zero-dependency tool belt for working with LLMs. The small utilities you end up re-writing on every project — token counting, cost estimation, retries, prompt templates, and text chunking — in one clean import.

CI PyPI Python License: MIT

  • 🪶 Zero required dependencies — pure standard library.
  • 🔌 Provider-agnostic — works with Anthropic, OpenAI, Gemini, or anything else.
  • 🧪 Fully tested across Python 3.9–3.12.

Install

pip install llmbelt

# Optional: exact token counts for OpenAI-family models
pip install "llmbelt[tiktoken]"

Command line

python -m llmbelt --help          # or the installed `llmbelt` command

llmbelt tokens "Hello, world!"               # -> token count
echo "long text" | llmbelt tokens -          # read from stdin
llmbelt cost --in 1500 --out 800 -m gpt-4o-mini   # -> $0.000630
llmbelt redact "email a@b.com"               # -> "email [EMAIL]"
llmbelt env check OPENAI_API_KEY             # exit 1 + message if unset
llmbelt env report OPENAI_API_KEY            # masked value report
llmbelt env keys                             # providers with a key set

Usage

Configure from the environment

from llmbelt import get_api_key, load_dotenv, EnvConfig, llm_settings_from_env

load_dotenv()                          # zero-dep .env loader (no override by default)
key = get_api_key("openai")            # reads OPENAI_API_KEY (+ known aliases)
settings = llm_settings_from_env()     # {api_key, model, base_url, temperature, ...}

# Declarative, typed config straight from the environment:
class Settings(EnvConfig):
    __prefix__ = "APP_"
    api_key: str                       # required (no default) -> clear EnvError if unset
    timeout: float = 30.0
    debug: bool = False

cfg = Settings.from_env()              # reads APP_API_KEY, APP_TIMEOUT, APP_DEBUG

Also included: env_str/int/float/bool/list, require_env, mask_secret, EnvNamespace, expand/resolve_layers (${VAR} interpolation), env_template (generate a .env.example), and snapshot_env/freeze_env.

Count tokens

from llmbelt import count_tokens, truncate_to_tokens

count_tokens("Hello, world!")              # exact if tiktoken installed, else estimated
count_tokens("Hello", model="gpt-4o")      # use a model-specific encoding

# Trim text to fit a budget (great before sending context to an API)
truncate_to_tokens(long_document, max_tokens=4000)

Estimate cost

from llmbelt import estimate_cost, Price

estimate_cost(input_tokens=1_500, output_tokens=800, model="gpt-4o-mini")   # -> USD

# Bring your own prices (the built-in table is approximate — always verify):
my_prices = {"my-model": Price(input_per_1m=2.0, output_per_1m=6.0)}
estimate_cost(1000, 500, "my-model", pricing=my_prices)

Retry with backoff

from llmbelt import retry

@retry(attempts=5, exceptions=(ConnectionError, TimeoutError))
def call_api():
    ...   # retried with exponential backoff + jitter on failure

# Works on async functions too — awaited, with non-blocking asyncio.sleep backoff
@retry(attempts=5, exceptions=(ConnectionError, TimeoutError))
async def call_api_async():
    ...

Extract JSON from a model reply

from llmbelt import extract_json

extract_json('Sure!\n```json\n{"ok": true}\n```')   # -> {"ok": True}
extract_json('The score is {"value": 0.9}.')         # -> {"value": 0.9}
extract_json("no json here", default=None)           # -> None (else raises ValueError)

Prompt templates

from llmbelt import PromptTemplate

t = PromptTemplate("Translate {text} into {language}.")
t.render(text="hello", language="French")   # "Translate hello into French."
t.render(text="hello")                      # KeyError: Missing template variables: ['language']

Fit a conversation into the context window

from llmbelt import count_message_tokens, trim_messages

messages = [
    {"role": "system", "content": "You are concise."},
    {"role": "user", "content": "..."},
    # ... a long history ...
]

count_message_tokens(messages)                      # total tokens of the chat
trim_messages(messages, max_tokens=8000)            # drop oldest turns, keep the system prompt

Manage a chat with a self-trimming history

from llmbelt import Conversation

chat = Conversation(system="You are concise.", max_tokens=8000)
chat.user("Hello!")
chat.assistant("Hi — how can I help?")

response = client.chat(messages=chat.messages)   # plug into any SDK; auto-trims as it grows

Redact PII before sending to an LLM

from llmbelt import redact, find_pii

redact("email jane@acme.com or call 555-123-4567")
# -> "email [EMAIL] or call [PHONE]"

find_pii("card 4111 1111 1111 1111")   # -> [("CREDIT_CARD", "4111 1111 1111 1111")]

Best-effort regex redaction — a guardrail, not a compliance guarantee. Extend DEFAULT_PII_PATTERNS for your own data.

Cache calls so you don't pay twice

from llmbelt import cached

@cached(ttl=3600)            # remember results for an hour; unhashable args are fine
def ask(prompt: str):
    ...                      # identical prompt -> served from cache, no API call

# works on async functions too
@cached()
async def ask_async(prompt): ...

Stay under rate limits

from llmbelt import RateLimiter

limiter = RateLimiter(rate=60, per=60)   # 60 requests per minute

@limiter                                  # decorator
def call_api(): ...

with limiter:                             # or a context manager
    call_api()

Chunk text for RAG

from llmbelt import chunk_text, chunk_by_tokens, split_text

chunks = chunk_text(document, chunk_size=1000, overlap=100)
# overlapping chunks so answers aren't split across a boundary

# Budget by tokens instead of characters (exact with tiktoken installed):
chunks = chunk_by_tokens(document, chunk_size=500, overlap=50)

# Smarter: break on paragraph/sentence/word boundaries instead of mid-word
chunks = split_text(document, chunk_size=1000, overlap=100)

Track spend across calls

from llmbelt import CostTracker

tracker = CostTracker()
tracker.add(input_tokens=1_500, output_tokens=800, model="gpt-4o-mini")
tracker.add(2_000, 1_200, "gpt-4o-mini")

print(tracker)          # "2 calls, 5,500 tokens, $0.0019"
tracker.summary()       # {"calls": 2, "input_tokens": ..., "cost_usd": ...}

API reference

Function Description
get_env / env_str / env_int / env_float / env_bool / env_list Typed environment readers (defaults, required, casting)
get_api_key(provider) / has_api_key / available_providers Resolve LLM provider API keys from env
load_dotenv / parse_dotenv / find_dotenv Zero-dependency .env loader
require_env / mask_secret / env_report Fail-fast validation + secret-safe reporting
EnvConfig Declarative typed config from the environment
EnvNamespace / collect_prefixed Group prefixed variables into one object/dict
expand / resolve_layers ${VAR} interpolation + layered resolution
llm_settings_from_env(provider) Assemble an LLM client config from env
env_template / env_template_from_config Generate a .env.example
snapshot_env / diff_env / freeze_env / FrozenEnv Snapshot, diff, freeze env config
count_tokens(text, model=None) Exact (tiktoken) or estimated token count
estimate_tokens(text) Dependency-free heuristic count
truncate_to_tokens(text, max_tokens, model=None) Trim text to a token budget
estimate_cost(input_tokens, output_tokens, model, pricing=None) USD cost estimate
CostTracker(pricing=None) Accumulate tokens + USD cost across many calls
retry(attempts, base_delay, backoff, jitter, exceptions, ...) Backoff retry decorator (sync and async)
PromptTemplate(template) Templating with missing-variable validation
chunk_text(text, chunk_size, overlap) Overlapping text chunks (by character)
chunk_by_tokens(text, chunk_size, overlap, model=None) Overlapping text chunks (by token budget)
split_text(text, chunk_size, overlap, separators=None) Boundary-aware chunks (paragraph/sentence/word)
extract_json(text, default=...) Parse the first JSON value out of an LLM reply
count_message_tokens(messages, model=None) Token count of a chat-format message list
trim_messages(messages, max_tokens, model=None, keep_system=True) Trim a conversation to a token budget
Conversation(system=None, max_tokens=None, model=None) Stateful chat history that self-trims to a budget
redact(text, patterns=None, mask="[{label}]") Mask PII/secrets in text
find_pii(text, patterns=None) List (label, match) PII detections
cached(maxsize, ttl) Memoize calls on an args hash (sync + async)
RateLimiter(rate, per, capacity=None) Token-bucket throttle (gate / context manager / decorator)

Development

git clone https://github.com/YoungAlpaccino/llmbelt
cd llmbelt
pip install -e ".[dev]"
pytest          # run tests
ruff check .    # lint

Publishing to PyPI (maintainer notes)

python -m build
twine upload dist/*

Before first publish: confirm the name llmbelt is free on PyPI. If taken, rename in pyproject.toml, the src/ folder, and imports (a single find-and-replace).


License

MIT — see LICENSE. Use it anywhere, including commercially.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llmbelt-0.7.0.tar.gz (42.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

llmbelt-0.7.0-py3-none-any.whl (36.2 kB view details)

Uploaded Python 3

File details

Details for the file llmbelt-0.7.0.tar.gz.

File metadata

  • Download URL: llmbelt-0.7.0.tar.gz
  • Upload date:
  • Size: 42.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.0

File hashes

Hashes for llmbelt-0.7.0.tar.gz
Algorithm Hash digest
SHA256 0e0e3364997b7e370179586cc0ce7f0266778ee8c6a73b3824d3254ae8b043a5
MD5 3328a51011f8cf94ba71a9e1a46e7efb
BLAKE2b-256 d82b7ae901dd8b50bf7ea0b7d60a9043750aae6b17743d08725e6c29e654cba4

See more details on using hashes here.

File details

Details for the file llmbelt-0.7.0-py3-none-any.whl.

File metadata

  • Download URL: llmbelt-0.7.0-py3-none-any.whl
  • Upload date:
  • Size: 36.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.0

File hashes

Hashes for llmbelt-0.7.0-py3-none-any.whl
Algorithm Hash digest
SHA256 1c7b8d47dabacc5fda07c8100043e027affc5448822febe26255564978a25249
MD5 b76762b3b9180aaa2c0ba54e37552559
BLAKE2b-256 19ba44dc7a712381b5bbaa61c2ca8c0e3bd3e1e1507e8bee9268a3cb4b8408e0

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page