llmbelt 🧰
A tiny, zero-dependency tool belt for working with LLMs. The small utilities you end up re-writing on every project — token counting, cost estimation, retries, prompt templates, and text chunking — in one clean import.
- 🪶 Zero required dependencies — pure standard library.
- 🔌 Provider-agnostic — works with Anthropic, OpenAI, Gemini, or anything else.
- 🧪 Fully tested across Python 3.9–3.12.
Install
pip install llmbelt
# Optional: exact token counts for OpenAI-family models
pip install "llmbelt[tiktoken]"
Command line
python -m llmbelt --help # or the installed `llmbelt` command
llmbelt tokens "Hello, world!" # -> token count
echo "long text" | llmbelt tokens - # read from stdin
llmbelt cost --in 1500 --out 800 -m gpt-4o-mini # -> $0.000630
llmbelt redact "email a@b.com" # -> "email [EMAIL]"
llmbelt env check OPENAI_API_KEY # exit 1 + message if unset
llmbelt env report OPENAI_API_KEY # masked value report
llmbelt env keys # providers with a key set
Usage
Configure from the environment
from llmbelt import get_api_key, load_dotenv, EnvConfig, llm_settings_from_env
load_dotenv() # zero-dep .env loader (no override by default)
key = get_api_key("openai") # reads OPENAI_API_KEY (+ known aliases)
settings = llm_settings_from_env() # {api_key, model, base_url, temperature, ...}
# Declarative, typed config straight from the environment:
class Settings(EnvConfig):
__prefix__ = "APP_"
api_key: str # required (no default) -> clear EnvError if unset
timeout: float = 30.0
debug: bool = False
cfg = Settings.from_env() # reads APP_API_KEY, APP_TIMEOUT, APP_DEBUG
Also included: env_str/int/float/bool/list, require_env, mask_secret,
EnvNamespace, expand/resolve_layers (${VAR} interpolation),
env_template (generate a .env.example), and snapshot_env/freeze_env.
Count tokens
from llmbelt import count_tokens, truncate_to_tokens
count_tokens("Hello, world!") # exact if tiktoken installed, else estimated
count_tokens("Hello", model="gpt-4o") # use a model-specific encoding
# Trim text to fit a budget (great before sending context to an API)
truncate_to_tokens(long_document, max_tokens=4000)
Estimate cost
from llmbelt import estimate_cost, Price
estimate_cost(input_tokens=1_500, output_tokens=800, model="gpt-4o-mini") # -> USD
# Bring your own prices (the built-in table is approximate — always verify):
my_prices = {"my-model": Price(input_per_1m=2.0, output_per_1m=6.0)}
estimate_cost(1000, 500, "my-model", pricing=my_prices)
Retry with backoff
from llmbelt import retry
@retry(attempts=5, exceptions=(ConnectionError, TimeoutError))
def call_api():
... # retried with exponential backoff + jitter on failure
# Works on async functions too — awaited, with non-blocking asyncio.sleep backoff
@retry(attempts=5, exceptions=(ConnectionError, TimeoutError))
async def call_api_async():
...
Extract JSON from a model reply
from llmbelt import extract_json
extract_json('Sure!\n```json\n{"ok": true}\n```') # -> {"ok": True}
extract_json('The score is {"value": 0.9}.') # -> {"value": 0.9}
extract_json("no json here", default=None) # -> None (else raises ValueError)
Prompt templates
from llmbelt import PromptTemplate
t = PromptTemplate("Translate {text} into {language}.")
t.render(text="hello", language="French") # "Translate hello into French."
t.render(text="hello") # KeyError: Missing template variables: ['language']
Fit a conversation into the context window
from llmbelt import count_message_tokens, trim_messages
messages = [
{"role": "system", "content": "You are concise."},
{"role": "user", "content": "..."},
# ... a long history ...
]
count_message_tokens(messages) # total tokens of the chat
trim_messages(messages, max_tokens=8000) # drop oldest turns, keep the system prompt
Manage a chat with a self-trimming history
from llmbelt import Conversation
chat = Conversation(system="You are concise.", max_tokens=8000)
chat.user("Hello!")
chat.assistant("Hi — how can I help?")
response = client.chat(messages=chat.messages) # plug into any SDK; auto-trims as it grows
Redact PII before sending to an LLM
from llmbelt import redact, find_pii
redact("email jane@acme.com or call 555-123-4567")
# -> "email [EMAIL] or call [PHONE]"
find_pii("card 4111 1111 1111 1111") # -> [("CREDIT_CARD", "4111 1111 1111 1111")]
Best-effort regex redaction — a guardrail, not a compliance guarantee. Extend
DEFAULT_PII_PATTERNSfor your own data.
Cache calls so you don't pay twice
from llmbelt import cached
@cached(ttl=3600) # remember results for an hour; unhashable args are fine
def ask(prompt: str):
... # identical prompt -> served from cache, no API call
# works on async functions too
@cached()
async def ask_async(prompt): ...
Stay under rate limits
from llmbelt import RateLimiter
limiter = RateLimiter(rate=60, per=60) # 60 requests per minute
@limiter # decorator
def call_api(): ...
with limiter: # or a context manager
call_api()
Chunk text for RAG
from llmbelt import chunk_text, chunk_by_tokens, split_text
chunks = chunk_text(document, chunk_size=1000, overlap=100)
# overlapping chunks so answers aren't split across a boundary
# Budget by tokens instead of characters (exact with tiktoken installed):
chunks = chunk_by_tokens(document, chunk_size=500, overlap=50)
# Smarter: break on paragraph/sentence/word boundaries instead of mid-word
chunks = split_text(document, chunk_size=1000, overlap=100)
Track spend across calls
from llmbelt import CostTracker
tracker = CostTracker()
tracker.add(input_tokens=1_500, output_tokens=800, model="gpt-4o-mini")
tracker.add(2_000, 1_200, "gpt-4o-mini")
print(tracker) # "2 calls, 5,500 tokens, $0.0019"
tracker.summary() # {"calls": 2, "input_tokens": ..., "cost_usd": ...}
API reference
| Function | Description |
|---|---|
get_env / env_str / env_int / env_float / env_bool / env_list |
Typed environment readers (defaults, required, casting) |
get_api_key(provider) / has_api_key / available_providers |
Resolve LLM provider API keys from env |
load_dotenv / parse_dotenv / find_dotenv |
Zero-dependency .env loader |
require_env / mask_secret / env_report |
Fail-fast validation + secret-safe reporting |
EnvConfig |
Declarative typed config from the environment |
EnvNamespace / collect_prefixed |
Group prefixed variables into one object/dict |
expand / resolve_layers |
${VAR} interpolation + layered resolution |
llm_settings_from_env(provider) |
Assemble an LLM client config from env |
env_template / env_template_from_config |
Generate a .env.example |
snapshot_env / diff_env / freeze_env / FrozenEnv |
Snapshot, diff, freeze env config |
count_tokens(text, model=None) |
Exact (tiktoken) or estimated token count |
estimate_tokens(text) |
Dependency-free heuristic count |
truncate_to_tokens(text, max_tokens, model=None) |
Trim text to a token budget |
estimate_cost(input_tokens, output_tokens, model, pricing=None) |
USD cost estimate |
CostTracker(pricing=None) |
Accumulate tokens + USD cost across many calls |
retry(attempts, base_delay, backoff, jitter, exceptions, ...) |
Backoff retry decorator (sync and async) |
PromptTemplate(template) |
Templating with missing-variable validation |
chunk_text(text, chunk_size, overlap) |
Overlapping text chunks (by character) |
chunk_by_tokens(text, chunk_size, overlap, model=None) |
Overlapping text chunks (by token budget) |
split_text(text, chunk_size, overlap, separators=None) |
Boundary-aware chunks (paragraph/sentence/word) |
extract_json(text, default=...) |
Parse the first JSON value out of an LLM reply |
count_message_tokens(messages, model=None) |
Token count of a chat-format message list |
trim_messages(messages, max_tokens, model=None, keep_system=True) |
Trim a conversation to a token budget |
Conversation(system=None, max_tokens=None, model=None) |
Stateful chat history that self-trims to a budget |
redact(text, patterns=None, mask="[{label}]") |
Mask PII/secrets in text |
find_pii(text, patterns=None) |
List (label, match) PII detections |
cached(maxsize, ttl) |
Memoize calls on an args hash (sync + async) |
RateLimiter(rate, per, capacity=None) |
Token-bucket throttle (gate / context manager / decorator) |
Development
git clone https://github.com/YoungAlpaccino/llmbelt
cd llmbelt
pip install -e ".[dev]"
pytest # run tests
ruff check . # lint
Publishing to PyPI (maintainer notes)
python -m build
twine upload dist/*
Before first publish: confirm the name
llmbeltis free on PyPI. If taken, rename inpyproject.toml, thesrc/folder, and imports (a single find-and-replace).
License
MIT — see LICENSE. Use it anywhere, including commercially.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file llmbelt-0.7.0.tar.gz.
File metadata
- Download URL: llmbelt-0.7.0.tar.gz
- Upload date:
- Size: 42.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0e0e3364997b7e370179586cc0ce7f0266778ee8c6a73b3824d3254ae8b043a5
|
|
| MD5 |
3328a51011f8cf94ba71a9e1a46e7efb
|
|
| BLAKE2b-256 |
d82b7ae901dd8b50bf7ea0b7d60a9043750aae6b17743d08725e6c29e654cba4
|
File details
Details for the file llmbelt-0.7.0-py3-none-any.whl.
File metadata
- Download URL: llmbelt-0.7.0-py3-none-any.whl
- Upload date:
- Size: 36.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1c7b8d47dabacc5fda07c8100043e027affc5448822febe26255564978a25249
|
|
| MD5 |
b76762b3b9180aaa2c0ba54e37552559
|
|
| BLAKE2b-256 |
19ba44dc7a712381b5bbaa61c2ca8c0e3bd3e1e1507e8bee9268a3cb4b8408e0
|