llm-router-ledger
Route LLM calls through one send_message() for text and one
create_embeddings() for vectors, and keep a JSONL ledger of every
request and response for offline cost reconciliation.
Provider support
| Status | Adapter | Providers |
|---|---|---|
| Supported | direct | Anthropic |
| Supported | OpenAI-compat | Azure OpenAI, DeepSeek, Local LM Studio, Local Ollama, MiniMax, OpenAI, OpenRouter, Qwen, Zhipu / GLM |
| Supported | via OpenRouter | ByteDance Seed, InclusionAI Ling, Nvidia Nemotron, Xiaomi MiMo |
| Planned | direct | Gemini |
- Every "Supported" row is live-smoke-verified end-to-end.
- Anthropic requires the optional
[anthropic]extra:uv pip install llm-router-ledger[anthropic]. - For the "via OpenRouter" families, use
provider: openrouterwith the appropriate model id. - The table above is about text. Embeddings are gated separately and verified on OpenRouter, Ollama and LM Studio only; see Embeddings.
Free chat models
Verified end-to-end via OpenRouter and configured in examples/llm_endpoints.example.yaml. Rates are USD per 1M tokens. Both endpoints declare an explicit 0.00 rather than omitting cost, so the ledger records their tokens the same way it does a paid endpoint.
| Model | In | Out | Context |
|---|---|---|---|
nvidia/nemotron-3.5-content-safety:free |
0.00 | 0.00 | 128000 |
nvidia/nemotron-3.5-lightning:free |
0.00 | 0.00 | 1000000 |
Free models share their capacity with everyone else using them, so a call fails with HTTP 429 when they are busy. One of the two above failed nine times in a row during verification.
nvidia/nemotron-3.5-content-safety:free is a safety classifier, not a general chat model. It answers every prompt with a verdict, so What is 17 * 23? returns User Safety: safe.
Install
uv pip install llm-router-ledger
Quickstart
Set OPENROUTER_API_KEY in .env and create llm_endpoints.yaml in the working directory. The fastest path is to copy examples/llm_endpoints.example.yaml to llm_endpoints.yaml in your working directory and edit it.
from llm_router_ledger import UsageTracker, send_message
tracker = UsageTracker(
log_path="logs/usage.jsonl",
project_id="my-blog",
)
result = send_message(
endpoint_name="openrouter-mimo-v2.5",
system="You are concise.",
user="Explain prompt caching in two sentences.",
tracker=tracker,
)
Or against a local Ollama server, with no API costs:
result = send_message(
endpoint_name="local-llama",
system="You are concise.",
user="Explain prompt caching in two sentences.",
tracker=tracker,
)
send_message()returns aChatResultwith.text,.usage, and.generation_id..usageaddscost,is_byok, andupstream_providerto the token keys when the provider reports them, plus flattened reasoning / cache detail keys (e.g.completion_reasoning_tokens,prompt_cached_tokens); see JSONL ledger schema.UsageTrackerappends pairedllm_request/llm_responseevents to the JSONL log, stamped withproject_id,run_tag,run_label, andpurposefor later grouping.- Prompt and response previews are redacted by default; pass
preview_lengthto opt in to storing truncated text, see JSONL ledger schema. - For multi-turn conversations, tool loops, or anything
system+usercan't express, passmessagesinstead; it replacessystemanduseroutright rather than merging with them:
result = send_message(
endpoint_name="openrouter-mimo-v2.5",
messages=[
{"role": "system", "content": [{"type": "text", "text": "You are concise."}]},
{"role": "user", "content": [{"type": "text", "text": "Explain prompt caching."}]},
{"role": "assistant", "content": [{"type": "text", "text": "..."}]},
{"role": "user", "content": [{"type": "text", "text": "Now in one sentence."}]},
],
tracker=tracker,
)
Each entry is {"role": ..., "content": [{"type": "text", "text": ...}]}, the OpenAI content-parts shape. It's kept even though only "text" parts are supported today, so adding image input later is additive rather than another break.
Embeddings
create_embeddings() embeds a list of texts and writes the same paired ledger events as send_message().
from llm_router_ledger import UsageTracker, create_embeddings
tracker = UsageTracker(
log_path="logs/usage.jsonl",
project_id="my-blog",
)
result = create_embeddings(
endpoint_name="openrouter-embed-bge-m3",
texts=["first passage", "second passage"],
tracker=tracker,
)
create_embeddings()returns anEmbeddingResultwith.vectors(one per input, in input order),.usage, and.generation_id..usageaddsdimensionsandembedding_countto the token keys, pluscost,is_byok, andupstream_providerwhen the provider reports them.completion_tokensis always 0. Embeddings bill input only.
Verified models
Via OpenRouter:
| Model | Dims | Context |
|---|---|---|
baai/bge-base-en-v1.5 |
768 | 512 |
baai/bge-m3 |
1024 | 8194 |
mistralai/mistral-embed-2312 |
1024 | 8192 |
nvidia/nemotron-3-embed-1b:free |
2048 | 32768 |
openai/text-embedding-3-large |
3072 | 8192 |
openai/text-embedding-3-small |
1536 | 8192 |
perplexity/pplx-embed-v1-0.6b |
1024 | 32000 |
qwen/qwen3-embedding-4b |
2560 | 32768 |
qwen/qwen3-embedding-8b |
4096 | 32768 |
Locally via Ollama: qwen3-embedding:0.6b, 1024 dims, 32768 context (ollama pull qwen3-embedding:0.6b). The same model at Q8_0 runs under LM Studio as text-embedding-qwen3-embedding-0.6b, downloaded from the Discover tab, so local runs on either server are directly comparable.
baai/bge-base-en-v1.5 is English only. Non-English input still returns vectors, with no error.
Prices are per endpoint in llm_endpoints.yaml, each with a pricing_url and pricing_checked date. See examples/llm_endpoints.example.yaml for the verified values, and llm-router-ledger stale for ones that need rechecking.
embedding_dimensions
An optional endpoint field declaring the vector width.
- Never sent on the wire, so a vector column or collection can be sized without first making a call. It is not OpenAI's
dimensionsrequest parameter and truncates nothing. - Enforced on the response: a different width raises
ProviderErrorinstead of returning vectors that would corrupt a fixed-width index. - OpenRouter re-routes between calls.
baai/bge-m3has been served by DeepInfra on one call and Parasail on the next. - Leave it unset to accept any width.
Provider gate
Embeddings are refused for providers not verified end-to-end, even where the chat adapter works: provider: openai raises NotImplementedError.
ollama and lmstudio are verified for embeddings. Other local servers are not, so they are refused despite serving the same OpenAI-compatible API.
Neither local server returns a response id, leaving provider_response_id empty. LM Studio additionally reports prompt_tokens and total_tokens as zero for embeddings, at any input size, so its rows record the vectors and their width but a token count of 0 rather than the true figure. Ollama reports real counts. Nothing is billed on either, so there is no invoice to reconcile against.
Smoke tests
python examples/smoke_test_openrouter_embeddings.py # free endpoint
python examples/smoke_test_openrouter_embeddings.py --endpoint openrouter-embed-qwen3-8b
python examples/smoke_test_ollama_embeddings.py # local, no cost
python examples/smoke_test_lmstudio_embeddings.py # local, no cost
Each takes --input-file, one text to embed per line, in place of the sample corpus.
Per-endpoint request params
Model-specific knobs belong in config, not in every caller. Give an endpoint an extra_body and it is sent on every call to that endpoint:
endpoints:
openrouter-deepseek:
provider: openrouter
model: deepseek/deepseek-chat
api_key_env: OPENROUTER_API_KEY
base_url: https://openrouter.ai/api/v1
extra_body:
reasoning:
enabled: false
- An
extra_bodypassed tosend_message()replaces the endpoint's value outright. The two layers are not merged, so a caller that wants both must combine them itself. An opaque vendor passthrough carries no merge rules to memorise as a result. - Known limitation:
provider: anthropicignoresextra_body, so the field has no effect there.provider: openrouterreaches Claude withextra_bodyintact.
Mirroring usage elsewhere
UsageTracker.subscribe() registers a callback that receives every ledger entry, so usage can be mirrored to another store without this library depending on it:
tracker.subscribe(lambda entry: my_container.upsert_item(entry))
- Each entry is written to the JSONL ledger before any subscriber runs.
- A callback that raises is logged and skipped. The entry is already in the ledger, the call that produced it is unaffected, and the remaining subscribers still run.
- Each subscriber receives its own copy of the entry.
- Callbacks are synchronous and run on the calling thread, so a slow one delays every call. Queue the work inside the callback if the destination is remote.
Recording calls this library did not make
An agent framework owns its own call path, so send_message() never
runs and the ledger never sees the tokens. UsageTracker records those
calls too, producing the same two events, with the same keys, that a
call through this library produces.
Record a call directly:
request_id = tracker.record_request(
model="xiaomi/mimo-v2.5",
provider="openrouter",
purpose="query-planning",
)
# ... something else makes the call ...
tracker.record_response(
request_id=request_id,
model="xiaomi/mimo-v2.5",
usage=raw_usage,
response_id=response.id,
provider="openrouter",
)
usage is the provider's usage mapping as reported. The three token
keys are lifted into the usage block and everything else is written
under usage_details, the same split send_message() performs.
response_id is routed the same way too: an id prefixed gen- lands in
generation_id, anything else in provider_response_id.
Or record a whole Pydantic AI run at once:
result = await agent.run("...")
tracker.record_run(
result.all_messages(),
purpose="query-planning",
provider="openrouter",
)
One request and response pair is written per model call, so a run that called a tool three times produces three pairs. The messages are read by duck typing, so this costs no dependency on pydantic-ai.
- Pass
provider. The provider name on the message is the framework's own, and isopenaifor every OpenAI-compatible server, so without it a local call is filed as an OpenAI one. - Pydantic AI's usage fields are translated to the names the adapters
already write, so rows from both sources join on one vocabulary. That
includes
finish_reason, recorded in the provider's own words rather than the framework's normalised ones. RequestUsage.costis recorded asusage_details.estimated_cost, never ascost: it is computed from a local price table rather than reported by the provider, andcostis reserved for what the provider said it billed. A provider's own reported cost does not survive the trip at all, since the framework keeps only integer usage fields, so reconcile these rows against the provider's export by response id.- A finished message list carries no record of why each call was made,
so one
purposeis stamped across the whole run and retries within it inherit it. Use a purpose scope where per-call purpose matters. - A run that raises produces no rows at all, unlike
send_message(), which logs the request before making the call. Both are correct, but it changes what an unpairedllm_requestmeans.
Setting a purpose an agent cannot pass
send_message() takes purpose per call, but by the time a framework's
request reaches the ledger there is no argument left to carry it. Set it
around the work instead:
from llm_router_ledger import purpose_scope
with purpose_scope("query-planning"):
result = await agent.run("...")
tracker.record_run(result.all_messages())
- The scope is a context variable, so it is per-task and per-thread: two agents running concurrently under asyncio each keep their own purpose.
- Scopes nest and the innermost wins. Entering a scope with
""is how a nested call records with no purpose rather than inheriting the one around it. - A
purposepassed to the call wins over the scope, and the scope wins overUsageTracker(default_purpose=...). The narrowest thing that was actually set is what reaches the ledger.
JSONL ledger schema
UsageTrackerwrites two events persend_message()orcreate_embeddings()call: anllm_requestbefore the call, and anllm_responseafter.- Both share a
request_idso they can be paired. Top-level fields on each event includeproject_id,provider,model,purpose,run_tag,run_label, andtimestamp. - The
llm_responseevent additionally carriesusage(withprompt_tokens,completion_tokens,total_tokens) and a response preview. - Previews are redacted by default:
system_prompt_preview,user_prompt_preview, andresponse_previeware written as"[REDACTED]"when the underlying text is non-empty,""when it genuinely is empty. Passpreview_length(a positive character count) toUsageTracker()to opt in to storing a truncated preview instead; the length and token counts are always recorded either way. - A failed call writes an
llm_errorevent sharing therequest_idof itsllm_request, carryingerror_type(the original SDK exception's class name),error_message, andstatus_codewhere the provider returned one. A third event type rather than anllm_responsewith an error field, because a failed call has no tokens and writing zeroes would corrupt anyone summing them. The SDK retries internally before raising, so onellm_errorstands for however many attempts it made. usage_detailson the response holds everything the provider reported beyond the three token keys, written only when non-empty.usagekeeps the same fixed three-key shape regardless of what lands inusage_details, across both modalities.- Embedding calls:
dimensionsandembedding_countalways, pluscost,is_byokandupstream_providerwhere available.
- Embedding calls:
Chat calls map provider fields onto ledger keys as follows. A key is written only when the provider reports a non-zero value for it.
| Provider reports | Ledger key | Observed on |
|---|---|---|
usage.prompt_tokens / completion_tokens / total_tokens |
usage.*, unchanged |
all |
Anthropic usage.input_tokens / output_tokens |
usage.prompt_tokens / completion_tokens |
Anthropic |
usage.cost, usage.is_byok |
usage_details.cost, .is_byok |
OpenRouter |
response provider |
usage_details.upstream_provider |
OpenRouter |
completion_tokens_details.reasoning_tokens |
usage_details.completion_reasoning_tokens |
OpenRouter, Qwen, Zhipu |
prompt_tokens_details.cached_tokens |
usage_details.prompt_cached_tokens |
OpenRouter, DeepSeek, Zhipu |
other keys in either *_tokens_details block |
same name, completion_ / prompt_ prefixed |
varies |
| anything else the provider reports | usage_details.unmapped.<key> |
see below |
usage_details.completion_reasoning_tokens and usage_details.prompt_cached_tokens are subsets of usage.completion_tokens and usage.prompt_tokens, not additions to them, so adding either to its parent double-counts. Verified against three paid OpenRouter endpoints: the reported cost matched the inclusive reading to the cent, while the additive reading overstated it by 21 to 58 percent.
Two keys are derived rather than reported: completion_tool_call_count, the number of tool calls on a turn that made any, and finish_reason, written only when the turn ended abnormally (e.g. length, truncated at max_tokens) in the provider's own vocabulary. A tool-call turn has no text, so it records response_length 0; the count is what distinguishes it from a model that answered with nothing.
usage_details.unmapped holds provider fields the library has no mapping for, so nothing a provider reports is silently discarded. Observed examples: DeepSeek's prompt_cache_hit_tokens and prompt_cache_miss_tokens, which duplicate prompt_cached_tokens; Qwen's completion_text_tokens and prompt_text_tokens; Azure's latency_checkpoint timing block; Anthropic's cache_read_input_tokens, cache_creation_input_tokens, cache_creation, service_tier and inference_geo. OpenAI reports nothing unmapped.
Treat unmapped as unstable. A key that later gains an explicit mapping moves out of unmapped and up a level, so read from it defensively.
Embedding calls additionally set modality: "embedding" on both events. The key is omitted entirely on text calls, so existing rows are unchanged and an absent modality means text. The response preview is empty and response_length is 0 for embeddings, since an embedding response carries no text. Neither the input text in full nor the vectors are ever written to the ledger.
Identifying a response for billing reconciliation: the response id is routed to one of two fields based on prefix:
generation_id: set when the id starts with"gen-"(OpenRouter convention). Use this when joining against OpenRouter's CSV export, which calls the columngeneration_id.provider_response_id: set for everything else. OpenAI, Azure OpenAI, Ollama, and most direct-provider endpoints return ids like"chatcmpl-..."that land here. Use this when joining against OpenAI-family billing exports or any provider-native log that exposes a chat completion id.
OpenRouter embedding ids are prefixed gen-emb-, so they route to generation_id and reconcile like any other OpenRouter call. Ollama returns no id at all, leaving provider_response_id empty; nothing is billed there, so there is nothing to reconcile against.
Exactly one of the two fields is populated per llm_response event; queries that join the ledger to billing data should COALESCE over both or branch on provider.
CLI
llm-router-ledger list # show configured endpoints
llm-router-ledger validate llm_endpoints.yaml # validate the YAML
llm-router-ledger stale --days 30 # endpoints with stale pricing
llm-router-ledger chat --endpoint openrouter-mimo-v2.5 --system "You are concise." --user "Hello." --log-path logs/usage.jsonl --project-id my-project
Env vars
| Variable | Purpose |
|---|---|
LRL_RUN_TAG |
Stamped on every JSONL event. |
LRL_RUN_LABEL |
Stamped on every JSONL event. |
LRL_CONFIG_PATH |
Default YAML path when load_config() is called with no argument. |
Development
git clone https://github.com/nirmalyaghosh/llm-router-ledger
cd llm-router-ledger
uv sync --extra dev
pytest tests/unit
Verify a local Ollama setup end-to-end with
python examples/smoke_test_ollama.py (see prerequisites at the top
of the script).
License
MIT. See LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file llm_router_ledger-0.2.1.tar.gz.
File metadata
- Download URL: llm_router_ledger-0.2.1.tar.gz
- Upload date:
- Size: 84.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.12.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fcd8b7ce1915e767597940e5f68e0f8b987e0f199592fd707dbb0374bd685c80
|
|
| MD5 |
191f7942dd0b0351d7ab5b7498d6bc24
|
|
| BLAKE2b-256 |
9440656ac002f527950dd282133256eb5113ee3036a8d039c17b3babc09895b1
|
File details
Details for the file llm_router_ledger-0.2.1-py3-none-any.whl.
File metadata
- Download URL: llm_router_ledger-0.2.1-py3-none-any.whl
- Upload date:
- Size: 50.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.12.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ef68e0e9985dae786458f250f5df7dc2b9963756cb2cec5c3d503c43dede5aa8
|
|
| MD5 |
bb3559cf4ec6fb21a1eb76080750695d
|
|
| BLAKE2b-256 |
28a9974c14041ed630f14f200867614f8d09ceefeb7cb1317604253f43803598
|