Skip to main content

Cut LLM token cost by reshaping the wire format: prompt-cache breakpoint placement and positional encoding for repeated structured output.

Project description

leanwire

Cut LLM token cost by reshaping the wire format, not the content. Two independent levers, both of which leave what the model actually decides alone.

Zero runtime dependencies. Works with the Anthropic SDK or raw HTTP.

pip install leanwire

1. leanwire.cache — stop re-paying for your transcript

A long agent conversation is re-sent on every turn. If cache_control is only on your system prompt and tools, the static prefix caches and the transcript is billed at full input price on every single call — while your aggregate cache-read numbers look great.

from leanwire.cache import CachePolicy

policy = CachePolicy(ttl="5m", model="claude-opus-4-8")

request = {"model": "claude-opus-4-8", "system": [...], "tools": [...],
           "messages": messages}
placement = policy.apply(request)      # your request is not mutated
response = client.messages.create(**placement.request)

apply() spends whatever breakpoint budget is left after your own tools/system markers (the API allows 4), spacing markers lookback blocks apart from the tail backwards so a valid read point always exists inside the 20-block lookback window — statelessly, with no need to remember where the last request put them.

It refuses to act rather than act wrongly: no budget left, no dict content blocks, or a prefix below the model's minimum cacheable size all produce a no-op with placement.skipped_reason set.

Find out if you have this problem in one loop

from leanwire.cache import CacheAudit

audit = CacheAudit()
for response in your_agent_run():
    audit.observe(response.usage)

print(audit.report)
40 calls | uncached 1,200,000 | read 2,000,000 | write 0 | out 40,000 | 62.5% of input served from cache
  [!] uncached input grows 10,000 -> 50,000 tokens across the run: the conversation
      transcript is being re-billed at full price every call while cache reads stay
      flat -- a static prefix is cached but the messages are not. Place a
      message-level breakpoint

Also detects nothing-cached, write-but-never-read (a timestamp or uuid in your prefix), and bulk prefix rebuilds (TTL expiry). Small per-turn writes are correct and are not flagged.

2. leanwire.codec — stop re-emitting field names

When a model returns N records sharing a schema, it re-emits every key N times.

from leanwire.codec import RecordCodec

codec = RecordCodec.infer(sample_records)     # or build Fields explicitly
codec.verify(sample_records)                  # raises unless round-trip is exact

schema = codec.json_schema()                  # put on output_config.format
prompt_hint = codec.legend()                  # field order + enum codes

records = codec.decode(response_rows)         # back to your original dicts

{"column_name": "loc_na", "score": 10, "criterion_met": true, "hallucination_risk": "low", ...} becomes ["loc_na", 10, true, "l", ...].

Lossless by construction and tested as such: fields that never vary leave the wire and are re-injected on decode, low-cardinality strings become single-character codes, and original key order is restored.

stats = codec.measure(records, token_counter)
print(stats)   # 40 records: 4,860 -> 2,489 tokens (48.8% smaller)

Measure before you promise. Savings depend entirely on how much of your payload is packaging versus free text. In our own testing the same codec gave 49% on records with short scalar fields and 25% on records dominated by long prose -- a 2x spread on identical code. measure() exists so you get a real number on your data rather than an estimate. Never quote a figure you have not run.

3. leanwire.accounting

from leanwire.accounting import cost_of
cost_of(response.usage, "claude-opus-4-8")     # -> Cost(input=..., cache_read=..., ...)

Current first-party prices, with the 1.25x (5m) / 2x (1h) cache-write and 0.1x cache-read multipliers applied.

Which lever applies to you

symptom lever
Long multi-turn agent, input tokens climbing per call cache
Cache reads look high but the bill still grows cache — run CacheAudit
Model returns many records with the same schema codec
Output is most of your spend codec
Single short calls, no repetition neither; measure before optimising

Caveats worth reading

  • Cache placement changes billing metadata only — the model sees a byte-identical prompt. It needs no accuracy evaluation.
  • The codec changes the output contract. It is lossless in encoding, but you are asking the model to emit a different shape, so evaluate that it still fills the fields correctly on your own data before rolling out.
  • Minimum cacheable prefix is model-dependent and not monotonic across generations (512 on Opus 5, 1024 on Opus 4.8, 4096 on Opus 4.6). Pass model= and a token_counter and the policy will skip rather than pay a write that never caches.

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

leanwire-0.1.0.tar.gz (20.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

leanwire-0.1.0-py3-none-any.whl (16.2 kB view details)

Uploaded Python 3

File details

Details for the file leanwire-0.1.0.tar.gz.

File metadata

  • Download URL: leanwire-0.1.0.tar.gz
  • Upload date:
  • Size: 20.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.2

File hashes

Hashes for leanwire-0.1.0.tar.gz
Algorithm Hash digest
SHA256 178fa84e1c1e7bed9497c00569232ad3db891150413fab6680d82a3fd6467c27
MD5 274ad6c0a7a7b5b84923e76a74b0fc4a
BLAKE2b-256 dcb87687bd77de7a0d0370698e9d486302c4414503333fde293e3cdbd36cbddf

See more details on using hashes here.

File details

Details for the file leanwire-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: leanwire-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 16.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.2

File hashes

Hashes for leanwire-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 c56efa1b7db806ddec1b49b2f4f6626a688ff1b83fb08f914a6d2f758460d9d6
MD5 add36e41f5672bf17e3323f95c8332c1
BLAKE2b-256 6ef0f0d4e653766c380d7d4e65fa6c837d190003bd2a2653e08d3f227b4eba0c

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page