batchwatch — Python client
Client for batchwatch.dev: crowdsourced measurement of queue time on LLM batch APIs.
Batch endpoints cost 50% of the synchronous ones, but "completes within 24 hours" is impossible to plan around. batchwatch measures what the queue actually does and answers one question: should I use batch for this job?
Standard library only. No dependencies, and none planned.
Install
Not published to PyPI yet. Until it is:
pip install git+https://github.com/batchwatch/client#subdirectory=python
or copy src/batchwatch/ into your project — it is two files.
Two lines
from batchwatch import Batchwatch
bw = Batchwatch(token="tk_...") # token optional; falls back to $BATCHWATCH_TOKEN
# 1. before you submit — does this belong in the queue?
if bw.should_batch("gpt-5.6-sol", max_wait="15m"):
job = client.batches.create(...)
else:
answer = client.chat.completions.create(...)
# 2. measure it, so the next person gets a better answer
with bw.track("gpt-5.6-sol", input_tokens=9720) as t:
result = wait_for(job)
t.done(output_tokens=result.usage.completion_tokens)
Get a key with no email and no card:
curl -X POST https://batchwatch.dev/v1/keys -d '{"label":"my pipeline"}'
It fails open, always
If batchwatch is down, slow, or broken, your job must not notice. That is the first requirement, ahead of collecting any data at all.
- Every submission runs on a daemon thread.
track()does no network I/O on your thread. - Two-second timeout by default (
BATCHWATCH_TIMEOUT). - Every batchwatch error is swallowed and logged at
DEBUGon thebatchwatchlogger. Nothing is printed unless you ask for it. should_batch()is the one synchronous call, because you are waiting for the answer. If it cannot answer, you get your owndefaultback — never a guess. The default isFalse, "run it synchronously": being wrong that way costs money, being wrong the other way blows a deadline.- An exception raised inside your own
withblock is recorded asfailedand re-raised untouched. We swallow our errors, never yours.
tests/test_fail_open.py proves it against a port nothing listens on and
against a socket that accepts but never answers.
It never sends your content
No prompts, no completions, no system prompts, no tool calls, no file names.
The request body is built from a fixed allowlist — provider, model, mode,
endpoint, request count, token counts, timestamps, status — and everything
else is dropped in _scrub() on the way out. There is no field to put text
in.
tests/test_no_content.py asserts it on the bytes a real HTTP server
received, and includes a positive control so the test cannot pass by the
client simply sending nothing.
output_tokens defaults to None, never 0
You know your input tokens. You cannot know your output tokens before the model has answered. So the default is absence, not zero.
Zero is not a harmless placeholder here: output costs five to six times as
much as input, so a saving computed on zero output is systematically too
low — measured at 3.4x too low on a real model — and nothing in the response
would tell you. If you know a ceiling, pass max_tokens instead and the
answer comes back labelled as a ceiling.
Spooling
When a measurement cannot be delivered, the completed record is appended to a
JSONL file and replayed later through POST /v1/calls/complete. Losing
measurements exactly when the network is bad means losing them exactly when
they are most interesting.
- Default path:
$BATCHWATCH_SPOOL, orbatchwatch-spool.jsonlin the system temp directory. SetBATCHWATCH_SPOOL=""or passspool=Noneto turn it off. - The spool is replayed automatically, at most once a minute, right after a
successful call — that is the moment we know the network is up. Call
bw.flush_spool()yourself from a shutdown hook if you want it drained on exit. - Spooling requires a token.
/v1/calls/completetakes your own timestamps, so it is closed to anonymous callers; without a key a spool file could never be sent, and writing one would just leak disk. Without a token, undeliverable measurements are dropped and logged atDEBUG. - The file is capped at 5 MB. Beyond that, measurements are dropped rather than filling your disk.
- A replayed measurement can arrive twice if the original
PATCHreached the server but the response did not. That is deliberate: a duplicate is visible in the dataset, a lost measurement is not. - Threads are handled. Two processes sharing one spool file may send a
record twice — give each process its own
BATCHWATCH_SPOOLif that matters.
Configuration
| Argument | Environment | Default |
|---|---|---|
token |
BATCHWATCH_TOKEN |
none (anonymous) |
base_url |
BATCHWATCH_URL |
https://batchwatch.dev |
timeout |
BATCHWATCH_TIMEOUT |
2.0 seconds |
spool |
BATCHWATCH_SPOOL |
<tempdir>/batchwatch-spool.jsonl |
enabled |
— | True |
enabled=False turns every network call into a no-op, which is what you want
in CI.
API
should_batch(model, max_wait=None, default=False, **kw) -> booladvice(model, max_wait=None, provider="openai", input_tokens=None, output_tokens=None, max_tokens=None, risk="p90") -> dict | Nonewait_now(model, provider="openai", mode="batch") -> dict | Nonetrack(model, provider="openai", mode="batch", requests=1, input_tokens=None, endpoint=None)— context managert.done(output_tokens=None, status="completed", ttfb_ms=None)t.failed()t.started(input_tokens=...)when the count is only known after submission
flush(timeout=5.0) -> bool— wait for outstanding submissions before exitflush_spool(timeout=None) -> int— send what is on disk, returns accepted
Tests
python -m pytest -q
25 tests, no network beyond loopback. They start real HTTP servers on
ephemeral ports rather than monkeypatching urllib: the thing under test is
network behaviour, so the network should be in the test.
Licence
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file batchwatch-0.2.0.tar.gz.
File metadata
- Download URL: batchwatch-0.2.0.tar.gz
- Upload date:
- Size: 32.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.16
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f4d5b76ac515236a2ed0b1c3de4c34a02908ec90779771343ac87fca2ad9b9bd
|
|
| MD5 |
40dc20139de514ada5ce8410dc3d08eb
|
|
| BLAKE2b-256 |
c79bc62cff77392eb23b8e9b4936660d9f44d1f97b1c4d3c3d671aabb99e6b2f
|
File details
Details for the file batchwatch-0.2.0-py3-none-any.whl.
File metadata
- Download URL: batchwatch-0.2.0-py3-none-any.whl
- Upload date:
- Size: 22.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.16
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
521f2a8fc48170df665a732792e8262a6a58305a07a082fa890ac407bf3ff269
|
|
| MD5 |
c897acfcacc131a5136f26936aeb2e30
|
|
| BLAKE2b-256 |
cbeca2578520d097a6306e93c36bbab6d9f984d64fb0cfb196c0b8b0140c962d
|