Skip to main content

jev-ultralightspeed

v0.1.0 · Apache-2.0 · no required dependencies

Faster, cheaper, same accuracy: hundreds of items a second against dozens, with agreement against human labels unchanged

You have a pile of text and one question about each. Fifty thousand support tickets to triage. A quarter of reviews to sort by sentiment. A month of logs to flag. A column to backfill on a table you already have. This answers the question for all of them, with TypeSafe's Jev, in the time it takes to get a coffee.

from jev_ultralightspeed import classify

answers = classify(tickets, "Does this message need a human to act on it today?")
urgent = [a.item for a in answers if a.yes]

15.9x the throughput of one request per item, 41% less money, and the same accuracy. Measured over 30,000 judgements against human labels, in one run, with a script in this repository.

Not for one item at a time. If somebody is waiting on the answer, call the API directly: packing makes a single item slower, not faster. This is for a queue.

Not affiliated with TypeSafe; the hedgehog is a parody and belongs to nobody.

The benchmark

30,000 judgements over the 1,347 completions in dinostomp's xstest-refusal pod, labelled compliance, refusal or partial by two human annotators, each completion seen about 22 times. Two arms, identical but for the shape of the requests. This is not Jev against another model: it is jev-1.13.0 against itself, same question, same items, same criteria.

regular jev jev + ultralightspeed
throughput 41.1 items/s 653.4 items/s
wall clock 12 min 11 s 46 seconds
requests 30,000 942
cost $0.729 $0.430
agreement with the human labels 89.3% (89.0 to 89.7) 89.2% (88.9 to 89.6)
same answer across an item's repeats 99.6% 98.0%
failed 0 0
retried 1 63

15.9x the throughput, 41% less money, and the accuracy is the same: 89.3% against 89.2%, with 30,000 judgements behind each figure and the intervals sitting on top of each other.

pip install "jev-ultralightspeed[fast]"
git clone https://github.com/collapseindex/dinostomp.git ../dinostomp
TYPESAFE_API_KEY=... python bench_eval.py            # about 13 minutes, about $1.20

Three things in that table are worth reading twice.

The baseline is not slow. It is one item per request with the same eight requests in flight, so the 15.9x is against a client that is already parallel. Against an actual sequential loop it is far larger and far less interesting.

Sixty-three retries, and nothing failed. Packing 32 deep with 8 in flight does hit transient limits at volume. The client backs off and sends them again, which is why this arm came in at 653 items/s rather than the 800 it reaches on short bursts. Retries are counted in usage.retries, so this is a number rather than an absence of complaints.

Aggregate accuracy is identical; individual answers are slightly less repeatable. Ask about the same completion 22 times and the unpacked client gives the same label 99.6% of the time, the packed one 98.0%. So pack freely when you want the total, and keep the pack size fixed when you are comparing item by item across runs.

Shorter items do better than this on every axis, because the per-item text is a smaller share of each request: a million synthetic support messages (138 tokens each) ran at 1,135 items/s, 14.7 minutes and $4.99 for the lot, over 31,400 requests with nothing retried. soak.py --items 1000000 reproduces that, and it is a different corpus, so it is a footnote here rather than a headline.

How

Nothing clever. Four things the obvious loop does not do, in order of how much they gave:

  1. Pack. Several items go in one request as item_1..item_N, each with its own question that names the item it judges. One round trip covers thirty-two items, and the shared overhead is paid once instead of thirty-two times. This is where both the speed and the token saving come from.
  2. Parallel. Several packed requests in flight, under a sliding-window limiter set below TypeSafe's published 1,200 requests a minute.
  3. One connection, kept open, multiplexed. With httpx and h2 installed, every request in flight shares a single HTTP/2 connection on one event loop, which measured about twice a thread per connection on HTTP/1.1; keeping it open between calls rather than rebuilding it per batch was worth another 2.5x. The transport idea is lifted from browser-use/jev-ultrafast, who got there first.
  4. Never ask twice. Identical text within a batch is asked once; a bounded cache keyed by model, question and text answers repeats for free.

Install

pip install "jev-ultralightspeed[fast] @ git+https://github.com/collapseindex/jev-ultralightspeed"
export TYPESAFE_API_KEY=...

The [fast] extra pulls in httpx and h2, which measured about twice the standard library path. Without it the client still works on the standard library alone:

pip install "git+https://github.com/collapseindex/jev-ultralightspeed"

Not on PyPI yet.

Use

from jev_ultralightspeed import Client

client = Client(pack=8, workers=8)     # the defaults are pack=8, workers=4
client.warm()                          # open the connections before the work arrives

answers = client.classify(
    messages,
    "Does this message need a human to act on it today?",
    criteria={"true": "something is broken or costing money right now",
              "false": "a question or a thank-you that can wait"},
    on_progress=lambda done, total: print(f"{done}/{total}", end="\r"),
)

for answer in answers:
    print(answer.label, round(answer.p, 2), answer.item[:60])

print(client.usage)     # 256 items in 1.16s (221.1/s, 32 requests, 156 tokens/item, $0.00166)

Pick-one questions work the same way:

answers = classify(tickets, "Which team should handle this?", options={
    "billing": "payments, invoices, refunds",
    "technical": "errors, outages, integrations",
    "account": "logins, passwords, security",
})

Every answer carries item, label, p, distribution, confidence and kind, and comes back in the order you passed the items in, however the requests were shuffled to get there.

Knobs

argument default what it does
pack 8 items per request. Higher is faster and cheaper, and slower per request. On short items: pack 8 does 320 items/s, pack 32 does 1,130.
workers 4 requests in flight.
transport auto http2 when httpx is installed, otherwise threads.
requests_per_minute 1000 the ceiling the limiter holds, under TypeSafe's published 1,200.
cache True answer repeats from memory, keyed by model, question and text.
model jev-latest passed straight through.
url the Jev endpoint point it at a gateway or a mock.

What it does not do

  • It does not change your question. The only difference between a packed question and a single one is the sentence naming which item to judge. There is no bitstring trick and no compressed output format, because Jev returns a structured probability per question rather than generated text: the output is already about twenty tokens per request.
  • It does not cache across processes. The cache lives in the client, in memory, bounded at 10,000 entries.
  • It does not hide failures. Retries cover 429, 500, 502, 503, 504 and 529 with backoff; anything else is raised with what the API said.
  • It is not an eval harness. It makes a judge fast, not trustworthy. See Related below.

Development

pip install pytest
python -m pytest tests -q        # 21 tests, no network, no key needed

TYPESAFE_API_KEY=... python bench_eval.py            # the table above, ~13 min, ~$1.20
TYPESAFE_API_KEY=... python bench.py --items 256     # pack and concurrency sweep, ~5 cents
TYPESAFE_API_KEY=... python soak.py --items 100000   # sustained load, ~50 cents

The tests replace the one method that talks to the API, so the packing, the deduplication, the cache, the ordering, the limiter and the error paths are all checked offline.

Three tools, one workflow, all Apache-2.0:

Costs are input tokens at TypeSafe's published $0.042 per million for jev-1.13.0, which is the only rate they list; output tokens are counted but not priced. Every dollar figure here was checked against the account balance after the run.

  • dinostomp is the harness the benchmark above was measured against: pods of labelled items, pre-registered thresholds, a checks registry and a findings ledger. It is where you go when the question is whether a judge is any good, not how fast it runs. The 1,347 labelled completions in the table are one of its audit pods.
  • jev-builder writes the request in the first place: paste your text, describe the question, and get something you can paste here.
  • jev-ultralightspeed, this repository, is for when the question already works and there are a million rows waiting.

If any of it saves you an afternoon, sponsorship keeps it maintained. Not required, and nothing here is gated.

Security

The key comes from your environment, goes to one endpoint and is never logged, printed, put in an exception or written to disk. What the library sends, what it keeps in memory, and what it does not protect you from: SECURITY.md.

Contributing

Issues and pull requests are welcome. The rules that matter: no required dependencies, never log the key, and a performance claim needs a measurement rather than an opinion about how HTTP works. See CONTRIBUTING.md.

License

Apache-2.0. Not affiliated with TypeSafe.

Release files for jev-ultralightspeed 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jev-ultralightspeed 0.1.0
File Size Uploaded
jev_ultralightspeed-0.1.0.tar.gz 23.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for jev-ultralightspeed 0.1.0
File Interpreter ABI Platform
jev_ultralightspeed-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 42.2 kB

Release files / jev_ultralightspeed-0.1.0.tar.gz

Download URL jev_ultralightspeed-0.1.0.tar.gz
Size 23.9 kB
Tags Source
SHA-256 checksum
How to use checksums
7e62b2bfd5a9f5a1ba1fdacb01692610f36b015f872d1030b846cec9f43852c4
BLAKE2b-256 checksum
How to use checksums
94c3f0ac3ceca10571e8f92347fcc88ca6f362e952882c90dc1332f6ae8566fa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.

Transparency log

Release files / jev_ultralightspeed-0.1.0-py3-none-any.whl

Download URL jev_ultralightspeed-0.1.0-py3-none-any.whl
Size 18.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cb8fb9606eaec731d305a1b5c97388c2a275f69d38a1bde4c7e1bf8bed532384
BLAKE2b-256 checksum
How to use checksums
d4e0dedfb9103e7e15e5ceb7023e380c122fcda65e6e4c98268cae23d88ce827
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.

Transparency log

Release history Release notifications | RSS feed

0.17.0

2 release files

0.16.1

2 release files

0.16.0

2 release files

0.15.2

2 release files

0.15.1

2 release files

0.15.0

2 release files

0.14.1

2 release files

0.14.0

2 release files

0.13.0

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.5

2 release files

0.10.4

2 release files

0.10.3

2 release files

0.10.2

2 release files

0.10.1

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page