Skip to main content

Jiamu Zhang1     Tianze Yang1     Yucheng Shi2     Liang Wu1

1 Nokia, Sunnyvale, CA      2 Tencent Hunyuan

Reverse the option order: the raw logit readout flips its answer, AnyJev L0 gives the same answer both ways
Qwen3-8B on a real BANKING77 item. Every number is a model output.

⚡ Serve it

Three commands take a model off the Hub and put a calibrated decision endpoint in front of it.

pip install "anyjev[hf]"

# 1. keep the blocks a decision needs — usually about two thirds
python -m anyjev.truncate Qwen/Qwen2.5-7B-Instruct 18 ./qwen-b18

# 2. serve it. L2 reads a hidden state, so the pooler hands one back untouched
vllm serve ./qwen-b18 --task embed \
  --override-pooler-config '{"pooling_type":"LAST","normalize":false,"softmax":false}'
from anyjev import Decider, Question
from anyjev.backends.vllm import VLLMBackend

d = Decider(VLLMBackend("http://localhost:8000", "./qwen-b18"), level="L2")
route = Question.choice("Which team should handle this?",
                        ["billing", "technical", "sales", "other"], name="route")

d.fit_head(route, states, labels, layers=[-1])   # 100–300 labels, one closed-form solve
d.decide(ticket, [route])["route"].distribution  # {"billing": 0.81, "technical": 0.07, ...}

Without labels, turn on the rotation budget — recommended for any K-option choice. L0 asks the model once per option rotation so that no option is favoured by its position. Most decisions do not need all K: read them one at a time, stop when the leader is far enough ahead, and the threshold can be calibrated so the answer matches the full cycle's a stated fraction of the time — measured against our own full-strength readout, so it needs no labels at all.

d = Decider(VLLMBackend("http://localhost:8000", "./qwen-b18"), adaptive_shifts=True)
d.calibrate_adaptive(route, unlabelled_tickets, target=0.01)   # a few hundred states, no labels
d.decide_batch(tickets, route)      # diagnostics: shifts_used, stop_threshold

7.2 rotations instead of 18 at a certified 1% disagreement rate, 2.2× the decisions per second on vLLM and 2.3×–2.7× on the transformers backend, accuracy unchanged (docs/rotation_budget.md). It is opt-in in 0.2.0 only because the tables in docs/ were measured before it existed; it becomes the default when they are regenerated.

An L2 deployment is a pooling server plus a few kilobytes of head. No logits, no parsing, no patched engine, and nothing generated. raw / L0 / L1 run the same way against a --task generate server. A head fit through transformers and served by vLLM answers the same as one fit and served on either alone — 99.0% identical answers, mean |dp| 0.0011 on BANKING77-20.

Measure it on your own box instead of trusting ours:

python -m anyjev.pipeline Qwen/Qwen2.5-7B-Instruct --labels-from banking20

That truncates, serves, fits a head, measures accuracy, ECE and ms per decision on held-out states, shuts the server down, and repeats at full depth so there is something to compare against. Timings are a median over --repeats passes with the spread printed next to them, because on a shared machine a single pass can report the same configuration as both faster and slower than the baseline.

Depth is usually a gain, not a trade. Cutting Qwen2.5-7B from 28 blocks to 18 left accuracy slightly higher and calibration better, and was faster: a middle block is a better feature space for a linear head than the last one, where the remaining blocks are busy turning the answer into tokens. --quantization fp8 is available and not recommended — it buys single-question latency and costs accuracy.

✨ What it is

Ask any open LLM a typed question — a choice, a yes/no, a score — and get a decision with a probability you can threshold, read from one prefill of its next-token distribution. Nothing is generated and nothing is parsed. Raw logits change their answer when you reorder the options and their confidence cannot be trusted; AnyJev fixes the first with zero labels and the second with a few hundred.

⚪ raw logits 🔵 L0
zero labels
🟢 L1
+ temperature
Labels required none none 100–500
Answer flips when options are reversed 0.230 0.073 0.077
Accuracy 0.747 0.803 0.807
Calibration error (ECE) 0.240 0.184 0.095
Auto-decidable at ≤5% error 7.7% 46.3% 52.0%

Qwen3-8B, BANKING77 20-way, 300 test items · every ablation

The last row is the point: accuracy moves six points, but the share of traffic you can safely automate goes 7.7% → 52.0%. With raw logits a "0.9" is not trustworthy enough to act on, so everything goes to a human. Once the probability means what it says, you can set a threshold.

📊 With labels: L2

A closed-form head per question, solved on 100–300 labels in seconds — no gradients, the model's weights untouched — and read from one prompt stopped partway down.

model L0, zero labels L2 block cost vs one forward
Qwen3-1.7B 0.494 0.730 18 / 28 0.70×
Qwen3-4B 0.564 0.786 24 / 36 0.69×
Qwen3-8B 0.647 0.771 24 / 36 0.68×
Qwen3-30B-A3B 0.630 0.799 40 / 48 —
Qwen3-32B 0.700 0.798 52 / 64 0.84×

LocalLLaMA/typed-decisions, 20 questions × 300 labels, 2,000 held-out decisions. Pooled ECE 0.03–0.05. Jev 0.727 and fine-tuned Laya 0.768 on the same set, as published by their authors · every cell

A 1.7B at 64% of its depth reaches the number Jev publishes; a 4B ties the fine-tuned 421M Laya. 100 labels already put the 8B head at 0.740. Heads for five Qwen3 models ship in anyjev-heads/, 23 heads per model in one 1.8–4.4 MB file.

A head maintains itself. Only its feature mean and scale move afterwards, re-estimated from unlabelled traffic, so it follows its question across rewordings and option orders on its own — a reworded question drops the Qwen3-8B head from 0.77 to 0.65–0.70 and 30 unlabelled requests bring it back to 0.74–0.75, against 0.77 for a full relabelled refit. New labels are needed only for a new question. How the routing works →

🧠 The levels

How one decision is read: ask a typed question, read it over every cyclic shift of the options, divide out the label prior estimated without labels, and return a decision that carries its level

Level Needs Does Does not
raw nothing restricted softmax over label tokens anything about bias or calibration
L0 nothing averages position bias out over the K rotations, divides out the label prior calibrate the uncertainty
L1 100–500 labels per question temperature scaling on top of L0 change the ranking
L2 100–300 labels per question a closed-form head on the hidden state partway down, one prompt per state transfer to another question or model

Every Decision carries its level, and require="L1" makes downstream code refuse to act on a weaker one. L0 costs K prefills for a K-option choice, or about 7 of 18 with the rotation budget on (adaptive_shifts=True, recommended — see Serve it). L2 costs less than one plain forward.

d.observe(q, state, label) collects labels as they arrive and solves the head by itself at 30, re-solving at 60, 120, … so day 0 runs at L0 with nothing and L2 arrives when the loop has fed it. The contract in full → · the method →

No GPU handy? python -m demo.jev_mode --backend fake runs the whole thing on a synthetic model in under a second.

Jev mode · 2048 and Minesweeper · NanoJev maze · when L0 helps · small models · research log, negative results included

Every number is regenerated from committed JSON (bash scripts/regen_docs.sh); a second run from a clean checkout reproduced every zero-label number bit for bit. Not affiliated with TypeSafe AI or Jev; rows published by their authors were not rerun here.

🧭 Roadmap

  • choice, noul and score from one prefill; L0 with zero labels; L1 artifacts
  • L2: a closed-form head per question, routing, label-free adaptation, observe
  • Shipped heads for five Qwen3 models; a packaged demo
  • 🚧 Speed optimization (ongoing): making every decision cheaper
  • L2 on served engines: vLLM, through an embed server's pooler or a truncated checkpoint (anyjev.pipeline, anyjev.truncate); SGLang not yet
  • Agent-loop evaluation: the same decisions inside a real agent, against the LLM they replace
  • Heads on the Hugging Face Hub, an interactive Space, a technical report
  • More models (Llama, Gemma, Mistral, DeepSeek), span readout beyond 26 options, conformal abstention

Dated plan and help-wanted files: ROADMAP.md.

🔍 Limitations

  • On typed-decisions, "accuracy" is agreement with a teacher LLM. The gold is the mean of three samples of one model; a fresh sample of that teacher agrees with it 0.735 of the time.
  • L2 is per question and per model. Heads fit on other questions do not help a new one, and only Qwen3 heads ship. It needs hidden states, which transformers and a vLLM embed server both provide; other engines do not yet.
  • Calibration cannot fix a model that cannot answer. On maze edges and Minesweeper no readout beats the trivial baseline.
  • L0 is not a free win everywhere. The batch prior costs accuracy when one label dominates (when L0 helps).

Also: at most 26 options in the letter readout (a span readout is on the roadmap, not in the code); coverage at 5% risk is a high-variance estimate at n = 300; the headline tables are Qwen models; every decision here is scored in isolation, not inside an agent loop.

🤝 Contributing and citation

Backends and bench providers are one file each; several are help wanted (ROADMAP.md, CONTRIBUTING.md). Changes: CHANGELOG.md. Credits: CREDITS.md.

@software{anyjev2026,
  title  = {AnyJev: Turn any LLM into a Jev-style decision model},
  author = {Zhang, Jiamu and Yang, Tianze and Shi, Yucheng and Wu, Liang},
  year   = {2026},
  url    = {https://github.com/nokia-applied-research/AnyJev}
}

Apache-2.0, see LICENSE. Datasets keep their own licenses, see THIRD_PARTY.md.

Release files for anyjev 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for anyjev 0.2.0
File Size Uploaded
anyjev-0.2.0.tar.gz 93.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for anyjev 0.2.0
File Interpreter ABI Platform
anyjev-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 162.9 kB

Release files / anyjev-0.2.0.tar.gz

Download URL anyjev-0.2.0.tar.gz
Size 93.0 kB
Tags Source
SHA-256 checksum
How to use checksums
a3f1b3195b8643a8019275fed5f2fffdb385c204ec58684dfaad011e8a7bfecb
BLAKE2b-256 checksum
How to use checksums
68b099b48656612de93e602df36f21bba37bb8e1aa87e68f4c1612924e309cd8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / anyjev-0.2.0-py3-none-any.whl

Download URL anyjev-0.2.0-py3-none-any.whl
Size 70.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0749a4c709069840906678c76efad3c29dafb538be44d91fa46f15cc1bd8486f
BLAKE2b-256 checksum
How to use checksums
4193f1be550c3bf105a469e96cffd65ab7c41373522d4dbf91e42f287dfd7a4a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.0

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page