AdmitPerf
Standardized admission control for LLM inference.
Bring your own infra. Your gateway already has raw metrics. AdmitPerf turns them into signals, policies decide on signals, and every decision is recorded in a form a report can compare across deployments.
It does not provision, serve, scrape, or generate load.
admitperf.core — the part that runs in your request path — imports nothing but
the standard library, opens no socket and spawns no subprocess. The CLI needs
click; core does not, and a test checks that by importing it with click blocked.
Three things
| Metrics | whatever you scrape. Raw names, hundreds of keys, any stack. Yours. |
| Signal | a named quantity, and how to read it from your metrics. Ours, so kv_pressure = 0.93 means the same thing in two deployments. |
| Policy | decides on signals. Four ship with AdmitPerf; yours is an equal citizen. |
Use it
from admitperf import Policy
from admitperf.core.signals import KV_PRESSURE, QUEUE_DEPTH
class KvWall(Policy):
name = "kv_wall"
def decide(self, metrics):
if KV_PRESSURE.read(metrics) >= self.threshold:
return self.reject("kv_pressure")
if QUEUE_DEPTH.read(metrics) > self.max_waiting:
return self.defer("queue_depth", retry_after_ms=200)
return self.admit()
In your gateway:
policy = KvWall(threshold=0.90, max_waiting=32, log="decisions.jsonl")
d = policy(raw_metrics, request_id=rid) # whatever you scraped
if not d.admitted:
return Response(d.status, retry_after=d.retry_after_ms)
d.status is derived from the reason — 503 for capacity, 429 for client-attributable
causes — so your refusals are comparable with anyone else's without you choosing a
code.
Your metrics, whatever they are called
A signal tries its sources in order. Add yours and it goes first:
from admitperf.core.signals import KV_PRESSURE
KV_PRESSURE.add_source("acme.cache.used_frac") # your name
KV_PRESSURE.add_source(lambda m: m["blocks_used"] / m["blocks_total"]) # computed
Out of the box it already reads vLLM (kv_cache_usage_perc, and the older
gpu_cache_usage_perc, so a version difference is a non-event) and DCGM — including
the 0–100 to fraction conversion, because every host getting that wrong differently
is how a shared metric name stops meaning anything.
Two rules it will not break
Absence is never zero. A signal nothing supplies reads None. A KV pressure of
0.0 claims the cache is empty, which looks like headroom — so the policy would
admit everything while appearing to work.
A value outside its range is skipped, not clamped. Map a 0–100 metric to a
fraction and you get None, not a cluster that appears permanently saturated.
What gets recorded
One JSON line per decision, holding every signal and the raw metrics — not just the one your policy read. That is what lets a report say "your signal never moved, but queue depth hit 61", and what makes replaying a different policy over your own production trace possible at all.
policy.outcome(rid, ttft_ms=418, ok=True) # optional; without it, no goodput
Shadow mode, and a counterfactual
Call it and ignore the verdict: recording is unconditional and enforcement is yours, so that is a complete shadow deployment with no flag.
policy = KvWall(threshold=0.90, enforce=0.5) # half the traffic governed
Both arms in one run under identical conditions, split deterministically by request id so a retry is treated the same way twice. In production it caps the blast radius.
Why
Across sixteen admission-primary papers surveyed, no two share a baseline, engine, workload, or SLO definition — and not one reports the observed range of the quantity its policy reads. So no reader can tell which published results describe a policy acting and which describe a policy that never got the chance.
The arithmetic matters as much as the literature. On one A100-40GB serving
Qwen2.5-7B, the KV pool holds roughly 384k tokens. At max_num_seqs=64 with a
2,168-token request, resident tokens cap near 139k — so kv_used_fraction cannot
exceed about 0.35, and a policy thresholded at 0.90 is unreachable at any arrival
rate. A run like that reports numbers indistinguishable from no policy at all.
That is the question AdmitPerf answers first: could the policy have fired at all?
What it gets you
The whole point: two logs, one without admission control and one with.
admitperf demo # no GPU, no cluster, 2 seconds
admitperf demo --repeats 5 # with an error bar
The policy refused 240 of 400 requests (60.0%).
p95 TTFT of served requests is 15.17x lower: 2730ms -> 180ms
goodput is DOWN: 0.480 -> 0.400 (of offered)
Goodput fell, so the policy refused requests the cluster could have
served. Faster tails bought at that price are not a win.
A 15× better tail, and the policy is worse. A KV threshold picked without reference to the SLO sheds requests the cluster could still have served in time. The same comparison against a queue bound derived from the SLO:
The policy refused 110 of 400 requests (27.5%).
p95 TTFT of served requests is 1.67x lower: 2730ms -> 1630ms
goodput is up: 0.480 -> 0.670 (of offered)
A report showing only latency would have picked the first policy. That is the argument
for comparing, and goodput divides by requests offered rather than admitted —
divide by admitted and refusing 95% of traffic reads as 1.00.
Running it again and again
One run of each policy has no error bar, and a gap smaller than the spread between
runs is not a result. --repeats writes one log per run, and compare reads them all:
admitperf compare --experiment demo
p95 TTFT of served requests is 1.78x lower: 2630 (2530-2830)ms -> 1480 (1430-1480)ms
goodput is up: 0.542 (0.530-0.573) -> 0.723 (0.703-0.755) (of offered)
ok the arms separate across 5 repeats
Four checks come before any of those numbers: the policy fired in every repeat,
every run faced the same load, outcomes were recorded, and the two arms' observed
ranges do not overlap. With one run per policy the last cannot be asked, and the page
says so instead of answering it. admitperf compare --check exits non-zero when any
fails.
Naming a measurement
Every decision carries the identity you gave it, so a result can be found from its log and vice versa:
KvThreshold(
threshold=0.90,
experiment="kv-wall-8k", # the question being asked
run="r1", # which repeat
notes="1xA100-40 · 8k prompts", # what a log cannot know
provenance={"threshold": "fitted"}, # fitted, inherited, or default
log="...",
)
Grouping is the claim: two policies in one experiment assert they faced the same conditions. That is declared rather than inferred from a directory, because otherwise moving a file silently changes what the result says.
Everything written lands in one place, and the path mirrors the identity:
experiments/<experiment>/<policy>/r<n>.jsonl
admitperf experiments # every experiment found here, and its policies
The CLI
Three verbs, and none of them is in a request path.
admitperf signals # what can your stack already feed?
admitperf watch http://host:8000/metrics --for 1h # record it, no code change
admitperf report trace.jsonl # the finding
admitperf experiments # what has been recorded here
admitperf compare --experiment kv-wall-8k # what did it buy?
admitperf dashboard # all of it, in a browser
Start with watch. It answers could a policy have fired here, and which signal
actually moved before you touch your gateway:
THE FINDING — UNKNOWN (3600 state samples)
signal min p50 p95 max n
kv_pressure 0.07 0.21 0.44 0.44 3600
queue_depth 0 3 47 61 3600
gpu_util 0.11 0.88 0.96 0.99 3600
prefix_hit_rate -- never supplied by your metrics
KV pressure never passed 0.44, so a policy thresholded at 0.90 could not have fired no matter how it was written. Queue depth hit 61. That is the finding, and it cost no integration.
admitperf report --check exits non-zero unless the log can support a claim, which
makes it usable in CI.
On a real fleet
Nothing here provisions or drives load — that is your infra's job. AdmitPerf's part is the last two commands.
# --- on the GPU box: your infra, your tooling ---
bash infra/setup/lambda_vllm.sh # serves Qwen2.5-7B on :8000
# --- from your laptop ---
ssh -L 8000:127.0.0.1:8000 -N ubuntu@$HOST # Lambda allows SSH only
admitperf watch http://127.0.0.1:8000/metrics --for 10m -o trace.jsonl &
vllm bench serve --base-url http://127.0.0.1:8000 \
--model Qwen/Qwen2.5-7B-Instruct \
--random-input-len 8192 --random-output-len 512 --ignore-eos \
--num-prompts 150 --request-rate 5
admitperf report trace.jsonl
--ignore-eos is not optional. Without it the model stops whenever it stops, and
on a filler prompt that can be 40 tokens instead of 512. The token count barely moves,
but duration collapses by an order of magnitude — and concurrency is rate × duration,
so the cache never fills and the whole run comes back inert for the wrong reason.
8192 input, not 2048. From the arithmetic above: at max_num_seqs=64 a short
request caps kv_used_fraction near 0.35. Requests need to be longer than about
6k tokens before KV binds before the scheduler does, and only then can a
KV-pressure policy fire at all.
On real hardware: not yet measured. The figures above come from the simulated
engine behind admitperf demo, which is honest about what it models — bounded slots, a FIFO
queue, KV pressure driven by resident tokens, Poisson arrivals and jittered service —
and is checked by tests, so the numbers in this README cannot go stale silently and
repeats cannot quietly become identical. Earlier figures from real runs were
removed rather than carried forward, because they could not be reproduced.
How it connects to the survey
The survey behind this found that of seven reporting items across sixteen admission-primary papers, six are fully populated and the seventh — signal liveness — is empty for every one of them.
So admitperf report checks a run against all seven, using the survey's own wording
for what each asks, and leads with the one nobody fills:
✓ 3. Signal liveness kv_pressure reached 0.907 against a threshold of 0.9
✓ 4. Repeats and spread 3 runs, with the spread reported
✓ 6. Metric definition goodput over OFFERED, so a refusal is a miss
✓ 7. Configuration from notes=
— 1. Capacity-relative load no measured serveable capacity recorded
— 2. Calibrated deadlines no unloaded-latency baseline recorded
— 5. Parameter provenance values recorded, but not where they came from
— is deliberately neither a pass nor a failure: it means this log does not carry
the fact. The survey's finding is that the corpus is silent on these, so silence
has to be its own state.
A policy also declares the taxonomy axes — unit, setting, slo_awareness,
signal_quantity, signal_structure — which is what lets a result here be placed
beside a published one. signal_structure is load-bearing rather than decorative: a
dual gate can hold its first signal above threshold for a whole run and never fire,
so one signal's range does not establish liveness for it.
The dashboard has a Terminology section defining every term a report can print, tagged where the wording is the survey's. It exists because two pairs cause nearly all the confusion here and both look interchangeable: offered versus admitted as a denominator, and INERT versus UNKNOWN.
Layout
src/admitperf/
core/ runs in YOUR request path. Stdlib only, no sockets. One class per file.
signal.py Signal signals.py the signals we ship
policy.py Policy decision.py Decision
verdict.py Verdict reasons.py reason -> status code
log.py Log
policies/ four baked-in policies, one per file
watch.py Watch report.py Report
comparison.py Comparison items.py the survey's seven items
experiment.py Experiment policy_runs.py PolicyRuns
discover.py finds logs terminology.py every term defined
demo.py a simulated run, no hardware needed
cli.py the CLI dashboard/ the Streamlit reader
infra/ one worked example of a host. NOT part of the package.
admitperf.core imports nothing but the standard library, opens no socket, and
spawns no subprocess — tests/test_layering.py enforces each, because that is what
makes it safe to install in a gateway.
make test · make lint
Metadata
Release files for admitperf 0.0.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| admitperf-0.0.2.tar.gz | 825.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| admitperf-0.0.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 904.1 kB
Release files / admitperf-0.0.2.tar.gz
| Download URL | admitperf-0.0.2.tar.gz |
|---|---|
| Size | 825.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7a2b004407f8aed00dcd30f7899bd9ebdac2fb557e34222330c7da8c5a0b4c32
|
|
BLAKE2b-256 checksum How to use checksums |
afa2855c989c59a3d52c18f1d5ca198a4a7b651826344408ee584c25fdba9141
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency logRelease files / admitperf-0.0.2-py3-none-any.whl
| Download URL | admitperf-0.0.2-py3-none-any.whl |
|---|---|
| Size | 78.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
81222db60e2c9bb6faa36bba62f8c08b91b36ad757b2ab7e3170c143e233db13
|
|
BLAKE2b-256 checksum How to use checksums |
7bb873bdd4f958f3775a3917a6d5122aee16404fae7cb0d02ec52a45ebe7659f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency log