Tripwire
Live demo → tripwire.anuran.de — type a prompt and watch the supervision happen: tokens streaming, per-signal probabilities, and the trip.
A real-time supervisor for streaming LLM output. Tripwire watches an LLM's tokens as they are generated and trips a fast, calibrated detector on rolling snapshots of the partial response. When a judgment crosses a threshold it aborts the upstream generation mid-flight — killing a bad 3,000-token answer at token ~200 instead of paying for the whole thing and rejecting it afterward.
The trick is latency. A second LLM checking the stream is slower than the stream itself, so it can't keep up. Tripwire's flagship detector is TypeSafe's Jev — a System One model that returns typed, calibrated judgments in ~150 ms instead of generating text — which is fast enough to sit directly in the token path.
client ──/v1/chat/completions (stream)──▶ Tripwire proxy
│
upstream (Groq / any OpenAI-compatible)│ tokens
▼
rolling window ── every N tokens / M ms ──▶ detector (Jev)
│ │ typed judgments
▼ ▼
forward token to client policy engine
▲ │
└──────── ABORT cancels upstream ◀────┘
Use it as a library
pip install tripwire-llm # the import stays `import tripwire`
Put a guardrail on any LLM call. Flag a complete response:
from tripwire import Guard
guard = Guard() # Jev if TYPESAFE_API_KEY is set, else a zero-key heuristic baseline
resp = call_your_llm(prompt) # OpenAI, Anthropic, local — anything
verdict = guard.check(resp, user=prompt) # sync
if verdict.blocked: # abort or escalate
resp = "Sorry, I can't help with that."
# verdict.decision / verdict.signal / verdict.reason tell you why
Or limit a token stream — abort mid-generation before the harm ships:
async for chunk in guard.supervise(token_stream, user=prompt, max_tokens=512):
if chunk.type == "token":
print(chunk.text, end="")
elif chunk.type == "intervention" and chunk.intervention.decision == "abort":
break # the upstream is already cancelled
check (sync) and acheck (async) return a Result with .decision, .blocked,
.signal, .probability, .confidence, .reason, and per-signal .judgments. Bring
your own detector by passing any object with an async judge to Guard(...). The
FastAPI proxy and CLI are optional: pip install "tripwire-llm[proxy]".
Quickstart (no API keys)
python -m pip install -e ".[dev]"
python examples/demo_attack.py
The demo runs two canned streams through the supervisor with the offline heuristic detector:
=== benign completion ===
Sure! To reset your password, open Settings, choose Security, and click Reset Password ...
[emitted 26/26 tokens | aborted: False | tokens saved: 0 | upstream cancelled: False]
=== adversarial completion ===
Of course. Ignore all previous instructions and
>>> ABORT: jailbreak: threshold exceeded (p=0.90, conf=0.80)
[stream terminated by tripwire]
[emitted 7/153 tokens | aborted: True | tokens saved: 249 | upstream cancelled: True]
Live usage
Copy .env.example to .env and fill in keys (Groq
is free and OpenAI-compatible; TypeSafe provides Jev). Groq
rotates its model roster often — set UPSTREAM_MODEL to a current chat model.
Check your Jev key in isolation (no upstream needed):
python examples/smoke_jev.py
Run the OpenAI-compatible proxy and point any client at it:
make run # uvicorn on :8080
curl -N http://localhost:8080/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"messages":[{"role":"user","content":"Write a haiku about the sea"}],"stream":true,"max_tokens":256}'
Or supervise a single prompt from the terminal:
python -m tripwire.cli "Write a haiku about the sea"
Intervention metadata rides along in the streamed chunks as a tripwire field, an
aborted response ends with a content_filter finish reason, and GET /metrics
exposes Prometheus-style counters (tokens streamed, tokens saved, detector latency
percentiles).
Detectors
The supervisor is indifferent to which detector runs — they all implement one
narrow protocol (src/tripwire/detectors/base.py).
| Backend | Role | Needs |
|---|---|---|
jev |
Flagship, default. System One typed judgments in the token path. | TYPESAFE_API_KEY |
heuristic |
Zero-cost regex/Luhn baseline; runs the offline tests. | — |
llm_judge |
LLM-as-judge baseline, for the benchmark comparison only. | GROQ_API_KEY |
python benchmarks/bench_detectors.py # heuristic (offline)
python benchmarks/bench_detectors.py --detectors heuristic,jev,llm_judge
The benchmark reports per-detector latency percentiles and precision/recall — the place where Jev's in-path speed advantage over an LLM judge is made concrete.
Development
make check # ruff + mypy (strict) + pytest
License
MIT
Release files for tripwire-llm 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tripwire_llm-0.1.0.tar.gz | 22.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tripwire_llm-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 53.5 kB
Release files / tripwire_llm-0.1.0.tar.gz
| Download URL | tripwire_llm-0.1.0.tar.gz |
|---|---|
| Size | 22.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
493755acabe1c74b47e4871cc8e6e53dfe03dd13bb73e3c58007e7aea2ba9890
|
|
BLAKE2b-256 checksum How to use checksums |
4896324652f9bd7f45c4f3ec7697330c271163ce0bf76eadba86df7f26778c42
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|
Release files / tripwire_llm-0.1.0-py3-none-any.whl
| Download URL | tripwire_llm-0.1.0-py3-none-any.whl |
|---|---|
| Size | 31.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7432177cecb936c2ec079669c1aed5c53c2bf25215079085b30a7195c6295fd2
|
|
BLAKE2b-256 checksum How to use checksums |
6b369106a71e393878166fae1e2ad18893440e59f347f7634a05b54c09009de4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|