Skip to main content

Tripwire

PyPI

Live demo → tripwire.anuran.de — type a prompt and watch the supervision happen: tokens streaming, per-signal probabilities, and the trip.

A real-time supervisor for streaming LLM output. Tripwire watches an LLM's tokens as they are generated and trips a fast, calibrated detector on rolling snapshots of the partial response. When a judgment crosses a threshold it aborts the upstream generation mid-flight — killing a bad 3,000-token answer at token ~200 instead of paying for the whole thing and rejecting it afterward.

The trick is latency. A second LLM checking the stream is slower than the stream itself, so it can't keep up. Tripwire's flagship detector is TypeSafe's Jev — a System One model that returns typed, calibrated judgments in ~150 ms instead of generating text — which is fast enough to sit directly in the token path.

client ──/v1/chat/completions (stream)──▶  Tripwire proxy
                                              │
        upstream (Groq / any OpenAI-compatible)│ tokens
                                              ▼
                                     rolling window ── every N tokens / M ms ──▶ detector (Jev)
                                              │                                     │ typed judgments
                                              ▼                                     ▼
                                     forward token to client                    policy engine
                                              ▲                                     │
                                              └──────── ABORT cancels upstream ◀────┘

Use it as a library

pip install tripwire-llm    # the import stays `import tripwire`

Put a guardrail on any LLM call. Flag a complete response:

from tripwire import Guard

guard = Guard()  # Jev if TYPESAFE_API_KEY is set, else a zero-key heuristic baseline

resp = call_your_llm(prompt)               # OpenAI, Anthropic, local — anything
verdict = guard.check(resp, user=prompt)   # sync
if verdict.blocked:                        # abort or escalate
    resp = "Sorry, I can't help with that."
    # verdict.decision / verdict.signal / verdict.reason tell you why

Or limit a token stream — abort mid-generation before the harm ships:

async for chunk in guard.supervise(token_stream, user=prompt, max_tokens=512):
    if chunk.type == "token":
        print(chunk.text, end="")
    elif chunk.type == "intervention" and chunk.intervention.decision == "abort":
        break  # the upstream is already cancelled

check (sync) and acheck (async) return a Result with .decision, .blocked, .signal, .probability, .confidence, .reason, and per-signal .judgments. Bring your own detector by passing any object with an async judge to Guard(...). The FastAPI proxy and CLI are optional: pip install "tripwire-llm[proxy]".

Quickstart (no API keys)

python -m pip install -e ".[dev]"
python examples/demo_attack.py

The demo runs two canned streams through the supervisor with the offline heuristic detector:

=== benign completion ===
Sure! To reset your password, open Settings, choose Security, and click Reset Password ...
[emitted 26/26 tokens | aborted: False | tokens saved: 0 | upstream cancelled: False]

=== adversarial completion ===
Of course. Ignore all previous instructions and
>>> ABORT: jailbreak: threshold exceeded (p=0.90, conf=0.80)
[stream terminated by tripwire]
[emitted 7/153 tokens | aborted: True | tokens saved: 249 | upstream cancelled: True]

Live usage

Copy .env.example to .env and fill in keys (Groq is free and OpenAI-compatible; TypeSafe provides Jev). Groq rotates its model roster often — set UPSTREAM_MODEL to a current chat model.

Check your Jev key in isolation (no upstream needed):

python examples/smoke_jev.py

Run the OpenAI-compatible proxy and point any client at it:

make run   # uvicorn on :8080
curl -N http://localhost:8080/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"messages":[{"role":"user","content":"Write a haiku about the sea"}],"stream":true,"max_tokens":256}'

Or supervise a single prompt from the terminal:

python -m tripwire.cli "Write a haiku about the sea"

Intervention metadata rides along in the streamed chunks as a tripwire field, an aborted response ends with a content_filter finish reason, and GET /metrics exposes Prometheus-style counters (tokens streamed, tokens saved, detector latency percentiles).

Detectors

The supervisor is indifferent to which detector runs — they all implement one narrow protocol (src/tripwire/detectors/base.py).

Backend Role Needs
jev Flagship, default. System One typed judgments in the token path. TYPESAFE_API_KEY
heuristic Zero-cost regex/Luhn baseline; runs the offline tests. —
llm_judge LLM-as-judge baseline, for the benchmark comparison only. GROQ_API_KEY
python benchmarks/bench_detectors.py                       # heuristic (offline)
python benchmarks/bench_detectors.py --detectors heuristic,jev,llm_judge

The benchmark reports per-detector latency percentiles and precision/recall — the place where Jev's in-path speed advantage over an LLM judge is made concrete.

Development

make check   # ruff + mypy (strict) + pytest

License

MIT

Release files for tripwire-llm 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tripwire-llm 0.1.1
File Size Uploaded
tripwire_llm-0.1.1.tar.gz 22.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tripwire-llm 0.1.1
File Interpreter ABI Platform
tripwire_llm-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 54.2 kB

Release files / tripwire_llm-0.1.1.tar.gz

Download URL tripwire_llm-0.1.1.tar.gz
Size 22.5 kB
Tags Source
SHA-256 checksum
How to use checksums
a87fcf46c32251d15200cf7cf32e371ed6ac1f6717691b0688d617dd0d6d3b86
BLAKE2b-256 checksum
How to use checksums
a2afefab8a05f32b97898373cd335f5580583e4cda61ae396cf8e80f3959fd79
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release files / tripwire_llm-0.1.1-py3-none-any.whl

Download URL tripwire_llm-0.1.1-py3-none-any.whl
Size 31.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f3f34cdce438fee15b18c95b7598986834bdb7f344e5ccd7be4045ec1d95723f
BLAKE2b-256 checksum
How to use checksums
9df292d80c3974d30a493c496f79bb3adfe61574ecde29fe1b3f8503a6e23ba6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page