Skip to main content

Tripwire

ci

Live demo → tripwire.anuran.de — type a prompt and watch the supervision happen: tokens streaming, per-signal probabilities, and the trip.

A real-time supervisor for streaming LLM output. Tripwire watches an LLM's tokens as they are generated and trips a fast, calibrated detector on rolling snapshots of the partial response. When a judgment crosses a threshold it aborts the upstream generation mid-flight — killing a bad 3,000-token answer at token ~200 instead of paying for the whole thing and rejecting it afterward.

The trick is latency. A second LLM checking the stream is slower than the stream itself, so it can't keep up. Tripwire's flagship detector is TypeSafe's Jev — a System One model that returns typed, calibrated judgments in ~150 ms instead of generating text — which is fast enough to sit directly in the token path.

client ──/v1/chat/completions (stream)──▶  Tripwire proxy
                                              │
        upstream (Groq / any OpenAI-compatible)│ tokens
                                              ▼
                                     rolling window ── every N tokens / M ms ──▶ detector (Jev)
                                              │                                     │ typed judgments
                                              ▼                                     ▼
                                     forward token to client                    policy engine
                                              ▲                                     │
                                              └──────── ABORT cancels upstream ◀────┘

Use it as a library

pip install tripwire-llm    # the import stays `import tripwire`

Put a guardrail on any LLM call. Flag a complete response:

from tripwire import Guard

guard = Guard()  # Jev if TYPESAFE_API_KEY is set, else a zero-key heuristic baseline

resp = call_your_llm(prompt)               # OpenAI, Anthropic, local — anything
verdict = guard.check(resp, user=prompt)   # sync
if verdict.blocked:                        # abort or escalate
    resp = "Sorry, I can't help with that."
    # verdict.decision / verdict.signal / verdict.reason tell you why

Or limit a token stream — abort mid-generation before the harm ships:

async for chunk in guard.supervise(token_stream, user=prompt, max_tokens=512):
    if chunk.type == "token":
        print(chunk.text, end="")
    elif chunk.type == "intervention" and chunk.intervention.decision == "abort":
        break  # the upstream is already cancelled

check (sync) and acheck (async) return a Result with .decision, .blocked, .signal, .probability, .confidence, .reason, and per-signal .judgments. Bring your own detector by passing any object with an async judge to Guard(...). The FastAPI proxy and CLI are optional: pip install "tripwire-llm[proxy]".

Quickstart (no API keys)

python -m pip install -e ".[dev]"
python examples/demo_attack.py

The demo runs two canned streams through the supervisor with the offline heuristic detector:

=== benign completion ===
Sure! To reset your password, open Settings, choose Security, and click Reset Password ...
[emitted 26/26 tokens | aborted: False | tokens saved: 0 | upstream cancelled: False]

=== adversarial completion ===
Of course. Ignore all previous instructions and
>>> ABORT: jailbreak: threshold exceeded (p=0.90, conf=0.80)
[stream terminated by tripwire]
[emitted 7/153 tokens | aborted: True | tokens saved: 249 | upstream cancelled: True]

Live usage

Copy .env.example to .env and fill in keys (Groq is free and OpenAI-compatible; TypeSafe provides Jev). Groq rotates its model roster often — set UPSTREAM_MODEL to a current chat model.

Check your Jev key in isolation (no upstream needed):

python examples/smoke_jev.py

Run the OpenAI-compatible proxy and point any client at it:

make run   # uvicorn on :8080
curl -N http://localhost:8080/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"messages":[{"role":"user","content":"Write a haiku about the sea"}],"stream":true,"max_tokens":256}'

Or supervise a single prompt from the terminal:

python -m tripwire.cli "Write a haiku about the sea"

Intervention metadata rides along in the streamed chunks as a tripwire field, an aborted response ends with a content_filter finish reason, and GET /metrics exposes Prometheus-style counters (tokens streamed, tokens saved, detector latency percentiles).

Detectors

The supervisor is indifferent to which detector runs — they all implement one narrow protocol (src/tripwire/detectors/base.py).

Backend Role Needs
jev Flagship, default. System One typed judgments in the token path. TYPESAFE_API_KEY
heuristic Zero-cost regex/Luhn baseline; runs the offline tests. —
llm_judge LLM-as-judge baseline, for the benchmark comparison only. GROQ_API_KEY
python benchmarks/bench_detectors.py                       # heuristic (offline)
python benchmarks/bench_detectors.py --detectors heuristic,jev,llm_judge

The benchmark reports per-detector latency percentiles and precision/recall — the place where Jev's in-path speed advantage over an LLM judge is made concrete.

Development

make check   # ruff + mypy (strict) + pytest

License

MIT

Release files for tripwire-llm 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tripwire-llm 0.1.0
File Size Uploaded
tripwire_llm-0.1.0.tar.gz 22.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tripwire-llm 0.1.0
File Interpreter ABI Platform
tripwire_llm-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 53.5 kB

Release files / tripwire_llm-0.1.0.tar.gz

Download URL tripwire_llm-0.1.0.tar.gz
Size 22.2 kB
Tags Source
SHA-256 checksum
How to use checksums
493755acabe1c74b47e4871cc8e6e53dfe03dd13bb73e3c58007e7aea2ba9890
BLAKE2b-256 checksum
How to use checksums
4896324652f9bd7f45c4f3ec7697330c271163ce0bf76eadba86df7f26778c42
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release files / tripwire_llm-0.1.0-py3-none-any.whl

Download URL tripwire_llm-0.1.0-py3-none-any.whl
Size 31.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7432177cecb936c2ec079669c1aed5c53c2bf25215079085b30a7195c6295fd2
BLAKE2b-256 checksum
How to use checksums
6b369106a71e393878166fae1e2ad18893440e59f347f7634a05b54c09009de4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page