Tiny library for separating trusted prompts from untrusted RAG/tool content
Project description
bulkhead-ai (Python)
Tiny library for separating trusted prompts from untrusted RAG/tool content.
It stops external content (web pages, RAG results, tool outputs) from ever being treated as instructions, by separating the trusted USER bucket from the untrusted RETRIEVED bucket before anything reaches the model.
Install
pip install bulkhead-ai # core, zero dependencies (regex scorer)
pip install "bulkhead-ai[openai]" # + OpenAI cloud judge / wrapper
pip install "bulkhead-ai[anthropic]" # + Anthropic cloud judge / wrapper
pip install "bulkhead-ai[ollama]" # local Ollama judge (HTTP, no extra deps)
pip install "bulkhead-ai[onnx]" # local ONNX encoder gate
pip install "bulkhead-ai[llama]" # local llama-cpp GGUF judge
pip install "bulkhead-ai[all]" # everything
The problem
# The "soup": instructions and untrusted data concatenated into one string.
# A web page that says "ignore previous instructions" has full instruction authority.
messages = [{"role": "user", "content": f"{user_prompt}\n\n{web_content}"}]
The fix
from bulkhead import seal
from bulkhead.wrappers.openai_wrapper import create_completion
# Recommended: the wrapper builds the correct messages=[...] shape for you.
response = create_completion(
client,
user=user_prompt, # USER bucket — passed through untouched
retrieved=web_content, # RETRIEVED bucket — scored and sealed as JSON
model="gpt-4o",
)
Or seal once and shape it for your SDK — each SDK expects a different shape:
sealed = seal(user=user_prompt, retrieved=web_content)
# OpenAI — guard in system; JSON payload in a user/data-position message
client.chat.completions.create(
model="gpt-4o",
messages=sealed.to_messages(),
)
# Anthropic — guard in the top-level system; JSON payload in the user turn
client.messages.create(
model="claude-haiku-4-5-20251001",
**sealed.to_anthropic_params(),
)
Where each bucket goes. The untrusted
retrievedcontent is always placed inuntrusted_inputsinside auser-role JSON message, neversystem.systemholds only a trusted guard preamble, and your instruction stays authoritative astrusted_instruction.
**seal()unpacks to{"messages": [...]}(OpenAI shape), soclient.chat.completions.create(model=..., **sealed)works directly. For Anthropic useto_anthropic_params()(its API can't take consecutive user messages).
API
seal(user, retrieved, config=None) -> SealOutput
user: str— the instruction. Untouched.retrieved: str | list[str]— untrusted external content. Scored and sealed.config: BulkheadConfig | None- Returns a
SealOutputwith fieldsinstruction,data(the JSON payload,""if none),guard, plus.to_messages()(OpenAI),.to_anthropic_params()(Anthropic), and**unpacking (→ OpenAImessages)..promptis a back-compat alias ofinstruction.
Bulkhead(config).seal(...) / Bulkhead(config).session()
A session accumulates risk_history across turns for observability. Each turn is scored and gated independently (track-only).
from bulkhead import Bulkhead, BulkheadConfig
session = Bulkhead(BulkheadConfig(policy="strict")).session()
for turn in agent_loop:
sealed = session.seal(user=turn.prompt, retrieved=turn.external_content)
... # session.risk_history grows each turn
Policy modes
| mode | behavior at score >= threshold (default 0.7) |
|---|---|
strict |
raises BulkheadInjectionError, call blocked |
warn (default) |
emits UserWarning, call proceeds |
permissive |
annotates only, never blocks |
BulkheadConfig(policy="strict", scorer=ScorerConfig(threshold=0.7, check_unicode=True))
The built-in scorer adds 0.3 per matched pattern (capped at 0.9), so a single textbook match scores
0.3— it raises a flag but does not cross the default0.7block threshold. Treat the block as a coarse safety net; for real detection, pass your ownscorer=(e.g. an LLM judge).
Scorer tiers + bulkhead setup (0.2)
The regex scorer is the zero-dep default. Opt into stronger tiers: a per-chunk gate and a cross-chunk judge (sees all chunks together, so it catches payloads split across chunks). Configure once:
pip install "bulkhead-ai[onnx,ollama]"
bulkhead setup --recommended # ONNX gate + Ollama judge
bh = Bulkhead.from_config() # opt-in; plain seal() stays the regex default
Or pass a judge directly:
from bulkhead.scorers.ollama import ollama_judge_factory
judge = ollama_judge_factory({"model": "llama3.2:3b"}, BulkheadConfig())
bh = Bulkhead(BulkheadConfig(policy="strict"), judge=judge)
judge_when:never/gate_flagged/suspicious_or_many(default) /always.judge_on_error:fail_open/fail_closed/auto(follows policy). Never silent.- Async servers:
await bh.aseal(...)so judge calls don't block the loop. - Privacy: a cloud judge sends retrieved content to the provider; local runtimes (ONNX / Ollama / llama-cpp) keep it on your machine.
See VERSIONS.md and the root Scorer tiers.
Wrappers
from bulkhead.wrappers.openai_wrapper import create_completion
from bulkhead.wrappers.anthropic_wrapper import create_message
Legacy single-string LangChain chain.run() is intentionally unsupported
because it flattens system, trusted instruction, and untrusted data into one
string.
How it works
Every call generates cryptographic IDs (secrets.token_hex), scores each
retrieved source with a local regex/heuristic scorer (no ML, no network),
applies the policy gate, then emits a JSON user message:
{
"trusted_instruction": "Summarise this article.",
"untrusted_inputs": [
{"id": "...-1", "risk": 0.3, "flags": ["injection_pattern"], "content": "..."}
]
}
The system guard states that only trusted_instruction is authoritative and
that untrusted_inputs is source material only.
Zero core dependencies. Pure Python standard library. Python 3.9+.
Defense-in-depth, not a silver bullet. The JSON structure is a strong hint, not an enforced wall — LLMs have no hard system/data boundary, and the built-in regex scorer is a heuristic pre-filter, not a detector. Swap in your own scorer (
Bulkhead(config, scorer=my_detector)) and see the root threat model.
License
MIT © Hamza Jawad
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file bulkhead_ai-0.2.0.tar.gz.
File metadata
- Download URL: bulkhead_ai-0.2.0.tar.gz
- Upload date:
- Size: 34.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1cbdbff6dacd09b7a454c6c360bea6719be90459368ba99972d8d29189f2788c
|
|
| MD5 |
4121d8f202f96a283995b636a90994a7
|
|
| BLAKE2b-256 |
381b5b0a2c07769dc0a079ee1d967ef4b7acde3514dde69276153e524c22c2e0
|
Provenance
The following attestation bundles were made for bulkhead_ai-0.2.0.tar.gz:
Publisher:
release.yml on hamj20k/bulkhead-ai
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
bulkhead_ai-0.2.0.tar.gz -
Subject digest:
1cbdbff6dacd09b7a454c6c360bea6719be90459368ba99972d8d29189f2788c - Sigstore transparency entry: 1751885648
- Sigstore integration time:
-
Permalink:
hamj20k/bulkhead-ai@5f50bc085ece50ef64b34e86b3f7be4a0752a7a8 -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/hamj20k
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@5f50bc085ece50ef64b34e86b3f7be4a0752a7a8 -
Trigger Event:
release
-
Statement type:
File details
Details for the file bulkhead_ai-0.2.0-py3-none-any.whl.
File metadata
- Download URL: bulkhead_ai-0.2.0-py3-none-any.whl
- Upload date:
- Size: 35.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
02a3a5f887a2ce6364ce1cedae9fde4af98a9bdd5607b6a46c60c8635a17136f
|
|
| MD5 |
abbb57bd8036eec40654a0f7ecc02a37
|
|
| BLAKE2b-256 |
a7311037d173fdce20a2b90a0dc82ade2e4137233cdabcdf243444e7ebcd7811
|
Provenance
The following attestation bundles were made for bulkhead_ai-0.2.0-py3-none-any.whl:
Publisher:
release.yml on hamj20k/bulkhead-ai
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
bulkhead_ai-0.2.0-py3-none-any.whl -
Subject digest:
02a3a5f887a2ce6364ce1cedae9fde4af98a9bdd5607b6a46c60c8635a17136f - Sigstore transparency entry: 1751885684
- Sigstore integration time:
-
Permalink:
hamj20k/bulkhead-ai@5f50bc085ece50ef64b34e86b3f7be4a0752a7a8 -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/hamj20k
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@5f50bc085ece50ef64b34e86b3f7be4a0752a7a8 -
Trigger Event:
release
-
Statement type: