🌊 LLM Stream Processor
A callback-driven, prefix-safe, lazy LLM stream sanitization library.
Real-time filtering, redaction, and control of streaming LLM outputs with sub-microsecond overhead.
✨ Features
- 🔒 Prefix-Safe Pattern Matching — Uses Aho-Corasick automaton to ensure no partial sensitive content leaks before full match confirmation
- ⚡ Ultra-Low Latency — Target <5μs per-token overhead, designed for real-time streaming
- 🔄 Sync & Async Support — Works seamlessly with both synchronous and asynchronous LLM SDKs
- 🎯 Flexible Actions — PASS, DROP, REPLACE, HALT, or CONTINUE_DROP/PASS based on pattern matches
- 📊 History Tracking — Optional input/output/action history for debugging and analytics
- 🔌 Runtime Updates — Dynamically register/deregister patterns without restarting streams
📦 Installation
pip install llm-stream-processor
For development:
git clone https://github.com/DevOpRohan/llm_stream_processor.git
cd llm_stream_processor
pip install -e .
🚀 Quickstart
from stream_processor import KeywordRegistry, llm_stream_processor, replace, halt
# Create a registry and register pattern callbacks
reg = KeywordRegistry()
reg.register("secret", lambda ctx: replace("[REDACTED]"))
reg.register("STOP", halt) # Halt stream on this keyword
@llm_stream_processor(reg, yield_mode="token")
def generate_response():
yield "The secret password is hidden. "
yield "Do not STOP here."
yield "This won't be seen."
# Consume the filtered stream
for token in generate_response():
print(token, end="", flush=True)
# Output: The [REDACTED] password is hidden. Do not
📖 API Reference
Core Classes
| Class | Description |
|---|---|
KeywordRegistry |
Register/deregister keywords and their callbacks, compiles to Aho-Corasick automaton |
StreamProcessor |
Low-level processor for character-by-character filtering |
ActionContext |
Context passed to callbacks with keyword, buffer, position, and history |
StreamHistory |
Tracks input/output/actions for debugging |
Decorator
@llm_stream_processor(registry, yield_mode='token', record_history=True)
| Parameter | Options | Description |
|---|---|---|
registry |
KeywordRegistry | Registry with registered patterns |
yield_mode |
'char', 'token', 'chunk:N' |
Output mode: per-character, per-token, or N-char chunks |
record_history |
True/False |
Enable/disable history tracking |
Action Helpers
| Function | Description |
|---|---|
drop() |
Remove the matched keyword from output |
replace(text) |
Replace matched keyword with custom text |
halt() |
Immediately abort the stream |
passthrough() |
Leave matched keyword unchanged (no-op) |
continuous_drop() |
Start dropping all content until continuous_pass |
continuous_pass() |
Resume normal output after continuous_drop |
🎯 Use Cases
PII Redaction
import re
from stream_processor import KeywordRegistry, llm_stream_processor, replace
reg = KeywordRegistry()
# Redact email-like patterns (register common domains)
for domain in ["@gmail.com", "@yahoo.com", "@outlook.com"]:
reg.register(domain, lambda ctx: replace("@[REDACTED]"))
# Redact SSN patterns
reg.register("SSN:", lambda ctx: replace("SSN: [REDACTED]"))
Content Moderation (Drop Segments)
from stream_processor import KeywordRegistry, llm_stream_processor, continuous_drop, continuous_pass
reg = KeywordRegistry()
# Drop internal "thinking" blocks
reg.register("<think>", continuous_drop)
reg.register("</think>", continuous_pass)
@llm_stream_processor(reg, yield_mode="token")
def llm_stream():
yield "Hello! <think>internal reasoning here</think>Here's my response."
print("".join(llm_stream()))
# Output: Hello! </think>Here's my response.
Async Streaming (OpenAI Pattern)
import asyncio
from stream_processor import KeywordRegistry, llm_stream_processor, replace
reg = KeywordRegistry()
reg.register("API_KEY", lambda ctx: replace("[HIDDEN]"))
@llm_stream_processor(reg, yield_mode="token")
async def stream_chat():
# Simulating async LLM response chunks
chunks = ["Your ", "API_KEY ", "is safe."]
for chunk in chunks:
yield chunk
await asyncio.sleep(0.1)
async def main():
async for token in stream_chat():
print(token, end="", flush=True)
asyncio.run(main())
🏗️ Architecture
Token Generator (sync/async)
│
▼
@llm_stream_processor ← Decorator API
│
▼
StreamProcessor ← Character-level engine
┌──────────────┐
│ Aho-Corasick │
│ Automaton │
└──────────────┘
│
▼
Lazy Buffer + Callbacks
│
▼
Re-packer (char/token/chunk)
│
▼
Consumer
For detailed architecture, see docs/ARCHITECTURE.md.
🧪 Development
# Install in editable mode
pip install -e .
# Run tests
python -m pytest tests/ -v
# Run the example
python -m examples.example
📚 Documentation
- Problem Statement: docs/PROBLEM.md
- Architecture & Design: docs/ARCHITECTURE.md
- Contributing Guide: CONTRIBUTING.md
🤝 Contributing
Contributions are welcome! Please read our Contributing Guide for details on:
- Code of Conduct
- Development setup
- Submitting pull requests
📄 License
This project is licensed under the MIT License - see the LICENSE file for details.
🙏 Acknowledgments
- Aho-Corasick algorithm for efficient multi-pattern matching
- Inspired by the need for real-time LLM output sanitization in production systems
Metadata
Release files for llm-stream-processor 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_stream_processor-0.1.0.tar.gz | 13.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_stream_processor-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 28.6 kB
Release files / llm_stream_processor-0.1.0.tar.gz
| Download URL | llm_stream_processor-0.1.0.tar.gz |
|---|---|
| Size | 13.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d59aa04a111a1d422d3b527c8f11bee8f99b0020bf93b2c8ff20ea75ae140e3b
|
|
BLAKE2b-256 checksum How to use checksums |
2361a0a381dd253d5dff2e8a7fe3ea6875bfd8f4e2af2c9b90ec6e374a36ac1c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Dec 20, 2025.
Transparency logRelease files / llm_stream_processor-0.1.0-py3-none-any.whl
| Download URL | llm_stream_processor-0.1.0-py3-none-any.whl |
|---|---|
| Size | 15.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
181f449f6074e1acc5d00d9a5214830fbfc1fd2007c425e57cc8912433147e96
|
|
BLAKE2b-256 checksum How to use checksums |
0cd8ab9f9b48918f16f8e7cd824de0844a4ba89f429442dc16c27476fbbaa705
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Dec 20, 2025.
Transparency log