Openlayer Guardrails
Open source guardrail implementations that work with Openlayer tracing.
Installation
pip install openlayer-guardrails
Usage
Standalone Usage
from openlayer_guardrails import PIIGuardrail
# Create guardrail
pii_guard = PIIGuardrail(
block_entities={"CREDIT_CARD", "US_SSN"},
redact_entities={"EMAIL_ADDRESS", "PHONE_NUMBER"}
)
# Check inputs manually
data = {"message": "My email is john@example.com and SSN is 123-45-6789"}
result = pii_guard.check_input(data)
if result.action.value == "block":
print(f"Blocked: {result.reason}")
elif result.action.value == "modify":
print(f"Modified data: {result.modified_data}")
With Openlayer Tracing
from openlayer_guardrails import PIIGuardrail
from openlayer.lib.tracing import trace
# Create guardrail
pii_guard = PIIGuardrail()
# Apply to traced functions
@trace(guardrails=[pii_guard])
def process_user_data(user_input: str):
return f"Processed: {user_input}"
# PII is automatically handled
result = process_user_data("My email is john@example.com")
# Output: "Processed: My email is [EMAIL-REDACTED]"
Toxicity Guardrail (Brazilian Portuguese)
Detects toxic content in Brazilian Portuguese using the ToxiGuardrailPT model.
pip install 'openlayer-guardrails[toxicity]'
from openlayer_guardrails import ToxicityPTGuardrail
# Create guardrail (default threshold=0.0; positive scores = safe, negative = toxic)
toxicity_guard = ToxicityPTGuardrail()
# Check inputs
result = toxicity_guard.check_input({"message": "Você é um idiota!"})
print(result.action) # GuardrailAction.BLOCK
# Check outputs with contextual scoring (sentence-pair encoding)
result = toxicity_guard.check_output(
output="Claro, aqui está a informação solicitada.",
inputs={"prompt": "Me ajude com meu trabalho."},
)
print(result.action) # GuardrailAction.ALLOW
Toxicity Guardrail (English)
Detects toxic content in English across six categories using unitary/toxic-bert.
pip install 'openlayer-guardrails[toxicity]'
from openlayer_guardrails import ToxicityENGuardrail
# Create guardrail (default threshold=0.5)
toxicity_guard = ToxicityENGuardrail()
# Check inputs
result = toxicity_guard.check_input({"message": "You are terrible and should die"})
print(result.action) # GuardrailAction.BLOCK
print(result.metadata["triggered_categories"])
# e.g. {'toxic': 0.98, 'severe_toxic': 0.72, 'insult': 0.89, 'threat': 0.81}
# Monitor only specific categories
guard = ToxicityENGuardrail(categories={"threat", "severe_toxic"})
Handling long texts
By default, all guardrails truncate inputs to 512 tokens for fast inference.
To evaluate the full text, enable chunking mode by setting max_length=None:
guard = ToxicityPTGuardrail(max_length=None) # or ToxicityENGuardrail(max_length=None)
In chunking mode, long texts are split into overlapping 512-token windows and each window is scored independently. The most toxic score across all windows is used. Latency scales linearly with the number of chunks.
Model Limitations
Prompt Injection Guardrail
| Property | Value |
|---|---|
| Model | meta-llama/Llama-Prompt-Guard-2-86M |
| Max tokens | 512 |
| Language | English |
| Parameters | 86M |
| Scope | Input-only (outputs are not checked) |
Texts longer than 512 tokens are truncated. Only the first 512 tokens are evaluated.
Toxicity Guardrail (PT-BR)
| Property | Value |
|---|---|
| Model | nicholasKluge/ToxiGuardrailPT |
| Max tokens | 512 |
| Language | Brazilian Portuguese |
| Parameters | 109M |
| Architecture | BERTimbau (bert-base-portuguese-cased) |
| Output type | Single scalar reward score (positive = safe, negative = toxic) |
| Reported accuracy | 70.36% (hatecheck-portuguese), 74.04% (told-br) |
| Scope | Input and output (output uses sentence-pair encoding for context) |
By default, texts longer than 512 tokens are truncated. Set max_length=None to enable chunking for full-text coverage. The model was trained on Brazilian Portuguese data and may not generalize well to European Portuguese or other languages.
Toxicity Guardrail (EN)
| Property | Value |
|---|---|
| Model | unitary/toxic-bert |
| Max tokens | 512 (chunking available via max_length=None) |
| Language | English |
| Parameters | 110M |
| Architecture | BERT (bert-base-uncased) |
| Output type | Multi-label probabilities across 6 categories |
| Categories | toxic, severe_toxic, obscene, threat, insult, identity_hate |
| Reported AUC | 0.98636 (Jigsaw Toxic Comment Challenge) |
| Scope | Input and output |
By default, texts longer than 512 tokens are truncated. Set max_length=None to enable chunking for full-text coverage. The model was trained on English data.
Metadata
Release files for openlayer-guardrails 0.7.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| openlayer_guardrails-0.7.1.tar.gz | 226.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| openlayer_guardrails-0.7.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 275.5 kB
Release files / openlayer_guardrails-0.7.1.tar.gz
| Download URL | openlayer_guardrails-0.7.1.tar.gz |
|---|---|
| Size | 226.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
83843dedd02dd1f319d6af94a80db643a80950e4f5b3b22c85f9f8c16475141b
|
|
BLAKE2b-256 checksum How to use checksums |
bd30c2dd95bd250f17cdeeb462a3912c9f1d0d90f40e8af00bd68bc072ffcc16
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / openlayer_guardrails-0.7.1-py3-none-any.whl
| Download URL | openlayer_guardrails-0.7.1-py3-none-any.whl |
|---|---|
| Size | 49.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c9bf02cc96ed593103deba436e06c6d3ecaa44670babf7372d71a7f83ca0b353
|
|
BLAKE2b-256 checksum How to use checksums |
59d5550ee041b24dee06dbf3e1680216b3d8abfbbfb9e65eb1a870c98408e03f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|