LLM-Rivotril
A lightweight Python framework to reduce LLM hallucinations, enforce guardrails, manage stateful memory, and monitor performance through a local web dashboard.
Features
- Guardrails — Block disallowed keywords, enforce allowed topics, limit output size, and validate JSON schemas on outputs.
- Semantic Guardrails (optional) — Match prompts against allowed topics using dense embeddings instead of exact keywords.
- Moderation Guardrail — Flag adversarial/unsafe content via OpenAI's moderation endpoint.
- Stateful Memory — Sliding-window conversation store to prevent context drift, with optional automatic disk persistence.
- Anti-Hallucination Verifier — Pluggable grounding checks, including keyword overlap, citation markers, and embedding-based faithfulness.
- Resilience — Built-in rate limiting, retry with backoff, and circuit breaker for LLM calls.
- Token-Saving Controls (optional) — Semantic (similarity-based) response cache, token-budget-aware memory trimming, and LLM-summarized history compaction.
- Telemetry & Dashboard — Built-in FastAPI dashboard with live request logs, token usage, latency, and success rate; optional per-token auth.
- Benchmark / Red-Team Evaluator — Labeled suite to measure guardrail and verifier accuracy without API costs.
- CLI — Launch the dashboard with a single command.
- Type-Safe Responses — Optional Pydantic response models via
instructor, including streamed structured output. - Async API —
run_async()for non-blocking execution. - Multi-Provider — OpenAI-compatible servers (Ollama, vLLM, ...), plus native Anthropic, Cohere, Gemini, Azure OpenAI, and AWS Bedrock adapters.
- Vector-Store Retrievers (optional) — pgvector, Qdrant, Weaviate, and Pinecone adapters for RAG beyond in-memory scale.
- Framework Integrations (optional) — Drop-in adapters for CrewAI, AG2/AutoGen, LangChain/LangGraph, and Google ADK, so guardrails/PII/RAG/observability apply inside those frameworks too.
Installation
pip install llmrivotril
For semantic (embedding-based) guardrails and verifiers:
pip install llmrivotril[semantic]
For local development:
git clone https://github.com/ailake-io/llmrivotril.git
cd llmrivotril
pip install -e ".[dev,semantic]"
This installs test, lint, type-check, and packaging tools (pytest, ruff, mypy, build, twine) plus every optional runtime dependency (all providers, RAG loaders, vector stores, Redis, framework integrations).
Quick Start
import os
from llmrivotril import RivotrilAgent, Guardrail
from llmrivotril.verifier import KeywordOverlapVerifier
agent = RivotrilAgent(
model="gpt-4o-mini",
api_key=os.getenv("OPENAI_API_KEY"),
guardrails=[
Guardrail(
name="safe-content",
allowed_topics=["AI safety", "machine learning"],
disallowed_keywords=["password", "secret"],
max_tokens=500,
)
],
verifier=KeywordOverlapVerifier(threshold=0.1),
)
response = agent.run("Explain what a guardrail is in AI safety.")
print(response)
Documentation
- Providers — OpenAI-compatible servers, Anthropic, Cohere, Gemini, Azure OpenAI, AWS Bedrock.
- Guardrails & Safety — Semantic guardrails, moderation, PII redaction, token budget.
- RAG — Local pipeline, vector-store retrievers (pgvector/Qdrant/Weaviate/Pinecone), grounding verification.
- Framework Integrations — CrewAI, AG2/AutoGen, LangChain/LangGraph, Google ADK adapters, and multi-agent setup notes.
- Async, Streaming & Function Calling
- Reliability — Resilience, response caching (including semantic cache), memory token budget & summarization, schema-repair fallback.
- Observability — Metrics persistence, local dashboard, benchmark, cost tracking.
- Configuration — Config files, environment variables, plugins, project scaffolding.
- Releasing — Build validation, TestPyPI, and the PyPI release workflow.
Interactive Demo
Run a complete walkthrough with mock LLM responses (no API key, no cost):
python examples/demo.py --mock --dashboard
Then open http://127.0.0.1:8767 to watch the dashboard update live.
Comparison: With vs. Without llmrivotril
See the same scenarios side-by-side:
python examples/comparison.py --mock
The comparison highlights how plain LLM calls return harmful or hallucinated
answers that are only caught manually afterwards, while llmrivotril blocks
them at runtime and records structured telemetry.
Project Structure
llmrivotril/
├── src/llmrivotril/
│ ├── agent.py # RivotrilAgent orchestrator (sync + async)
│ ├── config.py # Environment-variable and file configuration loader
│ ├── guardrails.py # Input/output guardrails
│ ├── moderation.py # OpenAI moderation-endpoint guardrail
│ ├── memory.py # Conversation memory store
│ ├── verifier.py # Hallucination / grounding checks
│ ├── providers.py # OpenAI / Azure / Anthropic / Cohere / Gemini / Bedrock adapters
│ ├── rag/ # RAG loaders, chunkers, retrievers, vector stores, and pipeline
│ ├── integrations/ # CrewAI / AG2 / LangChain / Google ADK adapters
│ ├── resilience.py # Rate limiter, retry, and circuit breaker
│ ├── semantic.py # Optional embedding-based guardrails/verifiers
│ ├── metrics.py # Telemetry collector
│ ├── server.py # FastAPI dashboard
│ ├── cli.py # Click CLI
│ ├── exceptions.py # Custom exceptions
│ ├── templates/
│ │ └── dashboard.html # Dashboard UI template
│ └── static/
│ └── tailwind.min.js # Bundled Tailwind CSS for offline dashboard
├── tests/
├── examples/
└── docs/
Development
Run lint, type checks, and tests:
pip install -e ".[dev,semantic]"
ruff check src tests examples scripts
ruff format --check src tests examples scripts
mypy src
pytest -v
Slow integration tests (e.g. loading sentence-transformers models) are skipped by
default. Run them with:
pytest -v --run-slow
For a reproducible provider/integration environment, apply the versions validated locally before installing the desired extras:
python -m pip install -c .github/constraints-runtime.txt \
-e ".[integrations,providers,semantic,qdrant]"
CI/CD
The GitHub Actions workflow runs linting, type checking, tests, and package builds on Python 3.10–3.13.
License
MIT License — see LICENSE.
Metadata
Release files for llmrivotril 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llmrivotril-0.1.1.tar.gz | 278.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llmrivotril-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 491.6 kB
Release files / llmrivotril-0.1.1.tar.gz
| Download URL | llmrivotril-0.1.1.tar.gz |
|---|---|
| Size | 278.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6ee9db3ee8ccd90940c8f2e9f7c24da04a64201fbc8d26546cbdde83debb38fc
|
|
BLAKE2b-256 checksum How to use checksums |
d92f1f4eb7a7422ad9dcb2ac5752bdd4014d2c3a956acd80356ca750ede5b8c9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|
Release files / llmrivotril-0.1.1-py3-none-any.whl
| Download URL | llmrivotril-0.1.1-py3-none-any.whl |
|---|---|
| Size | 213.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6c24ded74e9a29c210a4b049544e05442f8971b2b53c402cf3b19b052a017b51
|
|
BLAKE2b-256 checksum How to use checksums |
416c7c1c2026be489b8713d59dc685da72307b0dc089238e96c6443522387f46
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|