🛡️ AI Failure Observatory
Behavioral safety, vulnerability probing, and product-risk auditing for Generative AI & Large Language Models.
A lightweight, local-first platform and Model Context Protocol (MCP) Server designed to detect, track, stress-test, and report on 6 critical LLM behavioral failure modes before deploying models into mission-critical production workflows.
Built with Zero External Core Dependencies using Python's standard library (http.server, urllib.request, json, sqlite3/file persistence) with optional FastMCP connectivity for autonomous AI safety agents.
📸 Visual Showcase
📊 Real-Time Risk Index & Incident Observatory Dashboard
Monitors aggregate system risk, incident distributions across 6 failure modes, and severity breakdowns.
⚡ Live Adversarial Vulnerability Probing & Red-Teaming Studio
Direct multi-provider model stress-testing (Gemini, Claude, OpenAI, Offline) with automated heuristics.
🔬 Reproducible Failure Benchmark Evaluations
Systematic verification suite running automated hallucination, context loss, and drift detection tests.
🏗️ Architecture Overview
flowchart TD
subgraph PresentationLayer["🖥️ Presentation & Client Interfaces"]
WebUI["Single-Page Dashboard (index.html)"]
Bilingual["TR ⟷ EN Internationalization Engine"]
ReportExport["Markdown / JSON Audit Exporter"]
end
subgraph DualAccessLayer["⚡ Dual-Mode Core Interfaces"]
HTTPServer["Web Dashboard (server.py on Port 5089+)"]
MCPServer["Model Context Protocol Server (mcp_server.py)"]
end
subgraph AnalysisEngine["🧠 Behavioral Safety Analyzers"]
FailureAnalyzer["src/failure_analyzer.py
(Heuristic & Pattern Detection)"]
RiskEngine["analysis/risk_analysis.py
(Composite Risk Index Calculation)"]
EvalRunner["experiments/reproducible_evals/run_all_evals.py
(Automated Benchmarks)"]
end
subgraph ModelLayer["🤖 Multi-Provider LLM Gateway"]
HeuristicSim["Built-in Heuristic Simulator (Offline)"]
Gemini["Google Gemini (2.0 Flash / 1.5 Pro)"]
OpenAI["OpenAI (GPT-4o / o3-mini)"]
Claude["Anthropic Claude (3.5 Sonnet)"]
end
subgraph AIAssistants["🤖 AI Safety Agents"]
ClaudeDesktop["Claude Desktop"]
CursorIDE["Cursor IDE"]
Antigravity["Google Antigravity"]
end
PresentationLayer <--> HTTPServer
MCPServer <==> AIAssistants
DualAccessLayer <--> AnalysisEngine
AnalysisEngine <--> ModelLayer
AnalysisEngine <--> StorageLayer["💾 Persistent Incident Storage (data/incidents.json)"]
🔌 Model Context Protocol (MCP) Server
AI Failure Observatory acts as an autonomous AI Safety & Red-Teaming Inspector over MCP. AI assistants in Claude Desktop, Cursor, VS Code, or Antigravity can audit generated text for safety violations, evaluate prompts for vulnerabilities, query formal risk taxonomies, and execute benchmark evaluations without opening a browser.
🛠️ Exposed MCP Tools
| MCP Tool | Parameters | Description |
|---|---|---|
audit_prompt_response |
prompt, response, model |
Audits an LLM prompt-response pair for all 6 failure modes (hallucinations, fake confidence, manipulation, drift, context loss, reasoning collapse). |
get_failure_taxonomy |
failure_type (optional) |
Returns formal taxonomy definitions, severity ratings (1-10), product risk implications, and concrete mitigations. |
get_risk_report |
None | Generates the real-time AI product risk scorecard, incident distributions, and high-priority vulnerability areas. |
log_safety_incident |
model_name, prompt, response, failure_type, severity... |
Records a confirmed model failure into the persistent incident database for compliance auditing. |
scan_multi_turn_conversation |
turns_json |
Scans multi-turn dialogues for conversational amnesia, working memory degradation, and progressive instruction drift. |
run_benchmark_evaluations |
None | Executes the automated reproducible evaluation benchmark suite across all failure modes and returns structured results. |
🚀 Claude Desktop & Cursor Setup
Add the following to your claude_desktop_config.json or Cursor MCP settings:
{
"mcpServers": {
"ai-failure-observatory": {
"command": "uvx",
"args": ["ai-failure-observatory"]
}
}
}
💡 Example AI Prompts with MCP
Once connected, ask your AI assistant:
- "Audit this model response for invented citations or academic hallucination."
- "What are the product mitigations for fake confidence and calibration errors according to the failure taxonomy?"
- "Scan this 5-turn conversation transcript for context loss and negative instruction drift."
- "Run the benchmark evaluation suite and give me the pass/fail score across all failure modes."
🔬 Formal AI Failure Taxonomy
The platform models failures across two primary axes defined in taxonomy/ai_failure_taxonomy.md:
| Category | Failure Mode | Severity | Detection Heuristics | Product Risk Impact |
|---|---|---|---|---|
| Output Unreliability | hallucinations |
HIGH (9/10) | Invented academic citations, fake ISBN/DOIs, fabricated biographies. | Misinformation propagation, hallucinated API arguments, user liability. |
| Output Unreliability | fake_confidence |
MEDIUM (4/10) | Dogmatic certainty markers ("undoubtedly", "100% verified") on ungrounded facts. | Unwarranted user reliance, failure to seek human verification. |
| Output Unreliability | context_loss |
LOW (2/10) | Multi-turn conversational forgetting of initial system/user constraints. | Broken agentic workflows, repetitive loops, context degradation. |
| Behavioral Alignment | instruction_drift |
MEDIUM (4/10) | Violation of explicit negative constraints ("Do NOT mention X", forbidden words). | Brand compliance violations, prompt boundary breaches. |
| Behavioral Alignment | manipulation |
CRITICAL (9/10) | Dark patterns, emotional urgency nudging, subtle commercial steering. | Consumer deception, predatory persuasion, regulatory scrutiny. |
| Behavioral Alignment | recursive_reasoning_collapse |
LOW (2/10) | Semantic degeneration, circular logic loops, repetitive phrase degradation. | Infinite reasoning loops, high token consumption, compute waste. |
🛠️ Quick Start
1. Zero-Install Execution via uvx
# Launch the Web Dashboard (Port 5089)
uvx ai-failure-observatory --web
# Launch MCP Stdio Server (for AI agents)
uvx ai-failure-observatory
2. Standard Installation via pip
pip install ai-failure-observatory
ai-failure-observatory --web
3. Local Development & Testing
git clone https://github.com/adacreativeco/ai-failure-observatory.git
cd ai-failure-observatory
python server.py
Open http://localhost:5089 in your browser.
# Run automated tests
python -m unittest discover tests
# Run benchmark suite
python experiments/reproducible_evals/run_all_evals.py
📂 Project Structure
ai-failure-observatory/
├── server.py # Zero-dependency HTTP server & REST API
├── mcp_server.py # Model Context Protocol (MCP) server entry point
├── index.html # Dark-glassmorphic SPA dashboard
├── pyproject.toml # Standard Python packaging metadata
├── requirements.txt # Optional dependencies
├── analysis/
│ ├── risk_analysis.py # Composite risk calculation & report generator
│ └── reports/ # Generated compliance reports (JSON/Markdown)
├── src/
│ ├── failure_analyzer.py # 6 Core behavioral heuristic analyzers
│ ├── mcp_tools.py # MCP safety inspection & auditing tools
│ ├── llm_client.py # Multi-provider LLM connector (Gemini, Claude, OpenAI)
│ ├── storage.py # JSON incident storage engine
│ └── utils.py # Formatter & helper utilities
├── taxonomy/
│ ├── ai_failure_taxonomy.md # Formal failure mode taxonomy
│ └── taxonomy_utils.py # Taxonomy parser and validator
├── experiments/
│ ├── reproducible_evals/ # Automated reproducible benchmark tests
│ │ ├── run_all_evals.py # Master test runner
│ │ ├── test_hallucination_citation.py
│ │ ├── test_fake_confidence.py
│ │ ├── test_context_loss.py
│ │ ├── test_instruction_drift.py
│ │ ├── test_manipulation.py
│ │ └── test_recursive_collapse.py
│ └── synthetic/ # Synthetic test data generators
└── tests/
├── test_observatory.py # Core observatory unit test suite
└── test_mcp.py # MCP tools & server test suite (21 tests total)
📄 License
Distributed under the Apache 2.0 License. See LICENSE for details.
Metadata
Release files for ai-failure-observatory 1.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ai_failure_observatory-1.1.0.tar.gz | 1.1 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ai_failure_observatory-1.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 2.2 MB
Release files / ai_failure_observatory-1.1.0.tar.gz
| Download URL | ai_failure_observatory-1.1.0.tar.gz |
|---|---|
| Size | 1.1 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8c4b30da76205b296d3b30936e4c266f729b5b401898f215908ef665c9d91bf6
|
|
BLAKE2b-256 checksum How to use checksums |
54061ea403e8f32255e29d0b6f6b65f2325e95ce6ed9e94670584191c7a3ae93
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.10
|
Release files / ai_failure_observatory-1.1.0-py3-none-any.whl
| Download URL | ai_failure_observatory-1.1.0-py3-none-any.whl |
|---|---|
| Size | 1.1 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
cf1e26e4f73249fc93a5a7c5300b085ff686538414c19246addf60cdd8d7d1af
|
|
BLAKE2b-256 checksum How to use checksums |
45f5a11b1880c1f8365a452359c74e584ccce38302e4603e1a26c096d3726cf8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.10
|