ThreatVision AI: Technical Architecture, Mathematical Foundations, and System Reference Manual
Abstract
ThreatVision AI (threatvision-ai) is an open-source, modular computer vision framework designed for real-time threat detection, multi-object tracking, spatial analytics, and multi-channel incident response. Operating on streaming video input from live cameras, RTSP network streams, video files, or static images, the framework evaluates visual hazards and computes a calibrated threat score. ThreatVision AI is engineered as an operator assistance framework. Detections return explicit confidence bounds and threat score evaluations rather than absolute assertions, assisting human operators in physical security monitoring.
Table of Contents
- Introduction
- Computer Vision Background
- System Architecture
- Mathematical Foundations
- Installation Guide
- Quick Start Guide
- Package Structure
- Detection Modules Specification
- Threat Scoring Engine & Risk Matrix
- Web Dashboard & Command Center UI
- Cloud REST API & WebSocket Streaming Specification
- Python API Class Reference
- Configuration Management
- Command Line Interface (CLI) Reference
- Notification Channels
- Logging & Monitoring
- Performance Optimization & Hardware Acceleration
- Security, Data Privacy & Responsible AI
- Custom Plugin Development Guide
- Testing & Verification Strategy
- Production Deployment Guide
- Frequently Asked Questions (FAQ)
- Troubleshooting Matrix
- Version History & Changelog
- Future Development Roadmap
- Comprehensive Technical & Domain Glossary
1. Introduction
1.1 What is ThreatVision AI?
ThreatVision AI is an open-source Python framework for building camera-based physical security monitoring systems. It provides a modular pipeline—camera ingestion, object detection, multi-object tracking, behavior analysis, threat scoring, and human review—that developers assemble into monitoring applications ranging from a single-webcam prototype to a multi-site, multi-thousand-camera deployment.
Rather than shipping a single monolithic "detect everything" model, ThreatVision AI composes specialized, independently maintained detectors (person, vehicle, weapon-shape, fire/smoke, fall, crowd density, and others) behind a common interface, and combines their outputs through a transparent, configurable scoring engine.
1.2 Why Was It Created?
Camera-based security monitoring at scale faces a structural bottleneck: a human operator can attentively watch only a handful of video feeds at once, and attention degrades quickly during long shifts on low-signal footage. Existing options in this space historically fell into two categories:
- Closed, proprietary appliances — reliable but opaque, difficult to audit, and expensive to customize or extend.
- Research-grade computer vision code — flexible and transparent, but requiring significant engineering investment to turn into a deployable monitoring system with tracking, scoring, review workflows, and notifications.
ThreatVision AI was created to sit between these: a package that gives engineering teams inspectable, extensible building blocks while providing the operational scaffolding (dashboard, API, notifications, deployment tooling) needed to run a monitoring system in production.
1.3 Problems It Solves
| Problem | How ThreatVision AI Addresses It |
|---|---|
| Operators cannot watch every feed at all times | Automated detection surfaces candidate events for review rather than requiring continuous manual attention. |
| Raw model output is hard to act on | The threat scoring engine converts per-frame detections into zone- and time-aware incident scores. |
| Alert fatigue from high false-positive rates | Temporal persistence checks, zone rules, and confidence thresholds are combined and tunable per deployment. |
| Vendor lock-in and opaque decision logic | Fully open pipeline; every stage from detection to scoring is inspectable and replaceable. |
| Disconnected tooling | Single framework covering ingestion through notification, with a plugin system for custom components. |
| Difficult evaluation of detector quality | Built-in benchmarking, confusion-matrix, and PR-curve tooling. |
1.4 Goals & Principles
- Transparency — Every detection and score must be traceable to the model, frame, and rule that produced it.
- Human-in-the-loop by default — Surfacing information for human decision-making, not automating force or legal actions.
- Composability — Detectors, trackers, scorers, and notifiers are independently swappable.
- Operational Realism — Addresses the full lifecycle (deployment, logging, monitoring, performance tuning), not just model inference.
- Honesty about limitations — Documentation and defaults make failure modes and confidence limits visible rather than overstating reliability.
2. Computer Vision Background
2.1 AI, Machine Learning, and Deep Learning Hierarchy
Artificial Intelligence (AI)
└── Machine Learning (ML)
└── Deep Learning (DL)
├── Convolutional Neural Networks (CNNs for spatial features)
└── Vision Transformers (ViT for attention-based models)
Deep networks learn a hierarchy of representations: early layers respond to edges and textures, middle layers to parts (hand, blade shape), and later layers to whole objects. This hierarchy allows transfer learning across ThreatVision AI's different detectors.
2.2 Core Computer Vision Tasks
| Task | Core Question | ThreatVision AI Usage |
|---|---|---|
| Image Classification | "What is in this image?" | Scene-level checks (indoor/outdoor) |
| Object Detection | "What objects are here, and where?" | Bounding box localization for person, vehicle, weapon-shape, fire/smoke |
| Multi-Object Tracking | "Which object is the same across frames?" | Track persistence, trajectory analysis, loitering, and dwell time |
| Pose Estimation | "How is this person's body positioned?" | Keypoint estimation for fall and fight detection |
| Action Recognition | "What is happening over time?" | Temporal activity recognition (falling, fighting, motion energy) |
2.3 Object Detection Architectures
- Single-stage detectors (e.g., YOLO series) — Predict bounding boxes and class probabilities directly in a single forward pass. Faster, used for real-time streaming pipeline.
- Two-stage detectors (e.g., Faster R-CNN) — Propose candidate regions first, then classify each region. Higher accuracy, used for offline re-verification of borderline incidents.
3. System Architecture
3.1 High-Level Pipeline Architecture
Camera Sources (Webcam / RTSP / Video File / Cloud Stream)
│
▼
[Ingestion & Preprocessing Layer]
(Async Frame Decoding & Resizing)
│
▼
[Modular Detection Pipeline]
(Person, Weapon, Fire, Smoke, etc.)
│
▼
[Multi-Object Tracking Layer]
(Kalman Filter + Data Association)
│
▼
[Threat Scoring & Fusion Engine]
(Weights, Zone Multipliers, Persistence)
│
▼
[Human Review & Alert Dispatch]
/ │ \
▼ ▼ ▼
[Dashboard] [Cloud API] [Notifications]
3.2 Detailed Data Flow Lifecycle
- Ingestion: Reads raw frames from input stream at specified frame rate.
- Preprocessing: Resizes frames, normalizes color space, and applies optional frame-skipping.
- Detection: Runs active detectors in parallel, yielding raw
Detectionobjects. - Tracking: Assigns persistent track IDs and computes movement vectors.
- Scoring: Fuses detections, class severity, zone weights, and persistence into a single threat score $S \in [0, 1]$.
- Review Queue: Routes incidents above threshold to human review.
- Action: Dispatches alerts to Telegram, Discord, Slack, Webhooks, or REST endpoints upon confirmation.
4. Mathematical Foundations
4.1 Intersection over Union (IoU)
Given two bounding boxes $A = (x_{A1}, y_{A1}, x_{A2}, y_{A2})$ and $B = (x_{B1}, y_{B1}, x_{B2}, y_{B2})$:
$$\text{Area}(A) = (x_{A2} - x_{A1}) \times (y_{A2} - y_{A1})$$
$$\text{Area}(B) = (x_{B2} - x_{B1}) \times (y_{B2} - y_{B1})$$
$$I(A, B) = \max(0, \min(x_{A2}, x_{B2}) - \max(x_{A1}, x_{B1})) \times \max(0, \min(y_{A2}, y_{B2}) - \max(y_{A1}, y_{B1}))$$
$$\text{IoU}(A, B) = \frac{I(A, B)}{\text{Area}(A) + \text{Area}(B) - I(A, B)}$$
Numerical Example: Box $A = [0, 0, 10, 10]$ ($\text{Area} = 100$), Box $B = [5, 5, 15, 15]$ ($\text{Area} = 100$). Intersection is $[5, 5, 10, 10]$ ($\text{Area} = 25$). Union is $100 + 100 - 25 = 175$.
$$\text{IoU} = \frac{25}{175} \approx 0.143$$
4.2 Precision, Recall, and F1 Score
$$\text{Precision} = \frac{\text{TP}}{\text{TP} + \text{FP}}$$
$$\text{Recall} = \frac{\text{TP}}{\text{TP} + \text{FN}}$$
$$\text{F1} = 2 \times \frac{\text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}}$$
Numerical Example: $\text{TP} = 80, \text{FP} = 20, \text{FN} = 10 \implies \text{Precision} = \frac{80}{100} = 0.80$, $\text{Recall} = \frac{80}{90} \approx 0.889$, $\text{F1} \approx 0.842$.
4.3 Mean Average Precision (mAP)
$$\text{mAP} = \frac{1}{N} \sum_{i=1}^N \text{AP}_i$$
Where $\text{AP}_i$ is the area under the Precision-Recall curve for class $i$. Evaluated at $\text{mAP}@0.5$ and $\text{mAP}@[0.5:0.95]$.
4.4 Non-Maximum Suppression (NMS)
Sort detections $D$ by confidence descending. Select box $b_{max}$ with highest confidence, add to kept set $K$, and discard any box $b \in D$ where $\text{IoU}(b_{max}, b) > \tau_{\text{NMS}}$ (default $\tau_{\text{NMS}} = 0.45$).
4.5 Activation Functions: Sigmoid and Softmax
$$\text{Sigmoid}(x) = \frac{1}{1 + e^{-x}}$$
$$\text{Softmax}(z_i) = \frac{e^{z_i}}{\sum_{j=1}^K e^{z_j}}$$
4.6 Loss Functions: Cross-Entropy
- Binary Cross-Entropy (BCE):
$$\mathcal{L}_{\text{BCE}} = - [y \log(p) + (1 - y) \log(1 - p)]$$
- Categorical Cross-Entropy (CCE):
$$\mathcal{L}{\text{CCE}} = - \sum{c=1}^C y_c \log(p_c)$$
4.7 Kalman Filter Tracking Formulation
State vector $\mathbf{x}_k = [x, y, v_x, v_y]^T$.
- Predict Step:
$$\mathbf{\hat{x}}k = \mathbf{F} \mathbf{x}{k-1}$$
$$\mathbf{P}k = \mathbf{F} \mathbf{P}{k-1} \mathbf{F}^T + \mathbf{Q}$$
- Update Step:
$$\mathbf{K}_k = \mathbf{P}_k \mathbf{H}^T (\mathbf{H} \mathbf{P}_k \mathbf{H}^T + \mathbf{R})^{-1}$$
$$\mathbf{x}_k = \mathbf{\hat{x}}_k + \mathbf{K}_k (\mathbf{z}_k - \mathbf{H} \mathbf{\hat{x}}_k)$$
$$\mathbf{P}_k = (\mathbf{I} - \mathbf{K}_k \mathbf{H}) \mathbf{P}_k$$
4.8 Distance Metrics & Data Association
- Cosine Similarity:
$$\text{sim}(\mathbf{u}, \mathbf{v}) = \frac{\mathbf{u} \cdot \mathbf{v}}{|\mathbf{u}| |\mathbf{v}|}$$
- Euclidean Distance:
$$d(\mathbf{u}, \mathbf{v}) = \sqrt{\sum_{i=1}^n (u_i - v_i)^2}$$
- Association Cost Matrix:
$$\mathbf{C}{i,j} = \alpha (1 - \text{IoU}{i,j}) + (1 - \alpha)(1 - \text{sim}_{i,j})$$
5. Installation Guide
5.1 Package Installation
# Core CPU installation
pip install threatvision-ai
# Installation with GPU acceleration (PyTorch CUDA / ONNX Runtime GPU)
pip install "threatvision-ai[gpu]"
5.2 Local Source Setup
git clone https://github.com/Amit123103/multithread_detection.git
cd multithread_detection
pip install -e ".[dev]"
5.3 Containerized Execution (Docker)
docker build -t threatvision:latest .
docker run --gpus all -p 8000:8000 threatvision:latest
6. Quick Start Guide
from threatvision import ThreatVision
# Initialize security engine
tv = ThreatVision(camera=0, dashboard=True, save_incidents=True)
# Enable active threat detectors
tv.enable_person_detection(threshold=0.5)
tv.enable_weapon_detection(threshold=0.6)
tv.enable_fire_detection(threshold=0.55)
tv.enable_fight_detection(threshold=0.6)
tv.enable_smoke_detection(threshold=0.55)
# Start real-time analysis pipeline
tv.start(block=True)
7. Package Structure
threatvision-ai/
├── threatvision/
│ ├── __init__.py # Core exports
│ ├── engine.py # Main ThreatVision manager
│ ├── camera/ # Camera & RTSP streaming
│ ├── detectors/ # 11 Specialized Detectors
│ ├── models/ # YOLO, RT-DETR, ONNX, Fallback engine
│ ├── tracking/ # Multi-object tracker
│ ├── analytics/ # Threat engine & spatial rules
│ ├── storage/ # Incident manager (JSON/CSV)
│ ├── reports/ # ReportLab PDF exporter
│ ├── alerts/ # Alert data models
│ ├── notifications/ # Telegram, Discord, Slack, Webhooks
│ ├── api/ # FastAPI server & MJPEG streaming
│ ├── dashboard/ # Static Web UI assets
│ ├── plugins/ # Plugin SDK & registry
│ ├── cli/ # Command Line Interface
│ ├── config/ # Settings & YAML loader
│ ├── logging/ # Structured logging
│ └── utils/ # Geometry, draw, and metrics
├── tests/ # 32 Automated Unit Tests
├── pyproject.toml # Build metadata & dependencies
└── README.md
8. Detection Modules Specification
ThreatVision AI features 11 specialized detection modules:
- Person Detector — Human detection baseline for behavior analysis.
- Weapon Detector — Identifies visual patterns for handguns, rifles, and knives.
- Fire Detector — Detects open flames via HSV color analytics and deep model inference.
- Smoke Detector — Detects smoke plumes via chrominance and contrast analysis.
- Vehicle Detector — Detects cars, trucks, buses, and motorcycles.
- Accident Detector — Identifies vehicle collisions using bounding box IoU dynamics.
- Fight Detector — Flags physical altercations based on high-energy overlapping person bounding boxes.
- Fall Detector — Detects human falls when bounding box aspect ratio $\text{AR} = \frac{w}{h} > 1.25$.
- Intrusion Detector — Flags ray-casting point-in-polygon entry into restricted perimeter zones.
- Crowd Detector — Triggers density alerts when person counts in a region exceed threshold.
- Package Detector — Detects unattended backpacks, suitcases, and parcels.
9. Threat Scoring Engine & Risk Matrix
9.1 Threat Score Formula
$$S = \min\left(1.0, , w_1 C + w_2 W_{\text{class}} + w_3 W_{\text{zone}} + w_4 F_{\text{persist}}\right)$$
Where $C$ is model confidence, $W_{\text{class}}$ is class severity, $W_{\text{zone}}$ is zone multiplier, $F_{\text{persist}}$ is temporal persistence factor, and $w_1=0.4, w_2=0.3, w_3=0.2, w_4=0.1$.
9.2 Risk Category Mapping
$$L(S) = \begin{cases} \text{CRITICAL}, & S \ge 0.85 \ \text{HIGH}, & 0.60 \le S < 0.85 \ \text{MEDIUM}, & 0.30 \le S < 0.60 \ \text{LOW}, & 0.15 \le S < 0.30 \ \text{SAFE}, & S < 0.15 \end{cases}$$
10. Web Dashboard & Command Center UI
The Web Dashboard runs on FastAPI and HTML5/JS:
- Live Grid View: Real-time camera feed rendering with overlay HUD.
- Incident Timeline: Filterable event log with threat scores and screenshots.
- Telemetry Statistics: FPS, memory usage, CPU/GPU utilization charts.
- Review Queue: Confirmation controls for human operators.
11. Cloud REST API Specification
| Method | Endpoint | Description |
|---|---|---|
GET |
/health |
Server health status |
GET |
/statistics |
FPS, threat telemetry, CPU/RAM utilization |
POST |
/detect |
Upload frame for detection JSON payload |
GET |
/history |
Fetch incident log history |
GET |
/history/csv |
Download incident log as CSV |
GET |
/stream |
Live video MJPEG stream |
12. Python API Reference
from threatvision import ThreatVision, ThreatVisionConfig
config = ThreatVisionConfig()
tv = ThreatVision(camera=0, config=config)
tv.enable_person_detection(threshold=0.5)
tv.enable_weapon_detection(threshold=0.6)
tv.start(block=False)
stats = tv.get_statistics()
print(stats)
13. Configuration Management
Configurations can be loaded from YAML, TOML, or JSON:
camera:
width: 1280
height: 720
skip_frames: 0
analytics:
threat_threshold_low: 0.25
threat_threshold_high: 0.75
notifications:
enable_webhook: false
14. Command Line Interface (CLI)
# Run camera stream with web dashboard
threatvision camera --source 0 --port 8000
# Run offline video file evaluation
threatvision video sample.mp4
# Run single image evaluation
threatvision image frame.jpg
# Launch standalone dashboard server
threatvision dashboard --host 127.0.0.1 --port 8000
# Run benchmark test
threatvision benchmark --frames 100
15. Notification Channels
ThreatVision AI dispatches alerts to multiple destinations:
- Telegram Bot: API token & chat ID integration.
- Discord Webhook: Embedded alert messages with threat score details.
- Slack Webhook: Custom channel alert payloads.
- Generic HTTP Webhook: JSON payloads for integration with SIEMs.
- Audible Local Alarm: System alert beep on CRITICAL severity.
16. Logging & Monitoring
- Structured JSON Logs: Machine-readable logs for ELK / Loki pipelines.
- Performance Monitor: FPS and inference latency metrics via
PerformanceMonitor.
17. Performance Optimization
- GPU Acceleration: PyTorch CUDA and ONNX Runtime GPU support.
- Frame Skipping: Skip non-critical frames to optimize multi-camera throughput.
- Quantization: INT8 quantization support for edge devices.
18. Security & Responsible AI
- Human-in-the-Loop: High-severity automated actions require operator review by default.
- Data Minimization: Only metadata and short incident clips are logged.
- No Autonomous Weapon Integration: Strictly designed for monitoring and alerting.
19. Custom Plugin Development Guide
from threatvision import Plugin, register_plugin, Detection
import numpy as np
@register_plugin
class ThermalDetector(Plugin):
def __init__(self):
super().__init__(name="thermal")
def detect(self, frame: np.ndarray):
return [
Detection(label="hotspot", confidence=0.88, box=(50, 50, 150, 150), category="thermal")
]
20. Testing & Verification
The suite contains 32 automated tests:
pytest -o addopts=""
32 passed in 5.89s
21. Production Deployment Guide
Linux Systemd Service
[Unit]
Description=ThreatVision AI Service
After=network.target
[Service]
ExecStart=/usr/local/bin/threatvision camera --source 0
Restart=always
[Install]
WantedBy=multi-user.target
22. Frequently Asked Questions (FAQ)
- Q: Is GPU required? A: No, CPU inference with synthetic/ONNX fallback is supported out of the box.
- Q: Does ThreatVision AI perform facial recognition? A: No. ThreatVision AI tracks object shapes and bounding boxes, not personal biometric identities.
23. Troubleshooting Matrix
| Problem | Cause | Solution |
|---|---|---|
| Camera fails to open | Wrong index or RTSP URL | Verify camera index or test RTSP URL in VLC/ffprobe |
| Low FPS | High frame resolution or CPU bound | Enable frame skipping or use ONNX/CUDA backend |
Missing python-multipart error |
Form parser dependency missing | Run pip install python-multipart |
24. Changelog
- v2.4.1: Complete CI/CD fixes, Ruff import sorting, Mypy type annotation fixes,
python-multipartintegration, and full academic manual documentation. - v1.0.0: Initial release with person, vehicle, and weapon detection modules.
25. Roadmap
- Near-term: Multi-camera cross-view tracking and adaptive thresholding.
- Long-term: Edge device optimization for Jetson Orin and Raspberry Pi 5.
26. Comprehensive Technical & Domain Glossary
- IoU (Intersection over Union): Ratio of bounding box intersection area to union area.
- NMS (Non-Maximum Suppression): Post-processing step eliminating overlapping bounding box proposals.
- Kalman Filter: Recursive state estimation algorithm predicting object trajectory across frames.
- Threat Score: Calibrated scalar metric $S \in [0, 1]$ representing hazard severity.
License
ThreatVision AI is released under the MIT License.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file threatvision_ai-2.4.1.tar.gz.
File metadata
- Download URL: threatvision_ai-2.4.1.tar.gz
- Upload date:
- Size: 45.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
74ce5f017a8db92fd0a22a1d62f64ef5f6b7a6e626a45e96c2c9c5c50f92a790
|
|
| MD5 |
292fc9c6c1ba02f71ddf0995d1fa718b
|
|
| BLAKE2b-256 |
5ac6b4746c45f83e0ac67632f2e369752de6aaa815b8158c8bccd3101f1f3894
|
File details
Details for the file threatvision_ai-2.4.1-py3-none-any.whl.
File metadata
- Download URL: threatvision_ai-2.4.1-py3-none-any.whl
- Upload date:
- Size: 56.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8f6ca99b30ab9ec28eda6037441f2d6303a9aa67159feaf3d8c171e6b7147e91
|
|
| MD5 |
28d2abb7050dfc72113228add12189dd
|
|
| BLAKE2b-256 |
bdbf99b8ba5c0c355f8d80bc095564fac3c61424d5908d8c3da916cc8edce36a
|