ReproGuard
Pre-production risk scanner for Jupyter notebooks, Python scripts, and ML repositories.
ReproGuard analyzes notebooks, scripts, and whole repos for reproducibility risks, data leakage, PII & secret leaks, missing dependencies, GenAI/LLM configuration risks, and handoff readiness — before you share, review, or promote work toward production.
Positioning: SonarQube-style review for data science work. Not a replacement for MLflow, DVC, Databricks, DataHub, or data observability platforms.
Use Cases
| When | How ReproGuard helps |
|---|---|
| Before code review / PR | Catches out-of-order notebook execution, missing seeds, local file paths, and undefined variables — so reviewers see clean, reproducible work instead of debugging hidden state |
| Before sharing a notebook | Flags embedded PII (emails, phone numbers, card numbers), API keys, and base64 blobs in outputs before a notebook leaves your machine |
| Before production handoff | Verifies dependency files exist and are pinned, handoff docs are present, execution counts are clean, and no data leakage patterns (pre-split preprocessing, target leakage) remain |
| GenAI / LLM project audits | Statically checks LangGraph/CrewAI agent structure, LLM client config (model pinning, temperature, max_tokens, timeouts), prompt injection bait, system-prompt leakage, and llm.yaml manifests — no API keys needed |
| CI gate | Runs in GitHub Actions / GitLab CI / pre-commit with --fail-under, --fail-new, and SARIF upload for Code Scanning — blocks regressions, allows debt to be paid down gradually |
| Secret hygiene | --git-history catches secrets committed in the past; --fix gitignore prevents future .env commits |
| Open-source / release hygiene | Checks README presence, CI workflows, .env.example templates, DVC outputs on disk, and large model artifacts in git — before you publish a repo or cut a release |
What's New in v0.5
v0.5.1 patch — trust and habit fixes with no scoring changes: DATA001 and
DATA002 no longer flag paths in comments or docstrings; PY001 reports a
stable column for truncated-escape errors on every Python version; every text
scan ends with a top-three Next actions summary (fix + example location +
reproguard explain pointer); and scan --staged gates only staged files
for fast pre-commit runs (reproguard scan . --staged --fail-under 75).
Dogfood root scans exclude the site demo copies. Full details in the
changelog.
| Area | What's new | How to use |
|---|---|---|
| New check: DEP004 | Unpinned !pip install / %pip install in notebook cells — the #1 reproducibility gap in shared notebooks |
reproguard scan notebooks/ |
| New check: SEC013 | Literal credentials in URL query strings (?api_key=sk-…) — URLs leak into logs, browser history, error reports |
reproguard scan . --privacy |
| Stdout streaming | Stream machine formats to stdout for jq/CI piping instead of writing files | reproguard scan . -f json -o - | jq .score |
| GitHub Action | The reusable action now emits SARIF by default and uploads all reports | see GitHub Action |
| Scan history integrity | Same-tick scans no longer overwrite each other; trend deltas and --fail-trend are reliable |
nothing to do — fixed behavior |
What's New in v0.4
v0.4 adds nine checks (73 → 82), three new CLI commands, and faster, quieter scans:
| Area | What's new | How to use |
|---|---|---|
| CLI commands | init scaffolds config + CI workflow, explain <CODE> documents any check, list shows the inventory |
reproguard init --ci github · reproguard explain DOCKER001 |
| Incremental scans | --since <ref> scans only files changed since a git ref — fast PR scans |
reproguard scan . --since HEAD~3 |
| Score trend gate | per-project scan history + drop gate for CI drift detection | reproguard trend -n 10 · --fail-trend 5 |
| Dockerfile checks | DOCKER001 unpinned base image, DOCKER002 unpinned installs | docs/CHECKS.md |
| New checks | REPO008 missing license, REPO009 large data in git, HAND002 uncommitted changes, NB008 never-executed notebook, SEC012 credentials in URIs, PII006 deep PII (Presidio NER) | --presidio (extra: pip install reproguard[privacy]) |
| SBOM | CycloneDX 1.5 bill of materials from declared dependencies | --sbom bom.json |
| CI integrations | GitLab Code Quality report + GitHub Actions ::error annotations |
--format gitlab |
| Reports | interactive HTML (severity filters), Hugging Face model-card variant | --format html · --model-card-format hf |
| Faster scans | single-pass repo walk, indexed line lookups, cached reads, incremental notebook state analysis | ~2× faster walks, no flags needed |
| Quieter scans | 40+ false-positive fixes from field scans (corrupt notebooks, tool-cache dirs, @tool() variants, Windows globs, …) |
automatic |
v0.4.1 patch — scan --since <ref> no longer prints a stray None
line after the incremental-scan banner; Python 3.13 is now tested and
declared; CI enforces the coverage floor (≥ 85%) and gates the clean
example fixture on its score. Full details in the
changelog.
Every new check and flag is documented with runnable examples below and in the changelog.
What's New in v0.3 (history)
v0.3 doubled the check inventory (29 → 73) and added the GenAI risk category, custom rule engine, scan profiles, auto-fixes, git-history secret scanning, llm.yaml manifests, model cards, and a reusable GitHub Action. v0.3.4 automated the release pipeline (readiness gate + curated release notes).
Why This Exists
Data science projects often work on the author's machine but fail during review or production handoff because of:
- 📍 Local data paths —
C:\Users\...or/Users/...that don't exist on another machine - 📦 Missing dependencies — No
requirements.txtor unpinned packages causing environment drift - 🔄 Out-of-order execution — Notebook cells run in a non-linear order, hiding stateful assumptions
- 🔓 Hidden PII & secrets — Email addresses, API keys, or credentials buried in code or output
- 📊 Data leakage — Preprocessing before train/test split, test data in fit calls, target-like features
- 📝 Missing handoff docs — No clear statement of objective, data source, assumptions, or instructions
ReproGuard is local-first — scans happen on your machine without uploading data or notebooks to any third-party service.
Features
Detects 84 risk patterns across 7 categories
| Risk Category | What It Catches | Severity Range |
|---|---|---|
| Reproducibility | Missing dependency files, unpinned packages, no random seeds, out-of-order execution, stale outputs | LOW → HIGH |
| Data Leakage | Preprocessing before split, target-like feature columns, test data in fit calls, suspiciously high metrics, tabular data dumps, output tracebacks | LOW → CRITICAL |
| Privacy & Security | Email addresses, phone numbers, credit card numbers, AWS keys, hardcoded secrets, private key blocks, high-entropy credentials, base64 images, JSON blobs | LOW → CRITICAL |
| Data Dependency | Local machine paths, missing referenced data files, hardcoded paths in output | MEDIUM → HIGH |
| Handoff Readiness | Missing objective, data source, assumptions, or metric documentation | LOW |
| Execution | Notebook execution failures, kernel/dependency setup errors | HIGH → CRITICAL |
| GenAI | LLM client config (temperature, max_tokens, model pinning, timeouts), prompt injection bait, interpolated untrusted content, LangGraph/CrewAI agent structure, tool agency, system-prompt leakage, fine-tuning contamination, llm.yaml manifests | LOW → CRITICAL |
Output Formats
- Terminal — Color-coded summary with severity breakdown
- JSON — Structured data for programmatic consumption
- HTML — Interactive standalone report with severity filters
- SARIF 2.1.0 — Compatible with GitHub Code Scanning and VS Code
- GitLab Code Quality —
gl-code-quality-report.jsonfor GitLab CI merge-request annotations
Scoring
Penalty-based scoring from 0–100. Status: ready_with_caution (≥75), needs_review (50–74), or not_ready (<50 or any CRITICAL issue). All penalties and weights are configurable.
Installation
pip install reproguard
From source
git clone https://github.com/vipulgote1999/ReproGuard.git
cd ReproGuard
pip install -e .
Development install
pip install -e .[dev]
Privacy extra (Presidio support — future)
pip install reproguard[privacy]
Quick Start
Scan a single notebook:
reproguard scan examples/risky_customer_churn.ipynb
Scan an entire project directory:
reproguard scan .
Scan with HTML report and fail CI on low score:
reproguard scan . --format html --fail-under 50
Enable clean notebook execution:
reproguard scan notebook.ipynb --execute
Usage Examples
Basic scan
reproguard scan examples/risky_customer_churn.ipynb
Output:
ReproGuard score: 27/100 (not_ready)
Files scanned: 1 | Issues: 8
Critical: 1 | High: 3 | Medium: 2 | Low: 2
Severity Code Issue Location
──────── ──── ───── ────────
CRITICAL LEAK001 Possible preprocessing before… risky_customer_churn.ipynb:cell 6
HIGH DATA001 Local machine path detected risky_customer_churn.ipynb:line 2
HIGH LEAK002 Future/target-like column nam… risky_customer_churn.ipynb:cell 4
HIGH LEAK007 Exception traceback found in… risky_customer_churn.ipynb:cell 9
MEDIUM PII001 Email address detected risky_customer_churn.ipynb:cell 8
MEDIUM REP001 Non-deterministic code witho… risky_customer_churn.ipynb:cell 6
LOW NB001 Notebook cells were executed… risky_customer_churn.ipynb
LOW NB002 Notebook output exists witho… risky_customer_churn.ipynb:cell 9
Scan a GenAI project
ReproGuard detects LLM configuration risks, prompt-injection bait, and agent structure issues statically — no API keys or network access needed:
# GenAI-tuned weights: genai findings penalize 35x, handoff only 5x
reproguard scan . --profile genai
Real output from scanning a LangGraph + FastAPI agent repository:
ReproGuard score: 25/100 (not_ready)
Files scanned: 29 | Issues: 43
Points deducted by category: genai: -35, handoff: -0, privacy: -10, reproducibility: -30
Severity Code Issue Location
──────── ──── ───── ────────
HIGH AGENT004 Tool accepts free-form src/tools/research_tool.py:33
input without an allowlist
HIGH REPO001 '.env' file is present .env
and not gitignored
LOW AGENT002 Agentic graph without src/agents/base_agent.py:74
explicit recursion limit
LOW LLMC003 Model ID is not version- src/config/settings.py:19
pinned (gpt-4o-mini)
LOW LLMC004 High temperature on src/config/settings.py:24
agentic model (0.7)
Run --format json for machine-readable findings, or --model-card for a
Markdown handoff summarizing risks and recommended actions.
Custom rules
Add org-specific checks in .reproguard.yml — they behave like built-ins
(scored, reported, suppressible via checks.disabled/checks.ignore):
# .reproguard.yml
rules:
- code: NOSAMPLE # Flag sampling without a fixed random_state
title: Sample without seed
pattern: 'sample\(' # plain regex matched against source text
category: reproducibility
severity: medium
confidence: 0.8
file_patterns:
- 'src/**/*.py'
- '*.ipynb'
- code: TODO001 # Track tech-debt markers
title: TODO left in code
pattern: '# TODO'
severity: low
Generate reports
# All report formats
reproguard scan . --format all
# JSON only
reproguard scan . --format json
# SARIF for GitHub Code Scanning
reproguard scan . --format sarif
# Custom output directory
reproguard scan . --output-dir scan-reports
CI integration
# Fail the build if the score is too low
reproguard scan . --fail-under 50
echo $? # Exit code 1 when score < 50 or any CRITICAL issue
# Stream JSON to stdout for jq / other tooling (no files written)
reproguard scan . -f json -o - | jq '.score'
GitHub Action
Scan in CI with the reusable action (SARIF by default — feed it straight into Code Scanning):
steps:
- uses: actions/checkout@v4
- uses: vipulgote1999/ReproGuard/.github/actions/reproguard-scan@v0.5.1
with:
target: "."
fail-under: 50
profile: genai
format: sarif
- uses: github/codeql-action/upload-sarif@v3
if: always()
with:
sarif_file: .reproguard/reproguard-report.sarif
The action installs ReproGuard from PyPI, runs the scan, and uploads all
generated reports (.reproguard/) as an artifact. Inputs: target, format
(text/json/html/sarif/gitlab/all), fail-under, fail-new, baseline,
execute, profile, version.
Auto-fixes, model cards, git history
# Safe, idempotent remediations:
# gitignore — append .env protection to .gitignore
# clear-counts — reset notebook execution counts (clean handoff)
# seed — inject missing random/numpy seeds into notebooks
reproguard scan . --fix all # or pick one: --fix seed
# Markdown handoff artifact (risk summary + recommended actions)
reproguard scan . --model-card
# gitleaks-lite: scan git history diffs for committed secrets
reproguard scan . --git-history
# Validate llm.yaml model IDs against the OpenRouter catalog
reproguard scan . --manifest-online
# GenAI-tuned category weights (or 'profile: genai' in .reproguard.yml)
reproguard scan . --profile genai
# Parallel clean execution of notebooks (kernelspec-aware)
reproguard scan . --execute --parallel --max-workers 4
The LLM manifest validated by --manifest-online lives at the repo root:
# llm.yaml
models:
- id: gpt-4o-2024-08-06 # pinned — never 'gpt-4o' or '*-latest'
provider: openai
prompts:
- id: rag
path: prompts/rag.md # must exist on disk
Prompt files (prompts/*.md|txt|yaml|json) are scanned as whole prompts —
injection-bait phrasing, PII, and oversized prompts are flagged.
Regression detection (baseline diff)
Gate CI on new issues while existing debt is paid down gradually:
# First run: save the baseline
reproguard scan . --format json --output-dir .reproguard
# Later runs: block on new issues only
reproguard scan . --baseline .reproguard/reproguard-report.json --fail-new 0
reproguard scan . --baseline .reproguard/reproguard-report.json --fail-new-critical
New and resolved issues are printed in the terminal summary and recorded in the JSON report metadata. Exit codes: 1 when the gate trips, 2 for invalid baselines.
Scan with privacy disabled
reproguard scan . --no-privacy
Custom execution timeout
reproguard scan notebook.ipynb --execute --execution-timeout 300
Understanding Reports
Score interpretation
| Score | Status | Action Required |
|---|---|---|
| ≥ 75 | ready_with_caution |
Review minor issues before production |
| 50–74 | needs_review |
Address significant issues |
| < 50 | not_ready |
Blocking issues — must fix |
| Any CRITICAL | not_ready |
Immediate attention required |
| 0 files scanned | no_files |
No supported files found — check the scan path (exit code 2) |
Report files
Reports are written to .reproguard/ by default:
.reproguard/
├── reproguard-report.json # Structured data
├── reproguard-report.html # Styled HTML report
└── reproguard-report.sarif # GitHub Code Scanning compatible
Configuration
Create a .reproguard.yml in your project root:
# .reproguard.yml
exclude_paths:
- "archive/**"
- "tests/**"
exclude_dirs:
- "scratch"
fail_under: 50
checks:
disabled:
- "LEAK005" # Disable large tabular output check
- "PII004" # Disable base64 image check
Configuration is discovered by walking up from the scan path (like git). See docs/CONFIGURATION.md for the full reference.
Pre-commit Hook
# .pre-commit-config.yaml
repos:
- repo: https://github.com/vipulgote1999/ReproGuard
rev: v0.5.1
hooks:
- id: reproguard-scan
args: ["--fail-under", "75"]
The hook scans the entire repository on each commit and blocks the commit when the score falls below the threshold.
CI/CD Integration
GitHub Actions (with SARIF upload)
name: ReproGuard
on: [push, pull_request]
jobs:
scan:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.11"
- run: pip install reproguard
- run: reproguard scan . --format sarif --fail-under 50
- uses: github/codeql-action/upload-sarif@v3
with:
sarif_file: .reproguard/reproguard-report.sarif
GitLab CI
reproguard:
stage: test
script:
- pip install reproguard
- reproguard scan . --format html --fail-under 50
artifacts:
paths:
- .reproguard/
Project Structure
ReproGuard/
├── src/reproguard/
│ ├── cli.py # Typer CLI entry point
│ ├── scanner.py # Scan orchestrator
│ ├── models.py # Core data models & scoring
│ ├── config.py # .reproguard.yml loader (profiles, rules)
│ ├── python_analysis.py # Python source analysis
│ ├── notebook.py # Notebook parser + hidden-state analysis
│ ├── leakage.py # ML data leakage heuristics
│ ├── privacy.py # PII / secret scanning
│ ├── dependency.py # Dependency file analysis
│ ├── execution.py # Clean execution (parallel, kernel-aware)
│ ├── repo.py # Repository hygiene checks (REPO001-007)
│ ├── genai.py # LLM config / prompt / agent checks
│ ├── unsafe.py # Bandit-lite unsafe Python checks
│ ├── rules.py # Custom rule engine
│ ├── manifest.py # llm.yaml manifest validation
│ ├── fixes.py # --fix auto-remediation
│ ├── githistory.py # --git-history secret scan
│ ├── modelcard.py # --model-card generation
│ ├── report.py # JSON/HTML report generators
│ ├── sarif.py # SARIF 2.1.0 report generator
│ ├── plugin.py # Check registry & filtering
│ └── utils.py # Shared helpers
├── examples/
│ ├── risky_customer_churn.ipynb # Notebook with intentional issues
│ ├── clean_analysis.py # Clean script example
│ └── requirements.txt # Example dependency file
├── docs/
│ ├── ARCHITECTURE.md # Internal design & module docs
│ ├── CHECKS.md # Complete issue code reference
│ ├── CONFIGURATION.md # Configuration file reference
│ ├── GUIDES.md # Usage guides & integrations
│ └── ROADMAP.md # Future plans
├── pyproject.toml # Build & tool config
└── README.md # This file
Check Reference
| Category | Checks | Representative codes |
|---|---|---|
| GenAI | 27 | LLMC001–LLMC008, PROMPT001–PROMPT005, AGENT001–AGENT007, MAN001–MAN004, GEN001–GEN003 |
| Reproducibility | 20 | REP001, NB001–NB008, DEP001–DEP004, DOCKER001–DOCKER002, PY001, REPO002, REPO005, REPO008 |
| Privacy & security | 20 | PII001–PII006, SEC001–SEC013, REPO001 |
| Data dependency | 7 | DATA001–DATA002, LEAK006, REPO003, REPO004, REPO007, REPO009 |
| Data leakage | 5 | LEAK001–LEAK005, LEAK007 |
| Handoff | 3 | HAND001–HAND002, REPO006 |
| Execution | 2 | EXEC001–EXEC002 |
| Total | 84 | across 7 categories |
The per-check table lives in docs/CHECKS.md, which CI enforces against the live check registry — it cannot drift. This README is the PyPI description, so it keeps the summary rather than a table that goes stale the moment a check ships.
Query the registry from the terminal instead:
reproguard list # all 84
reproguard list --category genai # one category
reproguard explain LEAK001 # full record for one check
See docs/CHECKS.md for full details on every check.
Design Principles
- Local-first — Scans run entirely on your machine. No data or notebooks leave your environment.
- Explainable rules — Every issue has a code, severity, evidence, confidence score, and suggested fix. No black boxes.
- Low friction — CLI-first design with a single
reproguard scan <target>command. Pre-commit hook, CI integration, and GitHub Action out of the box. - Conservative scoring — The score is transparent (penalty-based, weighted by severity and category). You can customize all penalties and weights.
- Narrow wedge — Focused on catching pre-production risks before work enters heavier MLOps pipelines.
Limitations
ReproGuard v0.5 (alpha) uses heuristics. It flags likely risks but cannot prove every issue is real. Treat it as a review assistant, not a final governance decision.
- Leakage detection is heuristic — expect false positives (use
checks.ignoreto silence known-safe locations) - Dependency parsing is intentionally lightweight (regex-based for requirements/Pipfile, YAML for conda env files, TOML for pyproject/lock files — no full resolver)
- Privacy scanning uses regex rules by default (Presidio support planned)
- Clean notebook execution depends on local kernel and dependency availability
- Data file existence checks are limited to paths referenced from notebooks/scripts
Development
# Install dev dependencies
pip install -e .[dev]
# Lint
ruff check .
# Test
pytest
# Build release artifacts and validate metadata
python -m build
python -m twine check dist/*
Releases follow the procedure in docs/RELEASING.md — see the checklist there before tagging. Changes are tracked in CHANGELOG.md.
Roadmap
- v0.2 ✅ — correctness hardening (magic handling, conda/lock dependency parsing), path-scoped ignore rules, baseline diffing, extended coverage
- v0.3 ✅ — GenAI category (27 checks), repository hygiene, custom rule engine, profiles, auto-fixes, git-history secret scan, LLM manifests, model cards, parallel + kernel-aware execution, GitHub Action
- v0.4 ✅ —
init/explain/listcommands, incremental--sincescans, score trend +--fail-trend, Dockerfile checks, SBOM export, GitLab Code Quality, interactive HTML report, ~2× faster walks, 40+ false-positive fixes - v0.5 ✅ — DEP004 (unpinned in-notebook pip installs), SEC013
(credentials in URL query strings), stdout streaming (
-o -), GitHub Action emits SARIF by default - Next (v0.6): JupyterLab and VS Code extensions, prompt-suite drift
detection, more
--fixcodes, notebook diffs - v1.0+: ML pipeline scanning, differential scans, team dashboard, API
See docs/ROADMAP.md for the full roadmap.
License & Usage Rights
© 2026 Vipul Gote. All rights reserved.
ReproGuard is free to use for its intended purpose — pre-production risk scanning of data science notebooks, Python scripts, and ML repositories — but it is not open source. You may use the tool as-is for your own scanning work, but you may not:
- copy, reproduce, or clone the source code beyond what is needed to run it;
- modify or build derivative works from it;
- redistribute, sell, sublicense, or offer it as a hosted service;
- reuse its code, heuristics, or ideas to build a competing tool;
- claim authorship or remove the copyright notice.
All other use requires prior written permission from Vipul Gote. See LICENSE for the complete terms.
Release files for reproguard 0.5.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| reproguard-0.5.1.tar.gz | 6.9 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| reproguard-0.5.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 7.0 MB
Release files / reproguard-0.5.1.tar.gz
| Download URL | reproguard-0.5.1.tar.gz |
|---|---|
| Size | 6.9 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
9a3f447f104d85eceebe97dbea93787aa8ec41a628f390c2fb07d5fdabdd5f7b
|
|
BLAKE2b-256 checksum How to use checksums |
a327b2007e452e3a2c65b2a65faeeaa1fb23025ca8e75f0f9186451f0803fcc5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.
Transparency logRelease files / reproguard-0.5.1-py3-none-any.whl
| Download URL | reproguard-0.5.1-py3-none-any.whl |
|---|---|
| Size | 131.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3616ef9c325f6187357baef94332d996f4f0fa841b6319c427d86a3153c147fa
|
|
BLAKE2b-256 checksum How to use checksums |
93a5db6af3605a55bb7c8041e09caebbeac9c86f750b2aa6d0ce0ad091acb386
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.
Transparency log