Skip to main content

ReproGuard

Python License Copyright Ruff PyPI Status Tests

Pre-production risk scanner for Jupyter notebooks, Python scripts, and ML repositories.

ReproGuard analyzes notebooks, scripts, and whole repos for reproducibility risks, data leakage, PII & secret leaks, missing dependencies, GenAI/LLM configuration risks, and handoff readiness — before you share, review, or promote work toward production.

Positioning: SonarQube-style review for data science work. Not a replacement for MLflow, DVC, Databricks, DataHub, or data observability platforms.


Use Cases

When How ReproGuard helps
Before code review / PR Catches out-of-order notebook execution, missing seeds, local file paths, and undefined variables — so reviewers see clean, reproducible work instead of debugging hidden state
Before sharing a notebook Flags embedded PII (emails, phone numbers, card numbers), API keys, and base64 blobs in outputs before a notebook leaves your machine
Before production handoff Verifies dependency files exist and are pinned, handoff docs are present, execution counts are clean, and no data leakage patterns (pre-split preprocessing, target leakage) remain
GenAI / LLM project audits Statically checks LangGraph/CrewAI agent structure, LLM client config (model pinning, temperature, max_tokens, timeouts), prompt injection bait, system-prompt leakage, and llm.yaml manifests — no API keys needed
CI gate Runs in GitHub Actions / GitLab CI / pre-commit with --fail-under, --fail-new, and SARIF upload for Code Scanning — blocks regressions, allows debt to be paid down gradually
Secret hygiene --git-history catches secrets committed in the past; --fix gitignore prevents future .env commits
Open-source / release hygiene Checks README presence, CI workflows, .env.example templates, DVC outputs on disk, and large model artifacts in git — before you publish a repo or cut a release

What's New in v0.5

v0.5.1 patch — trust and habit fixes with no scoring changes: DATA001 and DATA002 no longer flag paths in comments or docstrings; PY001 reports a stable column for truncated-escape errors on every Python version; every text scan ends with a top-three Next actions summary (fix + example location + reproguard explain pointer); and scan --staged gates only staged files for fast pre-commit runs (reproguard scan . --staged --fail-under 75). Dogfood root scans exclude the site demo copies. Full details in the changelog.

Area What's new How to use
New check: DEP004 Unpinned !pip install / %pip install in notebook cells — the #1 reproducibility gap in shared notebooks reproguard scan notebooks/
New check: SEC013 Literal credentials in URL query strings (?api_key=sk-…) — URLs leak into logs, browser history, error reports reproguard scan . --privacy
Stdout streaming Stream machine formats to stdout for jq/CI piping instead of writing files reproguard scan . -f json -o - | jq .score
GitHub Action The reusable action now emits SARIF by default and uploads all reports see GitHub Action
Scan history integrity Same-tick scans no longer overwrite each other; trend deltas and --fail-trend are reliable nothing to do — fixed behavior

What's New in v0.4

v0.4 adds nine checks (73 → 82), three new CLI commands, and faster, quieter scans:

Area What's new How to use
CLI commands init scaffolds config + CI workflow, explain <CODE> documents any check, list shows the inventory reproguard init --ci github · reproguard explain DOCKER001
Incremental scans --since <ref> scans only files changed since a git ref — fast PR scans reproguard scan . --since HEAD~3
Score trend gate per-project scan history + drop gate for CI drift detection reproguard trend -n 10 · --fail-trend 5
Dockerfile checks DOCKER001 unpinned base image, DOCKER002 unpinned installs docs/CHECKS.md
New checks REPO008 missing license, REPO009 large data in git, HAND002 uncommitted changes, NB008 never-executed notebook, SEC012 credentials in URIs, PII006 deep PII (Presidio NER) --presidio (extra: pip install reproguard[privacy])
SBOM CycloneDX 1.5 bill of materials from declared dependencies --sbom bom.json
CI integrations GitLab Code Quality report + GitHub Actions ::error annotations --format gitlab
Reports interactive HTML (severity filters), Hugging Face model-card variant --format html · --model-card-format hf
Faster scans single-pass repo walk, indexed line lookups, cached reads, incremental notebook state analysis ~2× faster walks, no flags needed
Quieter scans 40+ false-positive fixes from field scans (corrupt notebooks, tool-cache dirs, @tool() variants, Windows globs, …) automatic

v0.4.1 patch — scan --since <ref> no longer prints a stray None line after the incremental-scan banner; Python 3.13 is now tested and declared; CI enforces the coverage floor (≥ 85%) and gates the clean example fixture on its score. Full details in the changelog.

Every new check and flag is documented with runnable examples below and in the changelog.


What's New in v0.3 (history)

v0.3 doubled the check inventory (29 → 73) and added the GenAI risk category, custom rule engine, scan profiles, auto-fixes, git-history secret scanning, llm.yaml manifests, model cards, and a reusable GitHub Action. v0.3.4 automated the release pipeline (readiness gate + curated release notes).


Why This Exists

Data science projects often work on the author's machine but fail during review or production handoff because of:

  • 📍 Local data paths — C:\Users\... or /Users/... that don't exist on another machine
  • 📦 Missing dependencies — No requirements.txt or unpinned packages causing environment drift
  • 🔄 Out-of-order execution — Notebook cells run in a non-linear order, hiding stateful assumptions
  • 🔓 Hidden PII & secrets — Email addresses, API keys, or credentials buried in code or output
  • 📊 Data leakage — Preprocessing before train/test split, test data in fit calls, target-like features
  • 📝 Missing handoff docs — No clear statement of objective, data source, assumptions, or instructions

ReproGuard is local-first — scans happen on your machine without uploading data or notebooks to any third-party service.


Features

Detects 84 risk patterns across 7 categories

Risk Category What It Catches Severity Range
Reproducibility Missing dependency files, unpinned packages, no random seeds, out-of-order execution, stale outputs LOW → HIGH
Data Leakage Preprocessing before split, target-like feature columns, test data in fit calls, suspiciously high metrics, tabular data dumps, output tracebacks LOW → CRITICAL
Privacy & Security Email addresses, phone numbers, credit card numbers, AWS keys, hardcoded secrets, private key blocks, high-entropy credentials, base64 images, JSON blobs LOW → CRITICAL
Data Dependency Local machine paths, missing referenced data files, hardcoded paths in output MEDIUM → HIGH
Handoff Readiness Missing objective, data source, assumptions, or metric documentation LOW
Execution Notebook execution failures, kernel/dependency setup errors HIGH → CRITICAL
GenAI LLM client config (temperature, max_tokens, model pinning, timeouts), prompt injection bait, interpolated untrusted content, LangGraph/CrewAI agent structure, tool agency, system-prompt leakage, fine-tuning contamination, llm.yaml manifests LOW → CRITICAL

Output Formats

  • Terminal — Color-coded summary with severity breakdown
  • JSON — Structured data for programmatic consumption
  • HTML — Interactive standalone report with severity filters
  • SARIF 2.1.0 — Compatible with GitHub Code Scanning and VS Code
  • GitLab Code Quality — gl-code-quality-report.json for GitLab CI merge-request annotations

Scoring

Penalty-based scoring from 0–100. Status: ready_with_caution (≥75), needs_review (50–74), or not_ready (<50 or any CRITICAL issue). All penalties and weights are configurable.


Installation

pip install reproguard

From source

git clone https://github.com/vipulgote1999/ReproGuard.git
cd ReproGuard
pip install -e .

Development install

pip install -e .[dev]

Privacy extra (Presidio support — future)

pip install reproguard[privacy]

Quick Start

Scan a single notebook:

reproguard scan examples/risky_customer_churn.ipynb

Scan an entire project directory:

reproguard scan .

Scan with HTML report and fail CI on low score:

reproguard scan . --format html --fail-under 50

Enable clean notebook execution:

reproguard scan notebook.ipynb --execute

Usage Examples

Basic scan

reproguard scan examples/risky_customer_churn.ipynb

Output:

ReproGuard score: 27/100 (not_ready)
Files scanned: 1 | Issues: 8
Critical: 1 | High: 3 | Medium: 2 | Low: 2

 Severity   Code     Issue                          Location
 ────────   ────     ─────                          ────────
 CRITICAL   LEAK001  Possible preprocessing before…  risky_customer_churn.ipynb:cell 6
 HIGH       DATA001  Local machine path detected     risky_customer_churn.ipynb:line 2
 HIGH       LEAK002  Future/target-like column nam…  risky_customer_churn.ipynb:cell 4
 HIGH       LEAK007  Exception traceback found in…   risky_customer_churn.ipynb:cell 9
 MEDIUM     PII001   Email address detected          risky_customer_churn.ipynb:cell 8
 MEDIUM     REP001   Non-deterministic code witho…   risky_customer_churn.ipynb:cell 6
 LOW        NB001    Notebook cells were executed…   risky_customer_churn.ipynb
 LOW        NB002    Notebook output exists witho…   risky_customer_churn.ipynb:cell 9

Scan a GenAI project

ReproGuard detects LLM configuration risks, prompt-injection bait, and agent structure issues statically — no API keys or network access needed:

# GenAI-tuned weights: genai findings penalize 35x, handoff only 5x
reproguard scan . --profile genai

Real output from scanning a LangGraph + FastAPI agent repository:

ReproGuard score: 25/100 (not_ready)
Files scanned: 29 | Issues: 43
Points deducted by category: genai: -35, handoff: -0, privacy: -10, reproducibility: -30

 Severity   Code     Issue                       Location
 ────────   ────     ─────                       ────────
 HIGH       AGENT004 Tool accepts free-form      src/tools/research_tool.py:33
                     input without an allowlist
 HIGH       REPO001  '.env' file is present      .env
                     and not gitignored
 LOW        AGENT002 Agentic graph without       src/agents/base_agent.py:74
                     explicit recursion limit
 LOW        LLMC003  Model ID is not version-    src/config/settings.py:19
                     pinned (gpt-4o-mini)
 LOW        LLMC004  High temperature on         src/config/settings.py:24
                     agentic model (0.7)

Run --format json for machine-readable findings, or --model-card for a Markdown handoff summarizing risks and recommended actions.

Custom rules

Add org-specific checks in .reproguard.yml — they behave like built-ins (scored, reported, suppressible via checks.disabled/checks.ignore):

# .reproguard.yml
rules:
  - code: NOSAMPLE      # Flag sampling without a fixed random_state
    title: Sample without seed
    pattern: 'sample\('   # plain regex matched against source text
    category: reproducibility
    severity: medium
    confidence: 0.8
    file_patterns:
      - 'src/**/*.py'
      - '*.ipynb'
  - code: TODO001       # Track tech-debt markers
    title: TODO left in code
    pattern: '# TODO'
    severity: low

Generate reports

# All report formats
reproguard scan . --format all

# JSON only
reproguard scan . --format json

# SARIF for GitHub Code Scanning
reproguard scan . --format sarif

# Custom output directory
reproguard scan . --output-dir scan-reports

CI integration

# Fail the build if the score is too low
reproguard scan . --fail-under 50
echo $?  # Exit code 1 when score < 50 or any CRITICAL issue

# Stream JSON to stdout for jq / other tooling (no files written)
reproguard scan . -f json -o - | jq '.score'

GitHub Action

Scan in CI with the reusable action (SARIF by default — feed it straight into Code Scanning):

steps:
  - uses: actions/checkout@v4
  - uses: vipulgote1999/ReproGuard/.github/actions/reproguard-scan@v0.5.1
    with:
      target: "."
      fail-under: 50
      profile: genai
      format: sarif
  - uses: github/codeql-action/upload-sarif@v3
    if: always()
    with:
      sarif_file: .reproguard/reproguard-report.sarif

The action installs ReproGuard from PyPI, runs the scan, and uploads all generated reports (.reproguard/) as an artifact. Inputs: target, format (text/json/html/sarif/gitlab/all), fail-under, fail-new, baseline, execute, profile, version.

Auto-fixes, model cards, git history

# Safe, idempotent remediations:
#   gitignore    — append .env protection to .gitignore
#   clear-counts — reset notebook execution counts (clean handoff)
#   seed         — inject missing random/numpy seeds into notebooks
reproguard scan . --fix all          # or pick one: --fix seed

# Markdown handoff artifact (risk summary + recommended actions)
reproguard scan . --model-card

# gitleaks-lite: scan git history diffs for committed secrets
reproguard scan . --git-history

# Validate llm.yaml model IDs against the OpenRouter catalog
reproguard scan . --manifest-online

# GenAI-tuned category weights (or 'profile: genai' in .reproguard.yml)
reproguard scan . --profile genai

# Parallel clean execution of notebooks (kernelspec-aware)
reproguard scan . --execute --parallel --max-workers 4

The LLM manifest validated by --manifest-online lives at the repo root:

# llm.yaml
models:
  - id: gpt-4o-2024-08-06     # pinned — never 'gpt-4o' or '*-latest'
    provider: openai
prompts:
  - id: rag
    path: prompts/rag.md      # must exist on disk

Prompt files (prompts/*.md|txt|yaml|json) are scanned as whole prompts — injection-bait phrasing, PII, and oversized prompts are flagged.

Regression detection (baseline diff)

Gate CI on new issues while existing debt is paid down gradually:

# First run: save the baseline
reproguard scan . --format json --output-dir .reproguard

# Later runs: block on new issues only
reproguard scan . --baseline .reproguard/reproguard-report.json --fail-new 0
reproguard scan . --baseline .reproguard/reproguard-report.json --fail-new-critical

New and resolved issues are printed in the terminal summary and recorded in the JSON report metadata. Exit codes: 1 when the gate trips, 2 for invalid baselines.

Scan with privacy disabled

reproguard scan . --no-privacy

Custom execution timeout

reproguard scan notebook.ipynb --execute --execution-timeout 300

Understanding Reports

Score interpretation

Score Status Action Required
≥ 75 ready_with_caution Review minor issues before production
50–74 needs_review Address significant issues
< 50 not_ready Blocking issues — must fix
Any CRITICAL not_ready Immediate attention required
0 files scanned no_files No supported files found — check the scan path (exit code 2)

Report files

Reports are written to .reproguard/ by default:

.reproguard/
├── reproguard-report.json      # Structured data
├── reproguard-report.html      # Styled HTML report
└── reproguard-report.sarif     # GitHub Code Scanning compatible

Configuration

Create a .reproguard.yml in your project root:

# .reproguard.yml
exclude_paths:
  - "archive/**"
  - "tests/**"
exclude_dirs:
  - "scratch"
fail_under: 50
checks:
  disabled:
    - "LEAK005"     # Disable large tabular output check
    - "PII004"      # Disable base64 image check

Configuration is discovered by walking up from the scan path (like git). See docs/CONFIGURATION.md for the full reference.


Pre-commit Hook

# .pre-commit-config.yaml
repos:
  - repo: https://github.com/vipulgote1999/ReproGuard
    rev: v0.5.1
    hooks:
      - id: reproguard-scan
        args: ["--fail-under", "75"]

The hook scans the entire repository on each commit and blocks the commit when the score falls below the threshold.


CI/CD Integration

GitHub Actions (with SARIF upload)

name: ReproGuard
on: [push, pull_request]
jobs:
  scan:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.11"
      - run: pip install reproguard
      - run: reproguard scan . --format sarif --fail-under 50
      - uses: github/codeql-action/upload-sarif@v3
        with:
          sarif_file: .reproguard/reproguard-report.sarif

GitLab CI

reproguard:
  stage: test
  script:
    - pip install reproguard
    - reproguard scan . --format html --fail-under 50
  artifacts:
    paths:
      - .reproguard/

Project Structure

ReproGuard/
├── src/reproguard/
│   ├── cli.py              # Typer CLI entry point
│   ├── scanner.py          # Scan orchestrator
│   ├── models.py           # Core data models & scoring
│   ├── config.py           # .reproguard.yml loader (profiles, rules)
│   ├── python_analysis.py  # Python source analysis
│   ├── notebook.py         # Notebook parser + hidden-state analysis
│   ├── leakage.py          # ML data leakage heuristics
│   ├── privacy.py          # PII / secret scanning
│   ├── dependency.py       # Dependency file analysis
│   ├── execution.py        # Clean execution (parallel, kernel-aware)
│   ├── repo.py             # Repository hygiene checks (REPO001-007)
│   ├── genai.py            # LLM config / prompt / agent checks
│   ├── unsafe.py           # Bandit-lite unsafe Python checks
│   ├── rules.py            # Custom rule engine
│   ├── manifest.py         # llm.yaml manifest validation
│   ├── fixes.py            # --fix auto-remediation
│   ├── githistory.py       # --git-history secret scan
│   ├── modelcard.py        # --model-card generation
│   ├── report.py           # JSON/HTML report generators
│   ├── sarif.py            # SARIF 2.1.0 report generator
│   ├── plugin.py           # Check registry & filtering
│   └── utils.py            # Shared helpers
├── examples/
│   ├── risky_customer_churn.ipynb  # Notebook with intentional issues
│   ├── clean_analysis.py           # Clean script example
│   └── requirements.txt            # Example dependency file
├── docs/
│   ├── ARCHITECTURE.md     # Internal design & module docs
│   ├── CHECKS.md           # Complete issue code reference
│   ├── CONFIGURATION.md    # Configuration file reference
│   ├── GUIDES.md           # Usage guides & integrations
│   └── ROADMAP.md          # Future plans
├── pyproject.toml          # Build & tool config
└── README.md               # This file

Check Reference

Category Checks Representative codes
GenAI 27 LLMC001–LLMC008, PROMPT001–PROMPT005, AGENT001–AGENT007, MAN001–MAN004, GEN001–GEN003
Reproducibility 20 REP001, NB001–NB008, DEP001–DEP004, DOCKER001–DOCKER002, PY001, REPO002, REPO005, REPO008
Privacy & security 20 PII001–PII006, SEC001–SEC013, REPO001
Data dependency 7 DATA001–DATA002, LEAK006, REPO003, REPO004, REPO007, REPO009
Data leakage 5 LEAK001–LEAK005, LEAK007
Handoff 3 HAND001–HAND002, REPO006
Execution 2 EXEC001–EXEC002
Total 84 across 7 categories

The per-check table lives in docs/CHECKS.md, which CI enforces against the live check registry — it cannot drift. This README is the PyPI description, so it keeps the summary rather than a table that goes stale the moment a check ships.

Query the registry from the terminal instead:

reproguard list                    # all 84
reproguard list --category genai   # one category
reproguard explain LEAK001         # full record for one check

See docs/CHECKS.md for full details on every check.


Design Principles

  1. Local-first — Scans run entirely on your machine. No data or notebooks leave your environment.
  2. Explainable rules — Every issue has a code, severity, evidence, confidence score, and suggested fix. No black boxes.
  3. Low friction — CLI-first design with a single reproguard scan <target> command. Pre-commit hook, CI integration, and GitHub Action out of the box.
  4. Conservative scoring — The score is transparent (penalty-based, weighted by severity and category). You can customize all penalties and weights.
  5. Narrow wedge — Focused on catching pre-production risks before work enters heavier MLOps pipelines.

Limitations

ReproGuard v0.5 (alpha) uses heuristics. It flags likely risks but cannot prove every issue is real. Treat it as a review assistant, not a final governance decision.

  • Leakage detection is heuristic — expect false positives (use checks.ignore to silence known-safe locations)
  • Dependency parsing is intentionally lightweight (regex-based for requirements/Pipfile, YAML for conda env files, TOML for pyproject/lock files — no full resolver)
  • Privacy scanning uses regex rules by default (Presidio support planned)
  • Clean notebook execution depends on local kernel and dependency availability
  • Data file existence checks are limited to paths referenced from notebooks/scripts

Development

# Install dev dependencies
pip install -e .[dev]

# Lint
ruff check .

# Test
pytest

# Build release artifacts and validate metadata
python -m build
python -m twine check dist/*

Releases follow the procedure in docs/RELEASING.md — see the checklist there before tagging. Changes are tracked in CHANGELOG.md.


Roadmap

  • v0.2 ✅ — correctness hardening (magic handling, conda/lock dependency parsing), path-scoped ignore rules, baseline diffing, extended coverage
  • v0.3 ✅ — GenAI category (27 checks), repository hygiene, custom rule engine, profiles, auto-fixes, git-history secret scan, LLM manifests, model cards, parallel + kernel-aware execution, GitHub Action
  • v0.4 ✅ — init/explain/list commands, incremental --since scans, score trend + --fail-trend, Dockerfile checks, SBOM export, GitLab Code Quality, interactive HTML report, ~2× faster walks, 40+ false-positive fixes
  • v0.5 ✅ — DEP004 (unpinned in-notebook pip installs), SEC013 (credentials in URL query strings), stdout streaming (-o -), GitHub Action emits SARIF by default
  • Next (v0.6): JupyterLab and VS Code extensions, prompt-suite drift detection, more --fix codes, notebook diffs
  • v1.0+: ML pipeline scanning, differential scans, team dashboard, API

See docs/ROADMAP.md for the full roadmap.


License & Usage Rights

© 2026 Vipul Gote. All rights reserved.

ReproGuard is free to use for its intended purpose — pre-production risk scanning of data science notebooks, Python scripts, and ML repositories — but it is not open source. You may use the tool as-is for your own scanning work, but you may not:

  • copy, reproduce, or clone the source code beyond what is needed to run it;
  • modify or build derivative works from it;
  • redistribute, sell, sublicense, or offer it as a hosted service;
  • reuse its code, heuristics, or ideas to build a competing tool;
  • claim authorship or remove the copyright notice.

All other use requires prior written permission from Vipul Gote. See LICENSE for the complete terms.

Release files for reproguard 0.5.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for reproguard 0.5.1
File Size Uploaded
reproguard-0.5.1.tar.gz 6.9 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for reproguard 0.5.1
File Interpreter ABI Platform
reproguard-0.5.1-py3-none-any.whl Python 3 none any Details

Total release size: 7.0 MB

Release files / reproguard-0.5.1.tar.gz

Download URL reproguard-0.5.1.tar.gz
Size 6.9 MB
Tags Source
SHA-256 checksum
How to use checksums
9a3f447f104d85eceebe97dbea93787aa8ec41a628f390c2fb07d5fdabdd5f7b
BLAKE2b-256 checksum
How to use checksums
a327b2007e452e3a2c65b2a65faeeaa1fb23025ca8e75f0f9186451f0803fcc5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.

Transparency log

Release files / reproguard-0.5.1-py3-none-any.whl

Download URL reproguard-0.5.1-py3-none-any.whl
Size 131.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3616ef9c325f6187357baef94332d996f4f0fa841b6319c427d86a3153c147fa
BLAKE2b-256 checksum
How to use checksums
93a5db6af3605a55bb7c8041e09caebbeac9c86f750b2aa6d0ce0ad091acb386
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.

Transparency log

Release history Release notifications | RSS feed

0.5.2

2 release files

This release

0.5.1 This release

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page