Skip to main content

RepoMind

Local-first code intelligence for Python repositories: deep static analysis, repository health and risk hotspots — with an optional LLM layer on top.

CI License: MIT Python 3.12+ Ruff Checked with mypy PRs Welcome

RepoMind answers one question: where should a team spend its refactoring budget?

It parses a Python repository, builds a module dependency graph, computes complexity metrics, finds dead code, and correlates everything with Git churn to rank the highest-risk parts of the codebase. The result is a health score, a ranked list of explainable findings, and shareable Markdown/JSON reports.

  • Deterministic first — the entire analysis works offline, with no API keys and no LLM.
  • Explainable — every finding carries the measured values and thresholds that produced it.
  • Fast — pure AST analysis plus a linear-time dependency graph; no code is executed.
  • Extensible — rules, parsers and reporters are separate layers with stable contracts.
  • Local-first — your code never leaves your machine unless you explicitly opt into a cloud model.

Features

Area What RepoMind detects
Complexity cyclomatic complexity, cognitive complexity, deep nesting
Maintainability long functions, wide parameter lists, oversized files
Design god objects (method count, class size, weighted methods per class)
Dead code unused imports, unreferenced private functions and methods
Dependencies import cycles (strongly connected components), external package usage, hub modules
History churn per file, authors, recency, and complexity × churn hotspots
Correctness files that fail to parse
Configuration path excludes, per-rule ignore entries, per-repo thresholds

Reports: rich terminal output, Markdown for pull requests, a self-contained HTML report, SARIF for GitHub code scanning, and JSON for automation (--fail-under and --fail-on-new turn RepoMind into a CI quality gate).

Quick start

Requires Python 3.12+.

# with uv (recommended)
uv venv
uv pip install -e .

# or with pip
python -m venv .venv
.venv\Scripts\activate        # Windows
# source .venv/bin/activate   # Linux/macOS
pip install -e .

After the first release the PyPI package will be repomind-analyzer (the plain repomind name is taken on PyPI by an unrelated project); it installs the same repomind command:

pipx install repomind-analyzer

Analyze a repository:

repomind analyze /path/to/repository
╭──────────────────────────────── RepoMind ─────────────────────────────────╮
│ Repository   /home/dev/RepoMind                                           │
│ Python code  50 files | 3,700 source lines                                │
│ Duration     0.46 s                                                       │
│ Git history  analyzed                                                     │
╰───────────────────────────────────────────────────────────────────────────╯

██████████  100/100 | grade A

No findings above the configured thresholds.
2 finding(s) suppressed by configuration (details in --format json).

Dependency graph
 Metric            Value
 Internal modules     50
 Internal imports    118
 Import cycles         0
 External packages    23

Change hotspots (last 5 commits on master)
 File                              Commits   Churn   Authors   Last change   Heat
 src/repomind/core/pyparser.py           3     423         1   2026-09-22     100
 src/repomind/cli/app.py                 2     500         1   2026-09-22      94
 src/repomind/core/engine.py             3     296         1   2026-09-22      70

Tip: use --format markdown --output report.md for a shareable report or --format json for machine-readable output.

Output above is trimmed and uncolored; borders and the score bar adapt to your terminal's encoding. The two suppressed entries are deliberate and explained in Dogfooding. A full HTML report generated from the deliberately flawed sample project is checked in at docs/example-report.html.

Generate shareable reports:

repomind analyze . --format markdown --output report.md
repomind analyze . --format json --output report.json
repomind analyze . --min-severity high --top 10
repomind analyze . --exclude migrations --exclude "*_pb2.py"
repomind analyze . --since main         # review only files changed since main
repomind analyze . --fail-under 70     # exit code 1 when the score drops
repomind analyze . --format sarif --output repomind.sarif
repomind analyze . --format html --output report.html
repomind rules                          # list every built-in rule

Adopting RepoMind in a legacy repository: accept today's findings once, then let CI fail only on new problems.

repomind baseline .                    # writes .repomind-baseline.json
repomind analyze . --fail-on-new       # exit code 1 only for new findings

The baseline file is plain JSON with one reviewable entry per accepted finding (rule, path, symbol, severity). Fingerprints ignore line numbers and measured values, so a known issue stays known when it moves or grows.

Useful flags:

Flag Description
-f, --format terminal (default), markdown, json, sarif or html
-o, --output write markdown/JSON/SARIF/HTML to a file instead of stdout
--min-severity hide findings below info, low, medium, high or critical
--top number of findings shown in the report
--history/--no-history force Git history analysis on/off (default: auto)
-x, --exclude glob pattern to skip; repeatable, matches any path segment
--config explicit configuration file
--baseline / --no-baseline use or ignore a baseline file (auto-detected by default)
--since only analyze Python files changed since a Git revision (committed + working tree)
--fail-on-new CI gate: exit code 1 for findings not accepted by the baseline
--fail-under CI gate: exit code 1 below the given health score

Configuration

Drop a repomind.toml at the repository root, or use [tool.repomind] in pyproject.toml:

[tool.repomind]
exclude = ["migrations", "*/generated/*"]
ignore = [
    "maintainability/too-many-parameters@src/app/cli.py",
    "dead-code/unused-import",
]
use_git_history = true
history_commits = 500

[tool.repomind.thresholds]
cyclomatic_warn = 10
cyclomatic_high = 15
cyclomatic_critical = 20
function_length_warn = 60
god_object_methods = 15
hotspot_min_commits = 5

Unknown keys are rejected loudly, so typos never silently change your analysis.

ignore entries silence known, accepted findings: rule-id suppresses a rule everywhere, rule-id@glob only for matching paths. Suppressed findings never disappear silently — every report shows how many were suppressed and --format json lists them in full.

GitHub Action

RepoMind ships a composite action that analyzes the repository, writes a SARIF report and uploads it to GitHub code scanning, so findings appear as alerts on the pull request diff:

name: repomind

on: [push, pull_request]

permissions:
  contents: read
  security-events: write

jobs:
  analyze:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: PeterCode-cmd/RepoMind@master
        with:
          fail-on-new: "true"   # optional: fail only on new findings
          fail-under: "85"      # optional: fail below this health score

Inputs: path, fail-under, fail-on-new, baseline, sarif-file, upload-sarif. The action runs with --no-history so results stay deterministic on shallow checkouts.

How the health score works

  1. Every finding adds penalty points: critical 8, high 4, medium 1.5, low 0.5, info 0.1.
  2. The penalty is normalised per 100 source lines: density = penalty / (loc / 100).
  3. score = 100 - density × 10 (clamped to 0–100), grade A < 1.0, B < 2.0, C < 4.0, D < 7.0, else F.

Size normalisation keeps large legacy repositories from being punished for their sheer size; the score measures density of problems, not their absolute count.

Architecture

repomind/
├── src/repomind/
│   ├── cli/           # Typer + Rich command line interface
│   ├── core/          # discovery, AST parser, complexity, graph, rules, engine
│   │   └── rules/     # one module per rule family, registered in default_rules()
│   ├── git/           # history analysis: churn, authors, hotspots
│   ├── models/        # shared dataclasses: findings, metrics, enums
│   ├── reporters/     # terminal, Markdown and JSON renderers
│   ├── semantic/      # (roadmap) embeddings + local vector search
│   ├── agents/        # (roadmap) LLM agents fed with prepared metrics
│   └── dashboard/     # (roadmap) Streamlit dashboard
├── tests/
│   └── fixtures/sample_project/   # deliberately flawed package used by tests
├── docs/
├── pyproject.toml
└── README.md

Pipeline: discover → parse → graph → history → rules → score → report. Rules never read files: they consume prepared metrics, graphs and history, which keeps them fast, deterministic and easy to test. See docs/architecture.md for details and extension guides.

Dogfooding

RepoMind analyzes itself in CI:

repomind analyze . --no-history --fail-under 95

The repository ships a repomind.toml with two deliberate decisions:

  • exclude = ["tests/fixtures"] — the sample project is intentionally broken (that is its job), so it must not count towards RepoMind's own health.
  • two ignore entries for src/repomind/cli/analyze.py — a Typer command is a declaration surface (one parameter per flag, rich help text, no logic), so parameter count and function length describe the CLI framework, not the code.

Everything else is clean: 100/100 (grade A), zero findings, zero import cycles. The score only became meaningful after the exclusions above; before them, RepoMind was grading its own test fixtures.

Benchmarks

First-run results on popular open-source projects (default thresholds, no configuration, 500-commit window):

Project Files Source lines Score Findings Wall time
psf/requests 37 8,029 66/100 C 120 18.6 s
pallets/flask 83 10,740 72/100 C 129 19.7 s
tqdm/tqdm 65 5,901 37/100 D 144 18.9 s
psf/black 351 118,690 86/100 B 688 22.2 s

The full write-up — methodology, notable findings, and why excluding tests can lower the score — is in docs/benchmarks.md.

Roadmap

  • MVP: CLI, static analysis, dependency graph, terminal + Markdown reports
  • Git history analysis with complexity × churn hotspots
  • JSON output and --fail-under CI gate
  • Configurable suppression (ignore) and decorator-aware dead-code detection
  • Baseline + --fail-on-new for incremental adoption in legacy repositories
  • Diff mode (--since) for pull-request reviews
  • SARIF output and a GitHub Action
  • Self-contained HTML report
  • Benchmarks against popular open-source projects (docs/benchmarks.md)
  • Semantic layer: local embeddings + natural-language questions about the code
  • Agent layer: Architect / Quality / Security / Maintainability / Documentation
  • Streamlit dashboard
  • GitHub URL input (clone + analyze) and PR comment bot
  • Additional languages (tree-sitter based)

Development

uv pip install -e ".[dev]"

ruff format .          # formatting
ruff check .           # linting
mypy                   # strict type checking
pytest --cov           # tests + coverage
pre-commit install     # optional: run checks on every commit

The test suite includes a deliberately flawed sample project (tests/fixtures/sample_project) that exercises every built-in rule, plus temporary Git repositories for history analysis.

Contributing

Contributions are welcome — see CONTRIBUTING.md. Good first issues: new rules, new reporters, language support, and the roadmap items above.

License

MIT — see LICENSE.

Release files for repomind-analyzer 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for repomind-analyzer 0.1.0
File Size Uploaded
repomind_analyzer-0.1.0.tar.gz 70.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for repomind-analyzer 0.1.0
File Interpreter ABI Platform
repomind_analyzer-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 138.4 kB

Release files / repomind_analyzer-0.1.0.tar.gz

Download URL repomind_analyzer-0.1.0.tar.gz
Size 70.3 kB
Tags Source
SHA-256 checksum
How to use checksums
1dc2b0a9ec0f633c3f0d1b2bd208a38d69300a9d319fa205b815466558ef2b1a
BLAKE2b-256 checksum
How to use checksums
817060e0470f321abf19bb02c3088d4da2114085c9d53e57510cf857358a9a21
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.

Transparency log

Release files / repomind_analyzer-0.1.0-py3-none-any.whl

Download URL repomind_analyzer-0.1.0-py3-none-any.whl
Size 68.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5eb2ffa004576d0694d9ff5787c21b73158567b1c7914a29e08b3f44ade999ae
BLAKE2b-256 checksum
How to use checksums
9361a4675b413bb32e0c8df95f3403f70928dc50c27f5a4d9e186d903a123c10
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.0

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page