Maintainability Agent
A deterministic CI gate + bounded remediation prompt generator for repos that use AI coding agents (Claude, Codex, Cursor, Copilot, Windsurf, …) — ships an invokable skill for Codex, Claude Code, and GitHub Copilot Chat so /maintainability-agent is one keystroke away in any of them.
pip install maintainability-agent # CLI + library
cp -r skills/maintainability-agent ~/.claude/skills/ # Claude Code skill
cp skills/maintainability-agent/copilot/maintainability-agent.prompt.md .github/prompts/ # Copilot Chat (VS Code)
# Codex picks up skills/maintainability-agent/ via its own convention.
Jump to Invokable Skill / Slash Command for the full install table.
v0.6.0 detects a helper written twice under two names — clone-instead-of-reuse, the most-cited complaint about AI-written code. Renamed copies defeat text matching, so bodies are compared structurally with identifiers anonymized. Each finding names the declaration to reuse, so the prompt says
toAtomicAmountatTradeTicket.tsx:862already does this rather than "there is duplication". Measured across the reference corpus, this is the first signal that separates AI-written applications from mature human-written OSS (median 1.49% vs 0.20% of production declarations). See docs/standard.md.v0.5.0 rebuilt the scoring engine. The old model counted findings absolutely, so it scored repo size rather than maintainability: it graded Django, pytest, black, tornado, click, httpx, attrs, lodash, svelte, axios and fastapi all at 0.0 / F while a 53-file toy repo scored 4.6 / A. Scores are now rates, normalized per dimension against what real code carries, and calibrated so the corpus median earns a B. See docs/standard.md.
Why this exists
AI-written code fails in recognizable ways: speculative refactors, duplicated helpers, broad rewrites for narrow bugs, stale comments that sound confident, tests that assert implementation details instead of behavior, architecture drift across modules. SonarQube / CodeClimate / Qlty / ESLint / Ruff / Radon all catch some of this. None of them ship a bounded prompt back to the agent that says "fix only these specific findings, do not refactor outside this scope."
That's the point of this tool:
- Run a deterministic local audit — file size, function size, approximate cyclomatic complexity, duplication, configurable risk patterns, ISO/IEC 25010-inspired 0–5 score.
- Emit Markdown, JSON, SARIF, a PR comment, and a baseline for incremental adoption.
- Generate an AI remediation prompt scoped to the actual findings — bounded, with explicit "don't rewrite the codebase" rules.
- Hand that prompt to your agent. Get a small, reviewable fix instead of a 600-line speculative cleanup PR.
- Drop the shipped portable invokable skill into Codex, Claude Code, or GitHub Copilot Chat so
/maintainability-agentis one keystroke away in any of them. See Invokable Skill below.
The remediation prompt is the differentiator. Every other tool in this space stops at "here's a list of findings."
Who it's for
- Teams running AI agents in the dev loop who are tired of unbounded agent rewrites and want a CI gate that actively constrains follow-up scope.
- Repos that want a maintainability gate without paying for SonarQube / CodeClimate / Qlty or sending code to a third party.
- Solo devs who want a single-binary deterministic audit they can pin in a Makefile, a pre-commit, or a local CI script.
Design principles
- Deterministic first, AI optional. The audit never calls an LLM by default. The remediation prompt is a generated artifact that you choose to hand to an agent.
- Bounded scope. The remediation prompt explicitly tells the agent to fix the listed findings only — not to embark on architecture cleanup.
- No vendor lock-in. All outputs (Markdown, JSON, SARIF, PR comment) are plain files. Pair this tool with mature analyzers (ESLint, Ruff, Radon, Semgrep, SonarQube, Qlty/Code Climate) — don't replace them.
- Pass-the-cost-of-disclosure. A finding that's "just a warning" never blocks CI alone. Hard gates are configurable + opt-in.
See docs/philosophy.md for the longer version.
Self-Audit
This repo eats its own dogfood — the tool is run against this codebase as part of CI, and the latest report is checked in at docs/self-audit.md:
| Metric | Value |
|---|---|
| Overall score | 5.0 / 5 (A+) |
| Files scanned | 74 |
| File warnings | 0 |
| File failures | 0 |
| Function warnings | 0 |
| Function failures | 0 |
| Duplicate blocks | 0 |
| Risk findings | 0 |
| Hard gate failures | 0 |
All five ISO/IEC 25010 categories (modularity, reusability, analyzability, modifiability, testability) score 5.0, against thresholds this repo deliberately sets stricter than the shipped defaults — a 250-line file warning versus the default 400.
That grade is maintained, not assumed, and since v0.5.0 it is also gated: an A+ requires every dimension to be clean, so it cannot be reached by averaging one bad dimension against four good ones. When the v0.4.0 work pushed this repo to 4.4 / 5 (B), the response was to split metrics.py along its real responsibilities rather than to publish a fix while advertising a stale A+. The score is the output of that work, not a claim that preceded it.
Regenerate with maintainability-agent --config maintainability-agent.json --output docs/self-audit.md (see the file's preamble for the path-sanitization step).
Install
python3 -m pip install maintainability-agent
maintainability-agent --root . --config maintainability-agent.json
Or run from a source checkout without installing:
python3 -m maintainability_audit --root . --config maintainability-agent.json
For an editable dev install, see CONTRIBUTING.md.
Quick Start
Copy the example config to your repo root as maintainability-agent.json:
cp maintainability-audit.example.json maintainability-agent.json
Run:
maintainability-agent \
--config maintainability-agent.json \
--format markdown \
--output maintainability-report.md
Fail CI on hard gates:
maintainability-agent \
--config maintainability-agent.json \
--fail-on-gate \
--output maintainability-report.md \
--prompt-output maintainability-remediation-prompt.md \
--comment-output maintainability-pr-comment.md
What It Analyzes
The deterministic scanner reads code from your repo (no LLM calls) and produces signals on:
- largest files (warn / fail thresholds configurable per-repo)
- function size and complexity — exact ranges for Python via
ast, brace-bounded for JS/TS/JSX/TSX/HTML - class size, against its own separate budget (
max_class_lines), on length alone - approximate cyclomatic complexity, plus cognitive complexity — nesting-weighted reading cost, so five guard clauses no longer score the same as five levels of nesting
- duplicate blocks (≥ N consecutive non-trivial lines, configurable)
- near-duplicate declarations — the same helper written twice under different names, compared structurally so renaming can't hide it, each paired with the original to reuse
- unreferenced private declarations — debris nothing in the repo can reach, scoped to internal names so a library's public surface is never flagged
- competing libraries for one concern — two HTTP clients or two schema validators mean two mental models; curated list, extensible via
idiom_groups - configurable risk patterns (regex matchers — TODO/FIXME,
eval(,exec(, custom) - expected files present (README, LICENSE, etc. — opt-in hard gate)
- expected test/lint commands declared in the config (opt-in hard gate)
- worktree-clean state at audit time (opt-in hard gate)
- ISO/IEC 25010-inspired 0–5 score per category + overall letter grade
The analyzer is intentionally conservative and dependency-free. It is built to under-report rather than over-report: a declaration it can't recognize costs one missed finding, never a cascade of false ones. Per-language accuracy, the known limitations, and why classes are graded separately are documented in docs/language-support.md.
Mature repos should pair this with native tools (ESLint, Ruff, Radon, Semgrep, SonarQube, Qlty / Code Climate) — not replace them. SARIF input from those tools can be folded into this tool's report via --sarif-input.
What It Produces
Each run can emit any combination of:
maintainability-report.md— the full Markdown report with summary, score, hotspots, duplicates, risk findings, external (SARIF) findings.maintainability-remediation-prompt.md— bounded AI prompt scoped to the run's findings.maintainability-pr-comment.md— short body suitable for agh pr commentpost.maintainability.sarif— SARIF 2.1.0 output for GitHub Code Scanning ingestion.maintainability-baseline.json— fingerprints of current findings, for--fail-on-newincremental adoption.- Per-tool agent instruction files (
AGENTS.md,CLAUDE.md,.cursor/rules/maintainability.mdc,.github/copilot-instructions.md,.windsurf/rules/maintainability.md,AI-MAINTAINABILITY.md) via--init-agent-standards.
AI Remediation Prompt
The runner can generate a bounded prompt for a human developer to give to Claude, Codex, or another coding assistant:
maintainability-agent \
--config maintainability-agent.json \
--output maintainability-report.md \
--prompt-output maintainability-remediation-prompt.md
The prompt is designed for AI-written or AI-assisted code reviews. It tells the assistant to:
- fix only the highest-value maintainability issues
- keep the patch small and reviewable
- preserve existing architecture and behavior
- add tests where behavior changes
- report false positives instead of rewriting blindly
This makes the CI artifact actionable without letting the audit turn into an unbounded refactor request.
PR and Baseline Workflows
PR-only audits, baseline grandfathering, AI-agent instruction generation, and reusable agent-standard file generation are covered in PR and Baseline Workflows.
Running Tests
PYTHONPATH=src python3 -m pytest
The full local verification sequence that matches CI — ruff, pip-audit, the 92% coverage gate, and the self-audit — is in CONTRIBUTING.md, along with the sandbox-friendly invocation for agents that disable plugin autoload.
Scoring Standard
The audit model is based on ISO/IEC 25010 maintainability — modularity, reusability, analyzability, modifiability, testability.
Scores are rates calibrated against real code, not counts. Every pressure is normalized against the median that a pinned 14-repo corpus of mature open-source projects (django, pytest, black, svelte, axios, requests, …) actually exhibits, so 2.5x means "two and a half times what well-maintained real code carries." The corpus median earns a B; A+ is gated, requiring every dimension clean rather than a good average.
The calibration is reproducible rather than asserted: python3 tools/calibration/measure.py --check re-measures the corpus and fails if the shipped constants have drifted, and tests/test_calibration_corpus.py re-derives them offline from checked-in measurements — no network, no trust required.
See docs/standard.md.
Documentation
- CLI reference
- Config schema
- Language support and detection accuracy
- Philosophy
- Changelog
- Analyzer adapters
- External quality tools
- IDE and agent integration
- PR and baseline workflows
- Roadmap
GitHub Action
This repo includes action.yml, so it can be used as a composite action:
- uses: marshallguillory86/maintainability-agent@v0.6.1
with:
config: maintainability-agent.json
changed-only: main...HEAD
Or copy .github/workflows/maintainability.yml into the target repo and adapt it.
IDE and Agent Integration
See docs/ide-agent-integration.md for VS Code tasks and integration notes for Copilot, Cursor, Codex, Claude Code, Windsurf, generic agents, local CI, and GitHub Actions.
Invokable Skill / Slash Command
For agents that support invokable skills, this repo ships a portable skill under skills/maintainability-agent/. The SKILL.md body is the source of truth; per-host adapters live under agents/ and copilot/.
| Host | Install destination | Invocation |
|---|---|---|
| Codex / OpenAI | wired via skills/maintainability-agent/agents/openai.yaml |
per Codex's skills convention |
| Claude Code | copy skills/maintainability-agent/ → ~/.claude/skills/maintainability-agent/ (user-scope) or <repo>/.claude/skills/maintainability-agent/ (project-scope) |
/maintainability-agent (or surfaced automatically when description matches) |
| GitHub Copilot (VS Code) | copy skills/maintainability-agent/copilot/maintainability-agent.prompt.md → <repo>/.github/prompts/maintainability-agent.prompt.md |
/maintainability-agent in Copilot Chat |
The copy commands are in the intro block at the top of this file. For non-invokable, always-on guidance, use --init-agent-standards (see docs/ide-agent-integration.md).
Local CI
For repos that do not use GitHub Actions, use:
examples/local-ci.sh
The local CI script enforces test coverage at >=92% and writes coverage.xml for SonarQube Cloud, Qlty, Codacy, or any other tool that can ingest Python coverage.
Get in Touch
- Bug reports / feature requests / general questions — open a GitHub Issue.
- Discussion / ideas — use the Discussions tab (enable in repo settings if not visible yet).
- Security vulnerabilities — see
SECURITY.mdand use the private security advisory flow. Do not post vulnerabilities in public issues.
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file maintainability_agent-0.6.1.tar.gz.
File metadata
- Download URL: maintainability_agent-0.6.1.tar.gz
- Upload date:
- Size: 76.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e072fecf7e69a353836164acc0070d5deedb7c363b695525bcbcc6d0975d5c42
|
|
| MD5 |
273cf3c03d23aef3ae3818d6f98823ba
|
|
| BLAKE2b-256 |
6dbb19a5814d1923c3b90f168607a5c7ebeff8853a1d290fd168a5ef85d05be8
|
Provenance
The following attestation bundles were made for maintainability_agent-0.6.1.tar.gz:
Publisher:
release.yml on marshallguillory86/maintainability-agent
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
maintainability_agent-0.6.1.tar.gz -
Subject digest:
e072fecf7e69a353836164acc0070d5deedb7c363b695525bcbcc6d0975d5c42 - Sigstore transparency entry: 2386514084
- Sigstore integration time:
-
Permalink:
marshallguillory86/maintainability-agent@751c623b56f9354b912cbc2973734db24c455905 -
Branch / Tag:
refs/tags/v0.6.1 - Owner: https://github.com/marshallguillory86
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@751c623b56f9354b912cbc2973734db24c455905 -
Trigger Event:
push
-
Statement type:
File details
Details for the file maintainability_agent-0.6.1-py3-none-any.whl.
File metadata
- Download URL: maintainability_agent-0.6.1-py3-none-any.whl
- Upload date:
- Size: 65.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e09fda5b37a2342b9725b7c249a519e9411ffad72a1845493915444de95baa2c
|
|
| MD5 |
9bf3cb29fdd9be385a135ce35a8d8a51
|
|
| BLAKE2b-256 |
625ff5954cf4ecc3b74d6d58d822030afa5c8f37cd0367c9640125b551ad6912
|
Provenance
The following attestation bundles were made for maintainability_agent-0.6.1-py3-none-any.whl:
Publisher:
release.yml on marshallguillory86/maintainability-agent
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
maintainability_agent-0.6.1-py3-none-any.whl -
Subject digest:
e09fda5b37a2342b9725b7c249a519e9411ffad72a1845493915444de95baa2c - Sigstore transparency entry: 2386514088
- Sigstore integration time:
-
Permalink:
marshallguillory86/maintainability-agent@751c623b56f9354b912cbc2973734db24c455905 -
Branch / Tag:
refs/tags/v0.6.1 - Owner: https://github.com/marshallguillory86
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@751c623b56f9354b912cbc2973734db24c455905 -
Trigger Event:
push
-
Statement type: