Maintainability Agent
A deterministic, offline maintainability audit whose output is a bounded work order for an AI coding agent — a copy-paste prompt, per finding, that says fix exactly these and refactor nothing else. Chat-primary; CLI for CI. Version 2.9.0.
Languages parsed: Python, Java, C, C++, C#, Swift, Fortran (free-form and fixed-form), and the JS/TS family — each by a scanner written for it, and measured with that language's own reading of what a branch is. What that means per language.
pip install maintainability-agent # CLI + library
pip install "maintainability-agent[mcp]" # optional local MCP server for chat / IDE hosts
cp -r skills/maintainability-agent ~/.claude/skills/ # Claude Code slash command
Executive summary
Agents write code faster than anyone can review it. Not measurably worse code — this project tested that claim about itself and retracted it — just more, arriving faster than trust can accumulate. Linters and quality dashboards catch some of the resulting slop. None of them ship the one thing that actually closes the loop: a bounded prompt back to the agent, scoped to the findings, that forbids unbounded rewrites.
That prompt is the product. Everything else — the scanner, the ISO/IEC 25010-inspired 0–5 score, the analyzer pool, the semantic and economic signals — exists to aim it. Remove the prompt and what's left is a worse version of tools that already ship.
What 1.0 guarantees:
- Deterministic and local. Same tree, config, pinned analyzer versions and history in → same evidence, findings and score out. The analysis performs no network access and invokes no LLM, and this agent does not transmit your source. Every report states what it examined and produced — files, declarations, findings, bounded work items — all computed with no model call: work a metered agent never has to do. Third-party analyzers it may spawn (eslint, jscpd, lizard) are not network-sandboxed — this process does not police whether they phone home; install them for an air-gapped run.
- One uniform rubric, readable in source, applied to every repository — so "better" and "worse" are not an argument. Calibrated against a query-selected corpus of mature open-source projects; the corpus median earns a B, and A+ is gated, not averaged.
- Honest about evidence. A score is withheld when too little was examined
to support one — a
--changed-onlydiff is not a repository grade, and a shallow clone is not an A. Every reported value names what measured it. - The complete work order is the report; chat/CLI is a bounded UI. The HTML/Markdown report carries the entire backlog with a deterministic copy-paste prompt for each item. The chat surface stays a tight summary so a host's payload cap can never truncate the prompt.
- One setup, three transports. Chat, MCP, and an interactive CLI TTY ask the same first-run questions. A surface that asks a subset is a bug.
Governing intent lives in docs/product-intent.md — that document is authoritative, and this README defers to it wherever the two differ.
Why this exists
The ratio of code-written to code-reviewed has collapsed. Unmaintainable code that used to accumulate over years can now accumulate in an afternoon: duplicated helpers, oversized files, speculative abstractions — the same slop hand-written codebases always accrued, now at machine speed.
The same speed is the way out. An agent pointed at specific, deterministic findings can fix them at the rate they appear. The loop this tool closes:
- Measure pressure points deterministically, with no LLM involved.
- Score with one uniform standard, so the verdict is not a debate.
- Emit a prompt scoped to those findings only, with explicit instructions not to refactor beyond them.
- Hand it to the agent. Review a scoped diff instead of a speculative rewrite.
Step 3 is the product; steps 1 and 2 are in service of it. Every other tool in this space stops at "here's a list of findings."
Who does the checking matters as much as what it checks. An author is never
the independent check on their own work, so a platform that generates code and
grades it is producing a self-assessment — a property of the arrangement, not a
criticism of any one of them. This writes nothing and runs no model, which is
what lets its verdict be evidence (the principle).
Since 2.1.0 the work order is also checked: --conformance compares the
returned diff against the paths the report named, --fail-on-regression
ratchets the dimension scores, and --attestation-output composes them into one
record (how that hole was closed).
The limit, stated in the same breath: those checks read the diff's shape,
never its correctness — whether the change works is not a claim this tool makes.
One pre-registered experiment has tested the bounded prompt. Generic prompting made 2 of 6 repositories worse; bounded prompting made 1 of 6 worse and improved 5 of 6, under this tool's own finding count. The registered hypothesis was narrower diffs, which did not hold, so the registered verdict stands at INCONCLUSIVE. Method, limits and raw data: docs/studies.md.
Who it's for
- Teams running AI agents in the dev loop, tired of unbounded cleanup PRs, who want a CI gate that actively constrains follow-up scope.
- Repos that want a maintainability gate without a SaaS analyzer or shipping code to a third party.
- Solo devs who want a single deterministic audit to pin in a Makefile, pre-commit, or local CI script.
The road to 1.0
1.0 was the line drawn under a long arc of subtraction — what remained after every claim the project could not stand behind was removed. A scoring engine rebuilt after it was found to grade repo size; a headline claim about AI authorship retracted against a matched control; a score that is withheld rather than guessed. That arc is the credibility, and it is recorded in full, with the figures quoted from their approved summaries, in the track record.
Install
python3 -m pip install maintainability-agent
maintainability-agent --root . --config maintainability-agent.json
Or run from a source checkout without installing:
python3 -m maintainability_audit --root . --config maintainability-agent.json
For an editable dev install and the full local-verification sequence, see CONTRIBUTING.md. Upgrading from 0.x? The scale and the evidence model changed on the way to 1.0 — see docs/migration-1.0.md.
Primary Surface: Chat / MCP
Drive the local MCP process from an IDE assistant or chat host. Call
audit_repository; unset action never audits. An unconfigured repository
(no repository maintainability-agent.json and no user config) returns
setup_needed and audit_ran: false — structured setup choices for analyzer
policy, history consent, economics, test-suite execution, and presentation, and
no report. A configured repository returns choice_needed (run or
reconfigure), also without a report. Answering setup does not start an audit.
action="run" returns the report and its bounded remediation prompt to the
conversation. record_history=None follows the persisted first-run consent and
always appends to an existing history; an explicit true or false wins.
python3 -m pip install "maintainability-agent[mcp]"
maintainability-agent mcp --allow-root /absolute/path/to/repository
Presentation is exactly three choices — chat, a Markdown file, or a single-file HTML report. Where to save is asked only after a file format is chosen, and no report file is written without that choice. See chat workflow help and IDE and agent integration.
Automation / CI: CLI
Use the CLI for scripts, repeatable automation, and CI gates. Copy the example
config to your repo root as maintainability-agent.json, then:
maintainability-agent \
--config maintainability-agent.json \
--fail-on-gate \
--output maintainability-report.md \
--prompt-output maintainability-remediation-prompt.md \
--comment-output maintainability-pr-comment.md \
--sarif-output maintainability.sarif
--fail-on-gate fails CI on hard gates only (a missing README, an undocumented
test command, a breached threshold) — never on a letter grade, and never on a
withheld estimate. For PR work, add --changed-only main...HEAD to audit the
diff; on a change too small to support a rate, the estimate and grade are
withheld and the scope is named as the reason.
Checking that the agent did the work
The work order tells an agent to fix exactly these findings and refactor nothing else. These flags check that the diff obeyed it — they read its shape, not its correctness:
maintainability-agent \
--config maintainability-agent.json \
--conformance main...HEAD \
--fail-on-out-of-scope \
--fail-on-regression \
--attestation-output maintainability-attestation.md
--conformance <revspec>compares the diff against the paths the report named, and answers two questions separately: did it stay in scope, and did it silence nothing. They stay separate because a change can obey the work order and still add a# noqato a finding inside it. A test added for a fix stays in scope even though the work order never named the test file.--fail-on-out-of-scopeturns that into a CI failure.--fail-on-regressionratchets the dimension scores against scan history, catching a change that improves one dimension while silently regressing another. It has three outcomes, not two — held, regressed, and not comparable, because two scans taken under different calibration cannot be differenced.--attestation-outputwrites the three into one per-change record: what was measured, what the agent was told to change, whether it stayed inside the work order, and what moved. It is reproducible and not signed — nothing here holds a key, and the document says so in its own text. A check nobody ran renders as not asked, never as passed.--transformation NAMEnames the class of work a scan followed and reports how this run of it compares with earlier ones. A report, never a gate. It says one run "moved further" than another, never that it was better — two runs of one codemod land on different code — and it measures the interval, not the transformation.
The roadmap has the history of why this was the hole under the product's central claim.
Before the commit
The gate above runs after the work is done. --staged runs while it is still in
the author's hands:
maintainability-agent --install-precommit-hook # writes .git/hooks/pre-commit
It scans the index, not the working tree — stage half a file with git add -p and keep typing, and what gets measured is what the commit will actually
contain. Clean, it prints nothing and exits 0. Breaching, it names each path,
line and remedy and exits 1; --format json gives an agent the same findings.
It is not a small audit. It produces no score, because a diff has no
population to draw a rate from — only threshold breaches, including a # noqa
this change added. It applies no repository gates: a missing README is a
property of the repository, not of your diff. And it runs nothing and writes
nothing, so it costs 0.17s here against the full audit's 266 — a hook slower
than the author's patience is one they uninstall.
--install-precommit-hook refuses to replace a hook it did not write, printing
the line to add to yours instead, and honours core.hooksPath.
What it analyzes
The deterministic scanner reads code from your repo (no LLM calls) and produces signals on:
- largest files (configurable warn/fail thresholds)
- function size and complexity — exact ranges for Python via
ast, brace-bounded for JS/TS/JSX/TSX/HTML — plus cognitive complexity (nesting-weighted reading cost) - class size, against its own budget (
max_class_lines) - duplicate blocks, and near-duplicate declarations compared structurally so renaming can't hide a copy, each paired with the original to reuse
- unreferenced private declarations (debris nothing can reach)
- competing libraries for one concern (two HTTP clients, two validators)
- configurable risk patterns (
eval(,exec(, TODO/FIXME, custom regex) - expected files / commands / clean-worktree (opt-in hard gates)
- TypeScript semantic facts from a recorded analysis or an
already-installed
tsc, including workspace projects (ADR 003) - test effectiveness from an opted-in suite run and parsed coverage (Class 5, default off — the one place the agent may execute the tree)
- an ISO/IEC 25010-inspired 0–5 estimate per category, and a verified grade — or a disclosed withholding when the evidence is thin.
The analyzer is intentionally conservative and under-reports rather than
over-reports: an unrecognized declaration costs one missed finding, never a
cascade of false ones. Pair it with native tools (ESLint, Ruff, Radon, Semgrep,
SonarQube, Qlty) rather than replacing them — their SARIF folds in via
--sarif-input. Accuracy and limits: docs/language-support.md.
Language support
Be clear-eyed: the tool does not support every language equally, and on an unrecognized language it under-reports rather than fails. Coverage has two layers.
Ten languages are parsed as of 2.7.0: Python (1.0), Java (1.0), C (1.1), C++ (1.2), C# (1.3), Fortran (free-form 1.4, fixed-form 1.6), Swift (2.4), COBOL (2.7), the JS/TS family, and HTML. Each has a scanner written for it and a documented list of what it misses — a language is claimed only when both exist.
Built-in scanner (always on, no dependencies) — reads function/class declarations, sizes and complexity for a fixed set of languages, and only these:
| Language | How it's measured | Fidelity |
|---|---|---|
Python (.py) |
ast — exact end_lineno |
Exact |
Java (.java) |
dedicated brace-bounded scanner | Bounded; under-reports some constructs |
C (.c, .h) |
dedicated brace-bounded scanner — functions, struct/enum/union |
Bounded; prototypes and macros are not declarations |
C++ (.cpp, .hpp, .cc, .cxx, .hh) |
dedicated brace-bounded scanner — functions, class members, namespaces, templates | Bounded; bodyless declarations are not definitions |
C# (.cs) |
dedicated brace-bounded scanner — methods, constructors, class/interface/struct/record/enum |
Bounded; properties are not declarations |
Swift (.swift) |
dedicated brace-bounded scanner — functions, initialisers, subscripts, class/struct/enum/protocol/actor |
Bounded; extension members carry their type, protocol requirements and computed properties are not declarations |
COBOL (.cbl, .cob, .cpy, and .CBL/.COB/.CPY) |
dedicated scanner — PROCEDURE DIVISION paragraphs, bounded by the start of whatever follows; fixed-form card columns read where the layout carries them | Bounded by the next header; level numbers and container programs/sections are not declarations |
Fortran, free-form (.f90, .f95, .f03, .f08, .F90, .F95, .F03, .F08, .pf) |
dedicated keyword-bounded scanner — modules, subroutines, functions, derived types | Bounded by end; measured with Fortran's own branch and nesting reading |
Fortran, fixed-form (.f, .for, .ftn, .F, .FOR, .FTN) |
the same scanner over card-column source; continuations joined, labelled DO loops understood |
Bounded by end or by the loop's label |
JS / TS / JSX / TSX (.js, .jsx, .mjs, .cjs, .ts, .tsx) |
brace/paren depth over a masked copy | Bounded by the declaration's own braces |
HTML (.html) |
same brace scanner (inline <script>) |
Bounded |
| TypeScript (semantic) | a recorded analysis or a locally-installed tsc, workspace projects included |
Type-level facts; unknown when no checker is present |
Any language not in that table — Go, Rust, Ruby, PHP, Kotlin, and the rest — is not parsed for declarations by the built-in scanner. Its files still count toward repo size, but the built-ins produce no function-size, complexity, duplication or dead-code findings for them, and the estimate leans on whatever evidence is available — which is why the report discloses its evidence tier and can withhold the grade.
So: first-class today is Python (and TypeScript for semantics); every other parsed language is bounded-but-real; anything outside the table is only as covered as the analyzer you point at it. Per-language accuracy and limits: docs/language-support.md. Which analyzer covers which language, how the opt-in pool extends coverage past the built-in set, and why COBOL's external tier is empty: docs/adapters.md.
What it produces
Any combination of: maintainability-report.md (or a single-file HTML report),
maintainability-remediation-prompt.md (the bounded prompt),
maintainability-pr-comment.md, maintainability.sarif (2.1.0, for GitHub
Code Scanning), maintainability-baseline.json (for --fail-on-new
incremental adoption), and per-tool agent instruction files
(AGENTS.md, CLAUDE.md, Cursor/Copilot/Windsurf rules) via
--init-agent-standards.
The report is the complete work order. The HTML and Markdown reports carry the whole backlog — every finding with a self-contained, deterministic copy-paste prompt telling the agent to fix only the listed items, keep the patch small, preserve architecture and behavior, add tests where behavior changes, and report false positives instead of rewriting blindly. Chat and CLI render a bounded view — a summary plus the top items and a pointer to the report — so a payload cap can never truncate the prompt.
Scoring standard
Based on ISO/IEC 25010 maintainability — modularity, reusability,
analyzability, modifiability, testability. Scores are rates calibrated
against real code, not counts: every pressure is normalized against the
median a pinned 112-repo corpus of mature projects actually carries —
eight of the ten parsed languages (django, angular, spring-framework,
ghidra, LAPACK, …); Swift and COBOL are parsed but unanchored. How that corpus was selected, what each language actually
reads, and what the numbers do not establish:
the calibration study, so 2.5x means
"two and a half times what well-maintained real code shows." The corpus median
earns a B; A+ is gated, requiring every dimension clean.
The corpus is selected by query, not taste — stars:>3000 created:<2021-01-01 pushed:>2026-01-01 across Python, TypeScript and
JavaScript, then filtered to repositories that contain code. The calibration is
reproducible, not asserted: python3 tools/calibration/measure.py --check
re-measures and fails on drift, and tests/test_calibration_corpus.py
re-derives the constants offline from checked-in measurements. See
docs/standard.md.
Self-audit
This repo eats its own dogfood — the tool runs against this codebase in CI, and a report is checked in at docs/self-audit.md, stamped with the exact source commit it was generated against (a provenance record, not a claim about HEAD). These figures mirror that stamped report row for row:
| Metric | Value |
|---|---|
| Maintainability estimate | 4.1 / 5 |
| Verified grade | B |
| Files scanned | 391 |
| File warnings | 120 |
| File failures | 0 |
| Function warnings | 65 |
| Function failures | 0 |
| Duplicate blocks | 0 |
| Risk findings | 0 |
| Hard gate failures | 0 |
Yes, a B — demoted from the A band because warning rates exceed the
A-grade ceilings, against thresholds this repo sets stricter than the shipped
defaults. The grade is gated (A+ needs every dimension clean; a
repo with production code and zero test files cannot earn an A-grade) and banded
from the evidence floor, so withholding evidence never buys a better letter.
Every threshold gate — file, function, duplication — is opted on for this
repo's own CI, so drifting below the bar fails the build rather than the README.
(CI note: actions/checkout defaults to fetch-depth: 1, which hides history
and costs roughly a grade — use fetch-depth: 0.)
Platform support
POSIX. Windows is not claimed, for a measured reason. A windows-latest probe
found most failures in two POSIX-only calls, os.fchmod and os.O_DIRECTORY. The
bounded, symlink-refusing write works through a file descriptor, so a validated
path cannot be swapped for a symlink before the write lands — Windows has no
equivalent and the portable rewrite is the hole. An unsupported platform never
buys green by weakening the supported ones.
Invokable skill / slash command
This repo ships a portable skill under
skills/maintainability-agent/, so
/maintainability-agent is one keystroke away in Claude Code, Codex and
Copilot Chat:
maintainability-agent --install-skill # writes ~/.claude/skills
Re-run it after every upgrade — a drifted skill teaches agents a dead workflow,
so a differing installed copy is refused with the list of differences. Per-host
install destinations, and --init-agent-standards for always-on guidance
instead of an invokable skill:
docs/ide-agent-integration.md.
GitHub Action
This repo ships action.yml, usable as a composite action:
- uses: marshallguillory86/maintainability-agent@v1.0.0
with:
config: maintainability-agent.json
changed-only: main...HEAD
fail-on-gate: "true"
Or copy .github/workflows/maintainability.yml into the target repo. For repos
not on GitHub Actions, examples/local-ci.sh enforces coverage and writes
coverage.xml.
Documentation
Start with the documentation index, which states each document's genre and what it is allowed to assert.
Governing — Product intent (authoritative: promises, and what it must never claim) · Architecture (layers, enforced invariants, known debt) · Philosophy (why AI-specific: volume, not pathology) · Decision register including ADR 001 (evidence and verification).
Reference — Maintainability standard · Studies and measured results · Report contract · CLI reference · Config schema · Language support · Analyzer adapters · IDE and agent integration · Chat workflow help · PR and baseline workflows · Roadmap · Changelog.
Running tests
PYTHONPATH=src python3 -m pytest
The full local verification that matches CI — ruff, pip-audit, the 92% coverage gate, and the self-audit — is in CONTRIBUTING.md.
Get in touch
- Bugs / features / questions — open a GitHub Issue.
- Discussion — the Discussions tab.
- Security — see
SECURITY.mdand the private advisory flow. Do not post vulnerabilities in public issues.
Support this work
This is a single-maintainer, MIT-licensed project — free to use, and built on a lot of unpaid hours. If it saves you or your team time, please consider sponsoring its continued development:
Sponsorship is optional and never gates a feature, a fix, or support — the whole tool stays free and open. It just helps keep the work going.
Acknowledgements
- Miles Parker — identified the friction-signal gap: that the "this keeps fighting me" signal a maintainer accumulates over months is exactly the evidence an LLM cannot hold across sessions, and that a tool positioned between the two should carry it. A contribution of insight rather than code, and it changed what this project is for.
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file maintainability_agent-2.9.0.tar.gz.
File metadata
- Download URL: maintainability_agent-2.9.0.tar.gz
- Upload date:
- Size: 995.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
23ab2ef7b04e4184b8da7f9b1715936e732859c1c6347a8dd6983aa3b5985667
|
|
| MD5 |
1d85c8d16eb466de6ab25b205fc868b3
|
|
| BLAKE2b-256 |
93f4c8a1af48132dcc9304effef0bf63427ba7aa74e7d0d517bac852edf2ed60
|
Provenance
The following attestation bundles were made for maintainability_agent-2.9.0.tar.gz:
Publisher:
release.yml on marshallguillory86/maintainability-agent
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
maintainability_agent-2.9.0.tar.gz -
Subject digest:
23ab2ef7b04e4184b8da7f9b1715936e732859c1c6347a8dd6983aa3b5985667 - Sigstore transparency entry: 2733471364
- Sigstore integration time:
-
Permalink:
marshallguillory86/maintainability-agent@7ce3b4468315728f5d9be1b5ae0e32b8a69f2653 -
Branch / Tag:
refs/tags/v2.9.0 - Owner: https://github.com/marshallguillory86
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@7ce3b4468315728f5d9be1b5ae0e32b8a69f2653 -
Trigger Event:
push
-
Statement type:
File details
Details for the file maintainability_agent-2.9.0-py3-none-any.whl.
File metadata
- Download URL: maintainability_agent-2.9.0-py3-none-any.whl
- Upload date:
- Size: 528.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3d8594567a7b2a0aee61c00a11c706a425a4e4c7635c4249e664a33634971dad
|
|
| MD5 |
8c2a775e58f32c0ff5b667b2b4a59d22
|
|
| BLAKE2b-256 |
bc073fafe6c4b28b1407679a387c8ff2c52d6ad829198936b48d1b2bced012d3
|
Provenance
The following attestation bundles were made for maintainability_agent-2.9.0-py3-none-any.whl:
Publisher:
release.yml on marshallguillory86/maintainability-agent
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
maintainability_agent-2.9.0-py3-none-any.whl -
Subject digest:
3d8594567a7b2a0aee61c00a11c706a425a4e4c7635c4249e664a33634971dad - Sigstore transparency entry: 2733471587
- Sigstore integration time:
-
Permalink:
marshallguillory86/maintainability-agent@7ce3b4468315728f5d9be1b5ae0e32b8a69f2653 -
Branch / Tag:
refs/tags/v2.9.0 - Owner: https://github.com/marshallguillory86
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@7ce3b4468315728f5d9be1b5ae0e32b8a69f2653 -
Trigger Event:
push
-
Statement type: