Skip to main content

🩻 ai-test-failure-analyzer


Root cause in seconds. Evidence, not intuition.

Feed it a real test result file from any of 24 supported frameworks — Playwright, Jest, Cypress, Newman, k6, JUnit, pytest, Vitest, WDIO, Mocha, Detox, Go test, RSpec, PHPUnit, NUnit, xUnit, Robot Framework, Allure, CTRF, SARIF, MSTest, Artillery, Gatling, Pact — and it traces back through your real git history, application logs, and config to surface the actual root cause, with a cited evidence chain and file:line precision. No guesses. No fixture noise. No repeating the obvious.

CI CodeQL npm PyPI License: MIT MCP server Agent Skill

ai-analyze running 10-phase analysis

🩻 A real analysis — evidence from git, app.log, and .env — no guesses, no fixture noise.


Why ai-test-failure-analyzer

Manual test failure investigation takes 30–60 minutes: open the test output, grep through logs, dig through git history, check recent deploys, ask Slack. And you can still point at the wrong thing — especially when the test file itself has an "intentional failure" comment or a fixture designed to trigger the analyzer.

This tool does it automatically in seconds:

  • Parses the test result file to extract failing tests with HTTP details
  • Scans git history for high-risk commits (endpoint renames, migrations, auth changes)
  • Scans application logs for ERROR/FATAL lines
  • Reads config files (.env, docker-compose)
  • Cross-correlates all evidence into clusters
  • Forms ranked, evidence-cited hypotheses with file:line precision
  • Never points to test fixtures or "intentional failure" comments as root causes

How it's different

ai-test-failure-analyzer Manual triage Generic LLM
Evidence source Real git/logs/config Human memory Training data
Fixture noise Blocked by Tier-1 gate No protection No protection
file:line precision Sometimes No
Works without source code ✅ API-only mode
Repeatable
CI-integrated

Supported frameworks

24 frameworks across all major ecosystems:

Framework Format Command
Playwright JSON reporter playwright test --reporter=json
Jest JSON jest --json --outputFile=results.json
Vitest JSON vitest run --reporter=json
Cypress Mochawesome JSON cypress run --reporter mochawesome
WebdriverIO JSON wdio run wdio.conf.js
Mocha JSON mocha --reporter json
Detox JSON detox test --reporter json
pytest JUnit XML pytest --junit-xml=results.xml
Newman (Postman) JSON newman run col.json --reporters json --reporter-json-export results.json
k6 Summary JSON k6 run --summary-export=results.json script.js
Go test JSON go test ./... -json > results.json
RSpec JSON rspec --format json --out results.json
PHPUnit XML phpunit --log-junit results.xml
NUnit XML standard NUnit output
xUnit XML standard xUnit output
MSTest TRX XML standard MSTest output
Robot Framework XML robot --output output.xml
Artillery JSON artillery run --output results.json
Gatling JSON standard Gatling simulation log
Pact JSON standard Pact verification output
Allure JSON allure generate results directory
CTRF JSON any CTRF-compliant reporter
SARIF JSON any SARIF-compliant scanner
Any JUnit-compatible XML TestNG, Karate, REST Assured, Insomnia CLI

Install

npm (global — JS/CI devs):

npm install -g ai-test-failure-analyzer
ai-analyze analyze playwright-report.json

npx (zero install):

npx ai-test-failure-analyzer analyze playwright-report.json

pipx (Python devs):

pipx install ai-test-failure-analyzer
analyzer analyze playwright-report.json

Claude Code skill:

/plugin install ai-test-failure-analyzer

Install skill to all agents (Claude, Cursor, Codex, Gemini, Windsurf):

ai-analyze install

Docker:

docker run --rm -v $(pwd):/workspace \
  ghcr.io/aks-builds/ai-test-failure-analyzer \
  analyze /workspace/results.json

Usage

CLI

ai-analyze analyze results.json
ai-analyze analyze results.json --mode api-only    # force API-only (no source scan)
ai-analyze analyze results.json --out report.md    # write Markdown report to file
ai-analyze analyze results.json --format json --out report.json   # JSON output
ai-analyze analyze results.json --format ctrf --out report.ctrf.json  # CTRF output
ai-analyze analyze results.json --create-issue     # file GitHub issue for top hypothesis
ai-analyze analyze results.json --enrich           # add AI root-cause annotations
ai-analyze analyze results.json --no-cache         # bypass result cache
ai-analyze watch results.json                      # re-analyze on every file change

CTRF output

ai-analyze analyze results.json --format ctrf --out analysis.ctrf.json

The CTRF output embeds ai root-cause annotations per test and includes a summary block:

{
  "results": {
    "tool": { "name": "ai-test-failure-analyzer", "version": "2.0.0" },
    "summary": { "tests": 42, "passed": 38, "failed": 3, "skipped": 1, "other": 0 },
    "tests": [
      {
        "name": "POST /api/clips → 404",
        "status": "failed",
        "duration": 312,
        "ai": "Root cause [95%]: endpoint moved. Evidence: git+config.",
        "extra": { "hypothesis_confidence": 95, "flakiness_score": 0.1 }
      }
    ]
  }
}

MCP server (Claude Code / Cursor)

Add to your MCP config:

{
  "mcpServers": {
    "ai-test-failure-analyzer": {
      "command": "ai-analyze",
      "args": ["serve-stdio"]
    }
  }
}

Then ask Claude: "Analyze the failures in playwright-report.json"

MCP HTTP (OpenAI / Gemini)

ai-analyze serve-http --port 8765

API-only mode

No source code? No problem. When your workspace has no src/, app/, lib/, or api/ directory — or when you pass --mode api-only — the tool switches to API-only mode.

It analyzes HTTP contract evidence directly from the test results:

ai-analyze analyze newman-results.json
# > API_ONLY mode — no workspace source detected, analyzing HTTP contract only
# Root Cause [95%] — POST /api/clips → 404 Not Found
#   Endpoint moved or removed. Check API changelog or versioning.
#   Evidence: response status 404 + URL /api/clips

Supports Newman, k6, Playwright (API tests), Jest, and any framework that records HTTP status codes.

CI integration

# .github/workflows/analyze-failures.yml
- name: Analyze test failures
  if: failure()
  run: |
    npx ai-test-failure-analyzer analyze test-results/results.json \
      --non-interactive \
      --out failure-analysis.md
- uses: actions/upload-artifact@v4
  if: failure()
  with:
    name: failure-analysis
    path: failure-analysis.md

Security

  • No shell injection: all subprocess calls use explicit argument lists
  • Path traversal protection: all paths resolved relative to workspace root
  • Size caps: 5 MB/file, 50 MB/scan, 200 commits max
  • Secrets redacted: .env token/secret/key/password values masked in reports
  • No outbound network from core analysis (GitHub issue creation is opt-in)

See SECURITY.md for the full threat model.

Repository layout

analyzer/                   Python package (MCP server + CLI + analysis)
  parsers/                  24 framework parsers (Playwright, Jest, Vitest, Cypress, WDIO,
                            Mocha, Detox, pytest, Newman, k6, Go, RSpec, PHPUnit, NUnit,
                            xUnit, MSTest, Robot Framework, Artillery, Gatling, Pact,
                            Allure, CTRF, SARIF, JUnit-compatible)
  evidence/                 Evidence collection (git, logs, config, flaky history)
  intelligence/             Flaky detector, clusterer, enricher
  render/                   Report rendering (Markdown, ANSI, CTRF)
  ui/                       User interfaces (CLI, TUI, Web)
  workspace_scanner.py      Phase 0 — mode detection, noise path discovery
  noise_filter.py           Evidence filtering and hypothesis deduplication
  orchestrator.py           10-phase analysis pipeline
  hypothesis.py             Confidence scoring and hypothesis formation
bin/cli.js                  Zero-dep Node wrapper (ai-analyze command)
skills/ai-test-failure-analyzer/SKILL.md  Claude Code agent skill
.claude-plugin/             Claude marketplace manifests
tests/analyzer/             pytest test suite
.github/workflows/          CI/CD (ci, release, publish, codeql)

Testing

pytest tests/analyzer -q    # Python: parsers, correlator, noise filter, workspace scanner
npm test                    # Node: CLI smoke tests

Contributing

See CONTRIBUTING.md.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ai_test_failure_analyzer-2.0.0.tar.gz (80.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ai_test_failure_analyzer-2.0.0-py3-none-any.whl (115.7 kB view details)

Uploaded Python 3

File details

Details for the file ai_test_failure_analyzer-2.0.0.tar.gz.

File metadata

  • Download URL: ai_test_failure_analyzer-2.0.0.tar.gz
  • Upload date:
  • Size: 80.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for ai_test_failure_analyzer-2.0.0.tar.gz
Algorithm Hash digest
SHA256 b3277d5ac192724029b441be53273f5d3e0bbbf60cf29cd7c748360741acf2bf
MD5 11d1dff5f47d24c6ee833300faa303d0
BLAKE2b-256 eafa7626110e766b10cb1e2a929758da3c87a2618f7e6e88f3f1e413b384a9d9

See more details on using hashes here.

File details

Details for the file ai_test_failure_analyzer-2.0.0-py3-none-any.whl.

File metadata

File hashes

Hashes for ai_test_failure_analyzer-2.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 8f898d60194cbbf605853b4ef9a57354e6a972bbb192c3661e8981934fe1a213
MD5 87952ce962ccf3a6f9910fe99df1d748
BLAKE2b-256 1709172f94711900bf28e0d26b5b8de041ce926fcbcba36ae45c655fedea122d

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page