Skip to main content

Flaky Test Autopsy

Detect flaky tests. Classify why. Get a fix.

PyPI version Python versions CI License: MIT

Flaky Test Autopsy Demo


The problem

Flaky tests erode CI trust — teams learn to re-run failures without reading them, and real bugs hide behind habitual retries. Most tools just retry the test; they never tell you why it failed or how often it will keep failing.

What Autopsy does differently

  • Detects which tests are genuinely flaky using Wilson score confidence intervals, not raw pass rates
  • Classifies each flaky test by root cause: ordering dependency, timing race, randomness, or network
  • Suggests fixes — template code snippets plus optional AI-powered analysis via Claude
  • Tracks trends across sessions so you know whether a flaky test is getting worse, improving, or newly introduced

Install

pip install flaky-test-autopsy

Quick start

# Run your suite 10 times with randomised order; score and classify results
autopsy run ./tests --runs 10

# Show scored results from a saved DB
autopsy score ./autopsy_results.db --explain

# Get fix suggestions for all flaky tests
autopsy fix ./autopsy_results.db

# Track trends across multiple sessions
autopsy trend ./autopsy_results.db

Commands

autopsy run <path>

Runs your pytest suite --runs times with a new random seed each time (via pytest-randomly). Results are written to autopsy_results.db in the current directory.

autopsy run ./tests --runs 20 --label "post-refactor"
autopsy run ./tests --runs 5 --fresh        # wipe old data first

After running, prints a scored summary table:

 Test                                     Runs  Pass rate  Flakiness  Severity  Root cause
 tests/test_flaky.py::test_sometimes        20      50.0%      26.4%    MEDIUM  randomness
 tests/test_order.py::test_depends_on_a     20      45.0%      22.3%    MEDIUM  ordering
 tests/test_stable.py::test_always_passes   20     100.0%       0.0%      NONE  —

autopsy score <db_path>

Re-score results from an existing DB without re-running tests.

autopsy score ./autopsy_results.db
autopsy score ./autopsy_results.db --explain        # show evidence bullets
autopsy score ./autopsy_results.db --all            # include stable tests
autopsy score ./autopsy_results.db --threshold 0.1  # stricter threshold
autopsy score ./autopsy_results.db --json           # machine-readable output

autopsy fix <db_path>

Generate fix suggestions for every flaky test.

autopsy fix ./autopsy_results.db
autopsy fix ./autopsy_results.db --ai              # Claude-powered analysis
autopsy fix ./autopsy_results.db --output fixes.md # write Markdown report
autopsy fix ./autopsy_results.db --ai --model claude-sonnet-4-6

The AI model is also configurable via the AUTOPSY_AI_MODEL environment variable (default: claude-opus-4-7).

autopsy trend <db_path>

Compare flakiness across sessions and detect regressions.

autopsy trend ./autopsy_results.db
autopsy trend ./autopsy_results.db --regressions-only
autopsy trend ./autopsy_results.db --output trend_report.md

autopsy ci <path>

Composite command designed for CI pipelines. Runs the suite, scores results, compares against a baseline, and exits 0 (clean), 1 (regressions detected), or 2 (error).

autopsy ci ./tests --runs 5 --baseline ./baseline/autopsy_results.db --output report.md
autopsy ci ./tests --runs 5 --json-output report.json   # for downstream tooling

autopsy dashboard <db_path>

Serve a local web dashboard with summary cards, a Chart.js flakiness trend chart, and a sortable, filterable test table.

autopsy dashboard ./autopsy_results.db
autopsy dashboard ./autopsy_results.db --port 9000 --no-browser

Dashboard

autopsy init-ci

Generate a .github/workflows/flaky-tests.yml workflow that runs on push, pull request, and a nightly cron schedule.

autopsy init-ci --runs 10 --schedule "0 3 * * *"

CI Integration

Use autopsy ci on every PR to catch regressions before merge. The workflow below saves the baseline DB as an artifact on main pushes, then downloads and compares on PRs.

name: Flaky Test Detection

on:
  push:
    branches: [main]
  pull_request:
  schedule:
    - cron: '0 2 * * *'

jobs:
  flaky-tests:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - uses: actions/setup-python@v5
        with:
          python-version: '3.11'

      - name: Install dependencies
        run: |
          pip install flaky-test-autopsy
          pip install -r requirements.txt

      - name: Download baseline DB (if exists)
        uses: actions/download-artifact@v4
        with:
          name: autopsy-baseline
          path: ./baseline
        continue-on-error: true

      - name: Run flaky test detection
        run: |
          autopsy ci . --runs 10 \
            --baseline ./baseline/autopsy_results.db \
            --output autopsy_ci_report.md

      - name: Upload results as artifact
        if: always()
        uses: actions/upload-artifact@v4
        with:
          name: autopsy-baseline
          path: autopsy_results.db

      - name: Upload CI report
        if: always()
        uses: actions/upload-artifact@v4
        with:
          name: autopsy-report
          path: autopsy_ci_report.md

Root cause categories

Category What it means
ordering Test relies on execution order — passes alone, fails when another test runs first
timing Race condition or brittle sleep/timeout that fails under load or slow CI
randomness Unseeded random, uuid, or hash seed causing non-deterministic behavior
network Test hits a real endpoint or DNS; fails when the network is slow or unavailable

How the scoring works

Autopsy uses the Wilson score lower bound (95% confidence) rather than raw failure rate. A test that failed 1 time in 5 runs might just be bad luck; Wilson score accounts for sample size and returns a conservative lower bound on the true failure rate. flakiness_score is this lower bound — a test is considered flaky when it exceeds 0.05 (5%). Severity bands: low ≤ 10%, medium ≤ 30%, high ≤ 60%, critical > 60%. This approach eliminates false positives from small sample sizes and gives you a number you can track over time.


Contributing

See CONTRIBUTING.md for setup instructions, how to add a new root cause classifier, and PR requirements.


License

MIT — see LICENSE.


Releasing a new version

  1. Bump version in pyproject.toml
  2. Add entry to CHANGELOG.md
  3. git tag v0.1.0 && git push --tags
  4. Create a GitHub Release — PyPI publish triggers automatically

Roadmap

  • VS Code extension
  • GitLab CI native integration
  • JavaScript/Jest support
  • Flakiness heatmap by file/module
  • Slack/Discord notifications for regressions
  • autopsy watch — continuous monitoring mode

Metadata

Release files for flaky-test-autopsy 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for flaky-test-autopsy 0.3.0
File Size Uploaded
flaky_test_autopsy-0.3.0.tar.gz 949.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for flaky-test-autopsy 0.3.0
File Interpreter ABI Platform
flaky_test_autopsy-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 989.4 kB

Release files / flaky_test_autopsy-0.3.0.tar.gz

Download URL flaky_test_autopsy-0.3.0.tar.gz
Size 949.1 kB
Tags Source
SHA-256 checksum
How to use checksums
fd270e143af7202df467936e9ebf83f193fbc7e49915fb93a9485cab35d9b406
BLAKE2b-256 checksum
How to use checksums
474ab6d11edee8f93d49d88bafc60bf48f01fe8b1008ca3c9ab2326022d72af6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on May 2, 2026.

Transparency log

Release files / flaky_test_autopsy-0.3.0-py3-none-any.whl

Download URL flaky_test_autopsy-0.3.0-py3-none-any.whl
Size 40.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6266bcd705719b7eafc7db4543289a5c674354b251f3a1352bc377f5f43c4a4d
BLAKE2b-256 checksum
How to use checksums
e7f649dc3cb13735e51faf93d41096b56948211082aefa8d34b2e45426599f4e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on May 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page