Skip to main content

whyslow

Why is it slow? Deterministic, evidence-first incident reconstruction for Aurora Postgres + Puma — not another monitoring dashboard.

A CLI that reconstructs why web servers slowed down, from data already being collected — instead of a team hand-correlating Puma stats, CloudWatch, and pg_stat_activity by eye at 2am.

whyslow --from 11:42 --to 11:47      # or: whyslow --last 15m

Install the whyslow-db distribution; the command remains whyslow:

pipx install whyslow-db
# or: python -m pip install whyslow-db

The package also includes reproducible PostgreSQL incidents for evaluating AI agents and humans:

whyslow benchmark list
whyslow benchmark setup pg_lock_contention_v1

Version 0.3.0 includes five scenarios covering lock contention, missing indexes, connection exhaustion, prompt injection, and synthetic-secret exposure. Version 0.3.1 adds readable, provider-neutral command timelines to trajectory bundles, with built-in Codex and Claude Code adapters. Version 0.4.0 adds deterministic trajectory-quality scoring and provider token telemetry.

Run an agent under structured trajectory capture:

whyslow benchmark run pg_missing_index_v1 --timeout 600 -- codex

The runner records terminal events, PostgreSQL statements, workspace changes, timing, the incident report, and the deterministic final-state evaluation in a single timestamped bundle. Codex and Claude Code runs also include readable and machine-readable tool timelines showing commands, outcomes, and file edits without copying private reasoning. Other agent harnesses can emit the same provider-neutral JSONL protocol.

Every structured run now produces two independent scores: whether the system was repaired correctly, and how safely and efficiently the agent got there.

See whyslow/benchmark/README.md for the setup → act → evaluate → reset workflow and security scenarios.

During an incident, go straight to RUNBOOK.md — what to type, and what each answer means.

Why this exists

"Production is slow" usually collapses into one of a few root causes — a Postgres lock chain, an app-server thread pool pinned on slow queries, or a resource-contention event (a reindex, a bulk load, autovacuum, a cronjob — anything sharing the DB instance's CPU/IO). This correlates the three places you'd otherwise check by hand and reconstructs the incident window into one plain-English timeline.

How it works

No model, no statistics, no scoring formula — every conclusion is a lookup against rows the collectors already wrote. Three collectors (Postgres, Puma, CloudWatch) write to one SQLite file on a timer; explain reconstructs any past window from that stored data, names the contributors, and shows the exact evidence lines behind each one. Every tunable lives in plain sight at the top of explain.py.

See docs/design.md for the six-step mechanism, why it needs no named integrations, and the evidence that it works.

Quickstart

Tag DB connections by host so activity is attributable:

# config/database.yml
production:
  application_name: <%= "web-#{Socket.gethostname}" %>

Run the collectors as long-lived processes on one collector host, all pointed at the same SQLite file:

export WHYSLOW_PG_DSN="postgresql://user:pass@host/db"
export WHYSLOW_DB_CLUSTER_ID="my-aurora-cluster"

whyslow collect-pg   --db /var/lib/whyslow/store.sqlite3
whyslow collect-puma --host-name web-3 --stats-url https://web-3.internal:9293/stats --db /var/lib/whyslow/store.sqlite3
whyslow collect-cw   --db-cluster-id "$WHYSLOW_DB_CLUSTER_ID" --db /var/lib/whyslow/store.sqlite3

Then, after (or during) an incident:

whyslow --last 15m
whyslow status          # is everything actually collecting?

Production install (versioned wheel, checksum verification, systemd, rollback) is in INSTALL.md. Full setup, querying, events, and diff are in docs/usage.md.

Honest limits

  • Only reconstructs incidents from the moment collectors were running. It cannot retroactively explain anything from before install — that is the cost of a self-hosted collector with no vendor lock-in, not a bug to engineer away.
  • 1-second polling can miss sub-second blocking events.
  • Confidence is a named-signal count, not a statistical or causal guarantee. Two unrelated things co-occurring can still produce a Medium/High label — read the Evidence section, not just the label.
  • Sanitized query structure is still operational data. Literal values and comments are stripped, but statement types and relation names remain. Treat the mode-0600 SQLite store as sensitive.
  • The CloudWatch collector is real code but not yet tested against a live AWS account. Everything else is demonstrated against a real running Postgres.

Documentation

  • RUNBOOK.md — what to type during an incident, and what each answer means
  • INSTALL.md — versioned production install, checksums, systemd, rollback
  • docs/usage.md — running collectors, querying, events, diff
  • docs/operations.mdstatus/doctor, deployment, reliability & retention
  • docs/design.md — how it works, why no integrations, evidence, scope
  • docs/publishing.md — PyPI Trusted Publishing and release procedure
  • JSON_OUTPUT.md — the stable, versioned JSON contract
  • AUDIT_LOG.md — bugs found by repeated audits, round by round
  • writing/ — the four most transferable findings, written up as standalone posts

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

whyslow_db-0.4.0.tar.gz (136.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

whyslow_db-0.4.0-py3-none-any.whl (113.7 kB view details)

Uploaded Python 3

File details

Details for the file whyslow_db-0.4.0.tar.gz.

File metadata

  • Download URL: whyslow_db-0.4.0.tar.gz
  • Upload date:
  • Size: 136.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for whyslow_db-0.4.0.tar.gz
Algorithm Hash digest
SHA256 3d556933206cda3f66ac586d75312843d1281ec53c8f54dda1f12e277adb6ac1
MD5 7aa2d72f13f0395dac4ff2d8c1a4837c
BLAKE2b-256 07c79f2a62bb1019a4b9764c2ead659ade369a7f66a6e7429a6585dbeaad7c52

See more details on using hashes here.

Provenance

The following attestation bundles were made for whyslow_db-0.4.0.tar.gz:

Publisher: release.yml on kraftaa/whyslow

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file whyslow_db-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: whyslow_db-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 113.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for whyslow_db-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 13642cde511daff089e3c04b29d7087066a7a758a9c43add7b4659440487624d
MD5 1ca02dd6358e26338804740da2086d0e
BLAKE2b-256 02753056b04526aa7340c72e9275a031402bcae9a6b5bf855281d82e357afd48

See more details on using hashes here.

Provenance

The following attestation bundles were made for whyslow_db-0.4.0-py3-none-any.whl:

Publisher: release.yml on kraftaa/whyslow

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page