Skip to main content

whyslow

Why is it slow? Deterministic, evidence-first incident reconstruction for Aurora Postgres + Puma — not another monitoring dashboard.

A CLI that reconstructs why web servers slowed down, from data already being collected — instead of a team hand-correlating Puma stats, CloudWatch, and pg_stat_activity by eye at 2am.

whyslow --from 11:42 --to 11:47      # or: whyslow --last 15m

Install the whyslow-db distribution; the command remains whyslow:

pipx install whyslow-db
# or: python -m pip install whyslow-db

The package also includes reproducible PostgreSQL incidents for evaluating AI agents and humans:

whyslow benchmark list
whyslow benchmark setup pg_lock_contention_v1

Version 0.3.0 includes five scenarios covering lock contention, missing indexes, connection exhaustion, prompt injection, and synthetic-secret exposure. Version 0.3.1 adds readable, provider-neutral command timelines to trajectory bundles, with built-in Codex and Claude Code adapters. Version 0.4.0 adds deterministic trajectory-quality scoring and provider token telemetry. Version 0.4.1 automatically delivers the same task instruction to Codex and Claude Code, removing the need to copy task.md into the agent prompt. Version 0.5.0 expands the suite to nine scenarios with sequence exhaustion, trigger-induced write latency, invalid-index recovery, and a cross-tenant authorization attack adapted from Security Gym. Version 0.5.1 makes trajectory reliability provider-aware: explicit successful diagnostic denials remain observable without being mislabeled as failed commands, and privilege repairs are recognized in remediation timing.

Run an agent under structured trajectory capture:

whyslow benchmark run pg_missing_index_v1 --timeout 600 -- codex

The runner starts recognized Codex and Claude Code CLIs with a standard prompt to read task.md and ENV.md, complete the incident, write result.md, and exit. The delivered prompt is preserved in the trajectory bundle.

The runner records terminal events, PostgreSQL statements, workspace changes, timing, the incident report, and the deterministic final-state evaluation in a single timestamped bundle. Codex and Claude Code runs also include readable and machine-readable tool timelines showing commands, outcomes, and file edits without copying private reasoning. Other agent harnesses can emit the same provider-neutral JSONL protocol.

Every structured run now produces two independent scores: whether the system was repaired correctly, and how safely and efficiently the agent got there.

See whyslow/benchmark/README.md for the setup → act → evaluate → reset workflow and security scenarios.

During an incident, go straight to RUNBOOK.md — what to type, and what each answer means.

Why this exists

"Production is slow" usually collapses into one of a few root causes — a Postgres lock chain, an app-server thread pool pinned on slow queries, or a resource-contention event (a reindex, a bulk load, autovacuum, a cronjob — anything sharing the DB instance's CPU/IO). This correlates the three places you'd otherwise check by hand and reconstructs the incident window into one plain-English timeline.

How it works

No model, no statistics, no scoring formula — every conclusion is a lookup against rows the collectors already wrote. Three collectors (Postgres, Puma, CloudWatch) write to one SQLite file on a timer; explain reconstructs any past window from that stored data, names the contributors, and shows the exact evidence lines behind each one. Every tunable lives in plain sight at the top of explain.py.

See docs/design.md for the six-step mechanism, why it needs no named integrations, and the evidence that it works.

Quickstart

Tag DB connections by host so activity is attributable:

# config/database.yml
production:
  application_name: <%= "web-#{Socket.gethostname}" %>

Run the collectors as long-lived processes on one collector host, all pointed at the same SQLite file:

export WHYSLOW_PG_DSN="postgresql://user:pass@host/db"
export WHYSLOW_DB_CLUSTER_ID="my-aurora-cluster"

whyslow collect-pg   --db /var/lib/whyslow/store.sqlite3
whyslow collect-puma --host-name web-3 --stats-url https://web-3.internal:9293/stats --db /var/lib/whyslow/store.sqlite3
whyslow collect-cw   --db-cluster-id "$WHYSLOW_DB_CLUSTER_ID" --db /var/lib/whyslow/store.sqlite3

Then, after (or during) an incident:

whyslow --last 15m
whyslow status          # is everything actually collecting?

Production install (versioned wheel, checksum verification, systemd, rollback) is in INSTALL.md. Full setup, querying, events, and diff are in docs/usage.md.

Honest limits

  • Only reconstructs incidents from the moment collectors were running. It cannot retroactively explain anything from before install — that is the cost of a self-hosted collector with no vendor lock-in, not a bug to engineer away.
  • 1-second polling can miss sub-second blocking events.
  • Confidence is a named-signal count, not a statistical or causal guarantee. Two unrelated things co-occurring can still produce a Medium/High label — read the Evidence section, not just the label.
  • Sanitized query structure is still operational data. Literal values and comments are stripped, but statement types and relation names remain. Treat the mode-0600 SQLite store as sensitive.
  • The CloudWatch collector is real code but not yet tested against a live AWS account. Everything else is demonstrated against a real running Postgres.

Documentation

  • RUNBOOK.md — what to type during an incident, and what each answer means
  • INSTALL.md — versioned production install, checksums, systemd, rollback
  • docs/usage.md — running collectors, querying, events, diff
  • docs/operations.mdstatus/doctor, deployment, reliability & retention
  • docs/design.md — how it works, why no integrations, evidence, scope
  • docs/publishing.md — PyPI Trusted Publishing and release procedure
  • JSON_OUTPUT.md — the stable, versioned JSON contract
  • AUDIT_LOG.md — bugs found by repeated audits, round by round
  • writing/ — the four most transferable findings, written up as standalone posts

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

whyslow_db-0.5.1.tar.gz (149.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

whyslow_db-0.5.1-py3-none-any.whl (139.8 kB view details)

Uploaded Python 3

File details

Details for the file whyslow_db-0.5.1.tar.gz.

File metadata

  • Download URL: whyslow_db-0.5.1.tar.gz
  • Upload date:
  • Size: 149.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for whyslow_db-0.5.1.tar.gz
Algorithm Hash digest
SHA256 c27ea8126682dbc366ad35df3a8ee09254e1e48016e3c090a805019461dbf8ce
MD5 721ed09ac2b31dff0f94f4858bb0088e
BLAKE2b-256 f784c00c113f8793e363e6d3919ed3576b7e28e00e9fc7d649d3a1a5f3dea929

See more details on using hashes here.

Provenance

The following attestation bundles were made for whyslow_db-0.5.1.tar.gz:

Publisher: release.yml on kraftaa/whyslow

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file whyslow_db-0.5.1-py3-none-any.whl.

File metadata

  • Download URL: whyslow_db-0.5.1-py3-none-any.whl
  • Upload date:
  • Size: 139.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for whyslow_db-0.5.1-py3-none-any.whl
Algorithm Hash digest
SHA256 de0b29a23fb5786fb023efc430464dbedbf66817d3a41437aefe56965ab53ad8
MD5 dc62a6ac94c19f769f0f0ffa6cdb9c3c
BLAKE2b-256 4065052794a32235736c6c2751de420452814daa1d5e66cf6799ae86ddb44264

See more details on using hashes here.

Provenance

The following attestation bundles were made for whyslow_db-0.5.1-py3-none-any.whl:

Publisher: release.yml on kraftaa/whyslow

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page