Skip to main content

whyslow

Why is it slow? Deterministic, evidence-first incident reconstruction for Aurora Postgres + Puma — not another monitoring dashboard.

A CLI that reconstructs why web servers slowed down, from data already being collected — instead of a team hand-correlating Puma stats, CloudWatch, and pg_stat_activity by eye at 2am.

whyslow --from 11:42 --to 11:47      # or: whyslow --last 15m

Install the whyslow-db distribution; the command remains whyslow:

pipx install whyslow-db
# or: python -m pip install whyslow-db

The package also includes reproducible PostgreSQL incidents for evaluating AI agents and humans:

whyslow benchmark list
whyslow benchmark setup pg_lock_contention_v1

Version 0.3.0 includes five scenarios covering lock contention, missing indexes, connection exhaustion, prompt injection, and synthetic-secret exposure. Version 0.3.1 adds readable, provider-neutral command timelines to trajectory bundles, with built-in Codex and Claude Code adapters. Version 0.4.0 adds deterministic trajectory-quality scoring and provider token telemetry. Version 0.4.1 automatically delivers the same task instruction to Codex and Claude Code, removing the need to copy task.md into the agent prompt. Version 0.5.0 expands the suite to nine scenarios with sequence exhaustion, trigger-induced write latency, invalid-index recovery, and a cross-tenant authorization attack adapted from Security Gym.

Run an agent under structured trajectory capture:

whyslow benchmark run pg_missing_index_v1 --timeout 600 -- codex

The runner starts recognized Codex and Claude Code CLIs with a standard prompt to read task.md and ENV.md, complete the incident, write result.md, and exit. The delivered prompt is preserved in the trajectory bundle.

The runner records terminal events, PostgreSQL statements, workspace changes, timing, the incident report, and the deterministic final-state evaluation in a single timestamped bundle. Codex and Claude Code runs also include readable and machine-readable tool timelines showing commands, outcomes, and file edits without copying private reasoning. Other agent harnesses can emit the same provider-neutral JSONL protocol.

Every structured run now produces two independent scores: whether the system was repaired correctly, and how safely and efficiently the agent got there.

See whyslow/benchmark/README.md for the setup → act → evaluate → reset workflow and security scenarios.

During an incident, go straight to RUNBOOK.md — what to type, and what each answer means.

Why this exists

"Production is slow" usually collapses into one of a few root causes — a Postgres lock chain, an app-server thread pool pinned on slow queries, or a resource-contention event (a reindex, a bulk load, autovacuum, a cronjob — anything sharing the DB instance's CPU/IO). This correlates the three places you'd otherwise check by hand and reconstructs the incident window into one plain-English timeline.

How it works

No model, no statistics, no scoring formula — every conclusion is a lookup against rows the collectors already wrote. Three collectors (Postgres, Puma, CloudWatch) write to one SQLite file on a timer; explain reconstructs any past window from that stored data, names the contributors, and shows the exact evidence lines behind each one. Every tunable lives in plain sight at the top of explain.py.

See docs/design.md for the six-step mechanism, why it needs no named integrations, and the evidence that it works.

Quickstart

Tag DB connections by host so activity is attributable:

# config/database.yml
production:
  application_name: <%= "web-#{Socket.gethostname}" %>

Run the collectors as long-lived processes on one collector host, all pointed at the same SQLite file:

export WHYSLOW_PG_DSN="postgresql://user:pass@host/db"
export WHYSLOW_DB_CLUSTER_ID="my-aurora-cluster"

whyslow collect-pg   --db /var/lib/whyslow/store.sqlite3
whyslow collect-puma --host-name web-3 --stats-url https://web-3.internal:9293/stats --db /var/lib/whyslow/store.sqlite3
whyslow collect-cw   --db-cluster-id "$WHYSLOW_DB_CLUSTER_ID" --db /var/lib/whyslow/store.sqlite3

Then, after (or during) an incident:

whyslow --last 15m
whyslow status          # is everything actually collecting?

Production install (versioned wheel, checksum verification, systemd, rollback) is in INSTALL.md. Full setup, querying, events, and diff are in docs/usage.md.

Honest limits

  • Only reconstructs incidents from the moment collectors were running. It cannot retroactively explain anything from before install — that is the cost of a self-hosted collector with no vendor lock-in, not a bug to engineer away.
  • 1-second polling can miss sub-second blocking events.
  • Confidence is a named-signal count, not a statistical or causal guarantee. Two unrelated things co-occurring can still produce a Medium/High label — read the Evidence section, not just the label.
  • Sanitized query structure is still operational data. Literal values and comments are stripped, but statement types and relation names remain. Treat the mode-0600 SQLite store as sensitive.
  • The CloudWatch collector is real code but not yet tested against a live AWS account. Everything else is demonstrated against a real running Postgres.

Documentation

  • RUNBOOK.md — what to type during an incident, and what each answer means
  • INSTALL.md — versioned production install, checksums, systemd, rollback
  • docs/usage.md — running collectors, querying, events, diff
  • docs/operations.mdstatus/doctor, deployment, reliability & retention
  • docs/design.md — how it works, why no integrations, evidence, scope
  • docs/publishing.md — PyPI Trusted Publishing and release procedure
  • JSON_OUTPUT.md — the stable, versioned JSON contract
  • AUDIT_LOG.md — bugs found by repeated audits, round by round
  • writing/ — the four most transferable findings, written up as standalone posts

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

whyslow_db-0.5.0.tar.gz (148.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

whyslow_db-0.5.0-py3-none-any.whl (139.3 kB view details)

Uploaded Python 3

File details

Details for the file whyslow_db-0.5.0.tar.gz.

File metadata

  • Download URL: whyslow_db-0.5.0.tar.gz
  • Upload date:
  • Size: 148.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for whyslow_db-0.5.0.tar.gz
Algorithm Hash digest
SHA256 bf22c672dd402e6c1ee373a408159c7aafe63cf247b3e789c648a873e2455ce0
MD5 37c65f4b26dfe7cbe1203e9f822a0617
BLAKE2b-256 f1f0d3e38f87eed153c317ef4eacccde21787111b3d0e5bd0e07b7fd4c16e897

See more details on using hashes here.

Provenance

The following attestation bundles were made for whyslow_db-0.5.0.tar.gz:

Publisher: release.yml on kraftaa/whyslow

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file whyslow_db-0.5.0-py3-none-any.whl.

File metadata

  • Download URL: whyslow_db-0.5.0-py3-none-any.whl
  • Upload date:
  • Size: 139.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for whyslow_db-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 0a7d6231c8c36c9b158eb8decb9fda08cbad8de14a74d98d89f57cb14de41b0f
MD5 524576d51c67089284901bc02b5181a2
BLAKE2b-256 fc9a68f4a6e787d57b041bd8c241f5046f72bd6e28cf1fc3ba8cdf42583848f0

See more details on using hashes here.

Provenance

The following attestation bundles were made for whyslow_db-0.5.0-py3-none-any.whl:

Publisher: release.yml on kraftaa/whyslow

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page