Skip to main content

whyslow

Why is it slow — and did the agent fix it safely?

Whyslow is an evidence-first PostgreSQL toolkit with two related jobs:

  1. reconstruct why an Aurora PostgreSQL + Puma service slowed down, using data already collected; and
  2. evaluate whether an AI agent or human repaired a disposable PostgreSQL incident correctly, efficiently, and within defined safety boundaries.
whyslow --from 11:42 --to 11:47      # or: whyslow --last 15m

Install the whyslow-db distribution; the command remains whyslow:

pipx install whyslow-db
# or: python -m pip install whyslow-db

Test database agents before trusting them

With Docker and Docker Compose installed, the package also includes nine reproducible PostgreSQL incidents for evaluating AI agents and humans. Each run creates a live but disposable database, lets a responder investigate and act, then independently checks what actually changed.

The benchmark is agent-agnostic. Codex, Claude Code, another CLI agent, a custom harness, or a human can work on the same task and be evaluated against the same hidden ground truth. Codex and Claude Code have built-in command timeline adapters; other harnesses can emit the provider-neutral JSONL format.

The nine scenarios cover operational failures and adversarial safety cases:

  • lock contention, missing indexes, connection exhaustion, sequence exhaustion, trigger latency, and invalid-index recovery;
  • prompt injection hidden in database evidence, synthetic-secret handling, and cross-tenant least-privilege repair.

Evaluators check incident recovery, data integrity, collateral damage, tenant isolation, least privilege, unsafe operations, and the observable commands the responder executed.

whyslow benchmark list
whyslow benchmark run pg_cross_tenant_access_v1 \
  --timeout 600 --reset-after -- claude

For a non-interactive Codex run, use codex exec. The generated benchmark workspace is intentionally not a Git repository, so Codex also needs --skip-git-repo-check:

whyslow benchmark run pg_cross_tenant_access_v1 \
  --timeout 600 --reset-after -- \
  codex exec --skip-git-repo-check --approve-for-me

The runner starts recognized Codex and Claude Code CLIs with a standard prompt to read task.md and ENV.md, complete the incident, write result.md, and exit. The delivered prompt is preserved in the trajectory bundle.

The runner records terminal events, PostgreSQL statements, workspace changes, timing, the incident report, and the deterministic final-state evaluation in a single timestamped bundle. Codex and Claude Code runs also include readable and machine-readable tool timelines showing commands, outcomes, and file edits without copying private reasoning. Other agent harnesses can emit the same provider-neutral JSONL protocol.

Every structured run produces two independent scores:

  • Final state: was the system repaired without breaking protected state or crossing the scenario's security boundary?
  • Trajectory: how reliably, efficiently, and safely did the responder get there?

This is a focused PostgreSQL regression pack, not a guarantee of production safety or a claim that an agent will behave safely outside the tested boundaries. It is RL-compatible in the narrow reset → act → evaluate sense; it evaluates responders but does not train models or make LLM API calls.

The database is disposable, but the agent process still runs on your host under its own sandbox and approval policy. Do not disable those protections merely because the benchmark database is isolated.

See the benchmark guide for the setup → act → evaluate → reset workflow and security scenarios.

During an incident, go straight to RUNBOOK.md — what to type, and what each answer means.

Why this exists

"Production is slow" usually collapses into one of a few root causes — a Postgres lock chain, an app-server thread pool pinned on slow queries, or a resource-contention event (a reindex, a bulk load, autovacuum, a cronjob — anything sharing the DB instance's CPU/IO). This correlates the three places you'd otherwise check by hand and reconstructs the incident window into one plain-English timeline.

How incident reconstruction works

The incident-reconstruction mode uses no model, statistics, or scoring formula: every conclusion is a lookup against rows the collectors already wrote. Three collectors (Postgres, Puma, CloudWatch) write to one SQLite file on a timer; explain reconstructs any past window from that stored data, names the contributors, and shows the exact evidence lines behind each one. Every tunable lives in plain sight at the top of explain.py.

See docs/design.md for the six-step mechanism, why it needs no named integrations, and the evidence that it works.

Quickstart

Tag DB connections by host so activity is attributable:

# config/database.yml
production:
  application_name: <%= "web-#{Socket.gethostname}" %>

Run the collectors as long-lived processes on one collector host, all pointed at the same SQLite file:

export WHYSLOW_PG_DSN="postgresql://user:pass@host/db"
export WHYSLOW_DB_CLUSTER_ID="my-aurora-cluster"

whyslow collect-pg   --db /var/lib/whyslow/store.sqlite3
whyslow collect-puma --host-name web-3 --stats-url https://web-3.internal:9293/stats --db /var/lib/whyslow/store.sqlite3
whyslow collect-cw   --db-cluster-id "$WHYSLOW_DB_CLUSTER_ID" --db /var/lib/whyslow/store.sqlite3

Then, after (or during) an incident:

whyslow --last 15m
whyslow status          # is everything actually collecting?

Production install (versioned wheel, checksum verification, systemd, rollback) is in INSTALL.md. Full setup, querying, events, and diff are in docs/usage.md.

Honest limits

  • Only reconstructs incidents from the moment collectors were running. It cannot retroactively explain anything from before install — that is the cost of a self-hosted collector with no vendor lock-in, not a bug to engineer away.
  • 1-second polling can miss sub-second blocking events.
  • Confidence is a named-signal count, not a statistical or causal guarantee. Two unrelated things co-occurring can still produce a Medium/High label — read the Evidence section, not just the label.
  • Sanitized query structure is still operational data. Literal values and comments are stripped, but statement types and relation names remain. Treat the mode-0600 SQLite store as sensitive.
  • The CloudWatch collector is real code but not yet tested against a live AWS account. Everything else is demonstrated against a real running Postgres.

Documentation

  • RUNBOOK.md — what to type during an incident, and what each answer means
  • INSTALL.md — versioned production install, checksums, systemd, rollback
  • docs/usage.md — running collectors, querying, events, diff
  • docs/operations.md — status/doctor, deployment, reliability & retention
  • docs/design.md — how it works, why no integrations, evidence, scope
  • docs/publishing.md — PyPI Trusted Publishing and release procedure
  • JSON_OUTPUT.md — the stable, versioned JSON contract
  • AUDIT_LOG.md — bugs found by repeated audits, round by round
  • writing/ — the four most transferable findings, written up as standalone posts

Metadata

Release files for whyslow-db 0.5.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for whyslow-db 0.5.2
File Size Uploaded
whyslow_db-0.5.2.tar.gz 150.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for whyslow-db 0.5.2
File Interpreter ABI Platform
whyslow_db-0.5.2-py3-none-any.whl Python 3 none any Details

Total release size: 291.4 kB

Release files / whyslow_db-0.5.2.tar.gz

Download URL whyslow_db-0.5.2.tar.gz
Size 150.8 kB
Tags Source
SHA-256 checksum
How to use checksums
bfeb02d4d45288d033eef2a33e154bd4f11cae818608ee1cf73110f82838bb90
BLAKE2b-256 checksum
How to use checksums
e2e5c6fbbb600486374272852f3fc4f1fb830323d72dc89f32706c99d0d6fdad
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 12, 2026.

Transparency log

Release files / whyslow_db-0.5.2-py3-none-any.whl

Download URL whyslow_db-0.5.2-py3-none-any.whl
Size 140.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4e3ddb28000deafc339b3afe55a19233f2a13def0af415d05fc5497b20d3458f
BLAKE2b-256 checksum
How to use checksums
e48b49d525cfd9bdc8af856562c191ee7203a7c6603a149f3cad26a21dee980b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 12, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.5.2 This release

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page