DB Snooper
DB Snooper generates compact, LLM-ready database context for SQL generation, query debugging, and schema exploration. Profiling alone drives state-of-the-art text-to-SQL accuracy (Automatic Metadata Extraction for Text-to-SQL). Supports SQLite, PostgreSQL, MySQL, MariaDB, and DuckDB. Requires Python ≥ 3.10.
Specification: spec/main.md
It inspects an existing database and produces a SQL profile (<database>/<schema>.sql): DDL, row counts, sampled rows, and per-column summaries. Use --per-table for one .sql per table.
AI agents and text-to-SQL pipelines can read this context instead of guessing table meanings.
Quick Start
Install with pip:
pip install db-snooper
Or run instantly with uvx (no install needed):
uvx db-snooper profile --db-type mysql --user user --password password --database db --schema sch --port 3306
This creates a profile at db/sch.sql.
What The Outputs Contain
The profile .sql file contains:
- Metadata with db-snooper version, UTC generation timestamp, SQL dialect, database name, and schema.
CREATE TABLEDDL, indexes, and constraints.- Total row counts.
- Deterministic sampled rows for small tables.
- Latest and random sampled rows for larger tables.
- Per-column null, non-null, distinct, numeric range, median, top-value, and shape summaries for larger tables.
- Catalog-derived estimates for very large tables (or metrics that are skipped on medium-large tables) from each engine's internal statistics — PostgreSQL
pg_stats, MySQLCOLUMN_STATISTICShistograms, and MariaDBmysql.column_stats— emitted with a≈/(catalog)marker so they are distinguishable from exact values. - Top-level key frequencies for JSON/JSONB columns and min/avg/max element counts for ARRAY columns (when row counts allow).
- Redacted values for sensitive column names containing
password,passwd,pwd,hash,salt,secret, ortoken. - A
-- skipped technical tables:line naming migration/framework tables excluded from the profile.
Database Examples
SQLite
db-snooper profile --db-type sqlite --database path/to/app.sqlite
PostgreSQL
Profile
db-snooper profile --db-type postgres --database app_db --schema sch --user readonly_user --host localhost --port 5432 --ask-password
MySQL
db-snooper profile --db-type mysql --database app_db --user readonly_user --host localhost --port 3306 --ask-password
MariaDB
db-snooper profile --db-type mariadb --database app_db --user readonly_user --host localhost --port 3306 --ask-password
DuckDB
db-snooper profile --db-type duckdb --database warehouse.duckdb --schema sch
Environment Variables
Connection values can come from environment variables instead of flags:
DB_SNOOPER_DB_TYPE=sqlite \
DB_SNOOPER_DATABASE=eval-dataset/student_club/student_club.sqlite \
db-snooper profile
Supported variables:
DB_SNOOPER_DB_TYPEDB_SNOOPER_DATABASEDB_SNOOPER_DB_HOSTDB_SNOOPER_DB_PORTDB_SNOOPER_DB_USERDB_SNOOPER_DB_PASSWORDDB_SNOOPER_SCHEMA
For server databases, --host defaults to localhost, --port defaults to the database default, and DB Snooper securely prompts for a password when DB_SNOOPER_DB_PASSWORD is not set.
Help
db-snooper -h
db-snooper profile -h
Table filters:
db-snooper profile --db-type sqlite --database app.sqlite --include-tables users,orders,line_items
Schema filter:
db-snooper profile --db-type postgres --database app_db --schema reporting --user readonly_user --port 5432 --ask-password
DB_SNOOPER_SCHEMA=reporting db-snooper profile --db-type postgres --database app_db --user readonly_user --port 5432 --ask-password
Profile options:
--small-table-threshold 50: tables with this many rows or fewer are sampled instead of column-profiled.--large-table-threshold 100000000: tables whose catalog row estimate is at/above this count are profiled from internal database stats only.COUNT(*), sampled rows, and per-column queries are skipped because they would be too slow on hundreds of millions of rows. Instead, each column is summarized from the engine's catalog statistics (approximate null fraction, distinct count, numeric min/max, and top values), marked with≈/(catalog).--sample-row-limit 50: maximum sampled rows for small tables.--include-tables table_a,table_b: only profile selected tables.--exclude-tables table_c: skip selected tables.--include-technical-tables: profile migration/framework tables (e.g.schema_migrations,alembic_version,flyway_schema_history,django_migrations) that are skipped by default.--per-table: generate one.sqlprofile for each table instead of a single schema profile.
Python API
Use the simple helpers when you have a SQLAlchemy URL:
from db_snooper import generate_profile
database_url = "sqlite:///eval-dataset/superhero/superhero.sqlite"
profile_sql = generate_profile(database_url)
Use the lower-level API when you already have a SQLAlchemy engine or need options:
from sqlalchemy import create_engine
from db_snooper import ProfileOptions, profile_database
engine = create_engine("sqlite:///eval-dataset/superhero/superhero.sqlite")
profile_sql = profile_database(
engine,
ProfileOptions(sample_row_limit=25, include_tables=frozenset({"superhero", "publisher"})),
)
Agent Skills
DB Snooper ships reusable agent skills that teach AI agents when and how to profile a database. Two skills are bundled:
| Skill | Triggers on | Command | Output |
|---|---|---|---|
db-snooper-profile |
profiling, schema/data context, table summaries, column distributions | db-snooper profile |
<db>/<schema>.sql |
db-snooper-context |
general text-to-SQL context | db-snooper profile |
<db>/<schema>.sql |
Join-path discovery (declared PK/FK plus inferred candidates) now lives in the separate schema-linker tool.
List the bundled skills:
uvx db-snooper skills list
Install the skills into an agent's discovery directory (no repo clone needed; the SKILL.md files ship inside the wheel):
# Default: opencode global (~/.config/opencode/skills)
uvx db-snooper skills install
# All three common discovery locations at once (opencode + Claude + agents)
uvx db-snooper skills install --target all
# Custom or project-local directory
uvx db-snooper skills install --dir ./.opencode/skills --force
Discovery directories:
--target opencode→~/.config/opencode/skills(default)--target claude→~/.claude/skills--target agents→~/.agents/skills--target all→ all three--dir PATH→ any custom path (overrides--target)
For zero per-user setup, commit the installed skill folders (for example .opencode/skills/) into your repository. Agents that walk the working directory discover them automatically.
License
The DB Snooper source code is licensed under the MIT License. See LICENCE.
Third-party Python dependencies remain under their own upstream licenses. See THIRD_PARTY_NOTICES.md for a dependency license summary.
The dataset files included under eval-dataset/ are derived from birdsql by The BIRD Team, and are used and redistributed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0).
These files are not covered by the MIT source-code license. They retain their original CC BY-SA 4.0 terms. Any derivative works that include these files must also be distributed under CC BY-SA 4.0.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file db_snooper-0.0.14.tar.gz.
File metadata
- Download URL: db_snooper-0.0.14.tar.gz
- Upload date:
- Size: 52.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.7.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d7fd5d4bac7bb5cd08d0fb0a92c0b07d5a48191339f63f97df8759cc990e2930
|
|
| MD5 |
7aae21980148f882def1ec1c4699b993
|
|
| BLAKE2b-256 |
e364f546a4f8d2b352139c921cc2dd86ab004640a3150af0fd0d59ddc2fa6e6b
|
File details
Details for the file db_snooper-0.0.14-py3-none-any.whl.
File metadata
- Download URL: db_snooper-0.0.14-py3-none-any.whl
- Upload date:
- Size: 38.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.7.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0a030620484fe94a8c4305dc3cd34d49650acb2fea3dfd368f8a672ba3048f80
|
|
| MD5 |
6a29f39964871c7bbb8fc9cfe435eeb0
|
|
| BLAKE2b-256 |
c1757371af1692724066c39eda7a42ee6398c8f1229d63fd0ad27d2b4b2f9ae0
|