athena-cost-guard
Estimate what an AWS Athena query will scan and cost — before you run it — and block queries that blow your budget.
Available on PyPI: pip install athena-cost-guard
Athena bills by the volume of data scanned from S3 (~$5/TB). Unlike BigQuery,
it has no built-in dry-run, so it's easy to fire off one unpartitioned SELECT *
and scan a terabyte by accident. athena-cost-guard gives you the pre-flight
check Athena is missing.
from athena_cost_guard import estimate
est = estimate(
"SELECT id FROM billing.line_items WHERE dt = '2026-08'",
region="us-east-1",
)
print(est.summary())
# tables: billing.line_items
# partitions matched: 1
# bytes scanned: ≤ 4.2 GB
# estimated cost: ≤ $0.0210 (@ $5.00/TB)
Guardrail mode — refuse to run anything over budget:
from athena_cost_guard import cost_guard, BudgetExceeded
@cost_guard(max_usd=1.00, region="us-east-1")
def run(sql):
return athena.start_query_execution(QueryString=sql, ...)
try:
run("SELECT * FROM billing.line_items") # unpartitioned full scan
except BudgetExceeded as e:
print(e) # estimated cost $52.31 exceeds budget $1.00 (scanning 10.5 TB)
print(e.estimate) # the full Estimate for logging
How it works
- Parse the SQL with
sqlglot(Athena dialect) to find the tables, referenced columns, and WHERE predicates on partition keys. - Prune partitions by calling Glue
GetPartitionswith a pushdown expression built from those predicates — so you only pay attention to the partitions the query would actually touch. - Size the surviving partitions by summing their S3 object bytes.
- Price the total using Athena's real billing rules (10 MB per-query minimum, rounded up to the nearest 10 MB, at your region's $/TB rate).
Install
pip install athena-cost-guard
# for column-aware (Tier-2) estimates, add the parquet extra:
pip install "athena-cost-guard[parquet]"
Note on the
[parquet]extra: the base install is intentionally dependency-light (sqlglot+boto3) and covers Tier-1 estimates and@cost_guard. Column-aware Tier-2 estimates (sample_columns=True) read Parquet footers viapyarrow, which only ships with the[parquet]extra. Without it, callingestimate(..., sample_columns=True)raises anImportErrorpointing you topip install "athena-cost-guard[parquet]"— the core estimator keeps working.
Requires AWS credentials with glue:GetTable, glue:GetPartitions, and
s3:ListBucket on the relevant tables/buckets (standard boto3 resolution:
env vars, shared config, or instance role).
Column-aware estimates (Tier-2)
By default (sample_columns=False) you get the Tier-1 upper bound: every
column in the matched partitions is assumed read. But Athena on Parquet only
scans the columns a query references, so pass sample_columns=True for a tight,
column-aware figure:
est = estimate(
"SELECT servicename, SUM(billedcost) FROM billing.line_items WHERE dt = '2026-08' GROUP BY 1",
region="us-east-1",
sample_columns=True, # needs the [parquet] extra
)
print(est.summary())
# bytes scanned: ~317.4 MB <- was ≤ 16.1 GB as a Tier-1 upper bound
# estimated cost: ~$0.0016
# column-aware: 2% of bytes referenced (from 8 sampled Parquet footer(s))
How it works: it reads the footers of up to sample_size (default 8) of the
largest Parquet files in the matched partitions — via ranged S3 GETs, never
downloading whole files — sums the compressed size of the columns the query
references, and scales the byte total by that fraction. The result is an
estimate (~), not an upper bound (≤). Falls back to the Tier-1 upper bound
for SELECT *, non-Parquet data, or unreadable footers (with a warning — never
silently wrong).
Proven against Athena's own numbers
On a real wide-table query selecting a handful of columns out of many, the
column-aware estimate tracked Athena's actual DataScannedInBytes closely:
| Data scanned | |
|---|---|
| Tier-1 upper bound | ≤ 16.1 GB |
Tier-2 estimate (sample_columns=True) |
~317 MB |
Athena actual (DataScannedInBytes) |
295 MB |
Within ~8%, and on the conservative (slightly-over) side — the safe direction for a budget guard.
Accuracy: read this
| Situation | Behaviour |
|---|---|
Column projection, sample_columns=True (Parquet) |
Modelled — tight column-aware estimate from footer sampling |
| Column projection, default | Upper bound (all columns assumed read) |
| Partition projection tables (no Glue partitions) | Warns, sizes the table root (wide upper bound) |
OR / function predicates on partition keys |
Not pushed down → those partitions are included |
| Iceberg / row-group stats pruning | Not modelled → upper bound |
Roadmap
- 0.2 — ✅ Tier-2 column-aware estimates via Parquet footer sampling.
- 0.3 — CLI (
athena-cost-guard "SELECT ..."), time-window partition pruning (literal date bounds), partition-projection support, Iceberg awareness.
Changelog
0.2.2
- Bug fix: a CTE named the same as a table it derives from no longer drops that table from the estimate (was collapsing the scan to the 10 MB floor — a silent under-count). CTE aliases are matched only against unqualified references.
- Bug fix: WHERE predicates are now attributed to the specific table they constrain. Previously only the first WHERE clause was read and predicates were matched by bare column name, so in multi-table queries a filter on one table could be pushed onto another (two same-key predicates could AND to zero partitions). Predicates that can't be unambiguously attributed are dropped, which only widens the estimate.
0.2.1
- Docs only: refreshed README with the
[parquet]-extra note and the Athena ground-truth validation table. No code changes.
0.2.0
- Tier-2 column-aware estimates (
sample_columns=True): reads Parquet footers (ranged S3 GETs) and scales the byte total by the fraction of referenced columns, turning the upper bound into a tight~estimate. - New optional
[parquet]extra (pyarrow); the core stayssqlglot+boto3. - Validated against Athena
DataScannedInBytes:≤ 16.1 GBupper bound →~317 MBestimate →295 MBactual.
0.1.0
- Initial release: Tier-1 upper-bound
estimate()(SQL parse → Glue partition pruning → S3 sizing → Athena pricing) and the@cost_guardbudget decorator.
Development
pip install -e ".[dev]"
pytest # parser & pricing tests need no AWS; estimate tests use fakes
License
MIT
Metadata
Release files for athena-cost-guard 0.2.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| athena_cost_guard-0.2.2.tar.gz | 19.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| athena_cost_guard-0.2.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 37.3 kB
Release files / athena_cost_guard-0.2.2.tar.gz
| Download URL | athena_cost_guard-0.2.2.tar.gz |
|---|---|
| Size | 19.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b23926b645078d8ebc7b640ca43f70fc33ccce9a79fdc0685a224f64c81271e8
|
|
BLAKE2b-256 checksum How to use checksums |
7652ef487228f8f03310a48018a30ae40d1afe491dfd2265a01166c2a07e15fc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.2
|
Release files / athena_cost_guard-0.2.2-py3-none-any.whl
| Download URL | athena_cost_guard-0.2.2-py3-none-any.whl |
|---|---|
| Size | 18.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b58e30be74d1ebe3286a44fc48cb1680b955a07a21c1f7102ed75f5772c205cb
|
|
BLAKE2b-256 checksum How to use checksums |
8979db42924bdace31eec2f914c041663df52ca0bc4830347da447b606314e1d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.2
|