Skip to main content

athena-cost-guard

PyPI Python License: MIT

Estimate what an AWS Athena query will scan and cost — before you run it — and block queries that blow your budget.

Available on PyPI: pip install athena-cost-guard

Athena bills by the volume of data scanned from S3 (~$5/TB). Unlike BigQuery, it has no built-in dry-run, so it's easy to fire off one unpartitioned SELECT * and scan a terabyte by accident. athena-cost-guard gives you the pre-flight check Athena is missing.

from athena_cost_guard import estimate

est = estimate(
    "SELECT id FROM billing.line_items WHERE dt = '2026-08'",
    region="us-east-1",
)
print(est.summary())
# tables:            billing.line_items
# partitions matched: 1
# bytes scanned:     ≤ 4.2 GB
# estimated cost:    ≤ $0.0210  (@ $5.00/TB)

Guardrail mode — refuse to run anything over budget:

from athena_cost_guard import cost_guard, BudgetExceeded

@cost_guard(max_usd=1.00, region="us-east-1")
def run(sql):
    return athena.start_query_execution(QueryString=sql, ...)

try:
    run("SELECT * FROM billing.line_items")   # unpartitioned full scan
except BudgetExceeded as e:
    print(e)            # estimated cost $52.31 exceeds budget $1.00 (scanning 10.5 TB)
    print(e.estimate)   # the full Estimate for logging

How it works

  1. Parse the SQL with sqlglot (Athena dialect) to find the tables, referenced columns, and WHERE predicates on partition keys.
  2. Prune partitions by calling Glue GetPartitions with a pushdown expression built from those predicates — so you only pay attention to the partitions the query would actually touch.
  3. Size the surviving partitions by summing their S3 object bytes.
  4. Price the total using Athena's real billing rules (10 MB per-query minimum, rounded up to the nearest 10 MB, at your region's $/TB rate).

Install

pip install athena-cost-guard

# for column-aware (Tier-2) estimates, add the parquet extra:
pip install "athena-cost-guard[parquet]"

Note on the [parquet] extra: the base install is intentionally dependency-light (sqlglot + boto3) and covers Tier-1 estimates and @cost_guard. Column-aware Tier-2 estimates (sample_columns=True) read Parquet footers via pyarrow, which only ships with the [parquet] extra. Without it, calling estimate(..., sample_columns=True) raises an ImportError pointing you to pip install "athena-cost-guard[parquet]" — the core estimator keeps working.

Requires AWS credentials with glue:GetTable, glue:GetPartitions, and s3:ListBucket on the relevant tables/buckets (standard boto3 resolution: env vars, shared config, or instance role).

Column-aware estimates (Tier-2)

By default (sample_columns=False) you get the Tier-1 upper bound: every column in the matched partitions is assumed read. But Athena on Parquet only scans the columns a query references, so pass sample_columns=True for a tight, column-aware figure:

est = estimate(
    "SELECT servicename, SUM(billedcost) FROM billing.line_items WHERE dt = '2026-08' GROUP BY 1",
    region="us-east-1",
    sample_columns=True,      # needs the [parquet] extra
)
print(est.summary())
# bytes scanned:     ~317.4 MB        <- was ≤ 16.1 GB as a Tier-1 upper bound
# estimated cost:    ~$0.0016
# column-aware:      2% of bytes referenced  (from 8 sampled Parquet footer(s))

How it works: it reads the footers of up to sample_size (default 8) of the largest Parquet files in the matched partitions — via ranged S3 GETs, never downloading whole files — sums the compressed size of the columns the query references, and scales the byte total by that fraction. The result is an estimate (~), not an upper bound (). Falls back to the Tier-1 upper bound for SELECT *, non-Parquet data, or unreadable footers (with a warning — never silently wrong).

Proven against Athena's own numbers

On a real wide-table query selecting a handful of columns out of many, the column-aware estimate tracked Athena's actual DataScannedInBytes closely:

Data scanned
Tier-1 upper bound ≤ 16.1 GB
Tier-2 estimate (sample_columns=True) ~317 MB
Athena actual (DataScannedInBytes) 295 MB

Within ~8%, and on the conservative (slightly-over) side — the safe direction for a budget guard.

Accuracy: read this

Situation Behaviour
Column projection, sample_columns=True (Parquet) Modelled — tight column-aware estimate from footer sampling
Column projection, default Upper bound (all columns assumed read)
Partition projection tables (no Glue partitions) Warns, sizes the table root (wide upper bound)
OR / function predicates on partition keys Not pushed down → those partitions are included
Iceberg / row-group stats pruning Not modelled → upper bound

Roadmap

  • 0.2 — ✅ Tier-2 column-aware estimates via Parquet footer sampling.
  • 0.3 — CLI (athena-cost-guard "SELECT ..."), time-window partition pruning (literal date bounds), partition-projection support, Iceberg awareness.

Changelog

0.2.2

  • Bug fix: a CTE named the same as a table it derives from no longer drops that table from the estimate (was collapsing the scan to the 10 MB floor — a silent under-count). CTE aliases are matched only against unqualified references.
  • Bug fix: WHERE predicates are now attributed to the specific table they constrain. Previously only the first WHERE clause was read and predicates were matched by bare column name, so in multi-table queries a filter on one table could be pushed onto another (two same-key predicates could AND to zero partitions). Predicates that can't be unambiguously attributed are dropped, which only widens the estimate.

0.2.1

  • Docs only: refreshed README with the [parquet]-extra note and the Athena ground-truth validation table. No code changes.

0.2.0

  • Tier-2 column-aware estimates (sample_columns=True): reads Parquet footers (ranged S3 GETs) and scales the byte total by the fraction of referenced columns, turning the upper bound into a tight ~ estimate.
  • New optional [parquet] extra (pyarrow); the core stays sqlglot + boto3.
  • Validated against Athena DataScannedInBytes: ≤ 16.1 GB upper bound → ~317 MB estimate → 295 MB actual.

0.1.0

  • Initial release: Tier-1 upper-bound estimate() (SQL parse → Glue partition pruning → S3 sizing → Athena pricing) and the @cost_guard budget decorator.

Development

pip install -e ".[dev]"
pytest            # parser & pricing tests need no AWS; estimate tests use fakes

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

athena_cost_guard-0.2.2.tar.gz (19.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

athena_cost_guard-0.2.2-py3-none-any.whl (18.2 kB view details)

Uploaded Python 3

File details

Details for the file athena_cost_guard-0.2.2.tar.gz.

File metadata

  • Download URL: athena_cost_guard-0.2.2.tar.gz
  • Upload date:
  • Size: 19.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.2

File hashes

Hashes for athena_cost_guard-0.2.2.tar.gz
Algorithm Hash digest
SHA256 b23926b645078d8ebc7b640ca43f70fc33ccce9a79fdc0685a224f64c81271e8
MD5 2dfda48d1d41bae872d98d9080894c3c
BLAKE2b-256 7652ef487228f8f03310a48018a30ae40d1afe491dfd2265a01166c2a07e15fc

See more details on using hashes here.

File details

Details for the file athena_cost_guard-0.2.2-py3-none-any.whl.

File metadata

File hashes

Hashes for athena_cost_guard-0.2.2-py3-none-any.whl
Algorithm Hash digest
SHA256 b58e30be74d1ebe3286a44fc48cb1680b955a07a21c1f7102ed75f5772c205cb
MD5 990facb576151b2fbe6787b9265b609d
BLAKE2b-256 8979db42924bdace31eec2f914c041663df52ca0bc4830347da447b606314e1d

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page