Skip to main content

athena-cost-guard

PyPI Python License: MIT

Estimate what an AWS Athena query will scan and cost — before you run it — and block queries that blow your budget.

Available on PyPI: pip install athena-cost-guard

Athena bills by the volume of data scanned from S3 (~$5/TB). Unlike BigQuery, it has no built-in dry-run, so it's easy to fire off one unpartitioned SELECT * and scan a terabyte by accident. athena-cost-guard gives you the pre-flight check Athena is missing.

from athena_cost_guard import estimate

est = estimate(
    "SELECT id FROM billing.line_items WHERE dt = '2026-08'",
    region="us-east-1",
)
print(est.summary())
# tables:            billing.line_items
# partitions matched: 1
# bytes scanned:     ≤ 4.2 GB
# estimated cost:    ≤ $0.0210  (@ $5.00/TB)

Guardrail mode — refuse to run anything over budget:

from athena_cost_guard import cost_guard, BudgetExceeded

@cost_guard(max_usd=1.00, region="us-east-1")
def run(sql):
    return athena.start_query_execution(QueryString=sql, ...)

try:
    run("SELECT * FROM billing.line_items")   # unpartitioned full scan
except BudgetExceeded as e:
    print(e)            # estimated cost $52.31 exceeds budget $1.00 (scanning 10.5 TB)
    print(e.estimate)   # the full Estimate for logging

How it works

  1. Parse the SQL with sqlglot (Athena dialect) to find the tables, referenced columns, and WHERE predicates on partition keys.
  2. Prune partitions by calling Glue GetPartitions with a pushdown expression built from those predicates — so you only pay attention to the partitions the query would actually touch.
  3. Size the surviving partitions by summing their S3 object bytes.
  4. Price the total using Athena's real billing rules (10 MB per-query minimum, rounded up to the nearest 10 MB, at your region's $/TB rate).

Install

pip install athena-cost-guard

# for column-aware (Tier-2) estimates, add the parquet extra:
pip install "athena-cost-guard[parquet]"

Note on the [parquet] extra: the base install is intentionally dependency-light (sqlglot + boto3) and covers Tier-1 estimates and @cost_guard. Column-aware Tier-2 estimates (sample_columns=True) read Parquet footers via pyarrow, which only ships with the [parquet] extra. Without it, calling estimate(..., sample_columns=True) raises an ImportError pointing you to pip install "athena-cost-guard[parquet]" — the core estimator keeps working.

Requires AWS credentials with glue:GetTable, glue:GetPartitions, and s3:ListBucket on the relevant tables/buckets (standard boto3 resolution: env vars, shared config, or instance role).

Column-aware estimates (Tier-2)

By default (sample_columns=False) you get the Tier-1 upper bound: every column in the matched partitions is assumed read. But Athena on Parquet only scans the columns a query references, so pass sample_columns=True for a tight, column-aware figure:

est = estimate(
    "SELECT servicename, SUM(billedcost) FROM billing.line_items WHERE dt = '2026-08' GROUP BY 1",
    region="us-east-1",
    sample_columns=True,      # needs the [parquet] extra
)
print(est.summary())
# bytes scanned:     ~317.4 MB        <- was ≤ 16.1 GB as a Tier-1 upper bound
# estimated cost:    ~$0.0016
# column-aware:      2% of bytes referenced  (from 8 sampled Parquet footer(s))

How it works: it reads the footers of up to sample_size (default 8) of the largest Parquet files in the matched partitions — via ranged S3 GETs, never downloading whole files — sums the compressed size of the columns the query references, and scales the byte total by that fraction. The result is an estimate (~), not an upper bound (). Falls back to the Tier-1 upper bound for SELECT *, non-Parquet data, or unreadable footers (with a warning — never silently wrong).

Proven against Athena's own numbers

On a real wide-table query selecting a handful of columns out of many, the column-aware estimate tracked Athena's actual DataScannedInBytes closely:

Data scanned
Tier-1 upper bound ≤ 16.1 GB
Tier-2 estimate (sample_columns=True) ~317 MB
Athena actual (DataScannedInBytes) 295 MB

Within ~8%, and on the conservative (slightly-over) side — the safe direction for a budget guard.

Accuracy: read this

Situation Behaviour
Column projection, sample_columns=True (Parquet) Modelled — tight column-aware estimate from footer sampling
Column projection, default Upper bound (all columns assumed read)
Partition projection tables (no Glue partitions) Warns, sizes the table root (wide upper bound)
OR / function predicates on partition keys Not pushed down → those partitions are included
Iceberg / row-group stats pruning Not modelled → upper bound

Roadmap

  • 0.2 — ✅ Tier-2 column-aware estimates via Parquet footer sampling.
  • 0.3 — CLI (athena-cost-guard "SELECT ..."), time-window partition pruning (literal date bounds), partition-projection support, Iceberg awareness.

Changelog

0.2.1

  • Docs only: refreshed README with the [parquet]-extra note and the Athena ground-truth validation table. No code changes.

0.2.0

  • Tier-2 column-aware estimates (sample_columns=True): reads Parquet footers (ranged S3 GETs) and scales the byte total by the fraction of referenced columns, turning the upper bound into a tight ~ estimate.
  • New optional [parquet] extra (pyarrow); the core stays sqlglot + boto3.
  • Validated against Athena DataScannedInBytes: ≤ 16.1 GB upper bound → ~317 MB estimate → 295 MB actual.

0.1.0

  • Initial release: Tier-1 upper-bound estimate() (SQL parse → Glue partition pruning → S3 sizing → Athena pricing) and the @cost_guard budget decorator.

Development

pip install -e ".[dev]"
pytest            # parser & pricing tests need no AWS; estimate tests use fakes

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

athena_cost_guard-0.2.1.tar.gz (16.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

athena_cost_guard-0.2.1-py3-none-any.whl (16.4 kB view details)

Uploaded Python 3

File details

Details for the file athena_cost_guard-0.2.1.tar.gz.

File metadata

  • Download URL: athena_cost_guard-0.2.1.tar.gz
  • Upload date:
  • Size: 16.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.2

File hashes

Hashes for athena_cost_guard-0.2.1.tar.gz
Algorithm Hash digest
SHA256 33457a7bf6569558f189288f54eb30a5fddb77f359258f1ad0f52631afbd9ddd
MD5 bf896bc7a056e271310f400b8222fb2b
BLAKE2b-256 9e29d8a3586639267beeab43f9303bff540ccb26cf54c245174260416af9fa07

See more details on using hashes here.

File details

Details for the file athena_cost_guard-0.2.1-py3-none-any.whl.

File metadata

File hashes

Hashes for athena_cost_guard-0.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 5742122cd6801b7b6bb0a564ea936ed228c5cb0061d21cd973c75d8a77221249
MD5 4a10c2c6e24d7f283ecb2987f77609a0
BLAKE2b-256 5332db3a5b5623247a4a63aee0055cdac21e8e4ff454f4eaaa9a454a1635a443

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page