Skip to main content

Cost and wall-time estimation for batched LLM jobs (per-iteration sampling with confidence intervals).

Project description

costscope

Cost + time estimation for batched LLM jobs.

Sample a handful of iterations, project the total cost and wall time with a confidence interval, confirm before spending the rest. Each "iteration" can be a single call or a multi-call pipeline. Works with OpenAI (chat completions + Responses API, including gpt-image-1), Anthropic, or a built-in synthetic backend for tests and demos.

Install

pip install -e .             # core
pip install -e '.[openai]'   # for OpenAI models
pip install -e '.[anthropic]' # for Claude models
pip install -e '.[dev]'      # with pytest

Requires Python 3.10+.

Usage

A single call per iteration (the classic case):

from costscope import CostEstimator

with CostEstimator(model="o1", total_iterations=500, sample_iterations=20) as ce:
    for prompt in prompts:
        response = ce.completion(messages=[{"role": "user", "content": prompt}])
        ...

Multiple calls per iteration — sample reflects the full pipeline cost:

with CostEstimator(model="claude-opus-4-7", total_iterations=500) as ce:
    for row in rows:
        with ce.iteration():
            facts = ce.completion(messages=[{"role": "user", "content": extract(row)}])
            summary = ce.completion(messages=[{"role": "user", "content": summarize(facts)}])

The first 20 iterations are billed normally and used to build a per-iteration cost and time distribution. After that you'll see something like:

┌────────────────────────────────────────────────────────────┐
│ Cost & Time Estimate                                       │
├────────────────────────────────────────────────────────────┤
│  Model:        claude-opus-4-7                             │
│  Sample:       20 of 500 iter  (actual $0.4321)            │
│  Per iter:     $0.0216  (σ $0.0042)                        │
│  Projected:    $10.81                                      │
│  95% CI cost:  $10.05 – $11.57  (±7.0%)                    │
│  Per iter time:  3.4s                                      │
│  Wall time:    28min  (sequential)                         │
│  95% CI time:  26min – 30min                               │
└────────────────────────────────────────────────────────────┘
  → Proceed? [y/N]:

Decline and subsequent .completion() calls raise EstimationCancelled.

Concurrency

If you plan to run iterations in parallel, pass concurrency=N so the wall-time projection accounts for it:

with CostEstimator(model="o1", total_iterations=500, concurrency=10) as ce:
    ...

Cost is unchanged; wall-time projection is divided by N.

Skip the prompt

  • auto_confirm=True — always proceed
  • threshold_usd=10.0 — auto-proceed when the upper bound is under the threshold
  • confirm_fn=... — supply your own confirmation callback

OpenAI Responses API

api="auto" (default) routes gpt-image-* and gpt-5* to the Responses API, leaving chat-style models on chat completions. Force one explicitly:

CostEstimator(model="gpt-5", api="responses", ...)

The adapter translates messages=input= and reads tokens from response.usage.input_tokens / output_tokens (and output_tokens_details.image_tokens for image generation).

Image generation (gpt-image-1)

with CostEstimator(model="gpt-image-1", total_iterations=200, concurrency=5) as ce:
    for prompt in prompts:
        ce.completion(input=prompt, tools=[{"type": "image_generation"}])

Image-output tokens are priced separately ($40/1M for gpt-image-1). See examples/image_generation.py.

Driving the SDK yourself

If you can't use ce.completion() (e.g. you call client.images.generate() directly, or stream), use the escape hatch:

with ce.iteration():
    resp = my_custom_call(...)
    ce.record(cost=compute_cost(resp), elapsed=measured_seconds)

Synthetic mode

For tests, demos, and dev loops where real API calls would cost money:

from costscope import CostEstimator, SyntheticConfig

cfg = SyntheticConfig(
    input_median=800, output_median=300, reasoning_median=2000,
    latency_median=1.2,            # simulate ~1.2s/call for time estimates
    image_output_median=4000,      # for image-gen models
    seed=42,
)

with CostEstimator(model="o1", total_iterations=500, synthetic=True, synthetic_config=cfg) as ce:
    ...

See examples/basic.py for a full runnable example.

Supported models (built-in pricing)

OpenAI o-series (o1, o3, o3-mini, ...), GPT-4o, GPT-5, gpt-image-1, Claude 4.x (Opus, Sonnet, Haiku). For other models, supply prices via SyntheticConfig.custom_prices or extend pricing.py.

Tests

pytest

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

costscope-0.2.0.tar.gz (14.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

costscope-0.2.0-py3-none-any.whl (12.7 kB view details)

Uploaded Python 3

File details

Details for the file costscope-0.2.0.tar.gz.

File metadata

  • Download URL: costscope-0.2.0.tar.gz
  • Upload date:
  • Size: 14.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.0

File hashes

Hashes for costscope-0.2.0.tar.gz
Algorithm Hash digest
SHA256 206284a498a37e8c973e81ad770bcd7bd9bef6d4b710f281173d8e7ae8d4cbb9
MD5 f466751aa2f495f9d52b921e36fea42f
BLAKE2b-256 792a57c6a3bd4e574896408e9d1e78143086928e3de25467b301ce7e45a0eca9

See more details on using hashes here.

File details

Details for the file costscope-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: costscope-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 12.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.0

File hashes

Hashes for costscope-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 8849ee95f62d0ee489190cb262cf57de227d57b1b76378b4f411f61e471a13fd
MD5 11a65321ae18c19e9e485d4523efb4ce
BLAKE2b-256 5804e4ea61dc77a96ba5fad35c85f9780b9dad85411ff7b6a3232438f1fb9e2e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page