Skip to main content

soif 💧

soif — French for thirst.

Estimate the water footprint of LLM prompts, the way you estimate their cost.

Every LLM answer evaporates real freshwater: data-center cooling towers (on-site) and the power plants feeding them (off-site) both consume it. Published per-prompt figures span two orders of magnitude — Google measured 0.26 mL per median Gemini prompt, while Mistral's lifecycle analysis reports 45 mL per 400-token Large 2 response. soif turns model + tokens + hosting assumptions into an honest low / mid / high water estimate with a fully documented, versioned methodology (METHODOLOGY.md).

  • Pure Python, zero runtime dependencies, MIT-licensed.
  • Library API, CLI, SDK response adapters (OpenAI / Anthropic), a Claude Code hook, and water-aware model routing for agent graphs.

📚 Docs: unchained-labs.github.io/soif

Install

pip install soif-llm                 # imports as `soif`; from a checkout: pip install .
pip install "soif-llm[tokenizers]"   # optional: exact token counts via tiktoken

(The PyPI distribution is soif-llm — the bare name was taken — but the module is soif.)

Quick start

import soif

est = soif.estimate("gpt-4o", prompt="Explain retrieval-augmented generation.")
print(est.humanize())
# ~1.15 mL of water (23.0 drops); range 0.113 mL - 12.46 mL

est.total_ml.mid        # millilitres, mid scenario
est.onsite_ml           # cooling-tower evaporation at the data center
est.offsite_ml          # water consumed generating the electricity
est.embodied_ml         # amortised manufacturing (chips, servers, buildings)
est.assumptions         # every default the estimate leaned on, spelled out

The accurate path is to feed real token usage from an API response — this captures actual output length, reasoning ("thinking") tokens, and cache hits:

response = client.chat.completions.create(...)   # OpenAI — or Anthropic messages.create
est = soif.from_response(response)

Reasoning models drink more — thinking tokens are output tokens:

soif.estimate("gpt-5", output_tokens=500, reasoning_effort="high")
soif.estimate("o3", output_tokens=500, reasoning_tokens=8000)   # from real usage

Control the hosting scenario:

soif.estimate("llama-3.1-70b", output_tokens=500,
              provider="aws", region="nordics",      # presets
              include_embodied=False)                # operational water only
soif.estimate("my-fine-tune", active_params_b=8, wue=0.2, pue=1.12, ewif=0.4)

CLI

soif estimate "why is the sky blue?" --model claude-sonnet-4-5
soif estimate -m gpt-4o -i 1200 -o 500 --json
soif compare gpt-4o gpt-4o-mini gemini-2.5-flash claude-haiku-4-5 -o 500
soif models

Agent graphs: metering and minimising water across a chain

Two primitives make water a first-class optimization target in agentic pipelines (LangGraph, hand-rolled DAGs, anything):

1. Meter — accumulate across nodes, with a soft budget:

meter = soif.Meter(budget_ml=50)

def summarize_node(state):
    resp = client.chat.completions.create(model=state["model"], ...)
    meter.record(soif.from_response(resp))
    if meter.over_budget:
        state["model"] = "gpt-4o-mini"   # degrade later hops
    return state

print(meter.summary())
# 7 call(s): ~18.2 mL of water (3.7 teaspoons); range ... — within budget (18.2/50.0 mL)

2. soif.optimize — route each node to the least-thirsty capable model:

from soif import optimize

optimize.pick_model(
    ["gpt-4o", "gpt-4o-mini", "gemini-2.5-flash", "claude-sonnet-4-5"],
    min_tier="small",              # capability floor for this node
    input_tokens=2000, output_tokens=300,
)
# -> 'gemini-2.5-flash'... whichever mid-scenario estimate is lowest

optimize.savings("claude-opus-4", "claude-haiku-4-5", output_tokens=500)
# {'baseline_ml': ..., 'alternative_ml': ..., 'saved_ml': ..., 'saved_pct': ...}

Model choice is the big lever (~30× between tiers); after that: shorter outputs, prompt caching, modest reasoning effort, and low-water regions/providers.

MCP server & agent stacks

  • MCP: soif-mcp exposes soif as MCP tools (estimate_water, estimate_from_usage, compare_models, pick_low_water_model, …) for Claude Code, Claude Desktop, Cursor, and any MCP client.
  • Agent guidance: AGENTS.md tells coding/orchestration agents how to use soif correctly (real usage over guesses, report ranges, quote assumptions).
  • Claude skill: .claude/skills/soif/ makes Claude answer water-footprint questions with soif automatically in this repo (copy it into any project's .claude/skills/).

Claude Code hook

Get a water read-out for every session, computed from the transcript's real token usage — see integrations/claude-code. In short, add to .claude/settings.json:

{
  "hooks": {
    "Stop": [
      { "hooks": [{ "type": "command", "command": "soif claude-hook" }] }
    ]
  }
}

After each turn: soif: this session used ~4.31 mL of water (0.9 teaspoons); range ... across 12 model call(s).

How it works (short version)

E_it       = tokens × Wh-per-token(model tier)        # server energy
E_facility = E_it × PUE                               # + cooling/power overhead
W_onsite   = E_it × WUE                               # cooling evaporation
W_offsite  = E_facility × EWIF                        # power-plant water
W_total    = (W_onsite + W_offsite) × lifecycle       # + embodied (optional)

Every factor is a (low, mid, high) triple propagated end-to-end, so the range reflects genuine uncertainty rather than false precision. Factors are versioned (soif.factors.FACTORS_VERSION) and calibrated against Google's measured Gemini numbers, Epoch AI's GPT-4o analysis, Mistral's Large 2 LCA, and Ren et al.'s methodology. Read METHODOLOGY.md before quoting numbers — these are estimates, not measurements.

factors.json — the factor set for other languages

The tables in src/soif/factors.py are the single source of truth for the whole soif project, but soif-app recomputes estimates in TypeScript. Rather than hand-port the numbers — which guarantees the two drift apart — the tables are serialised to factors.json:

python scripts/export_factors.py            # regenerate
python scripts/export_factors.py --check    # CI: fail if it drifts from the .py

The file carries the tier energies, token weights, tier boundaries, provider (WUE/PUE) and region (EWIF) tables, the lifecycle multiplier, and the full model registry with its matching rule. It also carries parity_vectors: estimates computed by this implementation that a port must reproduce, covering cached tokens, reasoning tokens, raw factor overrides, operational-only mode, and the unknown-model fallback. A port that matches every vector has demonstrated parity rather than claimed it.

factors.json is generated, never edited — CI fails the build if it disagrees with the Python. It is attached to each GitHub release, so downstream consumers can pin a factor set by URL. Stamp factors_version on anything you store, so historical estimates stay reproducible when the factors change.

Contributing

Factor updates (new disclosures, better WUE/PUE/EWIF data, new models) are the most valuable contributions — please include sources. pip install -e ".[dev]" && pytest && ruff check .

Changing a factor means regenerating the export: python scripts/export_factors.py. Bump FACTORS_VERSION in the same commit — stored estimates elsewhere are stamped with it, and a silent value change makes old rows irreproducible.

License

MIT

Metadata

Release files for soif-llm 1.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for soif-llm 1.1.0
File Size Uploaded
soif_llm-1.1.0.tar.gz 47.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for soif-llm 1.1.0
File Interpreter ABI Platform
soif_llm-1.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 70.4 kB

Release files / soif_llm-1.1.0.tar.gz

Download URL soif_llm-1.1.0.tar.gz
Size 47.6 kB
Tags Source
SHA-256 checksum
How to use checksums
719f1fc4f408a6598915458bfc2c5603180b2c3b21609c8bf261914c702f53f4
BLAKE2b-256 checksum
How to use checksums
e41f91f07df431fd94775b15bda02902b12a4013477092d7eaea9d978256f12e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.

Transparency log

Release files / soif_llm-1.1.0-py3-none-any.whl

Download URL soif_llm-1.1.0-py3-none-any.whl
Size 22.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6bc300a2ccd18630ee7492672e91ad9169e44e75dee34866f2a9c4bf88559c97
BLAKE2b-256 checksum
How to use checksums
fa9ae64bcc610879747ddad7949e13a6396052033d189866304d76200024a546
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.1.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page