soif 💧
soif — French for thirst.
Estimate the water footprint of LLM prompts, the way you estimate their cost.
Every LLM answer evaporates real freshwater: data-center cooling towers (on-site) and the
power plants feeding them (off-site) both consume it. Published per-prompt figures span two
orders of magnitude — Google measured 0.26 mL per median Gemini prompt, while Mistral's
lifecycle analysis reports 45 mL per 400-token Large 2 response. soif turns model +
tokens + hosting assumptions into an honest low / mid / high water estimate with a fully
documented, versioned methodology (METHODOLOGY.md).
- Pure Python, zero runtime dependencies, MIT-licensed.
- Library API, CLI, SDK response adapters (OpenAI / Anthropic), a Claude Code hook, and water-aware model routing for agent graphs.
📚 Docs: unchained-labs.github.io/soif
Install
pip install soif-llm # imports as `soif`; from a checkout: pip install .
pip install "soif-llm[tokenizers]" # optional: exact token counts via tiktoken
(The PyPI distribution is soif-llm — the bare name was taken — but the module is soif.)
Quick start
import soif
est = soif.estimate("gpt-4o", prompt="Explain retrieval-augmented generation.")
print(est.humanize())
# ~1.15 mL of water (23.0 drops); range 0.113 mL - 12.46 mL
est.total_ml.mid # millilitres, mid scenario
est.onsite_ml # cooling-tower evaporation at the data center
est.offsite_ml # water consumed generating the electricity
est.embodied_ml # amortised manufacturing (chips, servers, buildings)
est.assumptions # every default the estimate leaned on, spelled out
The accurate path is to feed real token usage from an API response — this captures actual output length, reasoning ("thinking") tokens, and cache hits:
response = client.chat.completions.create(...) # OpenAI — or Anthropic messages.create
est = soif.from_response(response)
Reasoning models drink more — thinking tokens are output tokens:
soif.estimate("gpt-5", output_tokens=500, reasoning_effort="high")
soif.estimate("o3", output_tokens=500, reasoning_tokens=8000) # from real usage
Control the hosting scenario:
soif.estimate("llama-3.1-70b", output_tokens=500,
provider="aws", region="nordics", # presets
include_embodied=False) # operational water only
soif.estimate("my-fine-tune", active_params_b=8, wue=0.2, pue=1.12, ewif=0.4)
CLI
soif estimate "why is the sky blue?" --model claude-sonnet-4-5
soif estimate -m gpt-4o -i 1200 -o 500 --json
soif compare gpt-4o gpt-4o-mini gemini-2.5-flash claude-haiku-4-5 -o 500
soif models
Agent graphs: metering and minimising water across a chain
Two primitives make water a first-class optimization target in agentic pipelines (LangGraph, hand-rolled DAGs, anything):
1. Meter — accumulate across nodes, with a soft budget:
meter = soif.Meter(budget_ml=50)
def summarize_node(state):
resp = client.chat.completions.create(model=state["model"], ...)
meter.record(soif.from_response(resp))
if meter.over_budget:
state["model"] = "gpt-4o-mini" # degrade later hops
return state
print(meter.summary())
# 7 call(s): ~18.2 mL of water (3.7 teaspoons); range ... — within budget (18.2/50.0 mL)
2. soif.optimize — route each node to the least-thirsty capable model:
from soif import optimize
optimize.pick_model(
["gpt-4o", "gpt-4o-mini", "gemini-2.5-flash", "claude-sonnet-4-5"],
min_tier="small", # capability floor for this node
input_tokens=2000, output_tokens=300,
)
# -> 'gemini-2.5-flash'... whichever mid-scenario estimate is lowest
optimize.savings("claude-opus-4", "claude-haiku-4-5", output_tokens=500)
# {'baseline_ml': ..., 'alternative_ml': ..., 'saved_ml': ..., 'saved_pct': ...}
Model choice is the big lever (~30× between tiers); after that: shorter outputs, prompt caching, modest reasoning effort, and low-water regions/providers.
MCP server & agent stacks
- MCP:
soif-mcpexposes soif as MCP tools (estimate_water,estimate_from_usage,compare_models,pick_low_water_model, …) for Claude Code, Claude Desktop, Cursor, and any MCP client. - Agent guidance: AGENTS.md tells coding/orchestration agents how to use soif correctly (real usage over guesses, report ranges, quote assumptions).
- Claude skill:
.claude/skills/soif/makes Claude answer water-footprint questions with soif automatically in this repo (copy it into any project's.claude/skills/).
Claude Code hook
Get a water read-out for every session, computed from the transcript's real token usage
— see integrations/claude-code. In short, add to
.claude/settings.json:
{
"hooks": {
"Stop": [
{ "hooks": [{ "type": "command", "command": "soif claude-hook" }] }
]
}
}
After each turn: soif: this session used ~4.31 mL of water (0.9 teaspoons); range ... across 12 model call(s).
How it works (short version)
E_it = tokens × Wh-per-token(model tier) # server energy
E_facility = E_it × PUE # + cooling/power overhead
W_onsite = E_it × WUE # cooling evaporation
W_offsite = E_facility × EWIF # power-plant water
W_total = (W_onsite + W_offsite) × lifecycle # + embodied (optional)
Every factor is a (low, mid, high) triple propagated end-to-end, so the range reflects
genuine uncertainty rather than false precision. Factors are versioned
(soif.factors.FACTORS_VERSION) and calibrated against Google's measured Gemini numbers,
Epoch AI's GPT-4o analysis, Mistral's Large 2 LCA, and Ren et al.'s methodology.
Read METHODOLOGY.md before quoting numbers — these are estimates,
not measurements.
factors.json — the factor set for other languages
The tables in src/soif/factors.py are the single source of truth for the whole soif
project, but soif-app recomputes estimates
in TypeScript. Rather than hand-port the numbers — which guarantees the two drift apart —
the tables are serialised to factors.json:
python scripts/export_factors.py # regenerate
python scripts/export_factors.py --check # CI: fail if it drifts from the .py
The file carries the tier energies, token weights, tier boundaries, provider (WUE/PUE)
and region (EWIF) tables, the lifecycle multiplier, and the full model registry with its
matching rule. It also carries parity_vectors: estimates computed by this
implementation that a port must reproduce, covering cached tokens, reasoning tokens, raw
factor overrides, operational-only mode, and the unknown-model fallback. A port that
matches every vector has demonstrated parity rather than claimed it.
factors.json is generated, never edited — CI fails the build if it disagrees with the
Python. It is attached to each GitHub release, so downstream consumers can pin a factor
set by URL. Stamp factors_version on anything you store, so historical estimates stay
reproducible when the factors change.
Contributing
Factor updates (new disclosures, better WUE/PUE/EWIF data, new models) are the most
valuable contributions — please include sources. pip install -e ".[dev]" && pytest && ruff check .
Changing a factor means regenerating the export: python scripts/export_factors.py.
Bump FACTORS_VERSION in the same commit — stored estimates elsewhere are stamped with
it, and a silent value change makes old rows irreproducible.
License
MIT
Metadata
Release files for soif-llm 1.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| soif_llm-1.1.0.tar.gz | 47.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| soif_llm-1.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 70.4 kB
Release files / soif_llm-1.1.0.tar.gz
| Download URL | soif_llm-1.1.0.tar.gz |
|---|---|
| Size | 47.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
719f1fc4f408a6598915458bfc2c5603180b2c3b21609c8bf261914c702f53f4
|
|
BLAKE2b-256 checksum How to use checksums |
e41f91f07df431fd94775b15bda02902b12a4013477092d7eaea9d978256f12e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.
Transparency logRelease files / soif_llm-1.1.0-py3-none-any.whl
| Download URL | soif_llm-1.1.0-py3-none-any.whl |
|---|---|
| Size | 22.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6bc300a2ccd18630ee7492672e91ad9169e44e75dee34866f2a9c4bf88559c97
|
|
BLAKE2b-256 checksum How to use checksums |
fa9ae64bcc610879747ddad7949e13a6396052033d189866304d76200024a546
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.
Transparency log