energy-bench
Measure the energy cost of local LLM inference, reproducibly, and pick the configuration that spends the least without losing quality.
Quantizing a model's weights cuts its memory roughly in half. Whether it cuts
its energy is a separate question: the answer depends on the model, the
task, and how many tokens you generate, and it sometimes comes out negative —
an INT4 model that draws more energy than the BF16 one because the dequantization
overhead outweighs the cheaper matmuls on a short generation. energy-bench
measures that, and turns the measurement into a recommendation.
The measurement core is the harness from a study that clocked inference energy on an L40S at a coefficient of variation of 0.83%: the NVML accumulated-energy counter, an idle-power baseline subtracted from every reading, and the NVML handle matched to the GPU by UUID so a co-tenant's job on a shared node is never measured.
Status: alpha (
0.1.x). The API can still move.transformersandvllmbackends are implemented; the numbers have not yet been validated against the original study's runs on real hardware.
Install
pip install energy-bench # core: the meter, no backend
pip install "energy-bench[transformers]" # + the transformers backend
pip install "energy-bench[transformers,bnb]" # + bitsandbytes, for int4 / int8
pip install "energy-bench[vllm]" # + the vllm backend
pip install "energy-bench[datasets]" # + build tasks from the Hub
Python 3.10+. nvidia-ml-py is the only hard dependency; everything else is an
optional extra, so installing the core never drags in a torch build.
The pieces
| what it is | |
|---|---|
Task |
a name, a frozen list of examples, and two pure functions — render(example) -> prompt and score(example, output) -> float | None. Three built in: mmlu, gsm8k, summarization. |
Backend |
loads one model and generates text. Owns the GPU; does no energy bookkeeping. TransformersBackend, VLLMBackend, or your own. |
EnergyMeter |
measures energy around a block of GPU work via NVML. Never loads a model, never imports torch. |
measure(backend, task, config) |
runs the protocol once — idle baseline, warm-up, N timed reps — and returns a CellReport. |
sweep / recommend |
measure across a grid of configs, then pick the cheapest one that holds quality. |
Measure one configuration
from energy_bench import measure, MeasureConfig, prepare_task
from energy_bench.backends.transformers import TransformersBackend
task = prepare_task("mmlu", n=150) # 150 examples, drawn with a fixed seed
with TransformersBackend("Qwen/Qwen3-8B", precision="int4") as backend:
report = measure(backend, task, MeasureConfig(reps=3), progress=print)
report.energy_j_net_mean # mean net energy over the reps, in joules
report.energy_cv # std / mean across reps -- the reproducibility check
report.energy_per_out_tok_mj
report.accuracy
report.save("qwen3-8b_int4_mmlu.json")
A CellReport carries the per-rep detail plus the aggregates
(energy_j_net_mean / _std, energy_cv, energy_per_out_tok_mj,
power_mean_w, accuracy, tokens_out, vram_bytes) and full provenance —
backend.info(), a device snapshot, the config, and the task's source and
seed. to_dict() / save() write the flat record.
Sweep configurations and get a recommendation
from energy_bench import sweep, recommend, SweepPoint, MeasureConfig
from energy_bench.backends.transformers import TransformersBackend
def make_backend(point):
return TransformersBackend("Qwen/Qwen3-8B", precision=point.precision)
report = sweep(
make_backend, task,
[
SweepPoint("bf16"),
SweepPoint("int4"),
SweepPoint("int4", max_new_tokens=8, label="int4+cap8"),
],
MeasureConfig(reps=3),
on_error="skip", # a config that breaks on this model is recorded, not fatal
)
report.matrix() # energy and quality, one row per cell
report.deltas("bf16") # energy saving and accuracy change vs bf16
rec = recommend(report, baseline="bf16", max_quality_drop=0.02)
rec.chosen # e.g. "int4"
rec.reason # "int4 saves 38.0% energy, quality +0.007, vs bf16"
compare_strategies({task_name: report, ...}) contrasts the per-task
recommendation with "always baseline" and "always cheapest" across a set of
tasks.
Bring your own task
A Task is just data plus two functions:
from energy_bench import Task, MeasureConfig, measure
sentiment = Task(
name="sentiment",
examples=[
{"id": "1", "text": "a joyless slog", "label": "negative"},
{"id": "2", "text": "an absolute delight", "label": "positive"},
],
render=lambda ex: f"Positive or negative? One word.\n\n{ex['text']}",
score=lambda ex, out: float(ex["label"] in out.strip().lower()),
metric="accuracy",
max_new_tokens=4,
)
Adding a backend
Subclass energy_bench.backends.base.Backend and implement generate,
vram_bytes, and device_uuid. Then check it against the contract:
from energy_bench.backends import assert_backend_contract
def test_my_backend():
assert_backend_contract(lambda: MyBackend("a-tiny-model"))
The vllm backend exists partly as proof the contract holds for a runtime
unlike transformers: vLLM batches internally (so generate hands over the
whole prompt list at once), pre-allocates its KV-cache pool (so vram_bytes
returns None), and runs its engine out of process (so device_uuid is
best-effort and the meter falls back to device 0).
How the measurement works, and what it can't claim
- Energy is the delta of NVML's accumulated-energy counter across the timed
block, at millijoule resolution. GPUs without that counter (pre-Volta, some
MIG) fall back to integrating sampled power, and the reading is marked
method="power_integration". - Idle power is measured with the model resident but not running, then
idle_w * durationis subtracted from every reading. If the idle estimate runs high the net energy can come out slightly negative — that is real noise, not clamped. - The GPU is matched by UUID. If no NVML device matches the backend's UUID the meter measures device 0 and warns.
- Tokens and scores are deterministic under greedy decoding, so only energy, power, duration and temperature vary across reps; their spread is the error bar.
- One GPU, one quantization method at a time. The numbers do not transfer across hardware or runtimes — that is the point of measuring rather than assuming.
License
Metadata
Release files for energy-bench 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| energy_bench-0.1.0.tar.gz | 41.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| energy_bench-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 77.6 kB
Release files / energy_bench-0.1.0.tar.gz
| Download URL | energy_bench-0.1.0.tar.gz |
|---|---|
| Size | 41.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bdb8806f25dcf8c54cc9404faadcefd64495261f15b64d305437aa86dbdea668
|
|
BLAKE2b-256 checksum How to use checksums |
a65ed3aa88cfee70e260d5509d2301cd220e303ec2e916252fee8ed27082ca66
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 30, 2026.
Transparency logRelease files / energy_bench-0.1.0-py3-none-any.whl
| Download URL | energy_bench-0.1.0-py3-none-any.whl |
|---|---|
| Size | 36.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
449fc84f72558d5eda642a077f30b081ea111d9cd02cc157b5f36b6a6dc19bb3
|
|
BLAKE2b-256 checksum How to use checksums |
d90b82db96840241ae40c590ff3b310a261ea0924715736b8821e38d4214f78f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 30, 2026.
Transparency log