Skip to main content

DecisionMetrics

Standalone telemetry and replay for models that choose among supplied actions. Record what the model chose, how long it took, what it cost, and whether the application actually applied the action.

Python 3.11+. The core uses only the standard library. OpenTelemetry is optional. The library has its own provider contract and no Vibecheck dependency.

Install and try it offline

With your Python 3.11+ environment active:

python -m pip install decision-metrics
decision-metrics demo --output output/demo.jsonl
decision-metrics summary output/demo.jsonl

For OpenTelemetry support, install decision-metrics[otel]. Version 0.1.0 is an alpha release; live provider compatibility and game performance need smoke tests.

The demo is synthetic and makes no model or game calls. Output files are created exclusively: use a fresh name for each run.

Record a decision and its action

from decision_metrics import ChoiceRequest, Harness, JsonlSink
from decision_metrics.providers import load_backend

backend = load_backend("examples/providers.toml", "strands-local")
try:
    with JsonlSink("output/run-01.jsonl") as sink:
        bench = Harness(backend, sink, metadata={"application": "my_game"})
        request = ChoiceRequest(
            state={"health": 5, "threat": "east"},
            options={"west": {"clearance": 150}, "east": {"clearance": 10}},
            instructions="Choose a movement that helps the player survive.",
        )
        decision = bench.choose(request, phase="combat", deadline_ms=850)
        if decision.expired():
            bench.record_action(decision, "expired", reason="combat_deadline")
        else:
            # Replace with the application's authoritative action execution.
            accepted = game.apply(decision.result.choice)
            bench.record_action(decision, "applied" if accepted else "rejected")
        bench.outcome("session_complete", metrics={"waves_completed": 9})
finally:
    backend.close()

game.apply is an application hook, not a library function. Pass the observation's actual time.monotonic() timestamp as observed_at when it predates the call. Action timestamps measure when the application records the result, which can be after a bridge acknowledgement; they do not reveal the exact simulation tick.

The game owns legal action generation, execution, freshness, resets, and scoring. DecisionMetrics records those facts without executing an action or retrying a stale model request. Action statuses are applied, rejected, expired, skipped, and unknown for delivery without a reliable acknowledgement.

Providers

examples/providers.toml contains separate entries for Jev, Clef, Clef-flash, Perplexity, Strands local/remote, Kev, and meraGPT. Download that file to examples/providers.toml before using the examples below, or clone the repository to get all sample files. They are not installed with the library.

Transport Configuration and credential environment
Jev System One provider = "jev"; JEV_KEY
Cloudflare Workers AI provider = "clef" or "clef-flash"; CLOUDFLARE_ACCOUNT_ID, CLOUDFLARE_AUTH_TOKEN
Perplexity Decisions provider = "perplexity"; PERPLEXITY_API_KEY
Local Strands / Kev provider = "strands" or "kev"; local HTTP endpoint
meraGPT System One provider = "meragpt"; MERAGPT_API_KEY; 2–10 options
Another compatible server provider = "systemone"; explicit endpoint and model

Hosted prices are dated defaults, with a source stored beside each rate. Override them in TOML when they change. Local entries record zero API charges while excluding hardware/energy costs. Remote Strands is a configuration template; replace its endpoint and record the actual server hardware before using it. No remote deployment is created by this library.

Use the same checkpoint for both Strands entries, and start its server with strict window validation. The v19 checkpoint has a 4,096-token window; the server's default truncation would change the benchmark input. The example's checkpoint, device, and strict-window fields describe the intended server configuration; they are not a runtime attestation. Strands serving instructions.

The HTTP transports are implemented and tested against fixtures and a local HTTP server. Live vendor compatibility, account access, and game performance still need smoke tests. OpenAI Decisions is pending a verified account/API contract.

Freeze inputs and replay them

A packet is one JSON object per line with state, ordered options, instructions, and optional phase and case_id. The sample cases are synthetic.

decision-metrics replay examples/packets.jsonl \
  --config examples/providers.toml --backend strands-local \
  --limit 3 --max-cost 0.50 --output output/strands-local-replay.jsonl
decision-metrics summary output/strands-local-replay.jsonl

Replay is sequential and records actions as skipped. It measures response validity, usage, latency, and distributions. It supplies no ground-truth action labels or gameplay win rate. The cost limit uses known API usage and can be crossed by the final request; unknown usage/prices remain explicit. Separate request limits still apply.

Use separate entries for deployment comparisons. Local response speed is part of the result; a faster local model can make more timely decisions in a live game. A second experiment with a shared request cadence can help isolate choice quality from deployment speed. Record cadence, hardware, precision, checkpoint, and server settings in run metadata.

For full episodes, set call and billing limits high enough to finish. An identical call ceiling can stop a faster controller earlier; budget stops belong in a separate outcome category from game deaths.

Bring another backend

Implement info, choose(request) -> DecisionResult, and close():

from decision_metrics import BackendInfo, DecisionResult

class MyBackend:
    info = BackendInfo("my_backend", "checkpoint-v1", "local_cpu")

    def choose(self, request):
        selected_id = my_model.select(request.state, request.options, request.instructions)
        return DecisionResult(choice=selected_id)

    def close(self):
        pass

Return Usage when available. Missing probabilities, confidence, and usage stay None; a generative or heuristic backend does not need to fabricate them. Native HTTP decision adapters require the promised complete option distribution. One legal option is recorded as forced without calling the backend.

Validation checks legal IDs, finite probabilities, complete distributions, sums, and agreement between the provider's selected choice and its argmax. Declared rounding precision permits bounded rounding error; raw values are retained. Valid usage from invalid responses is still accounted for. Provider confidence describes its decision distribution, not the probability of winning a game.

OpenTelemetry

Install the otel extra and wrap the local sink after your application configures its OpenTelemetry SDK providers/exporters:

python -m pip install 'decision-metrics[otel]'
from decision_metrics import Harness, JsonlSink
from decision_metrics.otel import OpenTelemetrySink

with JsonlSink("output/run-otel.jsonl") as local:
    sink = OpenTelemetrySink(local)  # Uses the application's tracer and meter.
    bench = Harness(backend, sink)
    # Call bench.choose and bench.record_action as above.

It emits decision.choose spans and these application-defined metrics:

Metric Unit
decision.requests, decision.actions, decision.deadline_misses count
decision.duration, decision.observation_age seconds
decision.tokens tokens, split by input/output
decision.api_cost USD, known API cost only

Metric attributes include provider, deployment, requested model, phase, and status. Decision/run IDs and packet hashes are trace attributes rather than metric dimensions. Raw prompt contents and probability vectors remain in JSONL. These names are DecisionMetrics conventions, not official OTel semantic conventions. Without an application SDK provider, the API is a no-op. Official Python instrumentation guide.

The local journal is written first. Export failures leave it intact and increment sink.export_failures. The library does not select a collector or initialize global SDK providers. Sink writing is synchronous; exporter overhead can affect the control loop, so use the same export setup across benchmark entries.

Events and summaries

Schema v1 uses run_start, decision, action, and run_outcome events. Each decision carries a UUID, common-packet SHA-256, candidate order, exact prompt, requested/returned model, deployment, timestamps, measured latency, deadline, result/error, usage, and estimated API cost. Probabilities and provider confidence remain in the result. Authentication headers and HTTP bodies from failures are not recorded. Disable prompt contents with Harness(..., record_inputs=False).

Summaries separate deployments and phases, preserve unknown usage/cost counts, exclude forced actions from model latency and requests, and reject duplicate decision/action IDs. Median and p95 use attempted requests; p95 is the nearest-rank percentile. An application outcome is whatever the application actually observed. No aggregate accuracy score is inferred from the model's own selections.

Development

git clone https://github.com/banjtheman/decision-metrics.git
cd decision-metrics
python -m pip install -e '.[otel]' opentelemetry-sdk
python -m unittest discover -s tests -v

The OpenTelemetry tests require opentelemetry-sdk; the core tests do not. Tests make no vendor calls. The project is MIT licensed.

See publishing instructions for release checks and PyPI Trusted Publishing setup.

Metadata

Release files for decision-metrics 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for decision-metrics 0.1.0
File Size Uploaded
decision_metrics-0.1.0.tar.gz 28.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for decision-metrics 0.1.0
File Interpreter ABI Platform
decision_metrics-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 49.0 kB

Release files / decision_metrics-0.1.0.tar.gz

Download URL decision_metrics-0.1.0.tar.gz
Size 28.0 kB
Tags Source
SHA-256 checksum
How to use checksums
47e6fcf1bfdf204ee95ad4ec2ec95ab73d12165aa830145aedff40dffb7fc722
BLAKE2b-256 checksum
How to use checksums
0edfaddd6f6d14e5f6754f07d882c2274548863594b5d5965ef66c0ac10ab455
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.11.8

Release files / decision_metrics-0.1.0-py3-none-any.whl

Download URL decision_metrics-0.1.0-py3-none-any.whl
Size 21.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6bbe7f58f58b07f0af8d83aed535cd8cab15ee745d019858c69e47b8e7cc35b7
BLAKE2b-256 checksum
How to use checksums
7853313a300af90c50e2a5b5969b41ffb0716f8fcae4d4cfb85dd5b068082fb8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.11.8

Release history Release notifications | RSS feed

0.2.0

2 release files

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page