DecisionMetrics
DecisionMetrics is a Python library for recording traces and metrics from decision models such as Jev, Cloudflare Clef, Perplexity Decisions, and Strands Decider.
A decision model chooses from a set of options supplied by your application. Those options might be support queues, agent tools, or game movements. DecisionMetrics records the inputs, available options, selected choice, and probabilities when the model provides them. It measures response latency, token usage, estimated API cost, and deadline misses, then links each decision to the action your application applied, rejected, or skipped.
Use these records to debug an application, compare models and local or hosted deployments, or replay the same inputs across providers. Save decision traces to JSONL and export spans and metrics through your application's OpenTelemetry setup.
Python 3.11+. The core uses only the standard library. OpenTelemetry is optional.
Install and try it offline
With your Python 3.11+ environment active:
python -m pip install decision-metrics
decision-metrics demo --output output/demo.jsonl
decision-metrics summary output/demo.jsonl
For OpenTelemetry support, install decision-metrics[otel]. This is an alpha
release.
The demo is synthetic and makes no model or game calls. Output files are created exclusively: use a fresh name for each run.
Application examples
| Application | Model decision | Action recorded |
|---|---|---|
| Support-ticket routing | Choose billing, technical support, or account access | Assign the ticket to an in-memory queue |
| Agent tool selection | Choose documentation search, a status check, or a follow-up question | Execute a local example tool and record its result |
| Game controller | Choose from available movements or upgrades | Record whether the game accepted the action |
Clone the repository to run the application examples:
git clone https://github.com/banjtheman/decision-metrics.git
cd decision-metrics
python examples/support_routing.py --output output/support-demo.jsonl
python examples/agent_tools.py --output output/tools-demo.jsonl
decision-metrics summary output/support-demo.jsonl output/tools-demo.jsonl
Both scripts use synthetic inputs and deterministic offline rules by default.
To measure a decision model instead, set its credentials and select a configured
backend. For example, with JEV_KEY set:
python examples/support_routing.py --backend jev --output output/support-jev.jsonl
python examples/agent_tools.py --backend jev --output output/tools-jev.jsonl
The actions stay within the local examples. The journals record the selected model's responses, timings, available token usage, and estimated API cost. See the examples guide for configuration and recorded fields.
Record a decision and its action
This example records a game controller's movement decision and whether the game
accepted it. Download the sample provider configuration
to examples/providers.toml and start a local Strands decision server before
running it. The same recording pattern applies to other applications and models.
from decision_metrics import ChoiceRequest, Harness, JsonlSink
from decision_metrics.providers import load_backend
backend = load_backend("examples/providers.toml", "strands-local")
try:
with JsonlSink("output/run-01.jsonl") as sink:
bench = Harness(backend, sink, metadata={"application": "my_game"})
request = ChoiceRequest(
state={"health": 5, "threat": "east"},
options={"west": {"clearance": 150}, "east": {"clearance": 10}},
instructions="Choose a movement that helps the player survive.",
)
decision = bench.choose(request, phase="combat", deadline_ms=850)
if decision.expired():
bench.record_action(decision, "expired", reason="combat_deadline")
else:
# Replace with the application's authoritative action execution.
accepted = game.apply(decision.result.choice)
bench.record_action(decision, "applied" if accepted else "rejected")
bench.outcome("session_complete", metrics={"waves_completed": 9})
finally:
backend.close()
game.apply is an application hook, not a library function. Pass the observation's
actual time.monotonic() timestamp as observed_at when it predates the call.
Action timestamps measure when the application records the result, which can be
after a bridge acknowledgement; they do not reveal the exact simulation tick.
Your application supplies the available actions and executes the selected one.
Record the execution result separately so the trace shows whether the decision
led to an action. Action statuses are applied, rejected, expired, skipped,
and unknown for delivery without a reliable acknowledgement.
Providers
examples/providers.toml contains separate entries for
Jev, Clef, Clef-flash, Perplexity, Strands local/remote, Kev, and meraGPT.
Download that file to examples/providers.toml before using the examples below,
or clone the repository to get all sample files. They are not installed with the
library.
| Transport | Configuration and credential environment |
|---|---|
| Jev System One | provider = "jev"; JEV_KEY |
| Cloudflare Workers AI | provider = "clef" or "clef-flash"; CLOUDFLARE_ACCOUNT_ID, CLOUDFLARE_AUTH_TOKEN |
| Perplexity Decisions | provider = "perplexity"; PERPLEXITY_API_KEY |
| Local Strands / Kev | provider = "strands" or "kev"; local HTTP endpoint |
| meraGPT System One | provider = "meragpt"; MERAGPT_API_KEY; 2–10 options |
| Another compatible server | provider = "systemone"; explicit endpoint and model |
Hosted prices are dated defaults, with a source stored beside each rate. Override them in TOML when they change. Local entries record zero API charges while excluding hardware/energy costs. Remote Strands is a configuration template; replace its endpoint and record the actual server hardware before using it. No remote deployment is created by this library.
Use the same checkpoint for both Strands entries, and start its server with strict window validation. The v19 checkpoint has a 4,096-token window; the server's default truncation would change the benchmark input. The example's checkpoint, device, and strict-window fields describe the intended server configuration; they are not a runtime attestation. Strands serving instructions.
The HTTP transports are implemented and tested against fixtures and a local HTTP server. Live vendor compatibility has not yet been validated.
Freeze inputs and replay them
A packet is one JSON object per line with state, ordered options,
instructions, and optional phase and case_id. The sample cases are synthetic.
decision-metrics replay examples/packets.jsonl \
--config examples/providers.toml --backend strands-local \
--limit 3 --max-cost 0.50 --output output/strands-local-replay.jsonl
decision-metrics summary output/strands-local-replay.jsonl
Replay is sequential and records actions as skipped. It measures response validity, usage, latency, and distributions. It supplies no ground-truth action labels or gameplay win rate. The cost limit uses known API usage and can be crossed by the final request; unknown usage/prices remain explicit. Separate request limits still apply.
Use separate entries for deployment comparisons. Local response speed is part of the result; a faster local model can make more timely decisions in a live game. A second experiment with a shared request cadence can help isolate choice quality from deployment speed. Record cadence, hardware, precision, checkpoint, and server settings in run metadata.
For full episodes, set call and billing limits high enough to finish. An identical call ceiling can stop a faster controller earlier; budget stops belong in a separate outcome category from game deaths.
Bring another backend
Implement info, choose(request) -> DecisionResult, and close():
from decision_metrics import BackendInfo, DecisionResult
class MyBackend:
info = BackendInfo("my_backend", "checkpoint-v1", "local_cpu")
def choose(self, request):
selected_id = my_model.select(request.state, request.options, request.instructions)
return DecisionResult(choice=selected_id)
def close(self):
pass
Return Usage when available. Missing probabilities, confidence, and usage stay
None; a generative or heuristic backend does not need to fabricate them.
Native HTTP decision adapters require the promised complete option distribution.
One legal option is recorded as forced without calling the backend.
Validation checks legal IDs, finite probabilities, complete distributions, sums, and agreement between the provider's selected choice and its argmax. Declared rounding precision permits bounded rounding error; raw values are retained. Valid usage from invalid responses is still accounted for. Provider confidence describes its decision distribution, not the probability of winning a game.
OpenTelemetry
Install the otel extra and wrap the local sink after your application configures
its OpenTelemetry SDK providers/exporters:
python -m pip install 'decision-metrics[otel]'
from decision_metrics import Harness, JsonlSink
from decision_metrics.otel import OpenTelemetrySink
with JsonlSink("output/run-otel.jsonl") as local:
sink = OpenTelemetrySink(local) # Uses the application's tracer and meter.
bench = Harness(backend, sink)
# Call bench.choose and bench.record_action as above.
It emits decision.choose spans and these application-defined metrics:
| Metric | Unit |
|---|---|
decision.requests, decision.actions, decision.deadline_misses |
count |
decision.duration, decision.observation_age |
seconds |
decision.tokens |
tokens, split by input/output |
decision.api_cost |
USD, known API cost only |
Metric attributes include provider, deployment, requested model, phase, and status. Decision/run IDs and packet hashes are trace attributes rather than metric dimensions. Raw prompt contents and probability vectors remain in JSONL. These names are DecisionMetrics conventions, not official OTel semantic conventions. Without an application SDK provider, the API is a no-op. Official Python instrumentation guide.
The local journal is written first. Export failures leave it intact and increment
sink.export_failures. The library does not select a collector or initialize
global SDK providers. Sink writing is synchronous; exporter overhead can affect
the control loop, so use the same export setup across benchmark entries.
Events and summaries
Schema v1 uses run_start, decision, action, and run_outcome events. Each
decision carries a UUID, common-packet SHA-256, candidate order, exact prompt,
requested/returned model, deployment, timestamps, measured latency, deadline,
result/error, usage, and estimated API cost. Probabilities and provider confidence
remain in the result. Authentication headers and HTTP bodies from failures are
not recorded. Disable prompt contents with Harness(..., record_inputs=False).
Summaries separate deployments and phases, preserve unknown usage/cost counts, exclude forced actions from model latency and requests, and reject duplicate decision/action IDs. Median and p95 use attempted requests; p95 is the nearest-rank percentile. An application outcome is whatever the application actually observed. No aggregate accuracy score is inferred from the model's own selections.
Development
git clone https://github.com/banjtheman/decision-metrics.git
cd decision-metrics
python -m pip install -e '.[otel]' opentelemetry-sdk
python -m unittest discover -s tests -v
The OpenTelemetry tests require opentelemetry-sdk; the core tests do not. Tests
make no vendor calls. The project is MIT licensed.
See publishing instructions for release checks and PyPI Trusted Publishing setup.
Metadata
Release files for decision-metrics 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| decision_metrics-0.1.1.tar.gz | 34.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| decision_metrics-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 56.0 kB
Release files / decision_metrics-0.1.1.tar.gz
| Download URL | decision_metrics-0.1.1.tar.gz |
|---|---|
| Size | 34.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a14da9d81b6fa1028ebc80b1fd29b67f293b338ae579631613ab94ef0e23bac8
|
|
BLAKE2b-256 checksum How to use checksums |
ae968c9eb5897c2f477fbeae57d1471b376c523f4fb206d3b0c23e02b5396e4b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.11.8
|
Release files / decision_metrics-0.1.1-py3-none-any.whl
| Download URL | decision_metrics-0.1.1-py3-none-any.whl |
|---|---|
| Size | 21.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1c70f76faae7701c46883e707960476332fe55edb53c656a32ca723dabfae0a3
|
|
BLAKE2b-256 checksum How to use checksums |
9ee5ffbbf6a42257cc88648b0f1c625777858804f8b55999c9c563de1e99fed1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.11.8
|