Skip to main content

neens-eval — pre-prod evaluation runner (PULL + PUSH)

The thin CI glue that turns a Neens pre-prod evaluation into a one-liner. Before a change ships, replay a golden dataset's frozen prompts against your candidate agent, capture the resulting OpenTelemetry traces, let Neens score them with your existing judges, and gate the build on regressions vs a baseline.

Two models are supported. In the PULL model (below) your harness runs your agent per prompt and emits traces Neens correlates. In the PUSH model (neens eval push) you register your agent as an agent_http connection and Neens calls it per prompt itself — no harness, no OTel wiring.

This runner does not try to be a universal agent-caller. In the pull model your harness runs your agent; the runner just:

  1. fetches the run's frozen prompts,
  2. runs your command once per prompt with the Neens correlation attributes injected into the environment,
  3. starts the run, polls until scoring finishes, fetches the gate, and
  4. returns an exit code (0 = gate passed) so CI can gate on it.

Requirement (be honest about it): your agent must already be OpenTelemetry-instrumented and exporting traces to the Neens OTLP receiver. Neens is your observability tool — it correlates the traces your agent emits; it does not instrument your agent for you.

Zero runtime dependencies (stdlib only): installing it adds nothing to a CI image beyond Python 3.10+.

Documentation: Pre-prod evaluations guide · Framework quickstarts (copy-paste OTel config for LangGraph, CrewAI, OpenAI Agents SDK, Pydantic AI, Claude Agent SDK, Vercel AI SDK) · neens.ai. A TypeScript/Node build with identical flags and exit codes is published as neens-eval.

Install

pip install neens-eval
neens --version

This installs the neens console script (also runnable as python -m neens_eval.cli …). Pin the version in CI (pip install "neens-eval==X.Y.Z") so a new release never changes a gate under you.

The CI one-liner

neens eval run \
  --base-url "$NEENS_BASE_URL" \
  --api-key  "$NEENS_API_KEY" \
  --run-id   "$RUN_ID" \
  -- python -m my_agent --answer   # <-- YOUR agent command, after `--`

Everything after -- is your command. It is invoked once per frozen prompt. The exit code is the gate verdict, so a GitHub Actions / GitLab step fails automatically when the gate fails:

- name: Pre-prod eval gate
  env:
    NEENS_BASE_URL: https://neens.example.com
    NEENS_API_KEY: ${{ secrets.NEENS_PROJECT_KEY }}   # a nk_live_… project key
    # Point YOUR agent's OTel exporter at the Neens receiver:
    OTEL_EXPORTER_OTLP_ENDPOINT: https://neens.example.com
    OTEL_EXPORTER_OTLP_TRACES_ENDPOINT: https://neens.example.com/v1/traces
  run: |
    neens eval run --run-id "$RUN_ID" -- python -m my_agent --answer

Create the run in the same step

If you don't pre-create the run, --create makes one against a dataset's golden version:

neens eval run --create \
  --dataset "$DATASET_ID" --version-label "$GIT_SHA" \
  --max-regressions 0 --min-pass-rate 0.9 \
  -- python -m my_agent --answer

Which baseline? — --baseline

--max-regressions counts items the baseline passed and the candidate failed, so the baseline decides whether it means anything:

--baseline Compares each golden item with
branch:main the newest completed run on that branch over the same frozen dataset version (resolved when the run is created, then pinned)
run:<preprod_run_id> one specific prior run
prod / prod:30d production scores over a window, matched by prompt text

Default in CI (0.6.0+): with no --baseline, a run created in GitHub Actions or GitLab CI asks for branch:<the repository's default branch> (GitLab CI_DEFAULT_BRANCH; GitHub repository.default_branch from the event payload). Outside CI, or when the default branch can't be detected, the server default applies (the 7-day production window). An explicit --baseline always wins, and a gates-only run (no --dataset) gets no default. Run the job on pushes to the default branch too, so pull requests have a run to compare with. The CLI prints the baseline it got:

Created run ppr_51c0…
Baseline: run ppr_9b3e… — branch:main: newest completed run on branch 'main' for this dataset version

A baseline that matched none of the run's items compared nothing: the summary prints regressions: not compared (no baseline) plus a WARNING, and a max_regressions policy rule is treated as a missing baseline (on_missing_baseline), never as "0 regressions".

Enforce eval gates in CI — --gates active

An eval gate is a failure mode Neens already caught in production, frozen into a golden set and a judge. --gates active makes the build replay every gate that is active (enforced in CI) in the project, with YOUR agent, and fail when a gate's pass rate drops below its threshold — so a failure you fixed cannot quietly come back. Paused gates are not enforced.

neens eval run --gates active -- python -m my_agent --answer
  • --dataset becomes optional: a gates-only run replays just the gates' golden sets. Pass --dataset too to gate the release on your own golden set in the same run.
  • --gates gate-abc,gate-def enforces exactly those gates (a paused one is reported, not enforced). NEENS_GATES is the environment default for --gates.
  • Commit, branch and pull request are detected from GitHub Actions (GITHUB_SHA, the PR head commit from the event payload, GITHUB_HEAD_REF/GITHUB_REF_NAME, refs/pull/<n>/merge) and GitLab CI (CI_COMMIT_SHA, CI_MERGE_REQUEST_IID, …); --git-sha, --branch, --pr-number and --pr-url override them. Without --version-label the candidate is labelled with the short commit. They are recorded on each gate's CI history in Neens.
  • The summary prints one row per gate:
  eval gates:
    GATE                      ENFORCED  PASS RATE  THRESHOLD  SCORED  VERDICT
    FM: Wrong refund amount   yes       33%        70%        3/3     FAIL
    FM: Leaks internal notes  yes       100%       70%        4/4     PASS
  eval gates:  1/2 passed, 1 enforced gate(s) failing the build

A gate whose items produced no score is NO DATA and fails the build — measuring nothing is never a pass. A project with no active gates (and no --dataset) is a run error (exit 2), not a green build.

Local dev loop — neens eval watch

neens eval watch is run --create on every file save. It watches your source tree, re-runs the gate against your local agent after each save, and prints what changed since the previous iteration: items that newly fail (▼), newly pass (▲) or still fail, the gate verdict, and the pass-rate / avg-score / failure deltas.

neens eval watch --dataset "$DATASET_ID" --version-label feat-refunds \
  --include "*.py" --include "prompts/*" \
  -- python -m my_agent --answer
  • Every iteration is a real pre-prod run and scores every golden item with your judges, on your LLM connection. That costs judge tokens on every save. Cap a session with --max-iterations N.
  • Iteration N is labelled <label>-watch.<session>.<N> (UTC start time as the session), so every run is traceable to the save that produced it.
  • Dependency-free polling watcher: --path (repeatable, default .), --include / --exclude globs, --debounce (default 1s). .git, node_modules, __pycache__, .venv/venv, dist, build and tool caches are always skipped.
  • Runs never overlap. Saves made while an iteration runs produce exactly one follow-up.
  • --once runs a single iteration and exits with the gate code (0/1/2), for a pre-push hook. Otherwise Ctrl-C stops the agent processes, cancels the in-flight run, and exits 130.
  • A failing iteration (backend error, scoring timeout) is printed, its run is cancelled, and the watcher keeps going. A run that cannot be created at all (bad key, unknown dataset, no golden version) exits 2.
  • --json writes one JSON object per iteration to stdout.

Every other flag (--baseline, --max-regressions, --min-pass-rate, --gate-policy and its overrides, --timeout, --concurrency, --poll-*, --no-gate, connection) means the same as for neens eval run.

How each prompt reaches your command

For every item the runner sets, and then runs your command:

Channel What your agent receives
stdin the prompt input (piped in)
$NEENS_ITEM_INPUT the same prompt input
$NEENS_EVAL_RUN_ID the pre-prod run id
$NEENS_DATASET_ITEM_ID the frozen prompt id
$NEENS_VERSION_LABEL the candidate version label
$OTEL_RESOURCE_ATTRIBUTES the three correlation attrs, merged into any existing value

Read the prompt from stdin or $NEENS_ITEM_INPUT — whichever suits your harness.

Under neens sweep run every invocation also gets the arm it belongs to:

Channel What your agent receives
$NEENS_EVAL_MODEL the arm's model — route your LLM calls to it, or every arm runs the same model
$NEENS_SWEEP_ID the sweep id
$NEENS_SWEEP_ARM the arm label
$NEENS_SWEEP_REP the repetition, 1..k
$OTEL_RESOURCE_ATTRIBUTES additionally neens.sweep_id, neens.sweep_arm, neens.sweep_rep

neens_eval.sweep_context() reads these back as {"sweepId", "arm", "rep", "model"} (or None outside a sweep).

How the OTel correlation attributes get injected

Neens links an incoming trace to its pre-prod item by three OpenTelemetry attributes:

neens.eval_run_id       # which pre-prod run
neens.dataset_item_id   # which frozen prompt
neens.version_label     # the candidate version being replayed

The runner injects them without any change to your agent code via the standard OTEL_RESOURCE_ATTRIBUTES environment variable (comma-separated key=value), which every OTel SDK merges into the trace resource. If you already set OTEL_RESOURCE_ATTRIBUTES (e.g. service.name=my-agent,deployment.environment=ci), the runner merges — it keeps your keys and adds/overrides only the three neens.* ones. So every span your agent emits during that invocation is tagged for correlation automatically.

Pointing your agent's traces at Neens

Your agent's OTel exporter must send to the Neens OTLP receiver (POST /v1/traces). Set the standard OTel env vars in your CI environment (the runner passes the environment through to your command):

export OTEL_EXPORTER_OTLP_ENDPOINT=https://neens.example.com
export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=https://neens.example.com/v1/traces
# If your OTel SDK sends auth via headers, use your Neens project key:
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer $NEENS_API_KEY"

Use whichever endpoint/protocol variables your OTel SDK honors — Neens accepts OTLP/HTTP (protobuf and JSON) at /v1/traces.

PUSH model — Neens calls your agent endpoint (no harness)

The pull model above needs your harness to run your agent per prompt. If instead you register your agent as an agent_http LLM connection in Neens, Neens can call that endpoint itself — once per frozen prompt — capture each response, score it, and gate. No local command, no subprocess, no OTel wiring: you just point the SDK at the run.

neens eval push \
  --base-url "$NEENS_BASE_URL" \
  --api-key  "$NEENS_API_KEY" \
  --dataset  "$DATASET_ID" \
  --version-label "$GIT_SHA" \
  --agent-connection "$AGENT_CONNECTION_ID" \
  --baseline prod:7d --max-regressions 0 --min-pass-rate 0.9

neens eval push always creates the run (runner_mode="push") against the dataset's golden version, triggers Neens to call --agent-connection per prompt, polls until scoring finishes, fetches the gate, and returns the same exit-code contract (0 = passed, 1 = failed, 2 = error). Under the inline queue the run is already terminal on trigger, so there is nothing to poll. Use it as a CI step exactly like neens eval run:

- name: Pre-prod eval gate (push)
  env:
    NEENS_BASE_URL: https://neens.example.com
    NEENS_API_KEY: ${{ secrets.NEENS_PROJECT_KEY }}   # a nk_live_… project key
  run: |
    neens eval push \
      --dataset "$DATASET_ID" --version-label "$GIT_SHA" \
      --agent-connection "$AGENT_CONNECTION_ID" --max-regressions 0

--no-gate and --json behave as in run (the push JSON carries no per-item items, since Neens drives the calls). Flags:

neens eval push --dataset ID --version-label LABEL --agent-connection ID [options]

target:
  --dataset ID                dataset to snapshot (its golden version, unless…)
  --dataset-version ID        …an explicit version
  --version-label LABEL       candidate version label (stamped on the run)
  --agent-connection ID       registered agent_http connection Neens calls per prompt (required)
  --name NAME                 run name (defaults to the version label)
  --baseline SPEC             branch:<name> | run:<preprod_run_id> | prod | prod:<range>
  --max-regressions N         gate: fail if regressions > N
  --min-pass-rate R           gate: fail if pass rate < R (0..1)

policy (see Gate-as-code policy below):
  --gate-policy FILE          gate-as-code policy file (env NEENS_GATE_POLICY)
  --max-cost-delta-pct PCT    --max-latency-delta-pct PCT
  --min-avg-score R           --max-abs-failures N

connection:
  --base-url URL              Neens API origin         (env NEENS_BASE_URL)
  --api-key KEY               nk_live_… project key    (env NEENS_API_KEY)
  --project-id ID             X-Neens-Project-Id      (env NEENS_PROJECT_ID; usually unneeded)

execution:
  --poll-interval SECONDS     status poll cadence (default 5)
  --poll-timeout SECONDS      overall scoring budget (default 1800)
  --no-gate                   a failed gate exits 0 (a run error still exits 2)
  --json                      machine-readable result on stdout

Gate-as-code policy

The server gate (--max-regressions + --min-pass-rate) stays authoritative for its two knobs. On top of it you can commit a gate-as-code policy file that enforces far richer conditions — cost/latency/step budgets, per-metric score floors, absolute failure caps — evaluated client-side from data Neens already returns (/comparison, /metrics, run detail). Point the runner at it with --gate-policy (or $NEENS_GATE_POLICY):

neens eval run --run-id "$RUN_ID" --gate-policy neens-gate.json -- python -m my_agent

The file is JSON (or YAML, only if pyyaml is installed). This example ships in the source distribution as neens-gate.example.json — copy it into your repo as a starting point:

{
  "gate": { "max_regressions": 0, "min_pass_rate": 0.9 },
  "rules": [
    { "cost":    { "max_avg_usd": 0.05, "max_delta_pct": 20 }, "severity": "warn" },
    { "latency": { "max_avg_ms": 3000, "max_delta_pct": 25 }, "severity": "warn" },
    { "steps":   { "max_avg": 8, "max_delta_pct": 50 }, "severity": "block" },
    { "max_abs_failures": 3, "severity": "block" },
    { "max_regressions": 0, "severity": "block" },
    { "min_pass_rate": 0.9, "severity": "block" },
    { "min_avg_score": 0.8, "severity": "block" },
    { "metric": "faithfulness", "min_avg": 0.8, "severity": "block" },
    { "metric": "toxicity_safety", "min_avg": 0.9, "severity": "warn" }
  ],
  "on_missing_baseline": "warn"
}

The gate block

The two server-gate knobs. The SDK passes these to the run at create time (--create / push), so the SERVER stays authoritative for regressions/pass-rate. CLI flags (--max-regressions, --min-pass-rate) override the block. On an existing run (--run-id) it's already baked in.

Rules — dimensions and where each signal comes from

Each rule object carries exactly one dimension key plus an optional "severity" ("block" default, or "warn").

Dimension Threshold keys Backend signal
cost max_avg_usd, max_delta, max_delta_pct avg from comparison.traits.candidate.avg_cost; delta/pct from comparison.deltas.cost. avg_cost averages only the sessions Neens could price, and is null when a side ran entirely on models with no rate — every cost rule is then reported skipped (unpriced) with the models named, never silently passed and never mistaken for a missing baseline. Set the rate in Settings → Model pricing.
latency max_avg_ms, max_delta_ms, max_delta_pct avg from traits.candidate.avg_latency; delta/pct from deltas.latencyMs
steps max_avg, max_delta, max_delta_pct avg from traits.candidate.avg_steps; delta/pct from deltas.steps
max_abs_failures int run.aggregate.failed — ABSOLUTE candidate failures
max_regressions int len(comparison.regressions) (fallback run.aggregate.regressions)
min_pass_rate float run.aggregate.passRate (fallback comparison.candidate.passRate)
min_avg_score float run.aggregate.avgScore
metric { "metric": "<key>", "min_avg"?, "max_avg"? } the /metrics rollup's avgScore for that metricKey

% deltas aren't returned by the backend — the SDK computes delta / before * 100 client-side, guarding before ∈ {null, 0}.

"New failures vs baseline" == regressions. There is no separate backend "new failure" signal, so the SDK doesn't invent one: use max_abs_failures for an absolute failure cap, max_regressions for the baseline-relative one.

metric rules use the NORMALIZED score (0..1, higher = better) — the /metrics avgScore is already normalized so higher is always better, safety metrics included. So a safety floor is a min_avg (e.g. toxicity_safety ≥ 0.9), never a raw ceiling.

on_missing_baseline

block | warn | pass (default warn). Governs baseline-relative rules (any *_delta / *_delta_pct budget, and max_regressions) when the baseline cohort is absent (the delta's before/after are null, or no baseline traits). A baseline that matched zero of the run's items (comparison.baseline.matchedItems == 0) is absent too, so max_regressions is never read as "0 regressions" when nothing was compared. Absolute rules (max_avg_*, max_abs_failures, min_pass_rate, min_avg_score, metric floors) always evaluate — they need no baseline.

  • block → a baseline-relative rule keeps its own severity (a block rule still fails the build).
  • warn → it's reported as a warning (never blocks).
  • pass → it's skipped entirely.

Block vs warn tiers

  • block (default) — a failing rule blocks the build (contributes to CI exit 1).
  • warn — a failing rule is reported but never blocks (exit stays 0 unless another block rule or the server gate fails).

The policy only ever ADDS a blocking condition on top of the server gate. With no policy, behavior is identical to before (server-gate only).

The server gate speaks cost too

The policy's top-level gate block — the one handed to create_run at run-create time — accepts the same cost family, spelled identically to the rules[].cost dimension, so a budget can live server-side where every consumer of GET /preprod-evals/{id}/gate sees it (the UI, another CI job, a model sweep), not just this CLI process:

{
  "gate": { "max_regressions": 0, "min_pass_rate": 0.9,
            "cost": { "max_avg_usd": 0.004, "max_delta_pct": -50 } },
  "rules": [ { "cost": { "max_avg_usd": 0.05 }, "severity": "warn" } ]
}

Two evaluation sites, one vocabulary. rules[].cost is evaluated client-side by this runner against the /comparison payload — it carries a severity and honours on_missing_baseline. gate.cost is evaluated server-side and is stored with the run; it has no severity, and a breach fails the gate. The same three keys are accepted in both places (an unknown one inside gate.cost is a GatePolicyError at parse time), and the SDK omits cost from the posted gate entirely when the policy declares none — it never sends "cost": null, which would fingerprint differently from an absent key server-side and void in-flight model sweeps.

The gate response then carries a costRules array ({rule, status, observed, threshold, reason}), absent when the gate declares no cost family. status is pass|fail|skipped, and only fail blocks: a run whose sessions all ran on models with no price reports skipped, naming the unpriced models — never a pass. Reading a null average as "$0, therefore under budget" is precisely the bug per-model pricing exists to kill, and an absent baseline (a delta with nothing to compare against) is reported as its own, differently-worded skip.

Override flags

Beyond --gate-policy, a few flags inject/override block rules (a flag replaces any same-dimension rule from the file):

--max-cost-delta-pct PCT     block if candidate avg cost rises more than PCT% vs baseline
--max-latency-delta-pct PCT  block if candidate avg latency rises more than PCT% vs baseline
--min-avg-score R            block if the run's aggregate avg score is below R (0..1)
--max-abs-failures N         block if absolute candidate failures exceed N

With neither a policy file nor any of these flags, policy=None — today's behavior exactly.

CI snippet

- name: Pre-prod eval gate (gate-as-code)
  env:
    NEENS_BASE_URL: https://neens.example.com
    NEENS_API_KEY: ${{ secrets.NEENS_PROJECT_KEY }}
  run: |
    neens eval run --run-id "$RUN_ID" \
      --gate-policy neens-gate.json \
      -- python -m my_agent --answer

--json additionally emits a policyEval object (per-rule dimension/severity/status/observed/ threshold/reason, plus blocked/warned) for a CI step to parse.

The build verdict is reported back to Neens

After neens eval run / neens eval push decides its exit code, it reports the build verdict to Neens (POST /preprod-evals/{id}/ci-result): passed, warned (a warn rule fired, or --no-gate let a breach through), blocked (exit 1) or errored (exit 2), with the exit code and every policy rule's rule/observed/threshold/severity/outcome/message, the policy file's name and SHA-256, and the SDK name + version. The run page and the Eval Gates strip then match the CI check — e.g. "blocked by policy: regressions 4 > 0 (all enforced gates passed)" — instead of showing only the server gate.

  • Best-effort. A failed report never changes the exit code; the CLI prints one line to stderr: neens: warning: could not report the build verdict to Neens (…); the exit code is unaffected.
  • No policy? The verdict still reflects the server gate and the exit status (rules: []).
  • --json carries buildVerdict: {"verdict": "blocked", "reported": true}.
  • neens eval watch (a local loop) never reports. Library callers can opt out with run_eval(..., report_ci_result=False).
  • A CI run created by an older CLI (< 0.6.0) shows build verdict unknown in Neens — never passed.

Multi-model sweeps — neens sweep

A sweep replays ONE frozen golden dataset version against N candidate model arms, k times each, under identical conditions (same items, same judges, same k, same gate) and reports pass^k per arm. It answers which model should we ship, so the CI story is estimate, then decide:

# 1. What would this cost? Writes nothing, spends nothing.
neens sweep preview \
  --dataset-id ds_checkout \
  --arm "sonnet=conn_sonnet" --arm "haiku=conn_haiku" --arm "local=conn_ollama" \
  --pass-k 3

# 2. Launch it, and wait for the per-arm verdict.
neens sweep start \
  --name "nightly model bake-off" --version-label "$GIT_SHA" \
  --dataset-id ds_checkout \
  --arm "sonnet=conn_sonnet" --arm "haiku=conn_haiku" \
  --pass-k 3 --budget-usd 10 --wait --json

# 3. Or poll it later.
neens sweep get msw-4f2c91ab0d3e

# 4. Get the ANSWER: which arm should we ship, and does it differ per agent?
neens sweep decide --sweep-id msw-4f2c91ab0d3e --bar 0.9

Each --arm is LABEL=CONNECTION_ID, where the connection is a registered agent_http LLM connection whose endpoint runs that model. Two arms on the same connection is rejected — that is the same model twice, not a comparison. Give one of --dataset-id (its frozen golden version), --dataset-version-id, or --scenario-suite-id.

neens sweep run — runner mode (no agent endpoint)

sweep start needs one publicly reachable agent_http endpoint per model, because Neens calls it. sweep run is the pull-mode twin of neens eval run: the CLI runs your own agent command once per arm x repetition x golden item, so it works on a laptop or in CI with no public URL and no tunnel.

neens sweep run --create --name "model choice" --dataset ds_checkout --version-label "$GIT_SHA" \
  --arm haiku=claude-haiku-4-5 --arm sonnet=claude-sonnet-4-6 --arm oss=openai/gpt-oss-20b \
  --k 3 --concurrency 2 --budget-usd 5 \
  -- python -m my_agent.eval_item

Here --arm is LABEL=MODEL, not LABEL=CONNECTION_ID. The model id is handed to your command as $NEENS_EVAL_MODEL; your agent must route its LLM calls to it. If every arm's traces show the same observed model, the verdict carries a prominent WARNING [arms_same_observed_model] — the agent is probably ignoring $NEENS_EVAL_MODEL.

What it does:

  1. Creates a runner sweep (--create) — or attaches to one created elsewhere (MCP, API) with --sweep-id ID. A push sweep is refused. The sweep id is printed first, so a CI log can link it.
  2. Runs your command for every arm x rep x item with the eval run environment plus $NEENS_EVAL_MODEL / $NEENS_SWEEP_ID / $NEENS_SWEEP_ARM / $NEENS_SWEEP_REP. Work is interleaved — rep 1 of every arm before rep 2 of any, arms round-robin within a rep — under one global --concurrency, so provider load and rate-limit drift never land on one arm. Each arm x rep's run is started as soon as all its items have executed; a transient error on that start (network, 5xx, 429) is retried up to 3 times with backoff.
  3. Polls until the sweep is scored, then prints the same per-arm table and verdict as sweep decide (pass rate with n and CI, pass^k, cost per case, regressions) plus any warnings.

Flags: --create --name --version-label --dataset-id/--dataset --dataset-version-id --scenario-suite-id --arm LABEL=MODEL --pass-k/--k --judge-deployment-id/--judge --budget-usd (create), --sweep-id (attach), --concurrency N (default 1), --timeout S (per invocation), --bar, --baseline-arm, --poll-interval (default 10, must be > 0), --poll-timeout (default 3600), --json (adds an invocations list, and failedStarts when a run could not be started), plus the connection flags.

Exit code sweep run situation
0 An arm cleared the bar — the decision is made.
1 The sweep was scored and no arm clears the bar.
2 Run error: void / failed / cancelled sweep, poll timeout, API error (including a budget refusal), a push sweep, or every agent invocation failed (the sweep is then cancelled). Also a child run that still could not be started after the retries: the other runs are driven to the end, the sweep is not cancelled, and the log prints the neens sweep run --sweep-id ID -- <command> that resumes it.
130 Interrupted (Ctrl-C or SIGTERM): the agent processes are killed and the sweep is cancelled server-side, so it does not sit waiting for traces.

A single failing invocation behaves exactly as in eval run: it is logged, its run is still started, and the judges score what arrived.

In GitHub Actions — no tunnel, no inbound endpoint:

- name: Model sweep (pass^k per model, no public endpoint)
  env:
    NEENS_BASE_URL: https://neens.example.com
    NEENS_API_KEY: ${{ secrets.NEENS_PROJECT_KEY }}   # a nk_live_… project key
    # Point YOUR agent's OTel exporter at the Neens receiver:
    OTEL_EXPORTER_OTLP_ENDPOINT: https://neens.example.com
    OTEL_EXPORTER_OTLP_TRACES_ENDPOINT: https://neens.example.com/v1/traces
  run: |
    pip install neens-eval
    neens sweep run --create --name "model choice ${{ github.sha }}" \
      --dataset "${{ vars.NEENS_GOLDEN_DATASET }}" --version-label "${{ github.sha }}" \
      --arm haiku=claude-haiku-4-5 --arm sonnet=claude-sonnet-4-6 \
      --k 3 --concurrency 2 --budget-usd 5 \
      -- python -m my_agent.eval_item

The step fails (exit 1) when no model clears the bar and errors (exit 2) on a void or failed sweep; cancelling the job (SIGTERM) cancels the sweep too.

neens sweep decide — the verdict

decide reads GET /model-sweeps/{id}/comparison and prints the server's verdict: the winning arm and why, a per-arm table (pass rate with its sample size and confidence interval, pass^k, cost per case, regressions), and one verdict per agent — because "Haiku is good enough" is routinely true for one agent and false for another.

Sweep msw-4f2c91ab0d3e — status=completed
  VERDICT: haiku clears the 90% bar at $0.000600 per case — 81% cheaper than sonnet.
  bar 90.0% (source: gate) | baseline: sonnet (incumbent) | 40 items x k=3
  ARM     MODEL              PASS RATE                    PASS^K    COST/CASE  REGRESSIONS
  sonnet  claude-sonnet-4-6  95.0% (n=40, CI 83.0-99.0%)  3/3 PASS  $0.003100  0
  haiku   claude-haiku-4-5   92.5% (n=40, CI 80.0-98.0%)  3/3 PASS  $0.000600  3
  local   gpt-oss:20b        no data                      0/3 —     unpriced   —

Flags: --bar 0.9 (the pass rate an arm must clear — omit it and the sweep's pinned gate.min_pass_rate is used, else the deployment default; the source is always reported), --baseline-arm ID (the incumbent the savings are measured against — defaults to the arm at position 0, never the best performer, because a baseline chosen after seeing the results is not a baseline), --require-winner, --wait, --json.

Two properties this command will not violate:

  • The ranking is the server's. decide renders verdict/agents[].verdict verbatim — it never computes a second ranking, because two rankings that can disagree is worse than one.
  • An unpriced arm never wins "cheapest." No rate means unpriced, never $0.00: it is still shown, still judged against the bar, and excluded from the cost ranking with a named reason.

Sweep exit codes

Exit code Situation
0 The sweep completed — including when an arm missed pass^k. A sweep is a model-selection decision, not a gate.
0 Launched without --wait: the verdict is not in yet.
1 --require-all-arms was passed and some arm did not hit pass^k (or the sweep has no verdict yet).
2 The sweep is void — the arms did not run under identical conditions, so the comparison is refused and voidReason is printed. A result we cannot trust must never look like a pass.
2 The sweep failed, or a transport/API error (including a budget refusal, which is a 422).

For sweep decide the same table holds, with --require-winner in place of --require-all-arms: a comparison that names no clearing arm still exits 0 (naming a winner is a decision, not a gate) unless --require-winner is passed, and a void / comparable: false comparison is exit 2 before any verdict is read — a void sweep still has per-arm numbers, and reading them first is exactly how an untrustworthy comparison gets to look like a pass.

--require-all-arms (and --require-winner) implies --wait: you cannot assert every arm hit pass^k — or require a winner — without having seen every arm finish.

An estimate whose arms include a model with no price in your tenant's table is printed as partial, naming the unpriced models — the total covers the priced arms only and is a floor, never the cost. Set a rate under Settings → Model pricing for a complete figure.

Gate exit-code contract

Exit code Meaning
0 Gate passed (or --no-gate was set and the run did not error)
1 Gate failed (regressions/pass-rate breached the configured gate, or an enforced eval gate failed or scored nothing)
2 Run error (no items, backend error, poll timeout, bad arguments)

--no-gate still fetches and prints the gate, but a failed gate exits 0 (report without blocking CI). A run error still exits 2: --no-gate waives the gate, not a broken run. --json emits a machine-readable result to stdout (human logs go to stderr) for a CI step to parse. neens eval watch --once follows this table; a continuous watch exits 130 on Ctrl-C, 2 when no run can be created, or the last iteration's code at --max-iterations.

Flags

neens eval run [target] [policy] [connection] [execution] -- <your agent command>
neens eval watch --dataset ID [--version-label LABEL] [--path P] [--include GLOB] [--exclude GLOB]
                 [--debounce S] [--max-iterations N] [--once] [policy] [connection] [execution]
                 -- <your agent command>          (see "Local dev loop" above)
neens --version            print the installed version

target:
  --run-id ID                 drive an existing run
  --create                    create a run first (needs --dataset + --version-label)
    --dataset ID              dataset to snapshot (its golden version, unless…)
    --dataset-version-id ID   …an explicit version
    --version-label LABEL     candidate version label (also tags traces)
    --name NAME               run name (defaults to the version label)
    --baseline SPEC           branch:<name> | run:<preprod_run_id> | prod | prod:<range>
                              (default in CI: branch:<the repo's default branch>)
    --max-regressions N       gate: fail if regressions > N
    --min-pass-rate R         gate: fail if pass rate < R (0..1)

eval gates (see "Enforce eval gates in CI" above):
  --gates active|ID,ID        replay + enforce eval gates; implies creating a run (env NEENS_GATES)
  --git-sha SHA  --branch NAME  --pr-number N  --pr-url URL
                              candidate provenance (default: detected from GitHub Actions / GitLab CI)

policy (see Gate-as-code policy above):
  --gate-policy FILE          gate-as-code policy file (env NEENS_GATE_POLICY)
  --max-cost-delta-pct PCT    --max-latency-delta-pct PCT
  --min-avg-score R           --max-abs-failures N

connection:
  --base-url URL              Neens API origin         (env NEENS_BASE_URL)
  --api-key KEY               nk_live_… project key    (env NEENS_API_KEY)
  --project-id ID             X-Neens-Project-Id      (env NEENS_PROJECT_ID; usually unneeded)

execution:
  --timeout SECONDS           per-item command timeout
  --concurrency N             parallel item invocations (default 1)
  --poll-interval SECONDS     status poll cadence (default 5)
  --poll-timeout SECONDS      overall scoring budget (default 1800)
  --no-gate                   a failed gate exits 0 (a run error still exits 2)
  --json                      machine-readable result on stdout

Library use

PULL model — drive your own agent over an existing run's prompts:

from neens_eval import NeensEvalClient, GatePolicy, run_eval

client = NeensEvalClient("https://neens.example.com", api_key="nk_live_...")
policy = GatePolicy.from_file("neens-gate.json")   # or GatePolicy.from_dict({...})
outcome = run_eval(
    "ppr_123",
    ["python", "-m", "my_agent", "--answer"],
    client=client,
    per_item_timeout=120,
    concurrency=4,
    policy=policy,                                    # client-side gate-as-code (optional)
)
if outcome.policy_eval:
    print("blocked:", outcome.policy_eval.blocked, outcome.policy_eval.reasons())
raise SystemExit(outcome.exit_code)

PUSH model — create + invoke a run against a registered agent endpoint (Neens does the calling):

from neens_eval import NeensEvalClient, run_push_eval

client = NeensEvalClient("https://neens.example.com", api_key="nk_live_...")
outcome = run_push_eval(
    client=client,
    dataset_id="ds_1",
    version_label="v2",
    agent_connection_id="conn_abc",          # a registered agent_http connection
    baseline={"kind": "prod_window", "range": "7d"},
    gate={"max_regressions": 0, "min_pass_rate": 0.9},
)
raise SystemExit(outcome.exit_code)

License

Apache-2.0 — see the LICENSE file shipped with this package.

Metadata

Release files for neens-eval 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for neens-eval 0.6.0
File Size Uploaded
neens_eval-0.6.0.tar.gz 150.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for neens-eval 0.6.0
File Interpreter ABI Platform
neens_eval-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 252.1 kB

Release files / neens_eval-0.6.0.tar.gz

Download URL neens_eval-0.6.0.tar.gz
Size 150.7 kB
Tags Source
SHA-256 checksum
How to use checksums
a0dc9e00bf5d10d149177cdd8cc47ebae8179231f974360e71711053642eb1c8
BLAKE2b-256 checksum
How to use checksums
ab530cad05b53cd021e698cb70ab67a439dbdbaaefad95e4460e9cb35a70caaa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / neens_eval-0.6.0-py3-none-any.whl

Download URL neens_eval-0.6.0-py3-none-any.whl
Size 101.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
fc52d84549ec8e0eede04caa239c6f4d4022cc85b6d7cd895f4514ab7c9f024e
BLAKE2b-256 checksum
How to use checksums
b4843e2a2e16a57de407d3da6f18ad45f72dbda2316775eb4686324fbd7bc4c2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.6.0 This release

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.1

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page