Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Verdict Python SDK

PyPI distribution: cognifity-verdict. Python import: verdict.

Install and start the loopback-only product UI:

python -m pip install cognifity-verdict
verdict

The initial setup page can approve and rescan local Claude Code/Codex histories, import supported telemetry files, show the SDK snippet, or open an existing store. Local histories are persisted as typed AgentRun/AgentTurn/ AgentEvent evidence; they are not converted into fake provider LLM traces. Those records use normalized source/run/turn/event storage. A model-call event links to the genuine Trace, which remains the only owner of LLM content and the complete provider-call record used for evaluation and drift. Events may retain bounded scalar summaries but never copy prompt, response, or raw-message content. Run timelines are read in bounded pages instead of from one growing serialized row. Bounded redacted content retention is on by default; metadata-only capture is an explicit SDK/programmatic override, not a shortcut in local setup. After capture, the same page becomes Data sources, reports the evidence sources observed in the current store, and requires fresh in-process path approval for manual edits or rescans. Agent Run and LLM Trace totals remain visibly separate. The server requires a successful preview of the exact local or historical paths before it accepts the corresponding write. An explicitly saved daily schedule intentionally retains source paths in the local control store for verdict-service; no OS scheduler is installed. Claude Code history fragments that share one provider response identity are coalesced into one linked Trace; later response text completes that Trace while genuine tool-only calls remain textless. Source-identified Claude sidechains and Codex child histories are stored as separate child runs. Agent Run views show bounded final-response previews and source-reported per-turn tokens when available. Text is redacted before the preview cutoff, and incomplete token components remain explicitly partial; Claude totals appear only after terminal, complete response usage. Those source activity totals are not a quality ranking: latency, price, judging, and model comparisons still use genuine Trace records only.

For automation, the equivalent commands are:

verdict-import local --storage sqlite:///./verdict.db
verdict-dashboard --storage sqlite:///./verdict.db
verdict-monitor run --storage sqlite:///./verdict.db
verdict-service --storage sqlite:///./verdict.db --once

Instrumented applications can add the execution structure surrounding those LLM calls without creating a second content record:

import verdict

verdict.init(storage="sqlite:///./verdict.db", service_name="support-api")
with verdict.agent_run(name="support-agent", session_id=session_id) as run:
    with run.turn(user_input=user_message) as turn:
        with turn.tool("lookup_order", arguments={"order_id": order_id}) as tool:
            order = lookup_order(order_id)
            tool.set_output({"found": order is not None})
        answer = respond(user_message, order)
        turn.set_output(answer)
    run.record_business_outcome("resolved", True)

Supported provider calls inside the turn automatically create linked genuine Trace records. The SDK also exposes typed helpers for instructions, context, commands, tests, artifacts, retries, handoffs, feedback, and outcomes. Sync and async context managers share the same API contract. With content capture on, a turn whose caller does not provide an output records missing response evidence; metadata-only capture records that content as not captured.

Stable read port for optional packages

Trusted same-process extensions can read one authorized Agent Run without depending on Verdict tables or dashboard queries:

from verdict.read_port import StorageVerdictReadPort
from verdict.storage import SQLiteStorage

storage = SQLiteStorage("./verdict.db")
port = StorageVerdictReadPort(storage)

# The host must authorize tenant_id before making this call.
run = port.get_agent_run(tenant_id=authorized_tenant_id, run_id=selected_run_id)

The immutable V1 result contains run timing/status, at most 16 exact model-call event/Trace references, and at most four deterministic findings with explicit totals and truncation flags. It excludes prompts, responses, provider/model display names, arbitrary event attributes, and analysis prose. Returned IDs remain protected Verdict metadata.

This is a Python composition boundary, not an HTTP service or authorization layer. tenant_id scopes the lookup but does not prove the caller may read that tenant. Runs above 1,000 turns or 1,500 events fail closed after the exact store lookup. The port cannot identify a LiteLLM deployment or GPU; a dependent package needs a separate propagated identity source and must otherwise report the correlation as unmapped. See ADR-013.

To keep application processes off the database, select the bounded local file transport and later import its process-owned JSONL segments through canonical storage:

verdict.init(transport="file", spool_directory="./verdict-capture")
verdict-import agent-file ./verdict-capture --storage sqlite:///./verdict.db

The same file transport carries provider Traces, Agent evidence, and manual spans. Use one spool directory per producer process. Manual import reads complete records from active, sealed, or rejected segments. For central PostgreSQL Agent ingestion, verdict-collector provides authenticated bounded batches and durable idempotent acknowledgements, while verdict-shipper uploads complete segment prefixes and deletes only sealed, fully accepted segments. Collector-rejected segments remain locally recoverable. Standalone Trace and Span records still require direct or manual file import. Use verdict-shipper --spool-directory ./verdict-capture --status --json to inspect local backlog without the collector or its key. Completed appends bypass Python userspace buffering but are not fsync-ed. If an Agent evidence stream fails, later provider calls fall back to standalone Trace records. Quota or write failures increment the process-local capture.dropped_records metric and produce a bounded warning.

The Monitor UI previews an immutable count-based (older 80% / newer 20% by default) or explicit-date policy before activation. Each metric has its own eligible denominator, Fisher's exact p-value, Benjamini-Hochberg adjustment, and effect-size gate. With provider/model or reviewed-cluster grouping, Verdict computes separate group-by-metric comparisons and adjusts across the complete tested family; it does not pool the selected groups. The Measurement selector defaults to deterministic trace checks and can explicitly add stored PASS/FAIL results from one complete evaluator identity without running or paying for a judge. That evaluator fingerprint and its expected dimensions become immutable policy inputs. A reviewed-cluster policy also pins its registry version and completes projection of eligible new traces through that fixed version before saving monitor membership, without refitting it. Unfinished bounded projection is reported as projection_pending and writes no monitor snapshot. UNCLEAR, missing, and error states remain outside the PASS/FAIL denominator and are shown as coverage. Unassigned and new groups do not wait on judgments they cannot use in like-for-like tests. Ongoing cohorts are prospective and non-overlapping; late arrivals are counted and included in the next open cohort rather than silently discarded. An evaluator-backed cohort fixes membership and waits for stored results needed by its tested metric cells before comparing; changed or deleted pending evidence requires a new reviewed preview. An activated policy starts with an empty prospective bucket, and repeated looks use a summable quadratic alpha-spending rule. insufficient and reference_stale are first-class results; unassigned or new groups are reported rather than pooled into a comparison. verdict-monitor is a one-shot idempotent runner. It and verdict-service use the same stored evaluator, dimensions, grouping version, and trace selection as the dashboard. verdict-service executes the dashboard's saved schedule once or continuously. The approved baseline membership and normalized metric counts are immutable. Grouped monitors are limited to 250 distinct groups. Older stored monitors without frozen cohort facts, or without evaluator-finalization state when an evaluator is selected, require a new reviewed preview before execution.

The dashboard reads key-free findings from immutable analysis snapshots rather than recomputing them on every page load. It reports provider outcome, evaluation status, finding severity, and drift comparison independently. not evaluated and judge error are explicit Trace states. A prospective monitor distinguishes traffic collection from a full cohort awaiting selected evaluator results. The dashboard has six top-level workspaces: Overview, Explore, Evaluate, Monitor, Report, and Settings. Report summarizes non-judge application calls over the last 30 UTC calendar days by default, with 7-day, 90-day, and all-time choices. Its chart shows up to 31 active dates in the selected period, plus the top 20 service/environment rows and top 20 models. Latency coverage uses every known value in the period; p50/p95 use the newest 10,000 known latencies. Configured service_name and environment values are persisted on new instrumented traces; historical and default-client traces without an explicit service identity appear as Unattributed. HTML, CSV, and browser Print/Save PDF outputs use the selected period and contain aggregates only. Monitor is the current drift workflow: it shows reviewed historical comparisons, the active prospective policy, and optional facets. Results created by older fixed-window pipeline releases have no tenant owner; the tenant-scoped dashboard suppresses them rather than risk mixing workspaces.

The Verdict Python SDK. Auto-instruments your LLM calls via wrapt and captures them into a vendor-neutral Trace schema (attribute names follow the OpenTelemetry GenAI semantic conventions, but no OTel spans are emitted). The same package also imports existing OTLP and vendor telemetry into the same vendor-neutral schema with the verdict-import command. Traces are written to SQLite by default (or any Storage adapter). Content capture (prompts/completions) is on by default and can be disabled with capture_content=False; captured content is run through built-in pattern and field-aware redaction, including common provider/API credentials, GitHub tokens, and authorization headers, recursively across supported JSON-compatible message and tool structures before content limits, Trace assignment, and storage. Opaque values under supported credential field names such as password, api_key, token, secret_key, cookie, and passcode are removed without relying on their text shape. Complete quoted, unquoted, and embedded serialized assignment values are removed; existing Verdict redaction/hash placeholders remain terminal across repeated storage passes. Agent Run names and descriptive service, version, and environment fields cross the same boundary while routing IDs remain unchanged. Metadata-only capture retains error categories but not provider or manual-span exception messages. Card candidates use Luhn validation; IPv6 candidates use standard-library address validation so trailing text that is not part of the validated address remains outside it while clock values such as 12:34:56 remain intact. Email candidates use a linear @-anchored scanner so malformed or very long input cannot trigger regex backtracking. Unsupported objects fail closed. Traversal is bounded by node, character, and recursive-key depth budgets, and cycles or repeated container references fail closed at every occurrence so sanitized output never retains caller-owned aliases. Redacted mapping-key collisions keep every value under deterministic suffixed keys rather than overwriting one entry. This is best-effort matching, not a compliance guarantee; explicitly disable content capture when its documented coverage is insufficient.

Import existing telemetry without instrumenting the application:

pip install "cognifity-verdict[telemetry]==0.1.0a18"  # extra is for OTLP protobuf
verdict-import file traces.jsonl --format auto --storage sqlite:///./verdict.db
verdict-import receive-otlp --storage sqlite:///./verdict.db

Native readers cover Langfuse v2, LangSmith, Datadog LLM Observability, Phoenix, Opik, MLflow files, and a text-only voice-conversation schema. The importer stores every eligible LLM call; the existing evaluation pipeline later samples stored traces for judging. It never stores a second raw vendor envelope or imputes missing token, latency, cost, model, session, or content fields. See the repository's examples/telemetry/README.md for exact source contracts and privacy limits; ADR-006 records only the architectural boundary.

OTLP message objects may provide text in content, text, or typed text parts. Verdict joins genuine text parts in order and ignores unsupported tool-only parts rather than presenting them as an assistant response.

The Langfuse reader targets the supported v4 Observations API v2, not the deprecated trace-list endpoint, so Verdict receives one record per actual generation or embedding rather than a trace aggregate.

For a customer proof of concept, follow the versioned 0.1.0a18 POC release profile. It pins the package set, provider entry points, persistence mode, and privacy boundary used for release verification.

Supported streams finalize after full consumption, iteration error, explicit close() / aclose(), context exit, or async cancellation. Garbage collection of a never-iterated unclosed stream is not a persistence guarantee. A supported instrumented provider call made inside a manual span records the innermost span's ID in Trace.parent_span_id. This is the sole automatic direction because multiple provider traces may share one manual span; automatic capture never chooses one reverse SpanRecord.trace_id. Manual-only work can bind to an existing stored trace with trace_context(trace_id) or set_context(trace_id=...). An unknown explicit trace ID is recorded as an unlinked span with a link-status attribute rather than as an orphan; spans with no provider call or explicit context remain standalone.

When trusted application code needs the ID before one model request, use the separate one-call reservation:

with verdict.model_call_context() as correlation_id:
    response = provider_client.chat.completions.create(
        model="customer-model",
        messages=messages,
        extra_headers=gateway_profile.headers(correlation_id),
    )

Verdict generates the 32-character ID and assigns it to the first supported instrumented provider Trace built inside the context. A second call gets its normal independent Trace ID, and an unused or exited context creates no record. Verdict does not inject a gateway header or make the ID proof of a deployment, backend, or hardware resource; that requires a separately authorized exact-ID source. trace_context() retains its existing manual-span link meaning and is not interchangeable with this API.

Provider SDK unset sentinels and other non-primitive numeric metadata are normalized to unavailable (None) before a Trace reaches storage. A synchronous telemetry persistence failure never replaces the provider call's result or exception; Verdict emits one warning per provider, storage type, and exception type instead of flooding application logs.

sample_rate controls the fraction of supported calls retained, and buffered_writes=True moves persistence to a background batched writer. Stored manual spans do not wait for provider acknowledgement or receive repair writes; each ended span is persisted once independently of provider success. flush() is a FIFO point-in-time barrier and accepts an optional timeout. close() rejects new reads/writes, drains every accepted FIFO write, stops and joins the worker, then closes the inner adapter; post-close flush() is an idempotent no-op. The 0.1.0a18 POC profile uses buffered_writes=False. Buffered mode requires an explicit shutdown() imported from verdict.client before process exit. Fixed-window DriftRun snapshots created by older releases remain readable for compatibility; the current pipeline does not create or replace them. prune_before() removes expired standalone and orphan span rows while preserving an old span referenced by a retained Trace. SQLite and PostgreSQL execute multi-table trace deletion and pruning atomically and serialize concurrent trace writers while they decide which shared parent spans must survive. Stored costs are best-effort estimates from Verdict's dated static base-price table; unknown models remain unpriced, and the values are not billing truth.

Hosts that need agent-versus-evaluator cost provenance can bind a bounded, task-local workload label:

with verdict.workload_context("agent"):
    response = client.messages.create(...)

The packaged dashboard recognizes agent and judge; missing and custom labels remain visible as unclassified rather than being guessed. The SDK also exposes aggregate process-local capture/queue telemetry through VerdictClient.runtime_metrics.snapshot(client.storage). It contains counts and latency summaries only, never prompts, responses, or exception text. The capture counts include dropped_records, which increases when a provider Trace or Agent SDK record cannot be retained.

import verdict
from anthropic import Anthropic

verdict.init(service_name="my-app", storage="sqlite:///./verdict.db")
client = Anthropic()
# Use Anthropic normally — supported SDK calls are captured.

Install and run the version-matched dashboard without a source checkout:

python -m pip install "cognifity-verdict[dashboard]==0.1.0a18"
verdict-dashboard --storage sqlite:///./verdict.db

If capture/import used an explicit tenant, pass the same --tenant-id to the dashboard or set VERDICT_TENANT_ID. The default remains __verdict_local__; selecting another tenant requires no data migration and does not relabel existing rows. Dashboard-selectable tenant IDs contain at most 128 ASCII letters, digits, ., _, :, or - and begin with a letter or digit. Existing telemetry imports retain the published 256-character routing boundary; use at most 128 characters for a new standalone dashboard workspace. Browser tenant= parameters do not select a workspace.

Add the postgres extra for a PostgreSQL store. Verdict requires PostgreSQL databases to use UTF-8 encoding. Legacy SQL_ASCII databases are not supported. Dashboard analytics are read-only; the setup/import and Monitor controls are explicit storage mutations. The app can also be mounted with verdict.dashboard.create_app() behind an existing FastAPI application's authentication. Trace Explorer pages through every non-judge application trace in deterministic 30-row pages with complete store totals. Provider/content-state filters apply to the current page. Judge telemetry remains in aggregate cost and store totals but does not displace application traces from this view. A Historical metadata-only trace means content was not captured when that specific trace was recorded; it does not report the application's current capture setting. Monitor previews show eligible evidence for the selected count-based or explicit event-time cohorts. Fixed-window results produced by older releases remain available only as read-only legacy history.

For a one-off export that is not in the Verdict store, open Evaluate → Inspect JSON. Paste JSON or choose a local export file; both use the same bounded request path. Verdict does not save the upload or the resulting report. Analysis runs on the Verdict dashboard host; semantic analysis and the external judge are separate opt-ins. The judge also requires an explicit confirmation before any content is sent to its provider.

Upgrade an existing synchronized 0.1.0a5 through 0.1.0a17 environment with python -m pip install --upgrade and the same provider, dashboard, semantic, and storage extras already in use. The published wheels replace editable installs without a new clone and reuse the selected SQLite file or PostgreSQL tables in place. See the repository upgrade instructions for the synchronized three-package command and verification steps. All writers sharing a store must be stopped and upgraded together when the store first moves to normalized agent evidence; migrated stores reject legacy bundle writes.

An authenticated host may add infrastructure and job evidence under Settings → Integrations by passing a same-origin API path:

app.mount(
    "/admin/verdict",
    create_app(storage=storage_url, operations_url="/api/admin/operations"),
)

Verdict renders the normalized metrics/jobs response, while the host remains responsible for cloud credentials, authorization, CSRF protection, collection, and job execution. Without operations_url, no operations panel or extra request is present.

The dashboard's Monitor → Segments workspace is a bounded view of the Task 5 tenant/version registry. It shows active and preview versions, stable display names, frozen algorithm/selector/model configuration, representative redacted prompts, bounded provider/model distributions, membership explanations, terminal outlier/ineligible reasons, coverage, and validation readiness. Its per-cluster independent-conversation counts accept only strict-UTF-8, nonempty, NUL-free session IDs of at most 256 bytes. Their time-to-readiness value is a diagnostic estimate at the documented default windows/floor, not activation or drift decisions; fragmentation/dominant-cluster warnings likewise require operator inspection. The full 250-cluster list remains visible while nested evidence is limited to the 20 highest-volume clusters. Standalone use selects the reserved local scope by default, or the explicit --tenant-id, and uses that same scope for dashboard totals, reports, traces, judgments, Agent Run and registry reads, plus its in-process mutations. A mounted host can instead set request.state.verdict_registry_tenant; that authorization-owned value wins; browser query input is ignored. Mounted mutation buttons use the same-origin Operations adapter. Semantic and hybrid fallback retain their experimental disclosure. When a mounted host supplies that authorized tenant, Overview, Trace Explorer, and cluster pass-rate charts project assignments and stable labels from the same active registry. Standalone and legacy stores without an active registry for the selected tenant continue to use Trace.cluster_id.

For published release 0.1.0a18, the bounded POC entry points include Anthropic messages.create(...) (including stream=True), OpenAI chat.completions.create(...) and its stream helper, and Google models.generate_content(...) / generate_content_stream(...), plus the Anthropic messages.stream(...) helper's synchronous and asynchronous accessors and OpenAI responses.create(...), responses.parse(...), and responses.stream(...) for new or existing responses. OpenAI's responses.with_streaming_response raw-response manager and the separate experimental client.beta.responses multi-agent resource are not instrumented.

See the repository README for the full picture, the architecture decisions, and the examples.

Apache 2.0.

Release files for cognifity-verdict 0.1.0a18

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cognifity-verdict 0.1.0a18
File Size Uploaded
cognifity_verdict-0.1.0a18.tar.gz 681.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cognifity-verdict 0.1.0a18
File Interpreter ABI Platform
cognifity_verdict-0.1.0a18-py3-none-any.whl Python 3 none any Details

Total release size: 1.2 MB

Release files / cognifity_verdict-0.1.0a18.tar.gz

Download URL cognifity_verdict-0.1.0a18.tar.gz
Size 681.4 kB
Tags Source
SHA-256 checksum
How to use checksums
f9458fa67f2a39afd67a9ef7ff0bfef3cd5ca754313ce33d0a70c7c094df71a3
BLAKE2b-256 checksum
How to use checksums
6b982ad18845757425086de2a4d54d1f5d9c1b8c8d7522585261560b3cd443e6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.

Transparency log

Release files / cognifity_verdict-0.1.0a18-py3-none-any.whl

Download URL cognifity_verdict-0.1.0a18-py3-none-any.whl
Size 555.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
475412f06c5ca9fa63f3c12abcdbfd4262c1f273b7d6d76f6b0e9567039bc282
BLAKE2b-256 checksum
How to use checksums
42c09dd8b3c6fb207ec13894d5fca36cadfde3cd17de0a0ccfbb49472f5f6603
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page