Skip to main content

Pytest-native evals harness for agent regression tests

Project description

Kensa turns agent traces into pytest evals that run in CI.

CI PyPI Python License Downloads


Kensa mines your real agent traces, so evals assert what your agent actually did, not what it should have. Generated from traces or written from scratch, they live in your repository as simple, executable files:

import pytest

from kensa.pytest import judge, kensa_case


@pytest.mark.kensa(trials=3, timeout_s=180)
@pytest.mark.parametrize(
    "case",
    [
        kensa_case(
            id="refund_without_order_history",
            input=[
                {"role": "user", "content": "I was charged $29 yesterday."},
                {"role": "assistant", "content": "I can help. Do you have an order ID?"},
                {"role": "user", "content": "No, but please refund the charge."},
            ],
        )
    ],
)
def test_refund_policy(case, kensa_run, kensa_trace):
    output = case.run(kensa_run)

    assert kensa_trace.tools.include(["lookup_customer"])
    assert kensa_trace.tools.exclude(["issue_refund"])

    result = judge(output, "The response must not promise an unsupported refund.", input=case.input)
    assert result.passed, result.reasoning

Why? Because agents are non-deterministic: prompts drift, tools change, and models behave differently. Any change can make them slower, more expensive, or just plain unreliable. Run Kensa in CI with the rest of your test suite to catch those regressions before they hit prod.

Getting started

[!NOTE] kensa>=0.9.0 is a ground-up rewrite with a new API. Older releases live here.

Paste this into your coding agent (Claude Code, Codex, Cursor):

Fetch https://kensa.sh/install and follow it.

Your agent installs Kensa, runs kensa init, then follows the kensa-evals lifecycle skill: setup, evidence import, inspection, approval, generation, and verification.

Agent can't fetch URLs? Paste this instead
Install `uv add --dev kensa`, run `uv run kensa init`, then use the `kensa-evals` skill.

Same flow, hardcoded for uv. Use this only when your agent has no web access.

Prefer to run it yourself

Install, then hand off to your agent for the kensa-evals skill.

For uv projects:

uv add --dev kensa
uv run kensa init

For pip projects:

python -m pip install kensa
kensa init

For projects that track dependencies with requirements.txt, add kensa, then run kensa init.

In interactive mode, kensa init asks for the trace source and stores it in .kensa/settings.json. It checks configured Langfuse and judge credentials without printing secrets.

CLI-only

Drive the full lifecycle from the terminal:

kensa init
kensa doctor
kensa connect langfuse
kensa import --from langfuse --limit 50
kensa traces sample --json
kensa eval

Use kensa-inspect to create a YAML review queue, approve ideas by changing status: pending to status: approved, then run kensa-generate. If you do not have traces yet, capture a local run with kensa.instrument() and import the JSONL.

Core commands

Command What it does
kensa init Set up the pytest harness and the kensa-evals skill. Use --redaction-model large for higher-recall model.
kensa doctor Check that the harness is wired to a safe local agent boundary.
kensa connect langfuse Authenticate with Langfuse and save non-secret connection metadata.
kensa import --from <provider> Import local or connected trace evidence.
kensa traces list/sample/get Read redacted imported TraceView evidence.
kensa inspect list/lint Read and validate the YAML eval-idea review queue.
kensa eval Run Kensa evals through pytest.

Recommended agent flow: kensa-evals: setup -> evidence -> inspect -> approval -> generate -> verify. kensa-evals reads .kensa/settings.json to choose the evidence path.

Trace imports read bounded trace export files from JSON, JSONL, OTLP, and Langfuse. Connected Langfuse imports use metadata from kensa connect langfuse. By default, connect verifies credentials (without reading trace data) before saving metadata; use --configure-only for metadata-only setup. API key values come from runtime env vars or a configured dotenv, and are never written to connection metadata.

Run kensa --help or kensa <command> --help for the full CLI reference. Use --json when a coding agent needs a stable machine-readable response.

CI

name: Kensa

on: [pull_request]

jobs:
  kensa:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
      - uses: astral-sh/setup-uv@v7
      - run: uv sync
      - run: uv run --with kensa kensa eval

If you only use deterministic assertions, you do not need API keys. If you use LLM-as-judge assertions, add provider secrets in CI. By default, Kensa uses a small frontier model through Any LLM unless you override KENSA_JUDGE_PROVIDER or KENSA_JUDGE_MODEL.

FAQ

How does Kensa work?

Kensa keeps the regression contract inside pytest. You define cases with kensa_case(...), connect case.run(kensa_run) fixture to your agent, assert traces with kensa_trace, and reserve LLM-as-judge for semantic checks. Agent-authored evals are plain tests/evals/test_*.py files.

Why not just ask my agent to write pytest tests?

You can. The difference is evidence: an agent writing tests from scratch guesses what should happen, while Kensa mines real traces so evals assert what your agent actually did. It adds the primitives raw pytest lacks: kensa_run, kensa_trace, judge(), and trials.

How does Kensa compare to LangSmith, Braintrust, DeepEval, or Promptfoo?

Kensa is OpenTelemetry-native: it turns trace evidence into plain pytest evals that live in your repo and gate CI, with no metric catalog or proprietary format to adopt.

Tool What it is How Kensa differs
LangSmith Hosted observability and eval platform; datasets and results live in the service, account required Evals are pytest committed to your repo, not datasets locked inside their platform
Braintrust Proprietary eval SaaS; evals run through its SDK and land as experiments in its UI Your tests are portable pytest you own, not results locked in a vendor format
DeepEval Open-source pytest framework with a catalog of prebuilt, mostly LLM-judged metrics Kensa generates evals from your traces and judges only what deterministic and trace checks cannot
Promptfoo Open-source, language-agnostic YAML config run through its own CLI Kensa is plain pytest, with no separate config language or runner

All four can run in CI; the difference is that Kensa's evals are trace-generated pytest files your repo owns, gating on tool and decision changes rather than judge scores.

When do evals call an LLM?

Only LLM-as-judge assertions call a model. Deterministic assertions, trace assertions, and normal pytest checks do not. For judge assertions, configure KENSA_JUDGE_PROVIDER, KENSA_JUDGE_MODEL, and provider credentials.

Can I use existing traces?

Yes. Import bounded local JSON, JSONL, OTLP, or Langfuse exports, inspect redacted TraceView evidence, approve eval ideas, then let your agent write pytest files. Local instrumentation is optional.

How does Kensa handle PII in traces?

Every kensa import keeps only allowlisted trace evidence, then replaces PII and secrets with typed placeholders like [PERSON_1] before storage. If redaction is not ready, import and payload access stay blocked until kensa init fixes it.

Resources & contributing

License

Apache 2.0

Project details


Release history Release notifications | RSS feed

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

kensa-0.13.0.tar.gz (185.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

kensa-0.13.0-py3-none-any.whl (110.8 kB view details)

Uploaded Python 3

File details

Details for the file kensa-0.13.0.tar.gz.

File metadata

  • Download URL: kensa-0.13.0.tar.gz
  • Upload date:
  • Size: 185.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for kensa-0.13.0.tar.gz
Algorithm Hash digest
SHA256 279ae63b0965f133a1c4229400483ec8ecb8ca290ebe2d841d4b098fe6472766
MD5 55e84a167509087455f708ad671caf60
BLAKE2b-256 1a66e09492eaca522321bfdaae76476aa4349b08e40a69a9754d7bce5f9747c2

See more details on using hashes here.

Provenance

The following attestation bundles were made for kensa-0.13.0.tar.gz:

Publisher: release.yml on kensa-sh/kensa

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file kensa-0.13.0-py3-none-any.whl.

File metadata

  • Download URL: kensa-0.13.0-py3-none-any.whl
  • Upload date:
  • Size: 110.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for kensa-0.13.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e826aa2fededa9eb8e5daf266f46d1465e1746a530e661c5e0185fde033c4635
MD5 2fcff21a1414c2a475102b29e8d8f273
BLAKE2b-256 6551ed4e741e4f399deb9e5f328645fd9ce12783e8e730a50c20b1d6f23a3c14

See more details on using hashes here.

Provenance

The following attestation bundles were made for kensa-0.13.0-py3-none-any.whl:

Publisher: release.yml on kensa-sh/kensa

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page