Skip to main content

kyvvu-engine

Runtime policy evaluation for AI agents — the Agent Security Kernel (ASK) as a stateful Python library.

kyvvu-engine evaluates policies against the full execution path of an AI agent. Given an intended behaviour — the smallest atomic unit of what an agent is about to do — it decides allow, warn, or block, and explains why.

Agent security is modelled as a pathwise problem: decisions depend not only on the step an agent is about to take, but on the full ordered history of what it has already done in the current task. This is "policies on paths," formalised in the paper Runtime Governance for AI Agents: Policies on Paths. kyvvu-engine is the reference implementation.

kyvvu-engine is used by the kyvvu SDK (Python agent integration), the Kyvvu platform, and is available as a standalone HTTP service via the SDK's kyvvu serve command for harnesses in other languages.


Contents


Installation

# Most users: install the SDK (includes the engine)
pip install kyvvu

# Engine only (no SDK, no agent integration — for embedding)
pip install kyvvu-engine

# Standalone HTTP server (for non-Python harnesses)
pip install "kyvvu-engine[serve]"

Quickstart

Two ways to get policies into the engine. The first needs no account, no API key and no network — evaluation itself is credential-free, and since #426 so is getting a policy set into it.

1. Offline — no account, no API key, no network

Write a manifest. This is the same format the manifests in Kyvvu/manifests are written in, so anything there runs locally too:

# my-policies.yaml
name: My first policy set
version: "1.0.0"
policies:
  - name: Code execution requires a preceding gate
    rule_type: step_requires_gate
    params:
      target_step_types: ["step.exec"]
    severity: critical
    enforcement_point: step_execution

Point a runner at it and evaluate:

from kyvvu_engine import KyvvuBlockedError, KyvvuRunner
from kyvvu_engine.schemas import Behavior, EvalContext, StepType

runner = KyvvuRunner(policy_file="my-policies.yaml", log_location="none")

# Load now, and check what is actually being enforced before trusting any
# verdict. Policies otherwise load lazily on the first evaluate(), so a
# load_report() read before that reports `unavailable` — nothing has been
# asked to load yet.
status = runner.ensure_policies()
print(status.load_state.value, status.policies_offered, status.policy_count, status.dropped)
# enforcing 1 1 []

ctx = EvalContext(
    agent_id="local-agent",
    task_id="task-1",
    # Declare the agent's tools. Not needed by the policy above, but the
    # OWASP set's allowlist policy reads this field (NOT the agent payload's
    # declared_tools) and fails closed: leave it unset there and every
    # step.resource / step.exec / step.credential blocks on "Tool calls must
    # be in the declared allowlist".
    agent_allowed_tools=["run_script"],
)

intended = Behavior(
    agent_id="local-agent",
    task_id="task-1",
    step_type=StepType.step_exec,
    step_name="run_script",
    input={"code": "rm -rf /"},
)

try:
    runner.evaluate(intended, ctx)
except KyvvuBlockedError as exc:
    print(exc)          # Step 'run_script' blocked (risk_score=1.00): …

Faster still, with the kyvvu SDK installed (pip install kyvvu):

kyvvu try                                  # bundled demo path, bundled frozen policy set
kyvvu try --manifest my-policies.yaml      # your manifest, still offline
kyvvu serve --policy-file my-policies.yaml # the same policies over HTTP, still offline

kyvvu.baseline.load_baseline_policies() returns a frozen 15-policy set bundled inside the SDK — pass it as KyvvuRunner(policies=…). It is a kyvvu feature rather than a kyvvu-engine one, because policy content carries the licence of the manifests it came from (Apache-2.0) while the engine is BSL-1.1. See The Bundled Baseline.

Check load_report(), not policy_count(). A hand-written policy set whose rule_type is misspelled, or whose params fail a rule's schema, is dropped — and a dropped set enforces nothing while looking configured. load_report() names which of the five PolicyLoadState values you are in and lists every dropped policy with its reason, and the same outcome is announced once on stderr, so this failure is loud rather than silent. That was the sharpest edge on the offline path before #441; it is now the noisiest.

2. Platform-connected

Credentials are needed only when policies are actually fetched — an agent's current, assigned policy set, with incidents and the audit trail on the other end of it:

from kyvvu_engine import KyvvuRunner
from kyvvu_engine.schemas import Behavior, EvalContext, StepType, Verb

runner = KyvvuRunner(
    api_url="https://platform.kyvvu.com",
    api_key="KvKey-…",
    agent_key="customer-support-agent",
)

ctx = EvalContext(agent_id="agent-123", environment="production")

# 1. Preflight: evaluate the intended behaviour before executing.
intended = Behavior(
    agent_id="agent-123",
    task_id="task-abc",
    step_type=StepType.step_model,
    verb=Verb.POST,
    step_name="chat_gpt-4o",
    input={"user_message": "Hi, my SSN is 123-45-6789"},
)

result = runner.evaluate(intended, ctx)
# action == "allow" → returns normally
# action == "warn"  → emits warnings.warn() and returns
# action == "block" → raises KyvvuBlockedError

# 2. Execute the step.
output = your_llm_call(intended.input)

# 3. Record the completed step. It becomes visible to future evaluate() calls.
runner.record(intended.model_copy(update={"output": output}))

# 4. Close the task when finished. No policies run here; this is cleanup.
runner.end_task("task-abc")

Mental model

The engine is a stateful decision machine.

               ┌──────────────────────────────────────────┐
  policies ──▶ │                                          │
               │              PolicyEngine                │
  intended ──▶ │       (zero-I/O, sub-ms decisions)       │ ──▶  allow | warn | block
  behaviour    │                                          │       + per-policy outcomes
  + context    │   internal state: per-task history       │       + aggregate risk score
               └──────────────────────────────────────────┘

Terminology

  • A behaviour is an atomic action an agent takes. It is the smallest governable unit.
  • A step is a triple {input, behaviour, output} — an executed behaviour with its context. The completed history of a task is an ordered sequence of steps, also called a path.
  • An intended behaviour is {input, behaviour, void} — a behaviour about to execute, with input but no output yet. This is what evaluate() inspects.
  • A policy is an instantiation of a rule function with specific parameters, severity, and enforcement point.
  • The organisational context (EvalContext) carries the agent's identity, classification, environment, user settings, and pre-fetched cross-task counts — everything rules may need that isn't in the behaviour or history itself.

What evaluate() does

  1. Reads the task's completed history from its internal tracker.
  2. Filters policies applicable to this agent and classification.
  3. Runs each applicable rule function, passing the flattened behaviour data, the policy's params, and a RuleContext that gives access to history and organisational context.
  4. Each rule returns a boolean. Per-policy outcomes are weighted by severity and aggregated into a risk score ∈ [0.0, 1.0]. The default aggregator is aggregate_max (worst-case severity wins); this is pluggable.
  5. The final risk score maps to an action: 0.0 → allow, (0, 1) → warn, 1.0 → block.

Two evaluation points

  1. Agent registration — once, before the first task begins, via evaluate_registration(). Agent-level policies run here (declared purpose, tool allowlist, classification).
  2. Every atomic step — before each step executes, via evaluate(). Step-level policies run here (path history, content, classification, rate limits).

Task completion is cleanup, not a decision point. end_task(task_id) evicts the task's history from memory and flushes buffered logs. It does not evaluate policies. To enforce task-end invariants, model them as policies on the task.end behaviour — templates emit a task.end Behavior that flows through the normal evaluate() path.

Runner action semantics

  • allow — evaluate() returns normally.
  • warn — evaluate() emits warnings.warn(...) and returns; incident webhook fires if configured.
  • block — evaluate() raises KyvvuBlockedError; incident webhook fires if configured. The caller can catch the exception to continue execution (retry, fall back, notify the user, abort the task).

Zero-I/O core

PolicyEngine never calls out, never queries a database, never fetches anything. KyvvuRunner is a wrapper that adds HTTP (fetching policies, flushing logs, firing incident webhooks). All network code is isolated in the io/ module; the core engine has no awareness of the network.

One engine per agent

Each KyvvuRunner is configured with a single agent_key and owns one PolicyEngine instance. Engines are per-agent by construction and are not designed to be shared across agents.


Atomic behaviours

Every action an agent takes is classified into one of 12 atomic behaviour types, all in the step.* namespace. Four mark the boundaries of a task; eight describe the agent's moves within one. Together with an HTTP-style verb (GET, POST, PATCH, DELETE, or none) they form the canonical vocabulary the engine operates on. TASK_BOUNDARY_STEP_TYPES names the four boundary types for callers that need to tell them apart.

Step type Valid verbs What it represents
task.start — A task begins.
task.end — A task completes normally.
task.error — A task terminates with an error.
task.idle — The agent is idle within a task (heartbeat / keepalive).
step.resource GET/POST/PATCH/DELETE Read or mutate an external resource (DB, file, API).
step.message GET/POST Receive (GET) or send (POST) a message — user input, UI events, outbound communication.
step.self GET/POST/PATCH/DELETE Read/write the agent's own internal state (memory, scratchpad, plan).
step.model POST Send a prompt to an LLM and receive a completion.
step.credential GET Retrieve a secret, token, or credential.
step.exec — Execute code (run a script, call a function, shell out).
step.gate — Cross a gate — a human approval, a policy check, a guardrail.
step.unknown — Uncategorisable behaviour (template fallback).

These combinations are enumerated in schemas.VALID_COMBINATIONS as (step_type, verb) pairs and enforced by Behavior's model validator. Any pair outside this set raises on construction. A serialized Behavior round-trips unconditionally: every key in the dump is an accepted input, and the model forbids unknown fields.

The step_execution enforcement point evaluates all twelve behavior types. step_execution and agent_registration are the only enforcement points the engine evaluates.

Task lifecycle events

  • task.start — emitted when the agent begins a task. Often the first behaviour in a path.
  • task.end — normal completion. evaluate() runs as usual; a policy matching on step_type == task.end can check whole-task invariants.
  • task.error — abnormal termination. History is evicted on end_task(task_id) identically to task.end. For forensic retention after errors, handle at the log-sink layer.
  • task.idle — emitted periodically when the agent is idle but the task isn't over. Keeps rate-limit and working-hours rules accurate across long pauses. Does not trigger cleanup.

Properties

Everything beyond the (step_type, verb) pair lives in properties, a nested dict that policies can inspect. Properties distinguish a step.resource reading customer-data from one reading product-data — the type is the same, the property is different.

Standard property groups:

  • target — the thing being acted on (domain, resource URI, table name).
  • auth — authentication access level (read, write, admin).
  • data — payload classification (sensitive fields, size, schema).
  • model — for step.model: provider, model id, parameters.
  • exec — for step.exec: runtime, isolation level, side-effect class.
  • guard — for step.gate: gate type (human_approval, policy_check, static_check).
  • message — for step.message: channel, sender, recipient.
  • usage — for step.model outputs: prompt_tokens, completion_tokens, cost_usd.

Custom groups are permitted; the engine passes them through unchanged, and rule functions read them via dot-path accessors (_get_prop(data, "target.table")).

Worked example

A step.resource GET reading customer data with realistic properties:

Behavior(
    agent_id="agent-123",
    task_id="task-abc",
    step_type=StepType.step_resource,
    verb=Verb.GET,
    step_name="read_customer_record",
    input={"customer_id": "CUST-9981"},
    properties={
        "target": {
            "system": "salesforce",
            "table": "customer-data",
            "object_id": "CUST-9981",
            "domain": "internal.crm.acme.com",
        },
        "auth": {"access_level": "read", "principal": "agent-123"},
        "data": {"classification": "pii", "fields": ["name", "email", "phone"]},
    },
)

The evaluation lifecycle

Three calls per step, one per task end.

1. evaluate(intended, context) → EvalResult

Called before a step executes. Reads the task's history, filters applicable policies, runs each rule, aggregates, returns. Does not modify history.

Outputs:

  • result.action — "allow", "warn", or "block".
  • result.risk_score — normalised [0.0, 1.0].
  • result.policies — one PolicyResult per evaluated policy.

2. Execute the step

The caller runs the tool, the LLM, the database write. The engine has no opinion about execution.

3. record(step) → Behavior

Called after the step executes. Assigns a monotonic step number within the task and appends the completed Behavior (with output populated) to the task's history. Future evaluate() calls in the same task_id see this step.

4. end_task(task_id)

Called when the task terminates. Cleanup only — no policy evaluation. Evicts the task's history from memory and, via the runner, flushes any buffered step logs.

Calling end_task() for an unknown task_id is a no-op. A new task_id is a fresh history with no relationship to any previous task — histories are keyed by task_id.

Memory management

History lives in memory, keyed by task_id. Two mechanisms prevent unbounded growth:

  • Normal termination: end_task(task_id) evicts history explicitly.
  • Abandoned tasks: runner.sweep_stale_tasks(), called periodically, evicts tasks older than KV_TASK_MAX_AGE_SECONDS (default 3600s). Tasks that crash before end_task() are cleaned up this way.

Wire sweep_stale_tasks() into a background thread or scheduler in production.

Interaction diagram

agent:  evaluate(intended) ───▶  engine: check policies against history + context
                           ◀───  EvalResult{allow/warn/block, risk_score, policies}
agent:  execute step, capture output
agent:  record(completed_step) ─▶  engine: append to history, assign step number
                                ◀─  Behavior{step=N, ...}
                           (repeat per step)
agent:  end_task(task_id) ─────▶  engine: evict history, flush logs

Agent registration

Before the first task, agents register themselves with the Kyvvu platform. Registration is where agent-level policies are evaluated — declared purpose, tool allowlist, owner domain, classification consistency.

Registration policies have enforcement_point: "agent_registration" and run against the agent's metadata rather than a Behavior:

result = runner.evaluate_registration(
    agent_data={
        "name": "customer-support-agent",
        "purpose": "Triage inbound customer tickets and draft responses",
        "owner": "support-team@acme.com",
        "declared_tools": ["zendesk_read", "llm_call"],
        "risk_classification": "limited",
    },
    context=EvalContext(
        agent_id="agent-123",
        environment="production",
        risk_classification="limited",
    ),
)

Semantics are identical to evaluate(): same EvalResult, same allow/warn/block, same runner behaviour (warn emits warnings.warn, block raises KyvvuBlockedError). The difference is which policies run — only those with enforcement_point=agent_registration.

Registration is typically called once at agent startup. A block at registration means the agent should not start at all — typically an illegally configured agent (empty purpose, disallowed tools, classification mismatch).


Rule functions

Rule functions are the unit of decidability. Each rule is a small pure Python function with the signature:

def rule(data: dict, params: dict, context: RuleContext) -> bool:
    """Return True if the policy passes; False if it is violated."""

A policy is an instantiation of a rule: policy = rule + params + (enforcement_point, severity, agent_id, risk_classification). The same rule backs many policies — field_matches_regex instantiated once for SSNs, once for credit cards, once for email domains.

Rule context

Every rule receives a RuleContext, the only surface through which rules read state beyond their own params:

  • context.agent_id, context.task_id, context.enforcement_point, context.now, context.hour
  • context.get_current_agent() → AgentRecord | None — agent metadata.
  • context.user_settings → dict | None — pre-fetched user preferences.
  • context.get_previous_step() → Behavior | None — last completed step.
  • context.get_all_steps_in_task() → List[Behavior] — full task history.
  • context.count_steps_of_type(step_type: str) → int — counter helper.
  • context.count_recent_nodes_across_executions(step_type, window_minutes, attribute_filter) → int — pre-fetched cross-task counts.

All surfaces are in-memory and pre-fetched. Rules perform no I/O.

Built-in rule functions

The engine ships with 27 built-in rules grouped into six categories. Each category lives in its own module (kyvvu_engine/rules/<category>.py) with a mirror test file.

Field rules (rules/field.py) — applicable to agent_registration and step_execution:

Rule What it checks
field_not_empty Named field has a non-empty value.
field_in_list Named field's value is in an allowlist.
field_matches_regex Named field matches a regex pattern.

Path & containment rules (rules/path.py) — step_execution only (the history rules replay accumulated history; path_within_root inspects the current target):

Rule What it checks
step_directly_preceded_by Previous step in history has a given type.
step_requires_predecessor Some earlier step in history has a given type.
step_preceded_by_without_intervening A required predecessor exists with no forbidden steps between.
step_requires_dedicated_predecessor Each target step has an unconsumed predecessor of the required type.
step_requires_gate A step.gate precedes this step.
sequence_forbidden A forbidden ordered sequence has not occurred.
step_not_after This step type is forbidden once a specified predecessor has occurred (permanently tainted).
history_contains The history contains a step matching type + optional verb + optional property filter.
current_is The intended behaviour matches type + optional verb + optional property filter.
path_within_root The target path resolves within project_root (allowlist containment); blocks targets that escape it. Lexical only — no symlink resolution.

Count rules (rules/count.py):

Rule What it checks
execution_max_steps Task has not exceeded a maximum step count.
max_consecutive_same_type No run of the same step type exceeds a limit.
cross_execution_rate_limit This agent has not exceeded N of this step_type in the last M minutes across tasks.
usage_budget Cumulative usage metric (tokens, cost) across task history has not exceeded budget.

Classification rules (rules/classification.py):

Rule What it checks
step_forbidden_for_classification This step type is not permitted for the agent's risk classification.
working_hours_only Current time is within a permitted window. Supports overnight wraparound and timezones.
step_name_in_allowlist This step's name is in the agent's declared tool allowlist.

Content rules (rules/content.py):

Rule What it checks
pii_in_request Step input does not contain PII matching configured regex patterns. Patterns are required; no defaults.
domain_allowlist Step's target domain is in an allowlist.

Flow rules (rules/flow.py):

Rule What it checks
conditional_successor_required If the previous step matches a trigger condition, the current step satisfies the configured successor constraints.
tainted_path_block If any prior step is tainted, certain downstream steps are forbidden.
all_of Compound: passes iff all sub-conditions pass.
any_of Compound: passes iff any sub-condition passes.
not Compound: passes iff the sub-condition fails.

Each rule exposes a description, parameter schema, and example parameters programmatically:

from kyvvu_engine import PolicyRule
metadata = PolicyRule.get_all_rules(enforcement_point="step_execution")
# → {"field_not_empty": {"description": "...", "enforcement_points": [...], "params_schema": {...}}, ...}

The table above is derived from this metadata.

Compound policies

The three compound rules accept sub-conditions as params. Compound rules recurse freely: all_of can contain any_of can contain not can contain a primitive.

Important: rule functions return True to pass and False to block. This means all_of returns True (passes) when all sub-conditions are met. If your intent is "block when conditions A, B, and C are all present," you need not(all_of(A, B, C)) — the all_of detects the dangerous combination, and the not inverts it into a block. Using bare all_of for a blocking trigger is a common authoring mistake: it would block every step where any condition is not met, which is the opposite of what you want.

Example: "If the agent has read customer-data AND product-data AND called a model, then POSTing a message requires a human-approval gate":

{
  "name": "PII + product data + model requires human approval",
  "rule_type": "all_of",
  "params": {
    "conditions": [
      {"rule_type": "current_is",
       "params": {"step_type": "step.message", "verb": "POST"}},
      {"rule_type": "history_contains",
       "params": {"step_type": "step.resource", "verb": "GET",
                  "property_filter": {"target.table": "customer-data"}}},
      {"rule_type": "history_contains",
       "params": {"step_type": "step.resource", "verb": "GET",
                  "property_filter": {"target.table": "product-data"}}},
      {"rule_type": "history_contains",
       "params": {"step_type": "step.model"}},
      {"rule_type": "not",
       "params": {"condition": {"rule_type": "step_requires_gate",
                                "params": {"target_step_types": ["step.message"],
                                           "target_verb": "POST",
                                           "gate_check_type": "human_approval"}}}}
    ]
  },
  "severity": "critical",
  "enforcement_point": "step_execution"
}

Incidents from a failed compound policy carry one incident with the condition tree in violation_details.

Rule-specific notes

  • step_requires_gate — the gate may be any distance earlier in history; this rule does not enforce gate freshness. For fresh-approval semantics, compose with step_directly_preceded_by.
  • step_not_after — once any forbidden predecessor has occurred, the target is blocked for the rest of the task (tainted-path semantics).
  • working_hours_only — accepts timezone: str (IANA name); falls back to UTC. Supports overnight windows (start_hour=22, end_hour=6).
  • pii_in_request — patterns param is required. Step input is serialised via json.dumps so nested dicts are scanned correctly.
  • usage_budget — sums a numeric property from completed steps in history and blocks when the cumulative value exceeds the budget. The first occurrence is always allowed; only subsequent steps see an accumulating total.

Writing your own rule function

Registering a rule

from kyvvu_engine import PolicyRule

@PolicyRule.register(
    name="step_name_forbidden",
    description="The step's name must not match a forbidden pattern.",
    params_schema={
        "patterns": {"type": "array", "required": True, "description": "Regex patterns"},
    },
    enforcement_points=["step_execution"],
    example_params={"patterns": ["dangerous_tool"]},
)
def check_step_name_forbidden(data, params, context):
    import re
    name = data.get("step_name", "")
    for pattern in params["patterns"]:
        if re.match(pattern, name):
            return False
    return True

The rule is immediately available as a rule_type in any policy definition. The Kyvvu platform UI discovers it via PolicyRule.get_all_rules() and renders a form from params_schema.

Rules must live in the appropriate module under kyvvu_engine/rules/ and must have a mirror test in tests/rules/. Tests use PolicyEngine directly:

# tests/rules/test_field_rules.py
from datetime import datetime
from kyvvu_engine import PolicyEngine
from kyvvu_engine.schemas import Behavior, EvalContext, StepType, Action

def test_step_name_forbidden_blocks_matching_name():
    engine = PolicyEngine()
    engine.load_policies([{
        "id": 1, "name": "no-dangerous", "enforcement_point": "step_execution",
        "rule_type": "step_name_forbidden",
        "params": {"patterns": [r"^dangerous_tool"]},
        "severity": "critical", "enabled": True,
    }])
    b = Behavior(
        agent_id="a", task_id="t", timestamp=datetime(2026, 1, 1),
        step_type=StepType.step_exec,
        step_name="dangerous_tool_v2",
    )
    result = engine.evaluate(b, EvalContext(agent_id="a", task_id="t", environment="prod"))
    assert result.action == Action.block

New rules without matching tests fail CI.

Worked example: token-usage / cost budget

To block an agent once it has spent a budget on LLM calls within a task, use usage_budget:

{
  "name": "Per-task $5 LLM budget",
  "rule_type": "usage_budget",
  "params": {"step_type": "step.model",
             "property_path": "usage.cost_usd",
             "budget": 5.0},
  "severity": "high",
  "enforcement_point": "step_execution"
}

This sums properties.usage.cost_usd across completed step.model behaviours in the current task; once the total exceeds 5.0, further model calls are blocked. The same pattern works for tokens (property_path: "usage.total_tokens", budget: 100000) or any numeric property templates emit.


The two-tier API

PolicyEngine — the pure core

Zero I/O. Zero logging config. Only dependency: Pydantic. For embedding, for running policies from an in-memory store, and for unit-testing policy logic.

Method Purpose
load_policies(policies: List[dict]) → None Replace the active policy set. Idempotent.
evaluate(intended: Behavior, context: EvalContext) → EvalResult Preflight a step.
evaluate_registration(agent_data: dict, context: EvalContext) → EvalResult Evaluate agent-registration policies.
record(step: Behavior) → Behavior Append a completed step to history; assigns step number.
end_task(task_id: str) → None Evict a task's history from memory.
get_history(task_id: str) → List[Behavior] Read the task's completed steps (snapshot).
evaluate_and_record(intended, context, output=None) → EvalResult Convenience: evaluate; if not blocked, record with the given output.
explain(intended, context) → str Human-readable per-policy evaluation trace.
policy_count() → int Number of loaded policies (diagnostic).
validate_rule_params(rule_type, params) → (bool, str | None) Check a rule name is registered.

KyvvuRunner — the I/O wrapper

PolicyEngine + HTTP + log buffering. For use when policies come from the Kyvvu platform.

Method Purpose
fetch_policies() → None Force a policy refresh (ignores TTL).
sweep_stale_tasks(max_age_seconds=None) → int Evict abandoned task buffers.
policy_status() → PolicyStatus Policy cache status: loaded, stale, source (api/disk_cache/none), timestamps, policy count, TTL remaining, plus the load_state / policies_offered / dropped fields below.
policy_count() → int How many policies are currently being enforced.
load_report() → PolicyLoadReport The engine's own accounting of the last load: offered, enforcing, dropped.
dropped_policies (property) Every offered policy that is not being enforced, with the reason.
ensure_policies() → PolicyStatus Load or refresh through the configured source (API or disk cache), respecting the TTL. Use this at startup, not fetch_policies().
load_policies(policies, source=..., stale=...) → PolicyStatus Load a policy set the caller obtained itself (own transport, own cache) with the same bookkeeping a fetch performs.
signal_policy_state() → None Announce the current load outcome on stderr, at most once per process per distinct outcome.
settings (property) The resolved KyvvuSettings.

All PolicyEngine methods are available on KyvvuRunner with the same names. KyvvuRunner.evaluate() additionally:

  1. Ensures policies are loaded (fetches if TTL expired).
  2. Emits warnings.warn() on warn.
  3. Fires the incident webhook on warn or block (if configured).
  4. Raises KyvvuBlockedError on block.

HTTP endpoints (the runner)

KyvvuRunner makes up to three kinds of HTTP requests. All endpoints are configurable. The log destination (KV_LOG_LOCATION) defaults to stdout (JSON-line output to the terminal for development). The incident webhook is off by default. Set KV_LOG_LOCATION=none (or empty) to disable log output entirely. Both endpoints accept stdout as a value for local debugging. See the Configuration Reference for the full list of environment variables.

Authentication: all requests carry Authorization: Bearer <api_key>. The instance identifier is sent as both ?instance={instance_id} in the query string and X-Kyvvu-Instance-Id: {instance_id} in a header. Both carry the same value.

1. GET /api/v1/policies — policy fetch

Called on first use and whenever the policy TTL expires (default 300 seconds).

GET {api_url}/api/v1/policies?agent_key={agent_key}&instance={instance_id}&enabled=true&limit=1000
Authorization: Bearer {api_key}
X-Kyvvu-Instance-Id: {instance_id}

Response: JSON array of PolicyDefinition dicts:

[
  {
    "id": 1,
    "name": "No PII to external LLMs",
    "enforcement_point": "step_execution",
    "rule_type": "pii_in_request",
    "params": {"patterns": ["\\d{3}-\\d{2}-\\d{4}"]},
    "severity": "critical",
    "enabled": true,
    "agent_id": null,
    "risk_classification": null
  }
]

Fields consumed at evaluation time: id, name, enforcement_point, rule_type, params, severity, enabled, agent_id, risk_classification.

Network failures are logged and swallowed. The runner falls back to the previously loaded policy set and retries after the TTL. The TTL clock is stamped on failure to prevent a down API from causing every evaluate() call to block on a re-fetch attempt.

HMAC verification (opt-in). When KV_POLICY_HMAC_SECRET is set on both the engine and the API, the API computes HMAC-SHA256(secret, response_body) and includes it in the X-Kyvvu-Policy-Signature header. The engine verifies the signature on receipt. If the signature is missing or invalid, the fetch is rejected and cached policies are kept. This prevents policy tampering by a compromised proxy or MITM within the internal network.

Disk cache (opt-in). When KV_POLICY_CACHE_PATH is set, the runner writes the fetched policies to disk after each successful fetch (atomic write via temp file + rename). On cold start, if the API is unreachable, the runner loads policies from this disk cache. A staleness warning is emitted if the cache exceeds KV_POLICY_CACHE_MAX_AGE_SECONDS (default 24h), but the cache is still used.

Fail-mode. When KV_POLICY_FAIL_MODE=closed, the runner blocks all step_execution behaviors if no policies could be loaded (from API or disk cache). Default is open (current behavior — allow all when no policies are available).

2. POST {log_location} — step log flush

Called on end_task() when KV_LOG_LOCATION resolves to an HTTP URL and steps are buffered.

POST {log_location}
Authorization: Bearer {api_key}
X-Kyvvu-Instance-Id: {instance_id}
Content-Type: application/json

{
  "agent_id": "agent-123",
  "task_id": "task-abc",
  "steps": [
    {
      "step_type": "step.model",
      "verb": "POST",
      "step_name": "chat_gpt-4o",
      "properties": {"model": {"provider": "openai", "name": "gpt-4o"},
                     "usage": {"total_tokens": 1250}},
      "meta": null,
      "input": {"user_message": "..."},
      "output": {"response": "..."},
      "timestamp": "2026-04-23T10:00:00+00:00"
    }
  ]
}

Payload redaction. For GDPR-sensitive environments, set KV_LOG_PAYLOADS=metadata_only. In this mode, each step's input and output fields are replaced with {"redacted": true, "keys": [...], "length": N} — shape preserved, content stripped. Default is full.

Response: {"steps_logged": N, "hash_tail": "..."} — only these two fields are consumed. HTTP errors are logged at WARNING and swallowed.

3. Incident sink — POST …/api/v1/incidents (kv)

Fired from evaluate() or evaluate_registration() when the action is warn or block, and routed through the incident sink. The incident sink follows the same location + format model as traces (incident_location / incident_format) and defaults to the trace sink — so if traces flow to the platform, incidents do too. The one asymmetry: for the kv format the Kyvvu HTTP path is …/api/v1/incidents (vs …/api/v1/logs/batch for traces). With incident_format=otlp a standalone kyvvu.incident span is emitted instead. The single entry point is KyvvuRunner.report_incident(payload). Disabled only when the sink resolves to none/off/empty.

Step-execution incident:

{
  "agent_id": "agent-123",
  "enforcement_point": "step_execution",
  "task_id": "task-abc",
  "step_name": "chat_gpt-4o",
  "step_type": "step.model",
  "action": "block",
  "risk_score": 1.0,
  "violations": [
    {
      "policy_name": "No PII to external LLMs",
      "severity": "critical",
      "details": {"matched_pattern": "\\d{3}-\\d{2}-\\d{4}"}
    }
  ],
  "timestamp": "2026-04-23T10:00:00+00:00"
}

Agent-registration incident: same shape with enforcement_point: "agent_registration" and no task_id / step_name / step_type.

Response: status code only; body is ignored. Errors are logged at WARNING and swallowed — incident reporting never blocks the agent.


Configuration

KyvvuRunner is configured via KyvvuSettings. Three equivalent patterns:

# Explicit kwargs
runner = KyvvuRunner(api_url="…", api_key="…", agent_key="…")

# Shared settings object
settings = KyvvuSettings(api_url="…", api_key="…")
runner = KyvvuRunner(settings=settings)

# Pure env-var driven
# export KV_API_URL=…  KV_API_KEY=…  KV_AGENT_KEY=…
runner = KyvvuRunner()

Precedence (highest to lowest): explicit kwargs → environment variables → .env in cwd → built-in defaults.

Authentication and identity

Setting Env var Default Purpose
api_url KV_API_URL http://localhost:8000 Base URL of the Kyvvu platform API.
api_key KV_API_KEY — Bearer API key. Required for policy fetch.
agent_key KV_AGENT_KEY — Stable agent identifier used to fetch policies.
instance_id KV_INSTANCE_ID auto-generated Identifier for this runner instance.

Local policy source (no credentials)

Setting Env var Default Purpose
policy_file KV_POLICY_FILE empty (disabled) One or more manifest YAML paths, comma-separated. Parsed by kyvvu_engine.manifests and enforced with no credentials and no network.
policies — None An inline policy set, in PolicyEngine.load_policies() shape. No env var: a policy set is not an environment-variable-shaped value. [] is a source that offers nothing (none_configured), not the absence of a source.

Either one makes api_key and agent_key unnecessary: they are required only when policies are actually fetched. See Where policies come from.

Log output

Setting Env var Default Purpose
log_location KV_LOG_LOCATION auto WHERE logs go. URL → HTTP POST, file path → JSONL, stdout → terminal, none/empty → disabled, auto → the platform when it is the policy source, else stdout (#357).
log_format KV_LOG_FORMAT kv HOW logs are formatted: kv (Kyvvu batch API), json, or otlp.
incident_location KV_INCIDENT_LOCATION unset → inherit trace sink WHERE incidents go. Same vocabulary as log_location; unset inherits the trace sink with the …/api/v1/incidents path for kv.
incident_format KV_INCIDENT_FORMAT unset → inherit log_format HOW incidents are formatted: kv, json, or otlp (standalone kyvvu.incident span).

Behaviour

Setting Env var Default Purpose
environment KV_ENVIRONMENT production Forwarded to EvalContext.environment.
log_payloads KV_LOG_PAYLOADS full full includes step input/output; metadata_only redacts them.

Cache and limits

Setting Env var Default Purpose
policy_ttl_seconds KV_POLICY_TTL_SECONDS 300 How long to cache fetched policies.
http_timeout_seconds KV_HTTP_TIMEOUT_SECONDS 10 Per-request HTTP timeout.
task_max_age_seconds KV_TASK_MAX_AGE_SECONDS 3600 Abandoned-task eviction threshold for sweep_stale_tasks().

Resilience

Setting Env var Default Purpose
fail_mode KV_POLICY_FAIL_MODE open open = allow all when no policies loaded; closed = block all step_execution behaviors.
policy_cache_path KV_POLICY_CACHE_PATH empty (disabled) File path for on-disk policy cache. Written after each successful fetch; loaded on cold start if API is down.
policy_cache_max_age_seconds KV_POLICY_CACHE_MAX_AGE_SECONDS 86400 Max age (seconds) of disk cache before a staleness warning. Cache is still used when stale.
policy_hmac_secret KV_POLICY_HMAC_SECRET empty (disabled) Shared secret for HMAC-SHA256 verification of the X-Kyvvu-Policy-Signature header on policy fetch responses.

Logging

Setting Env var Default Purpose
log_level KV_LOG_LEVEL INFO Log level for kyvvu / kyvvu_engine loggers.

Instance identification

Each runner instance gets a unique instance_id to disambiguate observability across horizontally scaled agents:

  • If KV_INSTANCE_ID is set (e.g. injected by Kubernetes as a pod name), a random 5-character suffix is appended to prevent collisions when orchestrators reuse names: KV_INSTANCE_ID=worker-3 becomes worker-3-a8f92.
  • If KV_INSTANCE_ID is unset, a random UUID is generated at runner construction time and remains stable for the runner's lifetime.

The instance_id is sent on every HTTP request as both a query parameter (?instance=...) and a header (X-Kyvvu-Instance-Id: ...).


Logging & Export

The engine supports pluggable exporters that control where and how task traces are sent.

Two settings, consistently named: KV_LOG_LOCATION (WHERE) and KV_LOG_FORMAT (HOW).

location format Exporter
URL kv HttpExporter (POST to URL/api/v1/logs/batch)
URL otlp OtlpExporter (POST to URL)
stdout any StdoutExporter (pretty JSON to stdout)
file path json FileExporter (JSONL)
none / empty any Disabled
# Local development
export KV_LOG_LOCATION=stdout

# Send to Kyvvu API
export KV_LOG_LOCATION=https://platform.kyvvu.com
export KV_LOG_FORMAT=kv

# Send to OTLP collector
export KV_LOG_LOCATION=https://otlp-collector:4318
export KV_LOG_FORMAT=otlp

# Write to file
export KV_LOG_LOCATION=/var/log/kyvvu.jsonl
export KV_LOG_FORMAT=json

See the Configuration Reference for the full env-var list.

Policy evaluation in audit trail

Every step recorded via record() automatically includes the policy evaluation result in meta["kyvvu.eval"] when a preceding evaluate() call was made:

{
  "meta": {
    "kyvvu.eval": {
      "action": "allow",
      "risk_score": 0.0,
      "policies": [
        {
          "name": "require_purpose",
          "rule_type": "field_not_empty",
          "severity": "medium",
          "violated": false
        }
      ]
    }
  }
}

This data flows through all exporters unchanged. The OTLP exporter maps it to span attributes (kyvvu.eval.action, kyvvu.eval.risk_score, kyvvu.eval.policies_checked, kyvvu.eval.policies_violated).

Manifest provenance in audit trail

When policies are materialised from a manifest assignment, the runner attaches their origin to the audit trail so a reviewer can answer "which manifest governed this?" without leaving the trace:

  • Per-policy — each entry in meta["kyvvu.eval"].policies carries manifest_id when the policy came from a manifest (omitted for hand-authored policies).
  • Per-task — the task.start step's meta["kyvvu.manifests"] lists the distinct manifest records in effect for the whole task. Each record carries manifest_id and manifest_name, plus the fuller origin fields (repo, path, sha, assigned_at, assigned_by) when the API supplied them.

The runner derives these records from the loaded policy set and forwards the in-effect manifests on EvalContext.manifests so rules and exporters can read them. This is populated automatically — no call-site changes are required.

OTLP integration

The OTLP exporter maps kyvvu data to OpenTelemetry spans:

  • Task → root span (kyvvu.task {task_id})
  • Step → child span (kyvvu.step {step_type})
  • Span attributes include kyvvu.step_type, kyvvu.step_name, kyvvu.verb, flattened properties, and eval results.

Span timing does not reflect actual execution timing (steps are recorded after the fact). The value is in the attributes and parent/child relationships.

# Quick test with Jaeger
docker run -d --name jaeger -p 16686:16686 -p 4318:4318 jaegertracing/all-in-one:latest
export KV_LOG_LOCATION=http://localhost:4318
export KV_LOG_FORMAT=otlp
# Run your agent, then open http://localhost:16686
# Search for service "kyvvu-engine"

Policy fetch resilience

The runner provides four opt-in mechanisms to harden policy delivery. All are backward compatible — when unconfigured, the runner behaves exactly as before.

Fail-open vs fail-closed

By default, the runner operates in fail-open mode: if no policies can be loaded, all steps are allowed. This keeps agents running during API outages.

For high-risk production deployments, set KV_POLICY_FAIL_MODE=closed. In this mode, if the engine has zero policies (no API, no disk cache), evaluate() raises KyvvuBlockedError with a synthetic no_policies_available violation. The agent must handle this — typically by pausing work until policies are restored.

Disk cache

Set KV_POLICY_CACHE_PATH=/var/lib/kyvvu/policy-cache.json to enable the on-disk policy cache.

  • Write: After each successful API fetch, policies are written to disk atomically (temp file + os.replace). Concurrent readers never see a partial file.
  • Read: On cold start, if the API fetch fails and the engine has zero policies, the runner loads from the disk cache. A staleness warning is emitted if the cache exceeds KV_POLICY_CACHE_MAX_AGE_SECONDS (default 24 hours).
  • The disk cache is a fallback only — in-memory policies from the API always take precedence.
  • When the API recovers, fresh policies replace the disk-cached set.

HMAC policy signing

Set KV_POLICY_HMAC_SECRET to the same value on both the API and the engine. The API computes HMAC-SHA256(secret, response_body) and sends it in the X-Kyvvu-Policy-Signature response header. The engine verifies the signature; if it is missing or invalid, the fetch is rejected and cached policies are kept.

This prevents a compromised proxy from silently modifying policies to weaken enforcement (e.g. disabling a critical rule) — even on networks where TLS is terminated upstream.

Policy status observability

runner.policy_status() returns a PolicyStatus object with programmatic fields:

Field Type Meaning
loaded bool True if policies have been loaded at least once.
stale bool True if last fetch failed and cache has exceeded TTL.
source str "api", "local", "disk_cache", or "none".
last_success datetime | None Wall-clock time of last successful fetch.
last_attempt datetime | None Wall-clock time of last fetch attempt (success or failure).
policy_count int Number of active policies (i.e. how many are actually enforced).
ttl_remaining_seconds float Seconds until cache expires and a re-fetch is triggered.
load_state PolicyLoadState Why the current set looks the way it does — see below.
policies_offered int How many policies were handed to the loader. Read against policy_count.
dropped list[DroppedPolicy] Every offered policy that is not enforced, with the reason.

Use this in health checks, observability dashboards, or agent startup logic to decide whether to proceed when policies are stale.

Knowing when nothing is being enforced

A policy count cannot tell four very different situations apart, and all four evaluate as allow. PolicyLoadState can:

State Meaning
enforcing At least one policy is loaded and being enforced.
none_configured The load succeeded and the source genuinely has no policies.
unavailable Nothing is loaded — no load attempted, or the fetch failed with no cache to fall back on.
stale_cache The fetch failed and a disk-cached set is enforcing; it may be older than the source's current set.
all_dropped The source offered policies and every one was dropped as unloadable. The configuration looks populated and the engine enforces nothing.

A dropped policy is a policy that is not being enforced, and a policy drops in two independent places: here, when its rule_type is not in the rule registry, and in PolicyStore, when PolicyDefinition validation rejects it. PolicyEngine.dropped and KyvvuRunner.dropped_policies merge both; PolicyStore.dropped keeps its narrower meaning for its own callers.

Whenever the effective enforcing set is empty — or part of the offered set was dropped — one message goes to stderr, naming which state it is:

KYVVU: 0 policies in effect — this agent is UNGOVERNED. state=all_dropped: the
policy source offered 15 policies and every one of them was dropped as
unloadable, so none of them is being enforced. Fix these and reload: …
KYVVU: enforcing 12/15 policies for agent_key='billing-bot' — 3 policies
dropped and NOT enforced (state=enforcing): …

stderr rather than a logger, because this has to reach someone running python agent.py who has configured no logging handler. Emitted at most once per process per distinct outcome, where "outcome" includes the agent_key and the source as well as the state, the counts and the dropped policies with their reasons — two different agents both reporting none_configured are two ungoverned agents, not one repeated message. A forked child starts with an empty ledger. The whole announcement is also incapable of raising: it runs on the evaluate path, where one of its callers turns any exception into an allow verdict, so a broken stderr must never let the report of non-enforcement cause non-enforcement.

stale_cache is a property of the load, not of which source the set came from: a long-lived agent whose API stops answering keeps enforcing the set it has and reports stale_cache at the next refresh attempt. PolicyStatus.stale is a different, narrower field and keeps its pre-#441 meaning (a failed fetch and a cache past its TTL), so the two can legitimately disagree.

A runner with no policy source fails at construction

KyvvuRunner.__init__ raises KyvvuConfigError when the resolved settings name no policy source at all, rather than deferring to the first evaluate(). KyvvuSettings.policy_source_kind() decides: the API (api_url and api_key and agent_key), else a configured policy_cache_path, else nothing. It judges configuration and performs no I/O — a cache file that has not been written yet is runtime state, failing construction on it would fail a first run and pass the second, and on an enforcement path that converts exceptions into "allow" it would turn a loud "0 policies" into a silent permit.

A cache-only configuration is a real source, not just a legal one: KyvvuRunner.ensure_policies() (and the lazy load inside evaluate()) dispatch on the resolved source and read the cache directly, so a runner with no API credentials enforces its cached policy set instead of raising agent_key is required to fetch policies. fetch_policies() keeps that raise — it is the API transport primitive and nothing else. kyvvu serve boots such a configuration too.

Where policies come from

KyvvuSettings.policy_source_kind() resolves one PolicySourceKind, and the first fully configured source wins:

Order Source Configured by Credentials
1 api api_url and api_key and agent_key yes
2 local policies=[…], or policy_file / KV_POLICY_FILE no
3 disk_cache policy_cache_path no (but written by an earlier fetch)
4 none nothing — refused at construction —

local sits above disk_cache because the two mean different things: naming a manifest file names the source, while a cache path names a place a fetch may once have written to. It sits below api so that a manifest dropped on a credentialed host cannot quietly demote it off the authenticated, current policy set — the credential-free path is for trying the engine, not for shadowing production policy.

One exception, and it is the documented precedence rather than a special case: a local source you pass by name outranks API credentials that were only ever in the environment. KyvvuRunner(policy_file="my-policies.yaml") is the credential-free entry point, and KV_API_KEY is exported on the machine of everyone who has ever used the platform — without this it would silently become network-dependent, and enforce nothing, for exactly those readers. Name both explicitly and api still wins; leave both to the environment and api still wins.

A local set is re-read once per policy_ttl_seconds, so editing a manifest takes effect in a long-running process without a restart.

When a local manifest cannot be read

Two different moments, because they are two different problems.

At construction, when policy_file is the only source: KyvvuRunner(…) raises KyvvuConfigError, naming the path, the reason and both the kwarg and the environment variable. A typo, a directory, an unreadable file, malformed YAML and a file that is not a manifest are all knowable when the configuration is supplied, and #427's thesis is that a static configuration error belongs there rather than on the first governed step. kyvvu serve and kyvvu-serve exit 1 instead of booting a server that would enforce nothing.

At load time, when there is something to fall back to: the loader never raises — it runs on the evaluate() path, where one caller (kyvvu-claude's hooks) converts any exception into an allow, so a raise there would let a bad path read as a permit. Instead it degrades and says so:

  • An inline policies list is kept, not discarded. The composition KyvvuRunner(policies=load_baseline_policies(), policy_file=<theirs>) degrades to the baseline when the user's half has a typo in it.
  • A configured policy_cache_path is read, the same fallback the API path has had since #441.
  • A set from an earlier good load keeps enforcing.

In every one of those the set is marked not current — load_state becomes stale_cache, not enforcing — and one line goes to stderr naming the path and the reason, even though something is still being enforced, because a silent substitution of one policy set for another is the failure mode this whole area exists to remove. A failed reload never refreshes ttl_remaining_seconds: retry timing and freshness are separate clocks.


Debugging and explainability

Set KV_LOG_LEVEL=DEBUG for full per-evaluation traces:

kyvvu_engine.engine DEBUG load_policies(): loaded 8/8 policies (0 dropped)
kyvvu_engine.engine.load DEBUG   policy id=1 name='no_pii' rule_type=pii_in_request severity=critical enforcement_point=step_execution
kyvvu_engine.engine.load DEBUG   policy id=2 name='domain_allowlist' rule_type=domain_allowlist severity=medium enforcement_point=step_execution
...
kyvvu_engine.engine DEBUG evaluate(): agent_id=agent-123 task_id=task-abc step_type=step.model verb=POST
kyvvu_engine.engine.eval DEBUG   policy 'no_pii': rule=pii_in_request → FAIL
kyvvu_engine.engine.eval DEBUG   policy 'domain_allowlist': rule=domain_allowlist → pass
kyvvu_engine.engine DEBUG evaluate(): agent_id=agent-123 step_type=step.model → action=block risk_score=1.00 (2 policies)

DEBUG-level output includes every policy loaded (on load_policies) and every policy's result (on each evaluate). If a policy does not appear here, the platform did not send it.

For structured JSON logging:

from kyvvu_engine import setup_logging
setup_logging(level="DEBUG", json=True)

For human-readable per-evaluation traces:

print(engine.explain(intended, context))
Evaluated 8 policies for step.model/POST "chat_gpt-4o" (task=task-abc step=5):
  ✓ domain_allowlist           (medium)   passed
  ✗ pii_in_request             (critical) FAILED: matched \d{3}-\d{2}-\d{4}
  ✓ step_requires_gate         (high)     passed
  ...
→ action=block (risk_score=1.00)

For compound rules, explain() renders the condition tree with pass/fail at each node.


Performance

The engine is designed for sub-millisecond evaluation on the hot path. Targets are indicative — actual numbers are machine-dependent and are measured per-release via tests/test_latency.py. Tests use absolute thresholds (e.g. p99 < 10 ms) as hard gates to catch catastrophic regressions while tolerating normal CI variance.

Scenario Target (p95)
Evaluate with 0 policies < 50 µs
Evaluate with 10 policies, empty history < 200 µs
Evaluate with 10 policies, 20-step history < 500 µs

End-to-end latency including KyvvuRunner.evaluate() is dominated by network I/O when a policy refresh or incident webhook fires; the engine-only numbers are the floor.

Run benchmarks locally:

pip install -e ".[dev]"

# Or, from the repo root, with uv (this repo is a uv workspace):
uv sync --all-packages --all-extras
pytest tests/test_latency.py -v -s

Running as a standalone service

For callers that aren't Python, kyvvu-engine runs as a local HTTP server. Install the SDK (which includes the engine) and use the kyvvu serve command:

pip install kyvvu

# Offline: no account, no API key, no network.
kyvvu serve --baseline
kyvvu serve --policy-file my-policies.yaml

# Platform-connected: fetches the agent's assigned policy set.
kyvvu serve --host 127.0.0.1 --port 8080 --agent-key my-agent

CLI arguments:

Flag Default Purpose
--host 127.0.0.1 Bind address.
--port 8080 Bind port.
--policy-file from KV_POLICY_FILE Manifest YAML to enforce locally. Repeatable. Needs no credentials.
--baseline off Enforce the frozen baseline set bundled in kyvvu. Needs no credentials.
--agent-key from KV_AGENT_KEY Agent key for policy fetch.
--api-url from KV_API_URL Kyvvu platform API URL.
--api-key from KV_API_KEY Bearer API key.

Credentials are required only for the fetching mode. A local source clears them for that process, so --baseline in a shell that exports KV_API_KEY still runs the bundled set rather than quietly fetching a different one. This answers the open "kyvvu serve — local server mode for offline/airgapped development. Is this documented and working?" checkbox from #201: it is.

kyvvu-engine's own kyvvu-serve entry point takes --policy-file too, but not --baseline: the frozen set is policy content and ships in the kyvvu SDK, not in the engine.

All KV_* environment variables and .env files work identically to KyvvuRunner.

Endpoints

Method Path Wraps Purpose
GET /health policy_status() Liveness probe — returns PolicyStatus JSON.
POST /evaluate evaluate() Preflight evaluation of an intended behaviour.
POST /register_agent evaluate_registration() Evaluate agent-registration policies.
POST /record record() Record a completed step. Returns {"step": <int>, "task_id": "<str>"}.
POST /end_task end_task() Close a task — evict history and flush logs. Returns {"status": "ok", "task_id": "<str>"}.

Example

curl http://127.0.0.1:8080/health
{"loaded": true, "stale": false, "source": "api",
 "policy_count": 8, "last_success": "2026-04-24T10:00:00+00:00",
 "last_attempt": "2026-04-24T10:00:00+00:00",
 "last_fetch_at": "2026-04-24T10:00:00+00:00",
 "last_fetch_succeeded": true, "instance_id": "worker-3-a8f92",
 "ttl_remaining_seconds": 280.5}
curl -X POST http://127.0.0.1:8080/evaluate \
  -H "Content-Type: application/json" \
  -d '{
    "intended": {
      "agent_id": "agent-123",
      "task_id": "task-abc",
      "step_type": "step.model",
      "verb": "POST",
      "step_name": "chat_gpt-4o",
      "input": {"user_message": "Hello"}
    },
    "context": {
      "agent_id": "agent-123",
      "task_id": "task-abc",
      "environment": "production",
      "risk_classification": "limited"
    }
  }'

Response:

{
  "action": "allow",
  "risk_score": 0.0,
  "policies": [
    {"policy_id": 1, "name": "pii_in_request", "severity": "critical",
     "violated": false, "violation_details": null}
  ],
  "blocked": false
}

When a policy blocks, blocked is true and action is "block". The server never returns a non-200 status for policy decisions — the caller reads blocked to decide whether to proceed.

The serve layer inherits the runner's sink configuration. Set KV_LOG_LOCATION=stdout (and KV_INCIDENT_LOCATION=stdout, or leave it unset to inherit the trace sink) to emit JSON-line output for local debugging without an API backend.

The Python SDK uses KyvvuRunner directly and does not need this server.


Multi-agent and branching patterns

The engine is framework-agnostic. It does not know LangGraph, AutoGen, or CrewAI exist. Multi-agent and branching patterns are handled entirely in the kyvvu SDK via behavioural templates — the engine evaluates policies against whatever Behavior objects templates emit.

The engine commits to two conventions for template authors:

Reserved meta keys. When a Behavior represents a step in a subtask, the template sets:

  • meta.parent_task_id — the task_id of the invoking parent task.
  • meta.parent_agent_id — the agent_id of the invoking parent agent.

Rules can read these via dot-path accessors. No rule primitives specific to multi-agent reasoning are required — the generic compound rules (all_of / any_of / not) plus history_contains cover the cases.

Cross-subtask aggregation. Policies that reason across sibling branches or parent/child tasks use EvalContext.cross_execution_counts, pre-fetched by the platform aggregating over parent_task_id. This is the same mechanism cross_execution_rate_limit uses.

The engine does not track branching paths as a DAG; histories are linear per task_id. If a DAG-aware history model is needed, it is a future-version change — sibling-subtask modelling is sufficient in the cases encountered so far.


Stability and versioning

Semantic versioning. The public API surface is:

  • PolicyEngine and its documented methods.
  • KyvvuRunner and its documented methods.
  • KyvvuSettings and its documented fields.
  • Behavior, EvalContext, EvalResult, PolicyResult, PolicyDefinition, PolicyStatus, AgentRecord, Action, StepType, Verb, TASK_BOUNDARY_STEP_TYPES.
  • PolicyRule and the built-in rule functions.
  • Aggregators: aggregate_max, aggregate_mean, aggregate_weighted_sum.
  • KyvvuBlockedError, KyvvuConfigError.
  • setup_logging.

Everything else (internal helpers, underscore-prefixed modules, deeper import paths) is private and may change between minor versions.

  • Before 1.0: minor versions may introduce breaking changes with a CHANGELOG entry.
  • From 1.0: breaking changes require a major-version bump.

See also


Licence

kyvvu-engine is source-available under the Business Source License 1.1 (BSL 1.1). It is not open source in the OSI sense.

  • Free use is permitted for development, testing, research, evaluation, and personal non-commercial purposes.
  • Production use requires a Kyvvu commercial subscription or a separate license agreement with Kyvvu B.V.
  • Each release converts to Apache License 2.0 four years after its publication date.

See LICENSE in this directory for the full terms.

Commercial licences: licensing@kyvvu.com

Release files for kyvvu-engine 0.11.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for kyvvu-engine 0.11.1
File Size Uploaded
kyvvu_engine-0.11.1.tar.gz 208.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for kyvvu-engine 0.11.1
File Interpreter ABI Platform
kyvvu_engine-0.11.1-py3-none-any.whl Python 3 none any Details

Total release size: 352.1 kB

Release files / kyvvu_engine-0.11.1.tar.gz

Download URL kyvvu_engine-0.11.1.tar.gz
Size 208.0 kB
Tags Source
SHA-256 checksum
How to use checksums
088cc8e2a1c5864a5fc66d91354467be83b77e647450c5ca8025b32607aed492
BLAKE2b-256 checksum
How to use checksums
6358a7370103417b09368b3e5ab73e0956cb70f4b2dfa16adbbb54a5d7281922
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / kyvvu_engine-0.11.1-py3-none-any.whl

Download URL kyvvu_engine-0.11.1-py3-none-any.whl
Size 144.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4067eca575e5a127a9bffdec8dd9de49c821c0a30e662534f249dc2d9bb76c1b
BLAKE2b-256 checksum
How to use checksums
8a9c1af244a6a140e082ecc1ccbe304750224c04d978fd265f9d949ce48f460d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.11.1 This release

2 release files

0.11.0

2 release files

0.10.0

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page