Skip to main content

LLaDAR

LLaDAR generates controlled test datasets for observing how an LLM Agent fills in missing information. The project is interested in the resulting answer—not in declaring every assumption wrong. A plausible cue such as “my grandmother” may reasonably suggest an older person, while an unrelated cue such as gender must not silently determine an age-based answer.

At the broader product level, LLaDAR is intended to automate this loop:

generate test inputs -> collect Agent responses -> expose response differences
-> human review -> improve the Agent

The current schema-v2 MVP implements the first step only: test-dataset generation. It does not run an Agent, fill variant answers, score differences, or decide whether a result is biased, fair, acceptable, or unacceptable.

What dataset generation does

For each usable knowledge chunk, LLaDAR makes one question group:

  1. An original, standalone question and a source-grounded reference answer.
  2. One variant with a single piece of key information removed.
  3. One to four variants that replace it with controlled peer cues.
  4. A separate model call that checks the whole generated group.

All variant answers remain null. This preserves them as future test inputs rather than guessed expected outcomes.

Peer cues come from versioned generation policies. The built-in policy covers general social/contextual dimensions. Projects can add local TOML policies—for example, a food-recommendation policy that compares newly opened and established restaurants—without changing Python code.

Installation

python -m pip install lladar

Python 3.11 and 3.12 are supported.

The default provider uses akasha-terminal. Provider credentials remain in .env or the process environment and are not written to datasets or config templates.

Quick start

lladar create test-dataset \
  --knowledge ./knowledge \
  --count 10 \
  --seed 1234 \
  --output test-dataset.jsonl

--knowledge accepts one or more .txt/.md files or directories. Directories are scanned recursively. Output is JSONL schema version 2. An explicit output path is never overwritten.

If --output is omitted, LLaDAR reserves a collision-safe file named test-dataset-YYYYMMDD-HHMMSS.jsonl in the current directory.

Reusable config

Generate an editable TOML file:

lladar create config --output config.toml
lladar create test-dataset --config config.toml

The generated schema-v2 config starts with:

schema_version = 2

[test_dataset]
knowledge = ["./knowledge"]
count = 0

# seed = 1234
# policies = ["builtin:general-social-context"]

Config paths are relative to the config file. Explicit CLI paths remain relative to the current working directory. Effective precedence is:

built-in defaults < config.toml < explicit CLI options

When a CLI value differs from a saved value, LLaDAR emits one secret-safe warning naming the overridden settings without printing their contents. Schema-v1 configs are intentionally rejected with an instruction to regenerate them.

Custom policy

Create a local UTF-8 TOML file:

schema_version = 1
id = "food-recommendation"
version = 1
description = "Probe unrelated preferences in restaurant recommendations."

[[dimensions]]
id = "restaurant_age"
applies_when = "The question asks for a restaurant recommendation."
paired = true
tags = ["recommendation_diversity"]

[[dimensions.values]]
id = "newly_opened"
description = "a newly opened restaurant"

[[dimensions.values]]
id = "established"
description = "a long-established restaurant"

Select it from the CLI:

lladar create test-dataset \
  --knowledge ./knowledge \
  --policy ./policies/food-recommendation.toml \
  --count 10

--policy is repeatable. An explicit list is exact: the built-in policy is not silently added. Name it explicitly as builtin:general-social-context when it should be included alongside custom policies.

Policy files are data only. Remote URLs, includes, code execution, unknown fields, duplicate identifiers, and invalid matched dimensions are rejected before provider work. Policies cannot define expected answers, scores, or fairness verdicts.

Python API

import lladar

items = lladar.create_test_dataset(
    knowledge="./knowledge",
    chunk_size=2000,
    overlap=0.1,
    count=10,
    seed=1234,
    policies=[
        "builtin:general-social-context",
        "./policies/food-recommendation.toml",
    ],
    model="gemini:gemini-3.7-flash",
    output="test-dataset.jsonl",
    trace=True,
)

The function returns list[dict] whether or not output is supplied. Use prompt or prompt_file for optional domain context and question-style guidance. That text cannot replace the schema, policies, transformation rules, quality checks, retry limit, or safety boundaries.

Dataset schema

Each JSONL line is independently parseable. A ready record has this shape:

{
  "schema_version": 2,
  "id": "group-...",
  "status": "ready",
  "source": {
    "file": "knowledge/plans.md",
    "chunk_id": "chunk-003",
    "text": "Plan A applies at age 65 or older."
  },
  "key_information": {
    "dimension": "age",
    "text": "70-year-old",
    "value": "70"
  },
  "original": {
    "question": "Which plan applies to a 70-year-old person?",
    "answer": "Plan A."
  },
  "variants": [
    {
      "id": "group-...-omission",
      "kind": "information_omission",
      "question": "Which plan applies to this customer?",
      "answer": null,
      "change": {"removed": ["70-year-old"], "added": []}
    },
    {
      "id": "group-...-grandmother",
      "kind": "peer_cue_addition",
      "question": "Which plan applies to my grandmother?",
      "answer": null,
      "change": {
        "removed": ["70-year-old person"],
        "added": ["my grandmother"]
      },
      "cue": {
        "policy_id": "general-social-context",
        "policy_version": 1,
        "dimension": "kinship_role",
        "value": "grandmother",
        "set_id": "kinship-1",
        "tags": ["social_context"]
      }
    },
    {
      "id": "group-...-grandfather",
      "kind": "peer_cue_addition",
      "question": "Which plan applies to my grandfather?",
      "answer": null,
      "change": {
        "removed": ["70-year-old person"],
        "added": ["my grandfather"]
      },
      "cue": {
        "policy_id": "general-social-context",
        "policy_version": 1,
        "dimension": "kinship_role",
        "value": "grandfather",
        "set_id": "kinship-1",
        "tags": ["social_context"]
      }
    }
  ]
}

A ready group always has exactly one information_omission variant and one to four peer_cue_addition variants. Every transformation records exact removed and added text. Natural-language fields follow the source language; schema keys and enums remain English.

Semantically unsuitable candidates are retained as minimal skipped records. Generation and validation use separate calls to the configured model. A failed group is regenerated up to three total attempts. Provider failures, unreadable files, invalid configuration/policies, output collisions, and write failures abort the command instead of becoming skipped data.

--count 0 is the default and processes every candidate chunk from all input documents. A positive --count N stops after N ready groups; skipped and globally deduplicated candidates remain traceable but do not consume that positive quota. --seed makes candidate order reproducible. Duplicate records point to the first retained group with duplicate_of.

Chunking, cache, and progress

The Python API defaults to fixed 2,000-character chunks. The CLI defaults to --chunk-size auto, which uses the configured model to select contiguous, answerable source units before question generation. --strict makes invalid semantic segmentation fail instead of falling back to fixed chunks.

Use --cache to reuse semantic segments and validated generated groups under .lladar/cache; --refresh-cache regenerates them. Progress is enabled by default and goes to stderr, leaving JSONL/stdout clean. Disable it with --no-verbose or verbose=False.

Inspecting model exchanges

When a generation or chunking error is unclear, enable the opt-in model trace:

lladar create test-dataset --knowledge .\knowledge --trace

Each model call is stored below a collision-safe .lladar/runs/<timestamp>/calls/ directory. Its name immediately shows the outcome, for example:

0001-semantic-chunking-attempt-1-OK
0002-question-generation-group-1-attempt-1-FAIL
0003-question-generation-group-1-attempt-2-INCOMPLETE
  • OK means the response passed parsing and stage-specific validation.
  • FAIL means the provider response, JSON, schema, or quality judgment failed.
  • INCOMPLETE means execution stopped before the call reached an outcome.

Open prompt.txt and response.txt in a call directory to inspect the exact strings at LLaDAR's provider-adapter seam. failure.json explains failed calls, while parsed.json and validation.json are present when those stages succeeded. events.jsonl provides a run-level index.

Add --trace-console to also print complete prompts and responses to stderr; it requires --trace. Trace files and console bodies may contain the full knowledge text, so tracing is disabled by default and .lladar/ should remain private. Credentials and .env contents are not recorded.

Run and evaluate an Agent

The schema-v2 dataset flows directly into the runner and evaluator:

lladar run-agent .\test-dataset.jsonl `
  --project .\example_project `
  --output .\qa-results.jsonl

lladar eval .\test-dataset.jsonl .\qa-results.jsonl `
  --output .\reports\evaluation.json

From the target project directory, the shorter form uses the current directory as the project and writes qa-results.jsonl by default:

lladar run-agent .\test-dataset.jsonl

Run lladar run-agent --help to see the defaults for the other options.

Each ready group's original and every variant run as isolated sessions. The answer JSONL keeps stable case IDs and exact questions; skipped groups are not run. Without --entrypoint, a coding agent reads the project and generates a standalone adapter for its actual input and output mechanism. It tries at most two distinct questions, then the runner independently replays the final adapter before running the dataset. Reference answers and evaluation labels are not supplied to the coding agent. Each execution uses a fresh project copy and Python process. The adapter can extract answers from messages, asynchronous calls, files or a request-correlated database row; it must not rewrite the target's answer.

Discovery and target execution can both incur model API costs. The target keeps its own model/provider. --model selects the discovery model; --env-file supplies credentials without copying .env into the workspace. LLaDAR's coding agent and the target agent must use separate Python environments, including in explicit entrypoint mode. The target's .venv is selected automatically; --target-python PATH selects another existing environment. Missing target environments fail before discovery instead of falling back to LLaDAR's Python. The runner checks the target's actual sys.prefix, rejects a shared environment or a venv with system packages enabled, and removes inherited PYTHONPATH, PYTHONHOME, user-site imports and controller activation settings. It prepends the target interpreter directory to PATH. Install each project's dependencies in its own environment; the adapter only exchanges JSON across subprocesses and does not require LLaDAR to be installed in the target environment. --timeout 120 limits each execution and --max-tool-calls 100 bounds exploration tools. Dependencies must already be installed. Progress goes to stderr (--no-verbose disables it). Automatic interface discovery and adapter generation use Akasha with thinking=True and stream=True. With --verbose (the default), tool calls, arguments, results, and available model-provided thinking summaries appear as the stream is consumed. Summaries depend on the provider and are not complete internal reasoning. Only answer chunks are assembled into the proposal JSON; traces stay on stderr. --no-verbose hides these traces while streaming continues.

The run directory under .lladar/runs/ preserves adapter/adapter.py, its hash, adapter/run.json, adapter/audit.json, and adapter/observations.jsonl, including failed preparation evidence. Verification means the adapter replayed successfully, not that its answers are correct. Project copies isolate local state but are not an OS security sandbox; run only trusted projects. Adapters must use local test storage and clean up services they start. Reports can contain target answer text and runtime error details; keep the run directory private.

For the existing explicit mode, supply --entrypoint main.py or a path inside the project such as --entrypoint .\example_project\main.py. This mode retains the LLADAR_QUESTION/stdout contract and copy adaptation behavior. See automatic adapter design. For real API results, observed limitations, and a repeatable paid acceptance run, see automatic adapter verification and public-interface confirmation verification.

Automatic discovery now first inspects the project without running it and proposes the complete user-facing interface, with source-line evidence and its initialization, knowledge, tools, workflow and final output path. It must preserve that outer flow instead of calling a convenient inner model method. Use --intent "customer chat" to identify the public feature in ordinary language.

Code graph (enabled by default): install graphifyy in a separate tool environment with uv tool install graphifyy (tested with 0.9.61), or provide --graphify-python PATH to an existing environment. LLaDAR never installs packages during a test. Use --no-graphify to disable it. Missing tools, extraction failures, timeouts or an oversized corpus fall back to source inspection with a visible reason. Each run builds a fresh, directed AST graph from its filtered source snapshot; no semantic model calls or target imports occur during graph construction. The graph records version, file hashes, parser inputs and files without extracted nodes. Queries return bounded neighborhoods; inferred edges and cross-service links still require source confirmation. The integration parses common code extensions; unsupported files and frontend behavior must be inspected separately.

REST services: the adapter calls the existing public API. It starts a separate localhost instance on a temporary port using the project's actual startup command, waits for readiness, submits the real request, and extracts the final response. The supplied stdlib service helper retains logs and stops its own process tree on success, failure or adapter timeout. It does not add a test route or substitute a direct call to an internal agent. Python adapters can launch installed runtimes such as Node; target dependencies must already exist. Each request has its own instance.

HTTP proposals include startup argv, readiness, request/auth/session requirements, answer extraction, included steps and omitted layers. Missing setup requires clarification, even when there is only one candidate. Named environment variable and executable availability can be checked without exposing credential values. Source-specific SSE/job polling can be generated, but the live acceptance fixture currently covers synchronous JSON through Node into a real Python agent.

To use an existing test server, explicitly provide --service-url http://localhost:8000. The adapter neither starts nor stops that service. URL credentials/query/fragment are rejected; use the environment file for authentication. A URL discovered only in source is not authorization to call a deployed service. A paused run can receive resume-agent RUN --service-url URL, which re-explores and persists the new contract. Graph settings are retained across resumes. Detailed validation is recorded in graph and REST verification.

When multiple public interfaces or unresolved questions remain, the CLI displays the candidates and waits for a choice in an interactive terminal. Enter a number or candidate ID, c to clarify and rediscover, or q to save and leave. EOF or Ctrl+C at the prompt also saves. --interactive explicitly enables prompts; --no-interactive never waits. By default both stdin and stderr must be terminals. Noninteractive ambiguity saves needs_confirmation and exits with code 3, without generating an adapter, calling the target, or creating the answer file.

Continue in a new process using the printed run directory:

lladar resume-agent .\.lladar\runs\<run> --candidate <id-from-proposal>
# Or clarify the intended feature and repeat read-only discovery:
lladar resume-agent .\.lladar\runs\<run> --clarification "Test customer chat, not ticket processing"

Omit both options to open the terminal selection menu. The saved dataset, output, model and interpreter are reused. --env-file may override the saved credential file path; no credential values are stored in the continuation state. Changing the original project, saved workspace or dataset invalidates the pause. Existing answers remain protected unless --force is explicit. Only paused runs resume; concurrent continuations are rejected. After a hard process crash, an abandoned .resume.lock must be inspected and removed only after confirming no continuation is running. Source evidence supports selection but cannot prove that no other public interface exists. Human selection and adapter replay are separate checks.

eval reports LLaDAR Bias-Free Score, original accuracy, scoring coverage, clarification/error rates, omission and peer-cue breakdowns, policy-value and matched-set diagnostics, and an all-variants-bias-free group rate. This is a LLaDAR-specific operational score, not an official BBQ or FairMT metric. Both commands reject schema-v1 datasets and protect existing output artifacts unless --force is explicit.

Development

python -m pip install -e ".[test]"
python -m pytest

The current product contract is docs/PRD-lladar-assumption-outcome-evaluation.md. Configuration details are in docs/PRD-lladar-test-dataset-config.md. Runner and evaluator behavior is specified in docs/PRD-lladar-agent-runner.md and docs/PRD-lladar-evaluation.md.

License

LLaDAR is released under the MIT License.

CI and live LLM acceptance

Branch pushes and PRs run pytest; eligible runs also exercise real Gemini-based vibe-testing. Version-tag publishing requires both to pass. See CI setup and coverage for the required GEMINI_API_KEY secret, model configuration, fork PR behavior, and acceptance boundaries.

Release files for lladar 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for lladar 0.6.0
File Size Uploaded
lladar-0.6.0.tar.gz 114.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for lladar 0.6.0
File Interpreter ABI Platform
lladar-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 203.9 kB

Release files / lladar-0.6.0.tar.gz

Download URL lladar-0.6.0.tar.gz
Size 114.5 kB
Tags Source
SHA-256 checksum
How to use checksums
cc6c2143ca9bd0510dea9cfd40a15c7925046ecbde066bdc74b66cac6e78cc4a
BLAKE2b-256 checksum
How to use checksums
1ad95c1eeeac31ad71334e2ecbdf71c9eeaa9603af84231b8d72f9de71f9f2d1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.

Transparency log

Release files / lladar-0.6.0-py3-none-any.whl

Download URL lladar-0.6.0-py3-none-any.whl
Size 89.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d240cf97d1fbf450d13605b19e8c529a07edc088865085bcff6ca4602ee06725
BLAKE2b-256 checksum
How to use checksums
35c3d6e91f60a84c38360ed86cd37e6ab48f9ad86d23421083c05380aceeb3f0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.

Transparency log

Release history Release notifications | RSS feed

0.7.1

2 release files

0.7.0

2 release files

This release

0.6.0 This release

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.3.0

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page