Skip to main content

A reproducible evaluation runner for tool-using Agent skills

Project description

yama

English | 简体中文

yama is a reproducible evaluation framework for tool-using Agent Skills: a declarative YAML Case describes what the model sees — the system prompt, Skill metadata, tool schemas, and multi-turn user messages — then yama calls the LLM through LiteLLM, drives the tool loop, and scores the transcript with deterministic hard checks or an LLM judge. See the design doc (Chinese) for the full Case schema and execution contract.

Quick start

# Run from a plugin root; collects __evals__/cases/**/*.yaml by default
cd dingding-simple
OPENAI_API_KEY=... uv run --project yama yama

# Or, from a workspace root (a directory with yama.toml), pick a plugin by name
OPENAI_API_KEY=... uv run --project yama yama --plugin dingding-simple

Writing a Case

A Case is a YAML file organized as context (what the model sees) → mocks (how tool calls get executed) → steps (user turns sent in order, each with assertions) → outcome (the pass bar for the whole Case):

description: The model reads the DSL before proposing creative directions

context:
  system_prompt: { default: true }     # use SYSTEM.md from the plugin root

  tools:
    - file: tools/read-dsl.yaml         # tool schema, resolved relative to __evals__

mocks:
  tools:
    read_dsl:
      respond:
        result:
          visualStyle: { name: Vintage Film }

steps:
  - id: request-directions
    user: Give me a few creative directions
    assert:
      hard:
        - tool_called: { name: read_dsl }
        - assistant_contains: creative direction

outcome:
  require:
    hard_checks: all_pass
  • description (optional) is a one-line summary of what the Case tests; it is shown in the CLI output and the HTML report.
  • context.tools is the tool schema sent to the model; mocks.tools is how the runner responds when a tool call arrives. The two sets of tool names must match exactly.
  • To use a Skill, declare it by name under context.skills and also declare {builtin: skill} under context.tools; the model reads the Skill body via skill(name, file?).
  • To simulate command-line tools, declare {builtin: bash} and configure per-command output under mocks.cli.
  • steps[].assert.hard are deterministic checks (nine types, including tool_called, tool_arguments, and assistant_contains); add assert.judge for LLM scoring.

For the full schema (every Skill/tool/mock form, the bash sandbox, message injection, judge configuration, and more), see the design doc (Chinese).

CLI usage

uv run --project yama yama --plugin dingding-simple --report
Flag Meaning
paths (positional, repeatable) Explicit Case YAML paths to run
--plugin-root PATH Use the given path as the single plugin root
--plugin NAME (repeatable) Select plugins by name from yama.toml
--all-plugins Run every plugin configured in yama.toml
--result-dir PATH Override the artifact root (default <plugin_root>/.yama/runs)
--concurrency N Maximum number of cases running in parallel (default 10; 1 runs cases one at a time)
--report [PATH] Also generate a single-file HTML report (default .yama/reports/latest.html)
--list Only print the collected Cases, without running them
--no-artifacts Write no artifact files at all
--json Emit machine-readable JSON instead of Rich tables

--plugin, --all-plugins, and --plugin-root are mutually exclusive. When none of them is given, yama walks up from the current directory to the nearest directory containing yama.toml and uses it as the workspace root (falling back to the cwd), running it as a single plugin root.

Cases run in parallel, up to 10 at a time by default (--concurrency changes the cap); runs within one case stay sequential, and results keep collection order regardless of completion order. While several cases run concurrently, a transient dashboard at the bottom of the terminal shows one row per case, updated in place — waiting, then a spinner with the run/step currently executing, then its ✓/✗ verdict — and each case's full run/step detail is printed above the dashboard as one block the moment it finishes, so blocks of different cases never interleave. With --concurrency 1 (or a single collected case) it streams instead — a spinner marks the step currently executing (on a terminal), and each case, run, and assertion result is printed the moment it completes. Either way a per-case summary table closes the output, and with --json the progress stream goes to stderr so stdout stays valid JSON.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

python_yama-0.3.0.tar.gz (283.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

python_yama-0.3.0-py3-none-any.whl (62.3 kB view details)

Uploaded Python 3

File details

Details for the file python_yama-0.3.0.tar.gz.

File metadata

  • Download URL: python_yama-0.3.0.tar.gz
  • Upload date:
  • Size: 283.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for python_yama-0.3.0.tar.gz
Algorithm Hash digest
SHA256 bfb08347f78be56e3a43152bc4f87df12b39711a729267c9cab76f050392d2e0
MD5 0b459cf41ff5819f7d0307ff0ff1a09a
BLAKE2b-256 3b2a67344ea8afb1e830533d295a7e3aa69a1e41c9567e51623bf6e618a095d3

See more details on using hashes here.

File details

Details for the file python_yama-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: python_yama-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 62.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for python_yama-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 597bfaf25e93079a19129d4cacfe3df477f630c78019b76955881ee25bdcab13
MD5 234fb8d8a092a36207eba56c028800f6
BLAKE2b-256 a9cf16d0aee2b321112a6dcbc73f49fcb64317dec7ea64511b9064a7aa8bec20

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page