Skip to main content

valcore

CI codecov PyPI Homebrew Python License: Apache 2.0

A small, self-contained local tool for developing, improving, and running agentic evaluations. Author evaluators in a web UI, run them over datasets from the command line, and gate CI on their accuracy.

Install

With Homebrew:

brew install duncankmckinnon/tap/valcore

Or as a uv tool:

uv tool install valcore

Either way you get an valcore command on your PATH.

Web UI

valcore serve opens a dark-themed web UI with four surfaces:

  • Overview — the landing page, summarizing what you have and pointing to the next step.
  • Evaluators — author, version, and validate LLM-as-judge evaluators.
  • Datasets — build and edit the datasets evaluators run over, by hand or generated from a description.
  • Runs — inspect completed runs, their metrics, and per-row scores, and compare runs against each other.

Quickstart

Store your gateway API key, then start the app:

valcore config set-key sk-...
valcore serve

serve starts the web UI and API on http://127.0.0.1:8000 and opens a browser (pass --no-browser to skip that, or --port to bind elsewhere). Author evaluators and datasets in the UI, then drive runs from the command line. Both can be written by hand or generated from a description; a generated result is an editable draft either way.

Seeding one from the other

An evaluator and a dataset have to agree on columns, so rather than retype that shape you can seed either one from the other and let the model fill in content you describe.

Generate a dataset from an evaluator version and the dataset always gets that version's required columns; you can add extra columns by naming them, and per-column notes say what each should contain. Suggested labels are optional — ask for them when you want the model to propose ground truth, and the label space comes from the evaluator.

Generate an evaluator from a dataset and it is drafted against that dataset's columns, with per-column notes saying how each factors into the assessment. The result is an editable draft, not a saved version, so you review and adjust it before keeping it.

A dataset needs no labels to be scored: an ordinary run just records the judge's output. Labels are only required for a validation run, which compares the judge against them to measure agreement.

Models and the gateway

valcore reaches models through the Pydantic AI Gateway. That is currently the only route: there is no direct-to-provider client and no per-provider API key, so a gateway key is required before anything that calls a model will run.

Model strings are always gateway/<provider>:<model>:

gateway/anthropic:claude-sonnet-5      # the default
gateway/anthropic:claude-opus-4-5
gateway/anthropic:claude-haiku-4-5
gateway/openai:gpt-5
gateway/google:gemini-2.5-pro

Valid providers are anthropic, openai, google, google-cloud, bedrock, and groq. A string that does not match this shape is rejected before any request is made, so a bare claude-sonnet-5 fails fast with a clear error rather than at call time.

The key is stored in ~/.valcore/config.toml (mode 0600) and exported as PYDANTIC_AI_GATEWAY_API_KEY when a command runs. An already-exported environment variable always wins over the stored key, which is what you want in CI:

export PYDANTIC_AI_GATEWAY_API_KEY=sk-...

Run valcore config set-key with no argument to be prompted without echoing the key.

Override the default model, highest precedence first: an explicit argument, VALCORE_DEFAULT_MODEL, model in config.toml, then the built-in default.

On other providers. Routing everything through one gateway keeps model access to a single credential and a single validated string format. It also means valcore inherits whatever the gateway supports and nothing else. Provider routing is confined to one module, so widening this later — direct provider clients, a self-hosted or OpenAI-compatible endpoint, local models — is a change to that resolution layer and the config schema rather than a change to how evaluators, datasets, or runs work.

Setup

valcore serve shows a setup card on the Overview page listing each key valcore knows about, whether it is currently set, and what it unlocks. Keys are never entered through the web UI — no secret crosses HTTP — and are always set from the CLI:

valcore config set-key                  # required: runs and generation
valcore config set-logfire-token        # optional: sends traces to Logfire
valcore config set-logfire-key          # optional: pushes datasets to Logfire

Without the gateway key, generation and runs are unavailable, and the UI shows why. Manual authoring, dataset upload, editing, hand-labeling, and every export still work with no key configured at all.

Commands

Command What it does
valcore serve Serve the web UI and API (--port, --host, --no-browser).
valcore list <evaluators|datasets|runs> List resources as a table or, with --json, as JSON.
valcore run <evaluator> <dataset> Run an evaluator version over a dataset.
valcore experiment <evaluator> <dataset> Run an evaluator version over a dataset via pydantic_evals.Dataset.evaluate.
valcore export <evaluator> Export an evaluator (and, with --dataset, a dataset) as a Python script or, with --format json, a portable eval package.
valcore import <file> Import a JSON eval package back into the local database.
valcore config set-key [KEY] Store the gateway API key in the config file.
valcore config set-logfire-token [TOKEN] Store the Logfire write token in the config file.
valcore config set-logfire-key [KEY] Store the Logfire API key in the config file.
valcore config get Show the current config (the key is masked unless --show-key).
valcore config path Print the path to the config file.
valcore config edit Open the config file in $EDITOR.
valcore logfire push <dataset> Push a dataset to Logfire's hosted dataset store.
valcore skills install Install the bundled agent skills (--claude, --copilot, …).
valcore skills list Show the bundled skills and where each is installed.
valcore skills uninstall Remove the bundled skills from the selected directories.
valcore version Print the installed valcore version.

Evaluators, versions, and datasets are addressable by name or by a unique id prefix; an ambiguous value is an error that lists the candidates. Pass --db PATH on the group to point at a SQLite database other than the default under ~/.valcore.

run accepts --version (defaults to the active version), --kind, --concurrency, --watch (one line per completed row), --json, and --min-accuracy. Progress goes to stderr and results go to stdout, so redirecting stdout yields clean JSON. The CLI talks to SQLite directly, so run works whether or not serve is up.

Portable eval packages

Both an evaluator and a dataset export two ways, as code or as JSON:

Code JSON
Evaluator standalone Python script pydantic_ai AgentSpec
Dataset module that builds a pydantic_evals.Dataset pydantic_evals Dataset

The JSON forms combine into an eval package: a pydantic_evals dataset plus a pydantic_ai AgentSpec, in one file by default or two with --split. Neither half is invented here — each is the serialization its own framework already defines — and a small valcore block beside them carries the prompt template, required columns, score field, and tool names that neither foreign format has a place for.

valcore export my-judge                                  # standalone Python script
valcore export my-judge --format json -o my-judge.json
valcore export --dataset my-data --format json -o my-data.json
valcore export my-judge --dataset my-data --format json --split -o pkg.json
valcore import my-judge.json

valcore export my-judge with no new flags still emits exactly the Python script it always has. --format json emits the package instead, --dataset folds a dataset into either form, and --split writes pkg.agent.json and pkg.dataset.json side by side rather than one bundle. import reads the JSON form back into your local database; a .py export is not importable.

Running a package

A valcore_judge.py companion module ships beside the JSON and needs no valcore install — it imports only stdlib json and pydantic, rebuilding the agent with Agent.from_spec. Register it as a custom evaluator type and the dataset runs under pydantic_evals:

from pydantic_evals import Dataset

from valcore_judge import ValcoreJudge

dataset = Dataset.from_file("my-data.json", custom_evaluator_types=[ValcoreJudge])
report = dataset.evaluate_sync(task)

What the formats can and cannot do

Three limitations are worth stating plainly:

  1. Reading a bundled package needs valcore_judge.py, because pydantic_evals.Dataset forbids unknown top-level keys and so refuses the bundle's agent and valcore blocks on its own.
  2. A dataset exported with an evaluator names ValcoreJudge in its evaluators, so a bare Dataset.from_file() — with no custom_evaluator_types — cannot read it either.
  3. AgentSpec has no tools field and silently ignores unknown keys, so Agent.from_spec(AgentSpec.from_file("pkg.agent.json")) builds a working agent with zero tools without complaint — the tool names live in the valcore block, which AgentSpec drops on load. Use valcore_judge.py, which restores them from source inlined into the module. If you load the bare spec and wonder why the judge behaves differently, this is why.

JSON is the only config format: the web UI validates and previews a package client-side with its built-in JSON.parse, adding no dependency to do it.

Agent skills

valcore ships a skill document that teaches a coding agent how to drive it — the data model, the author/validate/run/export loop, the evaluator-dataset compatibility rules, and a full CLI reference. Install it into whichever agent you use:

valcore skills install --claude          # ./.claude/skills/
valcore skills install --claude --copilot
valcore skills install                   # ./.agents/skills/, discoverable by any client
valcore skills install --claude --global # ~/.claude/skills/
Flag Destination
(none) or --agents .agents/skills/
--claude .claude/skills/
--copilot .github/skills/
--all all of the above

Flags are additive and nothing is implicit — --claude --copilot writes exactly those two directories and leaves .agents/ alone. Add --global for home-level directories instead of the current repository.

Copying is the default: an already-identical skill is skipped, and one you have edited prompts before being overwritten (--force to skip the prompt). Use --symlink to link to the packaged copy instead, so upgrading valcore upgrades the installed skill.

Run valcore skills list to see what is bundled and where each copy currently lives.

Using valcore in CI

run --json emits a single object with run metadata, metrics, and per-row scores, and --min-accuracy turns a validation run into a pass/fail gate:

valcore run my-evaluator my-dataset \
  --kind validation \
  --min-accuracy 0.9 \
  --json > run.json

Exit codes:

Exit Meaning
0 The run finished and, if --min-accuracy was set, accuracy met the threshold.
1 The run failed, or a domain error occurred (printed as error: <message> on stderr).
2 Accuracy fell below --min-accuracy.

--min-accuracy requires a categorical accuracy metric; numeric or unlabeled runs have no accuracy and error rather than silently passing.

Logfire

logfire is an optional extra — install it to get traces:

uv tool install 'valcore[logfire]'

With a Logfire token configured (see Setup), each valcore run opens a valcore.run span carrying the run's evaluator version, dataset, and concurrency, with a valcore.score_row child span per row; on close, the run span records its status and each agreement metric as attributes, so a Logfire query can filter runs by accuracy directly. The Pydantic AI Gateway already reports the LLM calls themselves — valcore adds only the surrounding run and row context around them, and deliberately does not re-report the calls, which would double-count tokens and cost.

valcore experiment <evaluator> <dataset> runs the same evaluation through pydantic_evals.Dataset.evaluate instead of run's own engine, so it also appears in Logfire's experiments view. It persists a run the same way run does, so it shows up on the Runs page too. Unlike run, it cannot be cancelled, because Dataset.evaluate has no cancellation.

valcore logfire push <dataset> publishes a dataset to Logfire's hosted dataset store. It needs a Logfire API key (see Setup) with the project:read_datasets and project:write_datasets scopes.

~/.valcore

All state lives under ~/.valcore (mode 0700). Set VALCORE_HOME to relocate it.

~/.valcore/              0700
  config.toml              0600  gateway key + defaults
  valcore.db             SQLite (plus -wal, -shm)
  logs/                    serve logs

Development

uv sync                    # install dependencies into a local venv
uv run pytest              # run the test suite

The web UI is a Vite + React SPA under web/:

cd web
npm install
npm run dev                # Vite dev server, proxying the API

To build the SPA into the wheel, run npm run build and copy web/dist/ into src/valcore/web_dist/ before uv build; the release workflow does this automatically on a v* tag.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

valcore-0.0.7.tar.gz (9.3 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

valcore-0.0.7-py3-none-any.whl (462.5 kB view details)

Uploaded Python 3

File details

Details for the file valcore-0.0.7.tar.gz.

File metadata

  • Download URL: valcore-0.0.7.tar.gz
  • Upload date:
  • Size: 9.3 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for valcore-0.0.7.tar.gz
Algorithm Hash digest
SHA256 80cd89a1ba5740a5b298bbd637bdcaa3bff3f4abb6892810c6f29e6e4eecfa86
MD5 d482dc969cc6c6cafc30ef20093662f1
BLAKE2b-256 da571920356a17ecb9a28f7c25413a759651ebc67d40295393f1cf55d944ae9e

See more details on using hashes here.

Provenance

The following attestation bundles were made for valcore-0.0.7.tar.gz:

Publisher: release.yml on duncankmckinnon/valcore

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file valcore-0.0.7-py3-none-any.whl.

File metadata

  • Download URL: valcore-0.0.7-py3-none-any.whl
  • Upload date:
  • Size: 462.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for valcore-0.0.7-py3-none-any.whl
Algorithm Hash digest
SHA256 bd435771e28a7297c541c6932c4f161fa9fd2956f398a206bf8031c090df367b
MD5 9765fdcc4be63092cc6503f99d803708
BLAKE2b-256 84790d0a14757d4589160190d77c33f255cdddf3990be95b856cce846bfab900

See more details on using hashes here.

Provenance

The following attestation bundles were made for valcore-0.0.7-py3-none-any.whl:

Publisher: release.yml on duncankmckinnon/valcore

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page