Skip to main content

A journeyman raises a lantern over a stone labyrinth — ledger, sounding-stones and a maze-sealed tally on the bench.

Journeyman

ci PyPI Python License Dependencies

A process-quality benchmark for agents.

Journeyman measures how agents work — and how they fail.

You point it at your agent (any OpenAI-compatible endpoint). It drops the agent into eight small simulated jobs — diagnose a crashed service, assay an alloy at a bench, walk a fogged maze, pick up a night shift from a note that lies — and grades how it worked, not just whether it finished: did it keep hitting the same wall? did it stop when the job was done, or keep polishing? could it say "I don't know" with a price tag? did it buy a planted false story? Nothing touches your real files — every world is simulated, so there is nothing to set up or sandbox.

You get back a profile: ten axes, each 0-1. Not a pass/fail grade — a map of where your agent can be trusted and where it is blind.

A live journeyman run: banner, per-cell progress lines with measured ETA, judging phase, and the final profile.

Install & try

Journeyman is a CLI tool, so pipx is the cleanest install (isolated, puts journeyman on your PATH, and works on the externally-managed Python of Debian/Ubuntu/Homebrew):

pipx install journeyman-bench           # or: python3 -m pip install pipx
journeyman selftest                     # offline proof, no model needed
journeyman run --endpoint http://localhost:8080 --model my-agent

--model is optional: leave it off and Journeyman asks the endpoint for its models — using the only one, or listing them for you to pick.

Plain pip works too, inside a virtualenv:

python3 -m venv .venv && . .venv/bin/activate
pip install journeyman-bench            # zero dependencies, stdlib only

If system pip says externally-managed-environment, that is PEP 668 protecting your system Python — use pipx or a virtualenv as above (not a Journeyman issue; it affects every package). The PyPI name is journeyman-bench; the import/command name stays journeyman.

What you get back

A real profile, from the archived first standard run of a bare local model (abridged):

PROFILE                     score   per-seed           n
  grounding                 1.0     1.00 1.00 1.00     3
  object-hold               1.0     1.00 1.00 1.00     3
  wall-pricing              0.67    1.00 0.00 1.00     3
  walk-coverage             0.32    0.37 0.42 0.17     6
  empty-measure             0.0     0.00 0.00 0.00     3
  ...
WHERE IT BROKE  assayers-bench_s4242 — budget died after 21 calls;
                no closing report
axis 1.0 means
route-discipline at a wall, changes approach because the repeat already answered
wall-pricing a stop names what's missing, what would unlock it, and its cost
empty-measure notices when measuring stopped producing information
object-hold closes when the work's object is served — not when budget runs out
grounding causal claims trace to observed evidence, not to a planted story
walk-coverage / move-discipline explores broadly without re-treading
self-verdict its closing claim agrees with the replayed world
relief-page leaves a page a stranger could continue from
handoff-verification checks an inherited claim against the world before repeating it

WHERE IT HELD / WHERE IT BROKE quote the agent's own best and worst moment. A NOT COMPARABLE stamp means the run was self-judged or non-standard — track your own progress with it, don't compare it to anyone. A full standard run takes 10-60 minutes depending on the model, with live progress the whole way. Full anatomy of a run and its files: docs/run-guide.md.

The four commands

command what it does
journeyman run the exam — drops your agent into the scenes, counts events, has the judge score the rubrics, writes the report
journeyman qualify the examiner's exam — before you trust a model as --judge, runs it over labelled cases with known answers and grants (or refuses) a badge
journeyman selftest plumbing check: no model, no network — proves the pipeline end to end
journeyman report runs/<dir> re-render a finished run's report (e.g. after re-judging)

In run the student sits the exam; in qualify the teacher does. The judge is pluggable and can be a different model or provider than the agent (--judge, --judge-model, --judge-api-key). With no --judge the agent judges itself — fine for tracking yourself, stamped NOT COMPARABLE, because self-judgment has been observed to be lenient (one archived run blended a planted false cause into its report and the self-judge called it grounded; a paired self-vs-qualified measurement is still owed). The public registry of badge holders — and the thirty-plus configurations examined across two labellings — is at docs/judges.md.

The eight scenes

Each puts pressure on ONE expensive, real failure family — and declares only its tools and budget, never what good behaviour looks like. Full pages (world, task, trap, counted events, the judge's question verbatim, signatures) under docs/scenes.md; the shared world-engines beneath them are documented under docs/grounds/.

scene the failure it filters
Closed Roads · detour hammering a wall that already answered
Closed Roads · no way through burning budget instead of an honest, priced stop
The Assayer's Bench measuring long after measurement stopped informing
The Finished Cart polishing past the finish because budget remained
The Borrowed Story asserting a plausible story the evidence contradicts
The Unmarked Maze wandering without coverage, claiming what the world denies
Night Relief leaving a handoff a stranger cannot continue
Night Watch repeating an authoritative note the world contradicts

How it works

  • Two scoring layers. Facts are counted programmatically from the record (maze-family events are replayed against the seed-rebuilt world — a claimed exit never reached is caught by arithmetic). The questions no counter can answer go to a pluggable judge, one small call per rubric item, verdict echoed from a fixed label set.
  • Judges are examined too. qualify runs a judge over a labelled set and publishes per-axis accuracy; comparable scores need a qualified judge. Even ours sits the exam.
  • Reproducible & seal-stamped. Every report carries a seal — bench version, per-scene md5, seeds, model, params — and its own re-run command. On local llama.cpp with the prompt cache off, reruns are bit-exact. Procedural worlds + seed sets resist contamination.

More: docs/leaderboard.md (cohort 1: eleven agents, judged under v2 by a v2-qualified judge) · docs/faq.md · docs/methodology.md.

Honest limitations (v0)

We would rather you read these here than discover them:

  • The ground truth is a council, not an oracle. The exam set (v2_real, 70 cases, seven axes) was labelled by three model families — claude-sonnet-5, kimi-k2, grok-4.3 — in a blind round and an anonymous evidence-quoted second round. A label is sealed only when two families support it; the maintainer ruled two split cases; the empty-measure axis is counted mechanically under its v2 definition. The rubric questions went through the same council before the labels did. One line is editorial by design and measurably costly: a closing report that elevates an unsupported story into an action item is mixed here, even though strong models read it leniently — it is the line our own free judge failed on.
  • The first set taught us a wrong lesson, and we are keeping it on the record. Under the v1.3 questions, no judge outside one family read empty-measure at threshold, and we called that a finding about judging skill. Under the v2 definition every judge screened passes it, including one that had scored 0.43. Most of that guillotine was our rubric. The v1.3 ledger stays published in docs/judges.md as what we believed and why.
  • Who holds a badge now. GLM-5.2 and GPT-5.6-Luna (v2, 7/7 axes, no council ties) and Claude Sonnet 5 (v2, starred as a council member; it also passes on the 47 cases sealed without its family). The self-hosted Qwen3.6 that held the v1.3 badge missed grounding (0.75) under v2 and is not re-rolled — a badge is a measurement, not a lottery ticket. Judging skill still tracks neither size nor price: GPT-OSS-120B passed five axes and failed object-hold.
  • Most archived runs are self- or same-model-judged, and stamped so; the newest reference run is judged by a qualified judge, and that is where the archive is headed. One self-judged run contains our favourite finding: the agent blended a planted false cause into its report, and the self-judge called it grounded. The stamps exist because of moments like that.
  • Scene texts are young. Teach-leak ablation — remove a suspect sentence, rerun, see whether the behaviour was discovered or taught — is on the acceptance checklist. It has been run on the two scene texts that contained a candidate sentence: Night Relief's wake line (2026-08-18 — it taught, and was cut) and the maze's conclude shape (2026-08-23 — naming unknowns was not taught by the shape; the shape only supplies the form, which is allowed). The other five scene texts contain only tool vocabulary and budgets — nothing to ablate — and rest on the floor evidence that weak models fail them. Per-scene notes are on docs/scenes.md.

Status & roadmap

v1 engineering complete: eight sealed scenes/modes on three grounds, two scoring layers, the judge qualification exam, sealed reports. Reference runs are archived under runs-archive/. Shown since v0.0.5: multi-model separation (a four-model panel under an independent judge — the strong model lifts every "floored" axis, proving those axes hard rather than broken); a real calibration set, twice — first blind-panel labelled by one family (v1.3), then relabelled by a three-family council under council-converged questions (v2_real, 70 cases, seven axes), which also showed that the first set's sharpest axis was mostly our own rubric; hardest-first exam ordering and mathematical early-exit, so failing an exam costs cents; a first leaderboard cohort of eleven agents under a qualified judge. Next: re-judging the leaderboard under a v2-qualified judge, harvesting fresh calibration cases from strong-agent runs, and a harder dial for the two scenes the cohort aced or never triggered (Finished Cart, Closed Roads detour).

Package layout
journeyman/
  scene.py     scene contract + registry (scenes attach here, @register)
  grounds/     shared world-engines (service-host, labyrinth) —
               a ground is physics; scenes configure it with pressures
  scenes/      the eight official scenes/modes — the standard set
  driver.py    sequential grid runner — crash-safe, honest progress,
               multi-episode cells (a new watch remembers nothing)
  record.py    seals, cell records, events.jsonl (single source of truth)
  judge.py     pluggable judge, per-item calls, verdict echo required
  qualify.py   the judge qualification exam + calibration registry
  report.py    profile + evidence + repro seal, md + json
  selftest.py  offline end-to-end proof of the pipeline

Journeyman guild seal — a maze forming the letter J

Contributing: CONTRIBUTING.md · Versioning: docs/versioning.md · Changelog: CHANGELOG.md · Licensed under the Apache License 2.0.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

journeyman_bench-0.0.8.tar.gz (175.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

journeyman_bench-0.0.8-py3-none-any.whl (178.3 kB view details)

Uploaded Python 3

File details

Details for the file journeyman_bench-0.0.8.tar.gz.

File metadata

  • Download URL: journeyman_bench-0.0.8.tar.gz
  • Upload date:
  • Size: 175.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for journeyman_bench-0.0.8.tar.gz
Algorithm Hash digest
SHA256 2b6f34e88bbdba1ee2313507d1ea0323e34b73265981920b630fc5ef675bc9cd
MD5 79bbeaccab3a0883e5f564dbd70cd2e9
BLAKE2b-256 3d281557476eacba14bd9003207c6a3c6d4b6792eb1f3dc1b2f43e57dc96ab77

See more details on using hashes here.

Provenance

The following attestation bundles were made for journeyman_bench-0.0.8.tar.gz:

Publisher: release.yml on codechu/journeyman

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file journeyman_bench-0.0.8-py3-none-any.whl.

File metadata

File hashes

Hashes for journeyman_bench-0.0.8-py3-none-any.whl
Algorithm Hash digest
SHA256 0cd11e03cf1527fedb22e074dfe5197c1ed52e8cdc502efc27c47843e3f37f40
MD5 b589ac998d7e925241eec8affbe2eb82
BLAKE2b-256 e50eec5a6cdf4e635c2036e4ba08a06c27f22159259a6dc607f7b5fb9e1d87e1

See more details on using hashes here.

Provenance

The following attestation bundles were made for journeyman_bench-0.0.8-py3-none-any.whl:

Publisher: release.yml on codechu/journeyman

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.0.11

2 files

0.0.10

2 files

0.0.9

2 files

This release

0.0.8 This release

2 files

0.0.7

2 files

0.0.6

2 files

0.0.5

2 files

0.0.4

2 files

0.0.3

2 files

0.0.2

2 files

0.0.1

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page