shōgym
A gym for running agents on tasks: bring any harness.
An environment describes a task, serves its tools over MCP, and verifies the recorded trajectory
with a pure function. A harness (Claude Code, Codex, Pi, Hermes, Prime Agent, or a plain loop)
drives the tools; shōgym never owns the model, the prompt, or the agent loop. The core is
zero-infrastructure: pip install shogym, point a harness at a served env, get a trace and a
score.
It is a substrate for questions about agents on tasks. How do agents improve themselves? What does adding a tool change? Does an agent's self-report match an honest score?
Under the hood, an environment (Env) is a task loader, an initial observation, a pure verifier,
and a set of MCP servers. Tools are MCP servers
(in-process, stdio, or HTTP); episodes are session-keyed for safe concurrent rollouts;
termination is a reserved terminate tool or the horizon; verification is a pure function over
the recorded trajectory.
Quickstart
Start with a harness. examples/ holds one directory per harness, each idiomatic to
that harness rather than squeezed into a shared abstraction. Every quickstart does the same three
things: serve a stream of tasks over one endpoint, swap the env with one variable, and
read the scores back out of the server's own durable rows (the harness never grades itself).
claude_code/: the reference implementation. One--mcp-configflag points theclaudeCLI at a queue of tasks;results.pyreads the scores back.pi/: Pi ships no MCP client, so this one adds a bridge extension, pinned to an exact version and scoped to the project's.pi/.hermes/: Hermes speaks MCP natively (with its[mcp]extra), configured by oneconfig.yamlunder an isolatedHERMES_HOME.codex/: Codex runscodex execwith the stream declared inline on the command line, or read from the project-scoped.codex/config.tomlchecked in here.prime_agent/: Prime Agent takes MCP as a Python-backed kernel skill, andserve.pyruns over HTTP in a second shell.
Each one defaults to automationbench over tasks [0, 1, 2]. One variable swaps the env for a
run, without editing a tracked file:
SHOGYM_ENV=tau2_banking_knowledge SHOGYM_TASKS=0,1 <the quickstart's command>
That one needs the tau2 extra, tau2's data, and OPENAI_API_KEY for the user simulator, and it
has 97 tasks (0 to 96). SHOGYM_ENV wins when it is set, so reach for the variable while you
are trying envs out and edit the ENV literal at the top of serve.py once you have picked. Any
quickstart runs any env in the catalogue below, and task index ranges differ per env. wordle_v1
needs no extra and no key, and is the cheapest place to start.
shōgym is pinned to Python 3.12 (requires-python = ">=3.12,<3.13", with a committed
.python-version) because the tau2-bench port needs it. With uv,
uv sync builds the 3.12 venv and uv run … runs against it.
Environments
Each env describes a task, serves its tools over MCP, and verifies a recorded
trajectory; an external harness drives it. The src/shogym/envs/
README covers the model and the shared env-README template.
wordle_v1: the reference environment. Wordle with aguesstool and a pure trajectory verifier. No extra deps.- tau2-bench: a port of
τ²-bench. Tool-using customer-service agents
(
tau2_mock,tau2_airline,tau2_retail,tau2_telecom,tau2_banking_knowledge), scored by tau2's own evaluator. Needs thetau2extra + data. yc_bench: a port of YC-Bench. Operate a simulated AI startup for a year via a singlerun_commandtool, scored on survival, funds, and tasks. Deterministic in-process sim (no data or key). Needs theyc_benchextra.hle: a port of Humanity's Last Exam. A single-turn, expert-level question answered via onesubmit_answertool and graded server-side (exact-match fast path, then an OpenAI model judge). shōgym's first model-graded verifier. Needs thehleextra,OPENAI_API_KEY, and gatedcais/hleaccess.browsecomp_plus: a port of BrowseComp-Plus. Answer reasoning-heavy queries against a fixed ~100K-doc corpus viasearch/get_document/submit_answer, graded by an LLM judge plus deterministic retrieval-recall and citation metrics. Needs thebrowsecomp_plusextra,OPENAI_API_KEY, Java 21, and gated dataset access.automationbench: a port of AutomationBench. Carry out a cross-application business workflow over a fully simulated world of ~47 SaaS apps via anapitool surface, scored end-state-only by a pure rubric. Deterministic and offline (no key). Needs theautomationbenchextra.frontier_bench: a port of Frontier-Bench. Operate a per-task Docker container through a shell (exec/read_file/write_file/done), scored by the task's own verifier over the container end-state. A CPU-only, single-container slice (5 tasks). Needs thefrontier_benchextra and a local Docker daemon (no key or data download).
The task server
Every quickstart is a thin wrapper around one object. A TaskStream publishes a whole queue over
a single MCP endpoint: it hands out tasks one at a time through get_task, routes the env's own
tools to whichever task is live, and seals and scores each one server-side.
import asyncio
from pathlib import Path
import shogym
from shogym.serve.stream import Immediate, TaskRef, TaskStream, build_stream_server
async def main() -> None:
stream = TaskStream(
shogym.make, # a factory, so each task gets a fresh env
[TaskRef("tau2_banking_knowledge", i) for i in (0, 1, 2)],
# Fresh per run: a stream refuses a directory another run recorded into.
prov_dir=Path("runs/banking-0001"),
feedback=Immediate(), # ending a task returns the env's own verdict
)
async with stream:
await build_stream_server(stream, name="shogym").run_async(transport="stdio")
asyncio.run(main())
Every dispensed task lands exactly one durable row under prov_dir, which
shogym.serve.stream.read_results reads back after the run.
For scores you intend to defend, construct EvalStream instead. It pins feedback=Never() and
refuses the argument outright, so a terminating call answers with the same fixed payload for every
env, task and outcome. It stamps feedback_regime="never" on every row it writes, and refuses to
resume a directory whose rows were recorded under any other regime. The agent is never told how it
did, so the harness cannot grade itself.
One episode, no harness
The quickstarts serve a stream of tasks. A single episode is smaller: shogym serve publishes
one task over MCP for any client to spawn and drive, and scores it off a local JSONL trace.
shogym serve wordle_v1 --task 17 --trace ./shogym_logs/run.jsonl
Or drive one in-process and read the terminal feedback:
import shogym
# `harness` is an async callable given a FastMCP client connected to the served env.
result = await shogym.evaluate("wordle_v1", task=17, harness=my_harness)
print(result.value("check_answer"))
License
Apache-2.0. Portions derived from llmgym (© TensorZero, Apache-2.0): see NOTICE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file shogym-0.0.1.tar.gz.
File metadata
- Download URL: shogym-0.0.1.tar.gz
- Upload date:
- Size: 6.8 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
eb779c4a6ff4dc9ff3db5621b4d3fc030b3c9fe45acf1485b14f88f12c6f478e
|
|
| MD5 |
1d6090467b8af45fcc83c4a52a263fd2
|
|
| BLAKE2b-256 |
d203097e6e65cd915a89ad2e570053cb9a7c2cf1945f9596323818df67977fff
|
File details
Details for the file shogym-0.0.1-py3-none-any.whl.
File metadata
- Download URL: shogym-0.0.1-py3-none-any.whl
- Upload date:
- Size: 504.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d770875895d9928e526ad777b193daeb207f19d0815662ef2b55457cf45f4c20
|
|
| MD5 |
f91ccf3c5176546405515d903b648510
|
|
| BLAKE2b-256 |
2a6b5cdcfa77293372f849f26c0d23bd59b5ed06a1508431852d96fd028ea91d
|