strands-env
A unified framework for building agent environments for RL training and evaluation with Strands Agents.
Features
An agent environment takes a task and runs the agent to completion over multiple turns, producing a rollout result — the trajectory, reward, and termination reason for that task. With strands-env, you can:
- Define Environments — Subclass
Environment, add@toolfunctions, plug inRewardFunction; typedTasksubclasses carry per-sample fields - RL Training — Token-level trajectories (TITO) for on-policy training with strands-sglang
- Benchmarking — CLI and
Evaluatorwith checkpointing, resume, and custom metrics
Install
pip install strands-env
For development:
git clone https://github.com/strands-rl/strands-env.git && cd strands-env
uv sync
Quick Start
Define an Environment
Subclass Environment and add tools as @tool-decorated functions:
import subprocess
import sys
from strands import tool
from strands_env.core import Environment
@tool
def run_python(code: str) -> str:
"""Run a Python snippet and return its output."""
proc = subprocess.run([sys.executable, "-c", code], capture_output=True, text=True, timeout=10)
return proc.stdout + proc.stderr
class CodingEnv(Environment):
def get_tools(self):
return [run_python]
Run It
from strands_env.core import Task
env = CodingEnv(model_factory=factory, reward_fn=reward_fn)
result = await env.rollout(Task(
message="Write Python to compute the 10th Fibonacci number, then run it.",
ground_truth="55",
))
result.rollout # Rollout(token_ids=..., loss_mask=..., ...); token-level rollout if using SGLang backend
result.final_response # "The 10th Fibonacci number is 55"
result.reward_result # {"reward": 1.0, "info": ...}
result.termination_reason # TerminationReason.TASK_COMPLETE
See the examples/ directory for complete, runnable demos.
Run Evaluations
python -m strands_env.eval \
--benchmark terminal-bench-2 \
--env examples.eval.terminal_bench.terminal_bench_env \
--backend sglang \
--base-url http://localhost:30000 \
--n-samples-per-prompt 4 \
--max-concurrency 8
Raise
--n-samples-per-promptfor more stable pass@k, and--max-concurrencyif you're using a hosted sandbox service.
Tip: For a non-agentic benchmark (no tool use), don't override
get_tools()— the base class returns[]by default.
Built-in Environments
Ready-to-use environments under src/strands_env/environments/. Each ships with its own README, system prompt, and requirements.txt.
| Environment | Description |
|---|---|
math |
Chat-only math reasoning with a symbolic-equivalence reward. |
agentcore_code |
Python / shell execution via AWS Bedrock AgentCore Code Interpreter. |
web_search |
Google search + Jina page scraping with optional LLM summarization, enlightened by OpenSeeker. |
harbor |
Run Harbor-format tasks in sandboxes. Supports training like SETA and evaluation like Terminal-Bench and SWE-bench. |
tau2_bench |
tau2-bench customer-service dialogues (airline/retail/telecom) driven by an LLM user-simulator. |
mcp_atlas |
MCP-Atlas benchmark runner across 36 MCP servers with 500 tasks. |
agent_world_model |
AgentWorldModel tasks with 1000 synthetic FastAPI + SQLite environments exposed as MCP tools. |
Documentation
- Evaluation Guide — CLI reference, hook files, custom evaluators
- RL Training Integration — integration with the slime RL training framework
Development
# Lint
ruff check src/ && ruff format --check src/
# Unit tests
pytest tests/unit/ -v
# Integration tests (requires running SGLang server)
pytest tests/integration/ -v --sglang-base-url=http://localhost:30000
Or if using Claude Code, just use /run-unit-tests and /run-integration-tests slash commands.
License
Apache License 2.0 — see LICENSE.
Metadata
Release files for strands-env 0.7.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| strands_env-0.7.2.tar.gz | 155.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| strands_env-0.7.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 296.4 kB
Release files / strands_env-0.7.2.tar.gz
| Download URL | strands_env-0.7.2.tar.gz |
|---|---|
| Size | 155.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a238fb82f4085872500f7cf417a4ac6e5d7e1b17394586ccddd46e82b276a259
|
|
BLAKE2b-256 checksum How to use checksums |
a4b75d89836d424b0cf3dec68b74f1a85ae71dbee1516fa9a2027c539e8b68b9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 4, 2026.
Transparency logRelease files / strands_env-0.7.2-py3-none-any.whl
| Download URL | strands_env-0.7.2-py3-none-any.whl |
|---|---|
| Size | 140.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e21a9f23c2944921885e684cfa032d35b6c425ab908565173c34a6f52609152b
|
|
BLAKE2b-256 checksum How to use checksums |
951fb0aae67277130c7020df993bb90fdfb47b4567f3b3c2e7026871cdda14d6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 4, 2026.
Transparency log