StateSet Agents is a production‑oriented RL stack for training and serving LLM‑backed agents that improve through multi‑turn interaction. The library provides:
- An improvement loop that turns your agent's own logs into a better agent:
ingest→improve(grade → curate) → fine‑tune. - Async‑first agent APIs (
MultiTurnAgent,ToolAgent) with Hugging Face and stub backends. - Environments for conversational and task‑oriented episodes.
- Trajectories and value/advantage utilities tailored to dialogue.
- Composable reward functions (heuristic, domain, multi‑objective, neural, LLM‑judge).
- A family of group‑based policy‑optimization trainers (GRPO, GSPO, GEPO, DAPO, VAPO) plus PPO and RLAIF.
- Offline RL algorithms for learning from logged conversations (BCQ, BEAR, CQL, IQL, Decision Transformer).
- Sim‑to‑Real transfer for training in simulation and deploying to real users (domain randomization, system identification, progressive transfer).
- Continual learning + long‑term planning utilities (replay/LwF/EWC, plan context injection).
- An MCP server so Claude Code/Desktop — or any MCP client — can drive the loop conversationally.
- Optional performance layers (vLLM generation, Rust acceleration, distributed training, HPO, FastAPI service).
If you want a framework that treats conversations as first‑class RL episodes (rather than single turns), this is it.
The improvement loop
Your agent already produces conversation logs. They are a training set.
pip install stateset-agents
# 1. Bring your own logs — OpenAI chat format or LangChain traces
stateset-agents ingest --format openai --input my_agent_logs.jsonl --output transcripts/
# 2. Grade every conversation, curate the best turns, get your next command
stateset-agents improve run \
--transcripts transcripts/ \
--reward customer_support \
--output improved/
# 3. Train on what worked (the exact command is printed in improved/next_steps.md)
python scripts/sft_from_curated.py --dataset improved/curated.jsonl --base-model <model>
No GPU? Step 3 is the only part that needs one. train-remote runs that
same job on rented compute:
stateset-agents train-remote --provider modal --gpu A100 \
--dataset improved/curated.jsonl --base-model <model>
improve writes three things: improve_summary.json (machine‑readable scores
and per‑reward breakdown), curated.jsonl (the turns above your threshold, ready
to train on), and next_steps.md (runnable training commands — regression‑tested
against the real CLI so they never drift).
Try it in five minutes, offline, no GPU and no API key:
bash examples/five_minute_demo.sh
It writes sample logs, runs the whole loop, and shows you the graded output.
Colab version: notebooks/improve_your_agent_5min.ipynb.
Or let an agent drive it:
pip install "stateset-agents[mcp]"
claude mcp add stateset-agents -- stateset-agents mcp
Seven MCP tools (list_rewards, ingest_transcripts, grade_transcript,
improve_run, improve_status, list_model_presets, dry_run_finetune) — see
docs/MCP_SERVER.md.
What's new
v0.25.0 (latest release — live on PyPI):
- Talk to your fine‑tuned model.
stateset-agents chat-remote --base-model X --adapter DIRrents a GPU pod, loads base + adapter, and holds a multi‑turn conversation (SSH‑piped, pod terminated on every exit path;--promptfor scripted mode). Live‑verified: a Muse‑Glimmer‑30B adapter resolved "I got double charged for it" to the order number from the previous turn. - Dead pods fail fast. SSH keepalives bound peer-loss detection to ~2 minutes (a pod restarting under a running job previously hung the executor indefinitely).
v0.24.0:
- Fine‑tune and call it, one command.
train-remotegains--eval-prompts FILE(the job generates base‑vs‑finetuned completions for held‑out prompts, greedy for comparability, returningeval_results.jsonbeside the adapter) and--container-disk-gb(pod disk sized to the checkpoint — a 63GB model needs ~160). Verified live on an H100: Muse‑Glimmer‑30B tuned on 140 support conversations answers held‑out order numbers in the trained persona while the base model stalls in its reasoning channel. - Vision‑exclusion that actually works. LoRA target inference is two‑pass — leaf names existing only in vision/audio stacks are dropped (peft matches names model‑wide), shrinking multimodal adapters ~13%; base‑eval generation now runs on GPU.
v0.23.1:
train-remotehandles large multimodal checkpoints — four fixes found and verified by trainingmeta-models/Muse-Glimmer-30Bend-to-end on a rented H100 (63GB BF16 download, LoRA on the text stack, 258MB adapter returned): multimodal-architecture fallback in the SFT loader, configurable RunPod pod disk (container_disk_gb), transformers-5.x-proofTrainingArgumentsconstruction, and vision-tower exclusion from LoRA target inference.
v0.23.0:
- Three new first-class starters.
stateset-agents qwen3-coder(Qwen/Qwen3-Coder-30B-A3B-Instruct, 256K ctx, Apache‑2.0),stateset-agents gpt-oss(openai/gpt-oss-20b, 128K, Apache‑2.0; 120B variant flagged multi‑GPU), andstateset-agents deepseek-v4(deepseek-ai/DeepSeek-V4-Flash, 1M ctx, MIT — GLM‑style QLoRA+vLLM path with MLA‑correct LoRA targets verified from the weight map).
v0.22.0:
- Architecture consolidation. The four flagship trainers (GSPO/DAPO/GEPO/VAPO) share one model-loading/checkpoint runtime (
training/trainer_runtime.py); the seven model starters are thin definition layers overtraining/starter_common.py; research modules moved tostateset_agents.experimental(old paths warn for one cycle). All RL math and public APIs unchanged — the full suite passes unmodified. - NVIDIA Nemotron 3.5 Lightning starter.
stateset-agents nemotron-3-5targetsnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16(hybrid Mamba‑2 + MoE, 30B total / 3B active, 256K ctx, OpenMDW‑1.1) with Mamba-aware LoRA targets.docs/nemotron_3_5_starter.rst - Quality gates that enforce themselves. Coverage and mypy-allowlist ratchets fail CI when floors fall behind reality; the security workflow's scanners now actually gate (45 HIGH/CRITICAL findings cleared);
make benchmark-loopscores the improvement loop against planted ground truth (precision 0.818 / recall 1.0); a weeklygpu-verifyworkflow re-proves the training job on rented hardware. - Slimmer repo.
dashboard/andmobile/moved to their own repositories.
v0.21.0:
- Muse Glimmer 30B first‑class starter.
stateset-agents muse-glimmertargetsmeta-models/Muse-Glimmer-30B— Meta's open agentic model (Aug 2026; dense 30B, 131K ctx, Apache‑2.0) — with the standard balanced/memory/quality QLoRA profiles,init --preset muse-glimmer, and amuse-glimmerpreset in the unified finetune driver.docs/muse_glimmer_starter.rst - RunPod provider for
train-remote.--provider runpodrents a GPU pod over SSH, runs the same packaged job every other provider runs, and copies the adapter back. GPU defaults are now per‑provider —RemoteJobSpec.gpuno longer hard‑codes a Modal‑specific name.
v0.20.0:
- Run the fine‑tune step without a GPU.
stateset-agents train-remoteruns the SFT job fromimproveon rented compute (--provider local|modal), closing the last gap in the improvement loop. The job itself is unchanged whichever provider runs it, and remote runs install a pinned published package rather than syncing your working tree.docs/CLI_REFERENCE.md
v0.19.0:
- MCP server.
stateset-agents mcp(pip install stateset-agents[mcp]) exposes the improvement loop as tools for any MCP client — Claude Code/Desktop or your own agent. Seven tools, stdio transport, dry‑run‑only training.docs/MCP_SERVER.md
v0.18.0:
- Bring your own agent's logs.
stateset-agents ingestconverts OpenAI chat‑format and LangChain conversation dumps into framework trajectories — logs from agents built anywhere plug straight into the loop. - The improvement loop in one command.
stateset-agents improve rungrades, curates, and emits verified‑runnable training commands (a regression test executes every suggestion against the real CLI parsers). - Flagship benchmark recipe.
make flagship-benchmark-all— a reproducible 3‑seed GSPO run on an 8B model with publish gates (benchmarks/FLAGSHIP.md).
v0.16.0 – v0.17.3 (correctness and distribution):
- RL‑core correctness overhaul. All five trainers fixed and behaviorally tested: DAPO freezes rollout‑time old log probs and honors µ inner updates; GEPO runs in log space; GSPO scores exactly the text it sampled and rescores vLLM rollouts; GSPO‑token regained its gradient path; VAPO clips values against rollout predictions with terminal‑token rewards and decoupled GAE; the GRPO loss path uses length‑normalized ratios. A cross‑trainer ratio‑invariant suite guards all of it.
- Convergence proof in CI. A nightly job trains a real (tiny) model and asserts the target‑token probability strictly increases — verified against zero‑signal and reversed‑reward controls.
- API hardening. Training‑lab routes auth‑gated behind
API_ENABLE_TRAINING_LAB(off outside development), bounded in‑memory state, fail‑closed production config validation, constant‑time key comparison, identity‑keyed rate limiting. - Distribution repaired. PyPI is current again, wheels ship the runtime config presets,
stateset-rl-coreis published so[rust]and[full]resolve, and CI is green across 4 Python versions + Windows. - Unified finetune driver.
examples/finetune_gspo.py --model <preset>replaced the per‑model script maze with a 12‑model preset registry.
Earlier highlights: v0.15.3 shipped Rust accelerator wheels (abi3‑py310) and model‑level Prometheus inference metrics; v0.15.0 added the getting‑started ladder; v0.13.2 shipped the whitepaper §11.7 three‑seed canonical benchmark (judge improvement +0.079, artifact).
Full breakdown in CHANGELOG.md.
Why group‑based optimization?
Traditional RLHF/PPO trains on one sampled response at a time. In long conversations this leads to high‑variance updates and brittle behavior.
StateSet Agents implements group‑relative methods:
- GRPO (Group Relative Policy Optimization): sample a group of trajectories per prompt, compute advantages relative to the group baseline, then apply clipped policy‑gradient updates.
- GSPO (Group Sequence Policy Optimization): a more stable sequence‑level variant (Alibaba Qwen team) that avoids token‑level collapse on long outputs and MoE models.
The result is steadier learning for dialogue tasks.
Core concepts
- Agent: wraps a causal LM and exposes
initialize()andgenerate_response().MultiTurnAgenthandles conversation history and state.ToolAgentadds function/tool calling.
- Environment: defines episode reset/step logic and optional reward hooks.
ConversationEnvironmentships with scenario‑driven multi‑turn conversations.TaskEnvironmentis for goal‑oriented tasks.
- Trajectory: a multi‑turn record of turns, rewards, and metadata (
MultiTurnTrajectory). - Rewards:
RewardFunctionsubclasses and factories; combined viaCompositeRewardor multi‑objective reward models. - Training: trainers in
stateset_agents.trainingimplement GRPO‑family updates, GAE/value heads, KL regularization, LoRA support, and optional distributed/vLLM execution.
Reward semantics
Reward functions can be evaluated per-step or only at episode end. Set
reward_type on your RewardFunction to control how the environment applies it:
RewardType.IMMEDIATEorRewardType.DENSE: compute per-step rewards only.RewardType.CUMULATIVEorRewardType.SPARSE: compute a final reward only.
If you pass a custom reward without reward_type, the environment assumes legacy
behavior and may compute both step and final rewards. For new rewards, always
set reward_type explicitly to avoid double counting.
Tool calling (ToolAgent)
ToolAgent lets a model request a tool via a JSON block, which the agent executes:
import asyncio
from stateset_agents.core.agent import AgentConfig, ToolAgent
def add(a: int, b: int) -> int:
return a + b
async def main():
agent = ToolAgent(
AgentConfig(model_name="stub://tools", use_stub_model=True),
tools=[
{
"name": "add",
"description": "Add two integers",
"parameters": {"a": "int", "b": "int"},
"function": add,
}
],
)
await agent.initialize()
# The model should respond with a JSON tool call like:
# {"tool": "add", "parameters": {"a": 1, "b": 2}}
print(await agent.generate_response("Please calculate 1 + 2"))
asyncio.run(main())
Installation
Core (lightweight, stub‑ready)
pip install stateset-agents # latest release (v0.23.0)
That's enough for the five-minute demo, the stub
backend, and the CLI. Training real models needs [training] below.
PyPI tracks the release tags. For unreleased work on master:
pip install "git+https://github.com/stateset/stateset-agents@master"
Training / real models
pip install "stateset-agents[training]"
Optional extras
pip install "stateset-agents[auto-research]" # Autonomous experiment loop + Optuna
pip install "stateset-agents[trl]" # TRL GRPO integration + bitsandbytes
pip install "stateset-agents[vllm]" # vLLM generation backend
pip install "stateset-agents[hpo]" # Optuna/Ray Tune HPO
pip install "stateset-agents[api]" # FastAPI service
pip install "stateset-agents[distributed]" # DeepSpeed / multi‑GPU helpers
pip install "stateset-agents[rust]" # Rust-accelerated GAE/advantage kernels (stateset-rl-core)
pip install "stateset-agents[full]" # Most extras in one go
This repository also contains an internal, unpublished Rust crate at the repo root (a StateSet commerce daemon) that is unrelated to the
stateset-rl-coreaccelerator behind[rust]above. Seedocs/RUST_CRATES.mdfor how the two Rust crates in this repo relate.
Model starter paths
One driver covers every supported model — --dry-run is the default, so nothing
trains until you ask:
pip install "stateset-agents[training,trl]"
python examples/finetune_gspo.py --list-models # 12 presets
python examples/finetune_gspo.py --model qwen3.5-0.8b # dry run: show the resolved config
python examples/finetune_gspo.py --model qwen3.5-0.8b --no-dry-run # actually train
Useful flags: --starter-profile {balanced,memory,quality} (the memory profile
uses 4‑bit quantization and smaller context/group sizes), --use-lora/--no-lora,
--use-4bit/--use-8bit, --use-vllm, --wandb, --export-merged,
--write-config PATH.
Seven models also ship a dedicated starter with tuned defaults and a hosting plan:
| Model | Dedicated entry point | Notes |
|---|---|---|
Qwen/Qwen3.5-0.8B |
stateset-agents qwen3-5-0-8b |
Cheapest path to a first run — see docs/QWEN3_FINETUNING_GUIDE.md |
google/gemma-4-31B-it |
stateset-agents gemma-4-31b |
Use --starter-profile memory on tighter GPU budgets |
moonshotai/Kimi-K2.6 |
stateset-agents kimi-k2-6 |
|
moonshotai/Kimi-K3 |
stateset-agents kimi-k3 |
Provisional — HF weights unpublished as of 2026‑07‑16; presets mirror Kimi‑K2.6 |
meta-models/Muse-Glimmer-30B |
stateset-agents muse-glimmer |
Meta's open agentic model (Aug 2026); dense 30B, 131K ctx, Apache‑2.0 |
zai-org/GLM-5.1 |
python examples/finetune_glm5_1_gspo.py |
754B MoE, QLoRA‑only + vLLM; docs/GLM5_1_HOSTING_PLAN.md |
zai-org/GLM-5.2 |
python examples/finetune_glm5_2_gspo.py |
754B MoE, QLoRA‑only + vLLM; docs/GLM5_2_HOSTING_PLAN.md |
Every CLI starter accepts the same flags: --json-output, --list-profiles,
--starter-profile NAME, --write-config PATH, --config PATH --no-dry-run.
The GLM starters are importable too (from stateset_agents.training.glm5_2_starter import get_glm5_2_config, run_glm5_2_config), as are the others.
Supported models
First-class starters ship for Qwen 3.5 0.8B, Gemma 4 31B IT, Kimi-K2.6, Kimi-K3 (provisional), Muse Glimmer 30B, GLM 5.1, and GLM 5.2. Reference examples and hosting plans cover Qwen 3.5 27B, Qwen 3, Qwen 2.5, Kimi-K2.5, Gemma 3 / Gemma 2 27B IT, Llama 3, Llama 2 7B, and Mistral 7B. Any HuggingFace causal LM compatible with AutoModelForCausalLM + TRL GRPO is supported through the generic flow.
See docs/SUPPORTED_MODELS.md for the full matrix, algorithm compatibility, and instructions for adding a new starter.
Experimental namespace
Research-grade modules (neural architecture search, multimodal processing,
long-term planning, few-shot adaptation, the intelligent orchestrator,
adaptive learning controller, and multi-agent coordination) live in
stateset_agents.experimental. They carry no API-stability guarantees
and may change or be removed in any release. The former
stateset_agents.core.<module> import paths still work for one deprecation
cycle and emit a DeprecationWarning.
Dashboard and mobile app (separate repos)
The React + Vite dashboard and Expo mobile app — working clients for the
simulator-backed /api/lab/* "Training Lab" router — live in their own
repositories:
stateset-agents-dashboard
and
stateset-agents-mobile
(extracted from this repo with full history, 2026-08-11). Both are runnable
locally but have no deployment path today — the router is gated behind auth
and the API_ENABLE_TRAINING_LAB flag (off by default).
API serving (/v1/messages)
export INFERENCE_BACKEND=vllm
export INFERENCE_BACKEND_URL=http://localhost:8001
export INFERENCE_DEFAULT_MODEL=moonshotai/Kimi-K2.5
# Optional: ask the backend to include token usage in streaming chunks when supported.
export INFERENCE_STREAM_INCLUDE_USAGE=true
curl http://localhost:8000/v1/messages \
-H "Content-Type: application/json" \
-d '{
"model": "moonshotai/Kimi-K2.5",
"max_tokens": 128,
"messages": [{"role": "user", "content": "Hello"}]
}'
OpenAI-compatible endpoint:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "moonshotai/Kimi-K2.5",
"max_tokens": 128,
"messages": [{"role": "user", "content": "Hello"}]
}'
Helm deployment
helm upgrade --install stateset-agents deployment/helm/stateset-agents \
--namespace stateset-agents
Quick start
1) Stub hello world (no downloads)
Runs without Torch/transformers and is ideal for CI or prototyping.
import asyncio
from stateset_agents import MultiTurnAgent
from stateset_agents.core.agent import AgentConfig
async def main():
agent = MultiTurnAgent(AgentConfig(model_name="stub://demo"))
await agent.initialize()
reply = await agent.generate_response([{"role": "user", "content": "Hi!"}])
print(reply)
asyncio.run(main())
2) Chat with a real model
import asyncio
from stateset_agents import MultiTurnAgent
from stateset_agents.core.agent import AgentConfig
async def main():
agent = MultiTurnAgent(
AgentConfig(
model_name="your-real-model-id",
max_new_tokens=128,
temperature=0.7,
)
)
await agent.initialize()
messages = [{"role": "user", "content": "What is GRPO?"}]
print(await agent.generate_response(messages))
asyncio.run(main())
For the zero-download onboarding path, run python examples/quick_start.py.
Train a multi‑turn agent with GRPO
The high‑level train(...) helper chooses single‑turn vs multi‑turn GRPO automatically.
import asyncio
from stateset_agents import (
MultiTurnAgent,
ConversationEnvironment,
CompositeReward,
HelpfulnessReward,
SafetyReward,
train,
)
from stateset_agents.core.agent import AgentConfig
async def main():
# 1) Agent
agent = MultiTurnAgent(
AgentConfig(
model_name="stub://quickstart",
use_stub_model=True,
system_prompt="You are a helpful customer support assistant.",
)
)
await agent.initialize()
# 2) Environment
scenarios = [
{
"id": "refund",
"topic": "refunds",
"context": "User wants a refund for a delayed order.",
"user_responses": [
"My order is late.",
"I'd like a refund.",
"Thanks for your help.",
],
}
]
env = ConversationEnvironment(scenarios=scenarios, max_turns=6)
# 3) Reward
reward_fn = CompositeReward(
[HelpfulnessReward(weight=0.7), SafetyReward(weight=0.3)]
)
# 4) Train
trained_agent = await train(
agent=agent,
environment=env,
reward_fn=reward_fn,
num_episodes=4,
profile="balanced",
training_mode="single_turn",
save_path="./outputs/refund_agent",
)
# 5) Try the trained model
resp = await trained_agent.generate_response(
[{"role": "user", "content": "My order was delayed, what can you do?"}]
)
print(resp)
asyncio.run(main())
More end‑to‑end scripts live in examples/complete_grpo_training.py and examples/production_ready_customer_service.py.
Continual learning + long‑term planning (optional)
Enable planning context and replay/LwF in the trainer with config overrides:
agent = MultiTurnAgent(
AgentConfig(
model_name="stub://quickstart",
use_stub_model=True,
enable_planning=True,
planning_config={"max_steps": 4},
)
)
trained_agent = await train(
agent=agent,
environment=env,
reward_fn=reward_fn,
num_episodes=4,
training_mode="single_turn",
# resume_from_checkpoint="./outputs/checkpoint-100",
config_overrides={
"continual_strategy": "replay_lwf",
"continual_kl_beta": 0.1,
"replay_buffer_size": 500,
"replay_ratio": 0.3,
"replay_sampling": "balanced",
"task_id_key": "task_id",
"task_schedule": ["task_a", "task_b"],
"task_switch_steps": 25,
},
)
context = {"conversation_id": "demo-trip", "goal": "Plan a 4-day trip to Kyoto"}
resp = await trained_agent.generate_response(
[{"role": "user", "content": "Can you draft a plan?"}],
context=context,
)
followup = await trained_agent.generate_response(
[{"role": "user", "content": "Great. What should we do next?"}],
context={"conversation_id": "demo-trip", "plan_update": {"action": "advance"}},
)
# To update the plan goal explicitly:
# context={"conversation_id": "demo-trip", "plan_goal": "Plan a 4-day trip to Osaka"}
Other training algorithms
All algorithms are available under stateset_agents.training when training deps are installed:
- GSPO: stable sequence‑level GRPO variant (
GSPOTrainer,GSPOConfig,train_with_gspo) - GEPO: expectation‑based group optimization for heterogeneous/distributed setups
- DAPO: decoupled clip + dynamic sampling for reasoning‑heavy tasks
- VAPO: value‑augmented group optimization (strong for math/reasoning)
- PPO baseline: standard PPO trainer for comparison
- RLAIF: RL from AI feedback via judge/reward models
Minimal GSPO sketch:
from stateset_agents.training import get_config_for_task, GSPOConfig, train_with_gspo
from stateset_agents.rewards.multi_objective_reward import create_customer_service_reward
base_cfg = get_config_for_task("customer_service", model_name="your-real-model-id")
gspo_cfg = GSPOConfig.from_training_config(base_cfg, num_outer_iterations=5)
trained_agent = await train_with_gspo(
config=gspo_cfg,
agent=agent,
environment=env,
reward_model=create_customer_service_reward(),
)
See docs/GSPO_GUIDE.md, docs/ADVANCED_RL_ALGORITHMS.md, and examples/train_with_gspo.py for full configs.
Scaffold a fine‑tuning project in 30 seconds
If you're building a fine‑tune for a client, start from a template instead of from scratch:
# See what's available
stateset-agents starter list
# Multi-turn customer support agent (the framework's differentiator)
stateset-agents starter customer-support ./my-client
# Single-turn math reasoner with verifiable rewards
stateset-agents starter gsm8k-math ./math-bench
# Agent that learns to invoke tools/APIs (weather, calculator, search stubs)
stateset-agents starter tool-calling-agent ./tool-agent
# Bare scaffold — edit everything
stateset-agents starter minimal ./hack
Each scaffold lands a runnable project: config.yaml, scenarios.jsonl (where applicable), reward.py, train.py, eval.py, serve.sh, plus a tailored README.md. From clone to running endpoint in three commands:
cd my-client
pip install -r requirements.txt
python train.py # trains on the bundled sample data
./serve.sh outputs/customer_support_v1 # serves via FastAPI gateway
Replace scenarios.jsonl with your client's data — same schema — and you're consulting.
Chat with your fine‑tune locally
# Interactive REPL — no API server needed, exits cleanly with /quit or Ctrl+D
stateset-agents chat --model Qwen/Qwen3.5-0.8B --checkpoint outputs/acme_v1
# With live reward grading — see scores after every assistant turn
stateset-agents chat --grade customer_support --history conversation.jsonl
The chat REPL is the fastest path from "did my fine-tune even load?" to "let me feel how it behaves on the queries that matter." The optional --history flag captures every turn to JSONL for later grading or replay; --grade shows live composite-reward scores so you can spot reward-function disagreements with your intuition in real time.
Curate good examples — build the next training set
After capturing many conversations, score them with the same reward function used during training, and curate the high-scoring ones as new training data:
# Grade every transcript in a directory + collect good examples into one JSONL
make grade-batch DIR=transcripts/ REWARD=customer_support \
CURATED=curated.jsonl THRESHOLD=0.7
# One-shot summary across all graded sessions
make grade-batch-summary GRADED_DIR=transcripts/graded
The curated file is idempotent across reruns — duplicate (prompt, response) pairs are skipped, so you can re-grade as your reward function evolves without polluting the curated set.
This closes the human-in-the-loop curation cycle: train → eval → chat → capture → grade → curate → train again.
Benchmark your fine‑tune
After training, you usually want a defensible number: did this actually improve over the base model, by how much, and is it reproducible? The framework ships a Phase‑0 benchmark pipeline that produces publication‑grade results across three tasks (GSM8K, the bundled customer‑support corpus, and the tool‑calling corpus).
Quick path: open one of the bundled Colab notebooks. The whitepaper §11.7 canonical result was produced by customer_support_3seed_judge.ipynb — judge improvement +0.079 with three-seed agreement on Qwen2.5-0.5B-Instruct (artifact).
| Notebook | Task | Runtime on A100 |
|---|---|---|
notebooks/customer_support_3seed_judge.ipynb |
Whitepaper §11.7 publication-gate notebook — 3 seeds × dual eval (rubric + LLM judge) | ~25 min |
notebooks/whitepaper_v1_comparative_trainers.ipynb |
TRL GRPO vs GSPO vs DAPO head-to-head on §11.7 protocol | ~45 min |
notebooks/whitepaper_v1_gsm8k_benchmark.ipynb |
GSM8K (single‑turn math) — binary reward | ~45 min |
notebooks/whitepaper_v1_gsm8k_benchmark_v2.ipynb |
GSM8K — dense-reward A/B variant | ~45 min |
notebooks/customer_support_4h.ipynb |
Multi‑turn customer support (single-seed) | ~3 h |
notebooks/vllm_speedup_benchmark.ipynb |
HF generate vs vLLM throughput sweep for §6.4 | ~20 min |
See notebooks/README.md for all ten core notebooks (the four above plus quickstart, tool-calling, curate, SFT-closure, and the standard GSM8K variant). Every notebook is JSON-validated and lint-checked in CI via scripts/lint_notebooks.py — pre-flighting against the eight foot-gun patterns from issue #16 (asyncio.run in Jupyter, abstract Agent base, flash-attn defaults, etc.).
CLI path (local A100 / H100):
# 6-second pipeline health check (no GPU)
make benchmark-smoke
# Run one configuration
make benchmark-phase0 TRAINER=gspo SEED=42
# Full matrix: 3 trainers × 3 seeds × 1 task = 9 runs
make benchmark-phase0-all
# Aggregate JSONs → markdown + CSV + PNG figures + gate report
make release-whitepaper-v1
The pipeline:
- Reproducibility.
set_all_seeds()covers Python random, NumPy, PyTorch (CPU + CUDA), and Transformers in one call. Every result JSON carries the git commit hash. - Schema. Each run produces a single JSON conforming to
benchmark_results/SCHEMA.md. Every published number traces back to a file. - Publication gates. 3 seeds, σ < 0.10, +0.03 improvement, single commit. Use
make benchmark-aggregate-strictin CI to enforce. - Figures.
make benchmark-plotproduces two whitepaper‑ready PNGs (pass@1 per trainer, improvement ranking) plus a matplotlib‑free text fallback. - One‑shot release.
make release-whitepaper-v1aggregates → plots → generates the whitepaper §11.7 markdown snippet → copies figures intodocs/figures/→ writes a release manifest. Six artifacts in one command.
See benchmark_results/README.md for the full pipeline reference.
Offline RL: Learn from logged conversations
Train agents from historical conversation logs without online interaction. Useful when:
- You have existing customer service transcripts
- Online training is expensive or risky
- You want to bootstrap before online fine‑tuning
Available Algorithms
| Algorithm | Best For | Key Innovation |
|---|---|---|
| BCQ | Conservative learning | VAE‑constrained action space |
| BEAR | Distribution matching | MMD kernel regularization |
| CQL | Pessimistic Q‑values | Conservative Q‑function penalty |
| IQL | Expectile regression | Implicit value learning |
| Decision Transformer | Sequence modeling | Return‑conditioned generation |
Quick Start
from stateset_agents.data import ConversationDataset, ConversationDatasetConfig
from stateset_agents.training import BCQTrainer, BCQConfig
# Load historical conversations
config = ConversationDatasetConfig(quality_threshold=0.7)
dataset = ConversationDataset.from_jsonl("conversations.jsonl", config)
# Train with BCQ
bcq_config = BCQConfig(
hidden_dim=256,
latent_dim=64,
num_epochs=100,
)
trainer = BCQTrainer(bcq_config)
await trainer.train(dataset)
Hybrid Offline + Online Training
Combine offline pretraining with online GRPO fine‑tuning:
from stateset_agents.training import OfflineGRPOTrainer, OfflineGRPOConfig
config = OfflineGRPOConfig(
offline_algorithm="cql",
offline_pretrain_steps=1000,
online_ratio=0.3, # 30% online, 70% offline
)
trainer = OfflineGRPOTrainer(config)
trained = await trainer.train(agent, env, reward_fn, offline_dataset=dataset)
See docs/OFFLINE_RL_SIM_TO_REAL_GUIDE.md for complete documentation.
Sim‑to‑Real Transfer
Train in simulation, deploy to real users. The framework provides:
Domain Randomization
Generate diverse training scenarios with randomized user personas:
from stateset_agents.training import DomainRandomizer, DomainRandomizationConfig
config = DomainRandomizationConfig(
persona_variation=0.3,
topic_variation=0.2,
style_variation=0.2,
)
randomizer = DomainRandomizer(config)
# Randomize during training
persona = randomizer.sample_persona()
scenario = randomizer.sample_scenario(topic="returns")
Conversation Simulator
Calibratable simulator with adjustable realism:
from stateset_agents.environments import ConversationSimulator, ConversationSimulatorConfig
simulator = ConversationSimulator(ConversationSimulatorConfig(
base_model="gpt2",
realism_level=0.8,
))
# Calibrate to real data
await simulator.calibrate(real_conversations)
# Measure sim‑to‑real gap
gap = simulator.compute_sim_real_gap(real_data, sim_data)
Progressive Transfer
Gradually transition from simulation to real interactions:
from stateset_agents.training import SimToRealTransfer, SimToRealConfig
transfer = SimToRealTransfer(SimToRealConfig(
transfer_schedule="cosine", # linear, exponential, step
warmup_steps=100,
total_steps=1000,
))
# Get current sim/real mixing ratio
sim_ratio = transfer.get_sim_ratio(current_step)
See docs/OFFLINE_RL_SIM_TO_REAL_GUIDE.md for complete documentation.
Hyperparameter optimization (HPO)
Install with stateset-agents[hpo], then:
from stateset_agents.training import TrainingConfig, TrainingProfile
from stateset_agents.training.hpo import quick_hpo
base_cfg = TrainingConfig.from_profile(
TrainingProfile.BALANCED, num_episodes=100
)
summary = await quick_hpo(
agent=agent,
environment=env,
reward_function=reward_fn,
base_config=base_cfg,
n_trials=30,
)
print(summary.best_params)
See docs/HPO_GUIDE.md and examples/hpo_training_example.py.
Custom rewards
Use the decorator for quick experiments:
from stateset_agents.core.reward import reward_function
@reward_function(weight=0.5)
async def politeness_reward(turns, context=None) -> float:
return 1.0 if any("please" in t.content.lower() for t in turns) else 0.0
Combine with built‑ins via CompositeReward.
Custom environments
Subclass Environment for task‑specific dynamics:
from stateset_agents.core.environment import Environment, EnvironmentState
from stateset_agents.core.trajectory import ConversationTurn
class MyEnv(Environment):
async def reset(self, scenario=None) -> EnvironmentState:
...
async def step(
self, state: EnvironmentState, action: ConversationTurn
):
...
Checkpoints
train(..., save_path="...")saves an agent checkpoint.- Load later:
from stateset_agents.core.agent import load_agent_from_checkpoint
agent = await load_agent_from_checkpoint("./outputs/refund_agent")
Auto‑Research
Run autonomous hyperparameter experiments overnight. The loop proposes configurations, trains with a time budget, evaluates on held‑out scenarios, and keeps only improvements.
# Quick test (no GPU)
stateset-agents auto-research --stub --max-experiments 5
# Real training with smart proposer
stateset-agents auto-research --proposer smart --improvement-patience 10
# From a config file
stateset-agents auto-research --config config.yaml
7 proposer strategies (perturbation, smart, adaptive, random, grid, bayesian, LLM), 5 search spaces, early abort on bad experiments, resume from checkpoint, W&B logging, and post‑run analysis with parameter importance.
# Load and analyze results after a run
from stateset_agents.training.auto_research import ExperimentTracker, compare_runs
tracker = ExperimentTracker.load("./auto_research_results")
tracker.print_summary()
print(compare_runs("./run_a", "./run_b"))
See docs/AUTO_RESEARCH_GUIDE.md for the full guide.
CLI
The CLI is a thin wrapper around the Python API:
stateset-agents version
stateset-agents doctor
stateset-agents train --stub
stateset-agents train --config ./config.yaml --dry-run false --save ./outputs/ckpt
stateset-agents evaluate --checkpoint ./outputs/ckpt --message "Hello"
stateset-agents serve --host 0.0.0.0 --port 8001
stateset-agents auto-research --proposer smart --max-experiments 50
For complex runs prefer the Python API and the examples folder.
Examples and docs
Start here:
docs/WHITEPAPER.md— the v0.13.4 technical whitepaper. Anchored to a specific git commit; every claim is verifiable via Appendix C.docs/WHITEPAPER_ERRATA.md— corrections published after each whitepaper revision.docs/PLATFORM_TOUR.md— a guided walk frompip installto a published v1.0 whitepaper revision (linear, journey-style).docs/COOKBOOK.md— copy-paste recipes for 8 common workflows (look up what you need).notebooks/README.md— a map of the ten bundled Colab notebooks: which to open when.benchmark_results/whitepaper_v1/— first-party result artifacts including the §11.7 canonical positive result.CHANGELOG.md— what changed in each release (latest releasev0.23.0).
Other entry points:
examples/getting_started/— start here afterpip install: five small examples (stub hello, custom reward, first GSPO fine-tune, LLM-judge eval, serve via FastAPI). All target the published PyPI version; the GPU-free three smoke-test the install end-to-end. Runmake getting-started-smoketo verify all three at once.examples/finetune_gspo.py– unified finetune driver:--model <preset>over the 12-model registry (--list-models), safe--dry-runby default,--no-dry-runto trainexamples/hello_world.py– stub mode walkthroughexamples/quick_start.py– stub-backed onboarding example with training + smoke testexamples/complete_grpo_training.py– end‑to‑end GRPO trainingexamples/train_with_gspo.py– GSPO + GSPO‑token trainingexamples/train_with_trl_grpo.py– Hugging Face TRL GRPO integrationexamples/auto_research_quickstart.py– autonomous experiment loop
Key docs:
docs/AUTO_RESEARCH_GUIDE.mddocs/RL_FRAMEWORK_GUIDE.md— canonical usage guidedocs/GSPO_GUIDE.mddocs/OFFLINE_RL_SIM_TO_REAL_GUIDE.mddocs/HPO_GUIDE.mddocs/CLI_REFERENCE.mddocs/ARCHITECTURE.md
Related Projects
- stateset-nsr - Neuro‑symbolic reasoning engine for explainable tools.
- stateset-api - Commerce/operations API that agents can drive.
- stateset-sync-server - Multi‑tenant orchestration and integrations.
- core - Cosmos SDK blockchain for on‑chain commerce.
- Public API docs: https://docs.stateset.com
Contributing
See CONTRIBUTING.md. Please run pytest -q and format with black/isort before opening a PR.
License
Business Source License 1.1. Non‑production use permitted until 2029‑09‑03, then transitions to Apache 2.0. See LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file stateset_agents-0.25.0.tar.gz.
File metadata
- Download URL: stateset_agents-0.25.0.tar.gz
- Upload date:
- Size: 904.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.10.16
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2c410ef25a30407b6ad5416011fc0299f0e41b64d43d2addab85138029f44c48
|
|
| MD5 |
bbcac363b7b65c4ce62f8fb323c27323
|
|
| BLAKE2b-256 |
70b23734b0c0838a730add1d778013ee62797c8a8c37598a2a9900ebc3c4737f
|
File details
Details for the file stateset_agents-0.25.0-py3-none-any.whl.
File metadata
- Download URL: stateset_agents-0.25.0-py3-none-any.whl
- Upload date:
- Size: 1.0 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.10.16
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
71213ffc264975356f773f030ff43a99685c4e59e982b9674ccac9f93f1d482c
|
|
| MD5 |
ea2c430e74852c4b21892d18b341a91a
|
|
| BLAKE2b-256 |
eb1a3de03f91f81af9b6edc98fdf5ad8da17cebece6f857f22accddcd0143910
|