Skip to main content

Python SDK and CLI for the Epsilab RL Environment Hub and Marketplace.

Project description

Epsilab Python SDK

Python SDK and CLI for the Epsilab RL Environment Hub.

What is Epsilab?

Epsilab is an open hub for RL environments. Search, run, and export training data from hosted environments, or publish your own with a single command.

  • Researchers and teams training models — search and run environments, export GRPO/DPO/SFT/KTO training data, batch evaluation
  • Environment builders — publish with epsilab deploy, immutable content-addressed releases, usage analytics
  • Open by default — public, shared, unlisted, or private visibility; auto-qualified releases

Installation

pip install epsilab                  # SDK + CLI only
pip install epsilab[training]        # + TRL, Together AI, torch
pip install epsilab[tinker]          # + Tinker (custom training loops)
pip install epsilab[fireworks]       # + Fireworks AI
pip install epsilab[all]             # everything

Quick Start

Deploy an environment

epsilab login
epsilab whoami
epsilab init my-environment
cd my-environment
# edit environment.py and tasks.json
epsilab deploy

The deploy command creates the namespace and listing when needed, builds the image, uploads it, registers immutable task and verifier releases, and makes the environment available for hosted sessions. No UUIDs or registry credentials are required.

Run an environment

epsilab run epsilab/bug-hunter
epsilab run epsilab/bug-hunter \
  --task bug-hunter-easy-train-001 \
  --action 'BUG_LINE: 6
FIX:
return total / len(numbers)'

The first form starts an interactive session. The second performs one action and prints its observation, reward, terminal state, and verification result.

from epsilab import Epsilab

# Credentials are loaded automatically from `epsilab login`
client = Epsilab()

# Browse the hub
listings = client.list_environment_listings(limit=20)
for l in listings:
    print(f"{l.slug:30s}  {l.title}")

# Create a session and interact
session = client.create_environment_session(
    listings[0].deployment_id,
    task_id="bug-hunter-easy-train-001",
    seed=42,
)
session = client.wait_for_session(session)
result = client.environment_step(
    session.session_id,
    {
        "action_type": "submit",
        "content": "BUG_LINE: 6\nFIX:\nreturn total / len(numbers)",
    },
    session_token=session.session_token,
)
print(f"Reward: {result.reward}, Done: {result.done}")

# Export training data
export = client.create_environment_export(
    deployment_id=listings[0].deployment_id,
    format="grpo",
)

Run a long-horizon coding agent

run_agent_episode separates model turns from environment steps. A model may reason for many turns before calling a tool; every request, reasoning chunk, assistant message, tool boundary, result, usage record, cancellation, and error is persisted immediately in the session trace.

from epsilab import AgentToolCall, AgentTurn, AgentUsage, Epsilab

client = Epsilab()

def call_your_model(context):
    # Invoke any provider here. Translate its response into AgentTurn.
    # Do not set a provider token limit if you want an uncapped generation.
    response = your_provider_call(context.history, context.observation)
    return AgentTurn(
        reasoning=response.reasoning,
        message=response.message,
        tool_calls=[
            AgentToolCall(call_id=call.id, name=call.name, arguments=call.arguments)
            for call in response.tool_calls
        ],
        usage=AgentUsage(
            input_tokens=response.input_tokens,
            output_tokens=response.output_tokens,
            cost_usd=response.cost_usd,
        ),
        provider=response.provider,
        model=response.model,
        provider_request_id=response.request_id,
    )

rollout = client.run_agent_episode(
    deployment_id="deployment-id",
    task_id="task-id",
    model_fn=call_your_model,
    max_turns=500,
    cancel_check=my_cancel_event.is_set,
)
print(rollout.stop_reason, rollout.turns_completed, rollout.session.total_reward)

The SDK has no token budget or per-generation cap. max_turns is the only runner limit (maximum 500 model calls); the environment's own immutable task horizon still governs tool actions. See examples/run_long_horizon_agent.py for a provider-neutral adapter skeleton.

Post-training with multiple environments

For generalist RL post-training, train across diverse environments simultaneously — coding, ops, business, etc.:

# Collect data from 3 envs, format as DPO, train locally
python examples/run_environment.py \
    --envs bug-hunter,refactor,test-writer \
    --algorithm dpo \
    --provider local

# Online GRPO across all available envs
python examples/grpo_training.py --envs all --provider tinker --steps 100

The example scripts support any training provider:

Provider Install Description
together pip install epsilab[training] Managed fine-tuning via Together AI
fireworks pip install epsilab[fireworks] Managed fine-tuning via Fireworks AI
tinker pip install epsilab[tinker] Custom training loops on remote GPUs
local pip install epsilab[training] TRL on local GPU (SFT/DPO/KTO/GRPO)

Batch evaluation with your policy

Batch creation provisions reproducible sessions; your code still supplies the agent actions. run_batch handles both parts so sessions cannot be left idle:

batch = client.run_batch(
    deployment_id="deployment-id",
    name="regression sweep",
    task_seed_pairs=[
        {"task_id": "task-001", "seed": 42},
        {"task_id": "task-002", "seed": 42},
    ],
    policy_fn=lambda observation, info: my_agent(observation),
)

for session in batch["sessions"]:
    print(session["task_id"], session["total_reward"])

For a batch created in the dashboard or CLI, call client.drive_batch(batch_id, policy_fn=my_policy).

Creating an Environment

An RL environment is a containerized task server. At minimum it needs:

File Purpose
environment.py Typed, deterministic OpenEnv environment logic
verifier.py Independent trajectory replay and reward verification
Dockerfile Builds the runtime container image
server.py Generated OpenEnv HTTP entry point
tasks.json JSON array defining the task set
.epsilab/project.json Local project linkage, populated on first deploy

Environment protocol

The generated server uses the OpenEnv contract and exposes:

GET  /health
POST /reset
POST /step
WS   /ws

/reset accepts a task ID and deterministic seed. /step accepts a typed action and returns an observation, optional reward, and terminal state. Most authors only edit environment.py; the generated server and verifier already implement the transport and replay contracts.

tasks.json format

[
  {
    "task_id": "find-the-bug-001",
    "prompt": "Find and fix the bug in the following code...",
    "difficulty": "easy",
    "split": "train",
    "max_steps": 10
  }
]

Creating an Application Tool

Application tools are reusable plugins (e.g. GitHub, Slack, Calendar) that environments can compose together. A tool needs:

File Purpose
plugin.py AppPlugin subclass defining the tool's identity and lifecycle
api.py API route handlers (FastAPI-style)
state.py Deterministic state model for the tool
cd my-tool/
epsilab deploy    # auto-detects tool structure

CLI Commands

Command Description
epsilab login Authenticate with your API key
epsilab logout Remove stored credentials
epsilab whoami Show current auth and profile status
epsilab init [name] Scaffold a working environment project
epsilab deploy Build, upload, register, and host an environment or tool
epsilab run <owner>/<name> Run an interactive or one-step hosted session
epsilab env list List your environment listings
epsilab env search [query] Search the hub
epsilab env verify Run local preflight checks
epsilab rl sessions List your RL sessions
epsilab rl trajectory <id> View step-by-step trajectory
epsilab namespace create <slug> Create a namespace

All commands support --json for machine-readable output and -v for verbose logging.

Configuration

The SDK resolves credentials in order:

  1. Explicit api_key= constructor argument
  2. EPSILAB_API_KEY environment variable
  3. ~/.epsilab/credentials.json (set by epsilab login)

Most users just need epsilab login — no env vars or code changes required.

Environment Variable Constructor Param Description
EPSILAB_API_KEY api_key Your API key (overrides stored credentials)
EPSILAB_API_BASE api_base API base URL (default: production)
EPSILAB_HTTP_TIMEOUT timeout_seconds Request timeout in seconds (default: 120)
max_retries Auto-retry count for 429/5xx (default: 3)
load_dotenv Also read a local .env file (default: false)

Error Handling

from epsilab import Epsilab, AuthError, InsufficientCreditsError, RateLimitError, ApiError

client = Epsilab()

try:
    session = client.create_environment_session("dep-id", task_id="task-001")
except AuthError:
    print("Invalid API key")
except RateLimitError as e:
    print(f"Rate limited. Retry after {e.retry_after}s")
except ApiError as e:
    print(f"API error: {e.status_code}")

The SDK retries automatically on rate limits (429) and transient server errors (500, 502, 503, 504) with exponential backoff and jitter.

Examples

Script What it does
examples/example.py Start here — discover envs, run one session, see rewards
examples/run_environment.py Collect data across multiple envs, train with SFT/DPO/KTO
examples/grpo_training.py Online GRPO with live environment rewards
examples/batch_evaluation.py Benchmark a policy across provisioned environment batches
examples/marketplace_example.py Consumer and publisher hub workflows

All examples use argparse — run with --help for options. Key flags:

python examples/run_environment.py \
    --envs bug-hunter,refactor,test-writer \
    --algorithm dpo \
    --provider local \
    --sessions-per-env 20

Documentation

Document Description
API Reference Full method reference for all SDK features
Evaluations (deprecated) Legacy evaluations, voice, routing — will be removed 2026-12-31

License

Apache 2.0 — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

epsilab-0.17.15.tar.gz (108.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

epsilab-0.17.15-py3-none-any.whl (129.7 kB view details)

Uploaded Python 3

File details

Details for the file epsilab-0.17.15.tar.gz.

File metadata

  • Download URL: epsilab-0.17.15.tar.gz
  • Upload date:
  • Size: 108.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for epsilab-0.17.15.tar.gz
Algorithm Hash digest
SHA256 93062dd37d0f41abafba8bccac8cb26a0da145235044c676a20ea33f8cae2fb6
MD5 e585663e77aaada7a36324425f29cdf5
BLAKE2b-256 1482b32f68c74c60a1fb7a57c8534d358074b38d1f673e2e90b58ebc86f9f43d

See more details on using hashes here.

Provenance

The following attestation bundles were made for epsilab-0.17.15.tar.gz:

Publisher: workflow.yml on EpsilabAI/epsilab-python

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file epsilab-0.17.15-py3-none-any.whl.

File metadata

  • Download URL: epsilab-0.17.15-py3-none-any.whl
  • Upload date:
  • Size: 129.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for epsilab-0.17.15-py3-none-any.whl
Algorithm Hash digest
SHA256 28ab015d729c09d92571ddd383fe83015556fc43fb14dbcd8b3597906e0aa957
MD5 261737e0a46fdfd5e2b3b63efc523f07
BLAKE2b-256 eae3cfbf71b6606701af433bf444c12ceb15e59389d8dfd518626d2496d857d0

See more details on using hashes here.

Provenance

The following attestation bundles were made for epsilab-0.17.15-py3-none-any.whl:

Publisher: workflow.yml on EpsilabAI/epsilab-python

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page