Skip to main content

🧬 AgentGenesis

An industrial-grade evaluation SDK for building, registering, and running agent-based coding challenges with dual-sandbox isolation.

PyPI Version Python Versions License Docker Required

🌐 Website • 🎮 Live Platform • 🇨🇳 简体中文


📖 Table of Contents


🎯 Overview

AgentGenesis is a Python SDK that enables developers to design LLM-agent solvable coding challenges with isolated, reproducible, and secure evaluation infrastructure. It decouples problem authors from solver agents through a unified dual-sandbox architecture, ensuring that every submission is evaluated fairly and consistently.

Whether you are:

  • 📝 A problem author designing novel agent benchmarks (multi-agent coordination, tool use, resilience testing...)
  • 🤖 A solver developer building LLM agents that tackle complex tasks
  • 🏢 A platform operator running large-scale evaluation workers

AgentGenesis provides the end-to-end toolkit you need.


💡 Why AgentGenesis?

Capability Description
🔒 Dual-Sandbox Isolation Judge and user code run in separate Docker containers; neither can see the other's internals or private data.
🏠 Local Zero-Difference Reproduction LocalEvaluator mirrors the cloud evaluation path exactly — eliminate "works on my machine" forever.
🚀 Fast Container Startup Template image pool with LRU garbage collection, keyed by pip dependencies + data directories.
📊 Token Metering & Quotas Per-submission LLM usage limits (chars/requests) with aggregate scaling across test cases.
🛠️ Built-in Tool-Calling Framework Async OpenAI-compatible agent loop with structured error taxonomy and self-correction.
📝 Revision Workflow Built-in problem registry with checksum-optimized artifact uploads and automatic revision fallback.

🏗️ Architecture

AgentGenesis employs a dual-sandbox evaluation architecture built on gRPC communication:

┌─────────────────────────────────────────────────────────────┐
│                    Evaluation Worker                         │
│  ┌─────────────────┐         ┌─────────────────────────┐   │
│  │  Judge Sandbox  │◄───────►│  User / Agent Sandbox   │   │
│  │  (run.py)       │  gRPC   │  (solution.py)          │   │
│  │  - Scoring      │ Bridge  │  - Tool Calling         │   │
│  │  - State Machine│         │  - LLM Agent            │   │
│  └─────────────────┘         └─────────────────────────┘   │
│           ▲                                                │
│           │ Cases / Results                                │
│           ▼                                                │
│  ┌─────────────────────────────────────────────────────┐  │
│  │              LocalEvaluator / Cloud Worker           │  │
│  │  - Case generation   - Parallel execution           │  │
│  │  - Event streaming   - Sandbox lifecycle            │  │
│  └─────────────────────────────────────────────────────┘  │
└─────────────────────────────────────────────────────────────┘

Evaluation Modes

  1. Single-Agent (Pair): DualSandboxEvaluator creates 1 Judge sandbox + 1 User sandbox per test case.
  2. Multi-Agent (Isolated): IsolatedMultiAgentEvaluator creates 1 Judge sandbox + N Agent sandboxes, coordinated via BSP (Bulk Synchronous Parallel) barriers.

Communication Protocol

  • gRPC SandboxBridge: Three core RPCs — CheckReady, SendMessage, RecvMessage.
  • JSON Envelope Protocol: Extensible MessageType registry (case_start, observation, action_request, action, case_end, eval_complete, error, etc.).

✨ Features

For Problem Authors

  • Rich Phase Configuration: Extensive per-phase config — resources, timeouts, dependencies, gateway quotas, visibility rules.
  • Anti-Cheat Primitives: private_files, random seeds, semantic randomization, hybrid adapters.
  • Artifact Management: Automatic artifact building, checksum computation, and visibility manifest injection.
  • Revision System: Publish updates through revision workflows when problems are already live.

For Solver Developers

  • Unified Tool-Calling API: Agent, LLMConfig, Tool, batch — async OpenAI-compatible loop.
  • Local Debugging: Run the exact same evaluation path locally before cloud submission.
  • Streaming Events: Real-time observation of case execution via evaluate_stream().

For Platform Operators

  • Worker Service: EvaluationService polls pending submissions, loads evaluators dynamically, and runs cases with health/metrics endpoints.
  • Template Pooling: LRU-cached Docker images slash cold-start time.
  • Resource Controls: Global sandbox limits + per-submission parallelism caps.

📦 Installation

Requires Python ≥3.10.

# Base SDK — problem authoring, local evaluation, registry client
pip install agent-genesis

# Server-side worker — Docker sandbox + gRPC transport
pip install "agent-genesis[server]"

# Development dependencies
pip install "agent-genesis[dev]"

Note: Server/worker mode requires Docker installed and running.


🚀 Quick Start

For Problem Authors

Create a new problem in problems/hello_world/:

# problems/hello_world/config.py
from agent_genesis import PhaseConfig

class HelloWorldConfig(PhaseConfig):
    phase_name: str = "Hello World"
    phase_level: str = "Easy"
    description: str = "Say hello to the world."
    evaluator_class: str = "DualSandboxEvaluator"
# problems/hello_world/register.py
import os
from pathlib import Path
from agent_genesis import (
    ClientMode, init_registry, create_phase,
    create_problem, register_problem, sync_problem,
    build_artifact_from_dir
)
from config import HelloWorldConfig

def main():
    api_key = os.environ.get("AGENT_GENESIS_API_KEY")
    backend_url = os.environ.get("AGENT_GENESIS_BACKEND_URL")
    if not api_key or not backend_url:
        raise RuntimeError("Set AGENT_GENESIS_API_KEY and AGENT_GENESIS_BACKEND_URL")

    init_registry(mode=ClientMode.USER, api_key=api_key, backend_url=backend_url)

    artifact = build_artifact_from_dir(Path(__file__).parent / "sandbox", Path(__file__).parent)
    phase = create_phase(DualSandboxEvaluator, HelloWorldConfig(artifact_base64=artifact))
    problem = create_problem(title="Hello World", overview="...", phases=[phase])
    register_problem(problem)
    print(sync_problem(problem.title))

if __name__ == "__main__":
    main()

Run registration:

export AGENT_GENESIS_API_KEY="your-api-key"
export AGENT_GENESIS_BACKEND_URL="http://your-backend"
python problems/hello_world/register.py

For Solver Agents

Implement a solution using the built-in tool-calling framework:

# answer/hello_world/solution.py
from agent_genesis.tool_calling import Agent, LLMConfig, Tool

llm_config = LLMConfig(
    model="deepseek-chat",
    base_url="https://api.deepseek.com",
    api_key=os.environ["LLM_API_KEY"],
)

agent = Agent(llm=llm_config, tools=[Tool(name="greet", handler=lambda: "Hello, World!")])
result = agent.run("Please greet the world.")
print(result)

For Local Evaluation

Test your problem or solution locally before cloud submission:

from agent_genesis.local import LocalEvaluator
from agent_genesis.tool_calling import LLMConfig

llm_config = LLMConfig(
    model="deepseek-chat",
    base_url="https://api.deepseek.com",
    api_key="your-api-key",
)

ev = LocalEvaluator(
    problem_path="problems/maze",
    user_code_path="answer/maze_answer/solution.py",
    llm_config=llm_config,
)

# Batch evaluation
result = ev.evaluate()
print(f"Passed: {result.passed_cases}/{result.total_cases}")

# Streaming with real-time visualization
for event in ev.evaluate_stream():
    print(event)

🎮 Problem Catalog

The platform includes diverse agent challenges testing different capabilities:

Category Problem Description
Multi-Agent Coordination werewolf Isolated multi-agent werewolf game with role-based strategy
microservice_avalanche Distributed transaction coordination across services
Tool Use & Planning maze Navigate random mazes using LLM agent with tool calls
tool_creator_challenge Dynamically create and use tools to solve queries
Parallel Execution parallel_weather Query 200 cities in <27s using parallel tool calls
short_circuit_scraper Fast-fail pattern with 10 endpoints under time pressure
Resilience & Retry resilient_scraper Exponential backoff with probabilistic failures
Semantic Analysis log_hunter Find hacker IPs in 800K tokens of access logs
interrupt_judge Determine when to interrupt user utterances
Structured Output structured_output Process 1000 questions with strict schema compliance
Shopping Agent sports_shopping Multi-constraint shopping with guardrails and time limits

Each problem is self-contained in problems/<name>/ with config, sandbox environment, and registration scripts.

Live Demo

🎮 Platform: Agent Genesis

  • Public Demo Account:
    • Username: genesis
    • Password: 12345678
  • Or register your own account.

Werewolf Game Demo


🧪 Testing

Run commands from the project root directory.

1) Default OSS Test Run (Recommended)

python -m pytest -q

Default pytest options exclude cross_module tests, so contributors can run the suite without private backend credentials.

Expected outcome:

  • ✅ passed: unit and integration tests executed locally
  • ⏭️ deselected: cross_module tests intentionally excluded by marker filter

2) Coverage Gate Run

python -m pytest agent_genesis/tests -q \
  --cov=agent_genesis \
  --cov-config=../.coveragerc \
  --cov-report=term-missing:skip-covered

The coverage threshold is enforced by .coveragerc (fail_under = 90).

3) Cross-Module Backend Run (Optional)

python -m pytest agent_genesis/tests -q \
  -m cross_module \
  -o addopts="-ra --strict-markers"

These tests require a live backend and environment variables:

Variable Description
BACKEND_URL Backend API base URL
INTERNAL_API_KEY Internal worker API key
AGENT_GENESIS_API_KEY User registry API key
CROSS_TEST_SLUG Test problem slug
CROSS_TEST_SUBMIT_ID Test submission ID
CROSS_TEST_SUBMIT_ID_CLAIMED Claimed submission ID
CROSS_TEST_USER_ID Test user ID
CROSS_TEST_KEY_ID Test key ID

Expected outcome:

  • ⏭️ skipped: environment-dependent fixtures are missing and tests self-skip
  • ✅ passed: backend and credentials are configured correctly

🤝 Contributing

We welcome contributions from the community! Please see our Contributing Guide for details on:

  • Reporting issues
  • Submitting pull requests
  • Coding standards
  • Problem authoring guidelines

For problem authoring in detail, refer to:


📄 License

AgentGenesis is licensed under the Apache License 2.0.


⭐ Star us on GitHub if you find AgentGenesis useful!

🇨🇳 查看简体中文文档

Metadata

Release files for agent-genesis 0.0.58

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agent-genesis 0.0.58
File Size Uploaded
agent_genesis-0.0.58.tar.gz 138.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agent-genesis 0.0.58
File Interpreter ABI Platform
agent_genesis-0.0.58-py3-none-any.whl Python 3 none any Details

Total release size: 251.8 kB

Release files / agent_genesis-0.0.58.tar.gz

Download URL agent_genesis-0.0.58.tar.gz
Size 138.9 kB
Tags Source
SHA-256 checksum
How to use checksums
31046fe758ff8cf9c775198d0e9c4f579c4e447773c21b88bfa8dfaf02a27f18
BLAKE2b-256 checksum
How to use checksums
d47f24f8d32f8e2156f70715fd3bfb41ea94ce665f8322e485cec7d388b328d2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.5

Release files / agent_genesis-0.0.58-py3-none-any.whl

Download URL agent_genesis-0.0.58-py3-none-any.whl
Size 113.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
093f1e20fdefc5c327c3e5d89df203735ea4e9bacdb47a20296a84c6118b5785
BLAKE2b-256 checksum
How to use checksums
d2609ff32e1d3ada92410f2691699890d9ab443111e4f246c8b5e58aa99f44f5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.5

Release history Release notifications | RSS feed

This release

0.0.58 This release

2 release files

0.0.57

2 release files

0.0.55

2 release files

0.0.54

2 release files

0.0.53

2 release files

0.0.52

2 release files

0.0.50

2 release files

0.0.49

2 release files

0.0.48

2 release files

0.0.47

2 release files

0.0.46

2 release files

0.0.45

2 release files

0.0.44

2 release files

0.0.43

2 release files

0.0.42

2 release files

0.0.41

2 release files

0.0.40

2 release files

0.0.26

2 release files

0.0.25

2 release files

0.0.24

2 release files

0.0.23

2 release files

0.0.22

2 release files

0.0.21

2 release files

0.0.20

2 release files

0.0.19

2 release files

0.0.18

2 release files

0.0.17

2 release files

0.0.16

2 release files

0.0.15

2 release files

0.0.14

2 release files

0.0.13

2 release files

0.0.12

2 release files

0.0.11

2 release files

0.0.10

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page