This release is a pre-release and may not be stable for production use.
SwarmBots
SwarmBots is a GPU-vectorized multi-agent reinforcement-learning benchmark for self-assembling modular robots. Policies control both movement and assembly: identical articulated units can move independently, connect into load-bearing structures, and disconnect during an episode. The swarm's physical structure is part of the control problem.
The task suite spans wall and bridge traversal, exploration under partial observability, climbing, navigation, and cooperative payload transport. Use your own learning algorithm with the benchmark's environment and evaluation API, or start with the included learning baselines.
| Wall traversal | Finding a hidden opening |
|---|---|
See the scenario catalog for task definitions.
Benchmark highlights
- Physical self-assembly. Agents control articulated limbs and connectors. Forming or releasing a connection changes how the units can move together and transmit forces.
- Partial observability. Hidden-wall and hidden-opening tasks require policies to act without direct access to obstacle geometry. Privileged simulator information is available for training critics, but excluded from evaluated actors.
- Variable assemblies. Registered tasks sample initial morphologies with four or five active units. An agent mask identifies active units within five padded slots.
- GPU simulation. Built on MuJoCo Warp, SwarmBots runs parallel worlds and keeps observations, actions, and rewards as PyTorch tensors on the simulation device.
Task suite
The full suite has 14 registered tasks, including an eight-task core suite for evaluation. Task names below expand to SwarmBots-<name>-v0.
| Task family | Tasks | Challenge |
|---|---|---|
| Obstacle traversal | WallEasy, WallMedium, WallHard, Bridge |
Cross walls from 0.2 to 0.4 m high or a narrow movable bridge. |
| Partial observability | POWallEasy, POWallMedium, FindOpening |
Cross a randomized hidden wall or explore to find a hidden opening. |
| Climbing and navigation | Climb, VerticalReach, MoveTo |
Climb onto a platform, reach an elevated goal, or move toward a sampled planar goal. |
| Payload transport | PayloadPlane, PayloadStep, DualPayloadPlane, MultiPayloadGoal |
Move one or more payloads, overcome a step, and deliver a variable set of payloads to assigned goals. |
The core suite covers medium fixed and hidden walls, bridge traversal, finding an opening, climbing, vertical reach, payload-over-step transport, and multi-payload goal transport. Access it through swarmbots.CORE_BENCHMARK_IDS; swarmbots.ALL_BENCHMARK_IDS exposes the full suite.
Each task defines its own rewards and, where applicable, a terminal success condition. The scenario catalog lists exact benchmark IDs, observations, and success criteria. Scenario parameters are customizable for new experiments; report modified tasks as custom variants.
Install
Python 3.11 or newer is required. CUDA is strongly recommended; CPU execution exists for development and tests but is not the benchmark's performance target. SwarmBots is currently versioned as an alpha. Once 0.1.0a1 is published to PyPI, install it explicitly with:
uv add "swarmbots==0.1.0a1"
Install from source:
git clone https://github.com/brn-dev/swarm-bots.git
cd swarm-bots
uv sync
The source checkout selects CUDA PyTorch on Windows and Linux, including matching Windows Triton, through its lockfile and default cuda dependency group. Subsequent syncs retain that setup. Follow GPU setup to configure the compiler, verify the compiled environment, or select CUDA when using the published package in another project.
Quick start
List the registered tasks and run a short smoke test:
.venv/bin/swarmbots list
.venv/bin/swarmbots smoke SwarmBots-WallEasy-v0 --device cuda --num-envs 64
On Windows, use .venv\Scripts\swarmbots.exe. No environment activation is required. For a CPU smoke test, use --device cpu --num-envs 2; CPU simulation is much slower than the intended CUDA path.
Add --compiled to verify PyTorch compilation as well as simulation. The ordinary smoke test uses eager operations; it does not check compiler setup.
Create an environment directly:
import torch
from swarmbots import make_env
env = make_env("SwarmBots-WallEasy-v0", num_envs=256, device="cuda", seed=42)
observations, info = env.reset(seed=42)
actions = {
"actuators": torch.zeros(env.action_space["actuators"].shape, device=env.device),
"connectors": torch.zeros(env.action_space["connectors"].shape, device=env.device),
}
observations, rewards, terminations, truncations, info = env.step(actions)
env.close()
The environment uses Gymnasium SAME_STEP autoreset. When a lane ends, the returned observation is already its next reset observation; the terminal observation is in info["final_obs"] and selected by info["_final_obs"].
Explicit resets and autoresets both settle physics before returning observations. The settling interval is outside the episode's control-step budget, so evaluation and training start from the same reset distribution.
Use your own policy
SwarmBots exposes vector environments with Gymnasium spaces and reset() / step() conventions. Data stays in PyTorch tensors; Gymnasium wrappers and training libraries that expect NumPy arrays require adaptation.
Observations are dictionaries of batched tensors:
local_obs: per-agent observations, shaped(worlds, agents, features).global_obs: task information visible to all agents.agent_mask: which padded agent slots are active.hidden_local_varsandhidden_global_vars: privileged state for centralized training or diagnostics. Do not pass these to an evaluated actor.
Actions contain per-agent actuators and connectors tensors. Rewards and done flags are team-level tensors with one value per simulated world.
Centralized, partially centralized, and decentralized actors are all allowed; actors may combine the permitted observations across agents within each world. The benchmark does not prescribe a learning algorithm or require decentralized execution. See the Python API for integration details.
Evaluate and compare policies
The built-in evaluator accepts a callable policy(observations, episode_starts) that returns the actuators and connectors action tensors. It passes only non-privileged observations and reports episode returns, lengths, and success rate where defined. Recurrent policies use episode_starts to reset their state per world.
Evaluate one seed of your policy with:
from swarmbots import evaluate_policy
result = evaluate_policy(
policy,
"SwarmBots-WallMedium-v0",
num_envs=256,
num_episodes=256,
seed=1000,
device="cuda",
action_mode="deterministic",
)
print(result.mean_return, result.success_rate)
Put neural-network policies in evaluation mode, select deterministic actions, and freeze observation normalization before calling the evaluator. action_mode records that choice; it does not change policy behavior.
For comparable results, protocol 0.1 specifies:
- Unmodified registered tasks with a 500-control-step episode limit.
- Seeds
1000through1004, with 256 worlds and exactly the first episode from each world per seed: 1,280 episodes per task. - Per-task mean return and success rate where defined, with mean and standard deviation across the five seed-level means, plus raw episode data and runtime metadata.
The evaluator requires num_episodes <= num_envs; additional seeds provide more samples. Reward scales differ between tasks, so report per-task scores. Evaluation samples the registered morphology pool; protocol 0.1 does not measure generalization to unseen morphologies. Report training budgets and training seeds separately.
The five-seed reporting example starts from a uniform-random sanity baseline and shows how to export episodes, settings, runtime versions, and aggregate statistics to JSON.
Optional learning baselines
The package includes PPO/MAPPO, multi-agent transformer (MAT), transformer-based SAC (TMASAC), and recurrent variants as starting points for benchmark experiments.
from swarmbots.learn import train
trainer = train(
"SwarmBots-WallMedium-v0",
"mappo",
num_envs=1024,
device="cuda",
total_timesteps=100_000_000,
run_dir="runs/wall-medium/mappo",
)
list_variants() lists the available presets, and as_benchmark_policy(trainer) adapts a trained actor for evaluation. See learning baselines for customization, checkpoint continuation, and custom training loops.
Record policy behavior
Record your policy with swarmbots record <benchmark-id> --policy module:function, or an included-baseline checkpoint with --checkpoint <path> --variant <variant>. See recording for the policy factory interface, checkout script, and video options.
Documentation
The documentation guide provides a starting point for exploring tasks, integrating policies, and reporting results.
- Scenario catalog
- Benchmark and reporting protocol
- Python API, compatibility, and custom policies
- GPU setup and compiled smoke test
- Policy recording and video options
- Optional learning baselines and presets
The API is alpha. Benchmark IDs and protocol versions are explicit so semantic changes can be introduced without silently invalidating results.
Citation
Please cite the software version used in your experiments. Machine-readable citation metadata is available in CITATION.cff.
For the benchmark and policy designs, also cite Dominik Baron (2026), SwarmBots: a GPU-accelerated multi-agent continuous control benchmark with transformer baselines, master's thesis, Johannes Kepler University Linz. The published thesis is available under the persistent identifier urn:nbn:at:at-ubl:1-108602.
Thesis source and supplementary material are available in the thesis repository.
License
SwarmBots is released under the Apache License 2.0.
Metadata
Release files for swarmbots 0.1.0a1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| swarmbots-0.1.0a1.tar.gz | 4.4 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| swarmbots-0.1.0a1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 4.9 MB
Release files / swarmbots-0.1.0a1.tar.gz
| Download URL | swarmbots-0.1.0a1.tar.gz |
|---|---|
| Size | 4.4 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c7e6bcf53eef1c2587ea7569e4ffbca91b0c6e566d4852a207940486cb334485
|
|
BLAKE2b-256 checksum How to use checksums |
3d1b69acb82d39388bbf523e86d45cbdc5ecbec9fa20daa43bc60db8265c1954
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency logRelease files / swarmbots-0.1.0a1-py3-none-any.whl
| Download URL | swarmbots-0.1.0a1-py3-none-any.whl |
|---|---|
| Size | 486.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3b8367a3b6e853a8ddda93c91859d5369227628df5d484598ce09ed27e3e5245
|
|
BLAKE2b-256 checksum How to use checksums |
d9eb9d91b6e21855e474eb0f55be76d5566d5cc8bf03f0b156023ad7560fa586
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency log