MARL-BattleGrounds
The JAX-native Benchmark for Heterogeneous and Competitive Multi-agent Reinforcement Learning.
Installation | Quick Start | Baselines | Guides | See Also | Citation
MARL-BattleGrounds (MARL-BGs) is an easy-to-use, JAX-native benchmark for heterogeneous and competitive team-vs-team multi-agent reinforcement learning (MARL). It is designed to be sample-efficient, computationally efficient and researcher-centric. Most MARL benchmarks are either fast but simple, like SMAX, where winning comes down to unit micromanagement, or rich but costly, like Dota 2, which took OpenAI Five ten months of training on large clusters. MARL-BGs fills the missing middle: rich team play, with fast and cheap experiments on one GPU.
- The Game. In Team Deathmatch, two teams of up to five agents fight on 52 hand-designed maps. Each agent is a Mage, Warrior, Hunter, Rogue or Priest. Each class has its own Basic ability, passive and Ultimate, built to help allies and counter enemies, so a team wins by working together. Obstacles block sight and attacks, so each agent sees only part of the map. Maps 0-11 form a curriculum, maps 0-41 are for training, maps 42-46 for validation and maps 47-51 for testing.
- The Evaluation. Eight hand-crafted endgame scenarios, each with a verified winning line, test skills such as agent modeling and multi-step planning. A monthly tournament retrains every entry on equal compute, plays it against the current "Big Nine" on the test maps, and rates it with draw-aware Elo and confidence intervals. Every entrant is released as a downloadable opponent. Up to 11,192 metrics per game, in 28 topics, show how a team won, such as how much the Priest healed each ally.
- The Tools. Bring your own learner through plain JAX
resetandstep, or start from six tuned baselines. Their trained Season 0 models, and every save from 0 to 250 million steps, load by name from Hugging Face. Training helpers cover self-play, opponent pools, curricula and reward shaping. One call evaluates any mix of policies, so cross-play and zero-shot coordination tests come out of the box. MARL-BGs also ships three strong scripted opponents, an LLM API that turns a language model into a team, the BattleClient for playing alongside or against your policies, and the analytical Replay Viewer for stepping through any game tick by tick.
Try it in your browser: the Colab notebook walks through the API, the speed, training a team and the trained Season 0 teams.
Installation
Start here:
| You want to | Do this |
|---|---|
| Look around in your browser, with nothing to install | Open the walkthrough in Colab: the API, the speed, training a team and the trained Season 0 teams |
| Play the game | The one command below |
| Use MARL-BattleGrounds in your own project | pip, below |
You do not need to install JAX first: each route installs the right JAX for your machine.
On Linux, an Apple Silicon Mac or Windows with WSL2, one command installs everything and opens the BattleClient. It picks the right JAX build for your GPU:
curl -L https://github.com/arm-2-lab/MARL-BattleGrounds/archive/refs/tags/v1.0.1.tar.gz | tar xz
cd MARL-BattleGrounds-1.0.1 && sh start.sh
In your own Python project (Python 3.12 to 3.14; 3.14 recommended):
pip install "marl-battlegrounds[all]" # CPU
pip install "marl-battlegrounds[all,cuda13]" # NVIDIA GPU, driver 580 or newer
With Docker or Podman, on any system (add --gpus all and the 1.0.1-cuda12 tag for an NVIDIA GPU):
docker run --rm -it -p 127.0.0.1:8765:8766 -p 127.0.0.1:8767:8768 -v "${PWD}:/work" ghcr.io/arm-2-lab/marl-battlegrounds:1.0.1
| Route | For | Tested |
|---|---|---|
One command, sh start.sh |
Linux (x86_64 and ARM), Apple Silicon Mac, Windows with WSL2 | Tested: Linux x86_64 with an NVIDIA RTX 5090; Linux x86_64 and ARM, and an Apple Silicon Mac, on the CPU. Not tested yet; it should work: WSL2 |
| pip | Linux, Apple Silicon Mac, WSL2; Python 3.12 to 3.14 | Tested: Linux x86_64 and ARM, and an Apple Silicon Mac, each with Python 3.12, 3.13 and 3.14. Not tested yet; it should work: WSL2 |
| Docker | Any system with Docker or Podman; NVIDIA GPUs on Linux and Windows | Tested: Linux with Docker (CPU, and an NVIDIA GPU with the CUDA 12 image), Linux with Podman, the ARM image on ARM Linux. Not tested yet; it should work: Docker Desktop on a Mac or Windows |
| Colab, Open in Colab | A browser | Not tested yet; it should work: Colab's free T4 GPU |
"Tested" means we ran exactly these instructions on that system for this release.
Docker runs everything: play, replays, training, evaluation, tournaments, analysis and notebooks. Its limits:
| Limit | What to do |
|---|---|
| No GPU in Docker on a Mac | Use the CPU image; Docker cannot reach Apple GPUs |
study (big multi-run jobs) is not tested in containers and stops when the container stops |
Start the container with -d; not tested yet |
| You cannot add packages to a running container | Build a two-line Dockerfile: FROM ghcr.io/arm-2-lab/marl-battlegrounds:1.0.1, then RUN uv pip install <package> |
| Docker Desktop (Mac, Windows) limits the container's memory | Raise the memory limit in Docker Desktop's settings for large training runs |
Every route, the GPU builds, Windows, clusters and offline use: the install guide.
Quick Start
make, reset and step use JaxMARL's call order (reset(key), step(key, state, actions)), and compile under jit, vmap and lax.scan.
import jax
import marl_battlegrounds as marl_bgs
env = marl_bgs.make("tdm", map_id=0, num_envs=4) # four games at once
key = jax.random.key(0)
observation, state = env.reset(key)
for _ in range(50):
key, act_key, step_key, reset_key = jax.random.split(key, 4)
actions = env.sample_actions(act_key, state) # legal random actions
observation, state, reward, done, info = env.step(step_key, state, actions)
observation, state = env.reset_done(reset_key, state) # restart finished games
On one RTX 5090 with 1,024 parallel 5v5 games, the simulator alone runs 50,416 game steps per second (random legal actions, compilation excluded), and recurrent MAPPO trains at 29,878 game steps per second in pure self-play.
Train
One command trains recurrent IPPO with the Season 0 learner settings (250 million game steps) against itself, its past copies and the three scripted opponents, and keeps its best checkpoint on the validation maps. What Training Costs gives the time and memory on one GPU.
python -m marl_battlegrounds.experiments.train method=ippo pool=scripted \
'validation_opponents=[tdm-alpha,tdm-beta,tdm-gamma]' output_dir=runs/my-ippo
Evaluate
import marl_battlegrounds as marl_bgs
results = marl_bgs.evaluate("random", "tdm-alpha", num_episodes=32, num_envs=32)
print(results.head_to_head()) # games, wins, draws, losses, point margin
Your method is always Team A, and paired games swap the spawn ends, so a map's layout cannot favor either side.
Watch And Play
python -m marl_battlegrounds replay path/to/game.marlbg-replay.json
import marl_battlegrounds as marl_bgs
marl_bgs.battle_client(team_b="tdm-alpha") # play in your browser against a scripted team
Baselines
| Method | Name | Memory | Critic | Reference | Source |
|---|---|---|---|---|---|
| Recurrent MAPPO | mappo |
Recurrent | Centralized | Yu et al., 2022 | ppo.py |
| Feedforward MAPPO | ff_mappo |
None | Centralized | Yu et al., 2022 | ppo.py |
| Recurrent IPPO | ippo |
Recurrent | One per agent | de Witt et al., 2020 | ppo.py |
| Feedforward IPPO | ff_ippo |
None | One per agent | de Witt et al., 2020 | ppo.py |
| QMIX | qmix |
Recurrent | QMIX mixer | Rashid et al., 2018 | qmix.py |
| PQN-VDN | pqn_vdn |
Recurrent | Sum of agent values | Gallici et al., 2024; Sunehag et al., 2017 | pqn.py |
Every save of the six Season 0 models, from 0 to 250 million steps, loads by name from
Hugging Face, for example
marl_bgs.load_method("MARL-BattleGrounds/tdm-season-0-recurrent-ippo@final")
(Released Models).
Your own method can be anything that turns permitted observations into legal actions: a network, a planner, a language model or a mix.
Guides
| Guide | What It Covers | Website |
|---|---|---|
| Getting Started | Install, your first games, maps, rosters, what agents see | start |
| Game Rules | Classes, abilities, scoring, Red Zone, scenarios | game |
| Your Method | Plug in your own model or learner, with JAX's jit, vmap and scan | your-method |
| Training | The six learners, opponents, curriculum, rewards, checkpoints | training |
| Released Models | Load, play, train against or start from the six trained Season 0 models | models |
| Hydra Configs | Settings files, overrides, sweeps, studies and the experiment stages | configs |
| Evaluation | Fair games, validation and test maps, scenarios, saved results | evaluate |
| Metrics | What is measured and how to read it | metrics |
| Tournaments | Round robins, Elo ratings and tiers | round-robin |
| Tournament Rules | The leaderboard rulebook | leaderboard-rules |
| Replays | Save games and watch them in the analytical Replay Viewer | record |
| BattleClient | Play against or beside your own agents in the browser | battleclient |
| LLM Agents | Run a language model as a team | llm |
| Performance | Speed and memory, and how to measure them on your machine | performance |
| Method Sources | Where the built-in learners come from | method-sources |
| API Reference | Every public name on one page | api |
Reproduce The Paper
The paper's commands, frozen settings and result tables are in
scripts/paper_experiments/.
See Also
JAX-native environments:
- JaxMARL: Multi-Agent RL Environments and Algorithms in JAX.
- Gymnax: classic RL environments in JAX.
- Jumanji: combinatorial and planning environments in JAX.
- Pgx: board games in JAX.
- Brax: physics simulation for robot learning in JAX.
- XLand-MiniGrid: open-ended grid worlds for meta-RL in JAX.
- Craftax: an open-ended survival game in JAX.
JAX-native algorithms:
- Mava: JAX implementations of popular MARL algorithms.
- PureJaxRL: JAX implementation of PPO, and demonstration of end-to-end JAX-based RL training.
Citation
If you use MARL-BattleGrounds in your work, please cite the paper:
@article{hawili2026marlbattlegrounds,
title = {{MARL-BattleGrounds}: A Novel {JAX}-native Benchmark for Heterogeneous and Competitive Multi-agent {RL}},
author = {Hawili, Ulixes Tariq and Ramamoorthy, Subramanian and Altmann, Yoann and Carlucho, Ignacio},
journal = {arXiv preprint},
month = oct,
year = {2026}
}
Contribute
Pull requests are welcome. Start with CONTRIBUTING.md, and open an issue for bugs, ideas, questions and leaderboard submissions.
License
Licensed under the Apache License, Version 2.0.
Metadata
Release files for marl-battlegrounds 1.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| marl_battlegrounds-1.0.1.tar.gz | 4.5 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| marl_battlegrounds-1.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 9.4 MB
Release files / marl_battlegrounds-1.0.1.tar.gz
| Download URL | marl_battlegrounds-1.0.1.tar.gz |
|---|---|
| Size | 4.5 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e2511113234986ade37c2822a3e901afe911b12efeeb0c2b7fb98cb95e40c6c1
|
|
BLAKE2b-256 checksum How to use checksums |
895d0daed71816337fa18d9d746af1109ff94d15f6ab98f85f622b553dd65448
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.
Transparency logRelease files / marl_battlegrounds-1.0.1-py3-none-any.whl
| Download URL | marl_battlegrounds-1.0.1-py3-none-any.whl |
|---|---|
| Size | 4.9 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
12e2c08ef1dcaead0869a224ad6aa0f08f87a8f1e710c7446afff5f3d8b03d4a
|
|
BLAKE2b-256 checksum How to use checksums |
6dd6215e618898d23a73d5d50902ba84175a5c28989b85390bdd7e7174bcc6a1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.
Transparency log