Skip to main content

MARL-BattleGrounds

The JAX-native Benchmark for Heterogeneous and Competitive Multi-agent Reinforcement Learning.

PyPI Python License Open In Colab Hugging Face Website

Installation | Quick Start | Baselines | Guides | See Also | Citation

Two teams fight a Team Deathmatch game in the analytical Replay Viewer

MARL-BattleGrounds (MARL-BGs) is an easy-to-use, JAX-native benchmark for heterogeneous and competitive team-vs-team multi-agent reinforcement learning (MARL). It is designed to be sample-efficient, computationally efficient and researcher-centric. Most MARL benchmarks are either fast but simple, like SMAX, where winning comes down to unit micromanagement, or rich but costly, like Dota 2, which took OpenAI Five ten months of training on large clusters. MARL-BGs fills the missing middle: rich team play, with fast and cheap experiments on one GPU.

  • The Game. In Team Deathmatch, two teams of up to five agents fight on 52 hand-designed maps. Each agent is a Mage, Warrior, Hunter, Rogue or Priest. Each class has its own Basic ability, passive and Ultimate, built to help allies and counter enemies, so a team wins by working together. Obstacles block sight and attacks, so each agent sees only part of the map. Maps 0-11 form a curriculum, maps 0-41 are for training, maps 42-46 for validation and maps 47-51 for testing.
  • The Evaluation. Eight hand-crafted endgame scenarios, each with a verified winning line, test skills such as agent modeling and multi-step planning. A monthly tournament retrains every entry on equal compute, plays it against the current "Big Nine" on the test maps, and rates it with draw-aware Elo and confidence intervals. Every entrant is released as a downloadable opponent. Up to 11,192 metrics per game, in 28 topics, show how a team won, such as how much the Priest healed each ally.
  • The Tools. Bring your own learner through plain JAX reset and step, or start from six tuned baselines. Their trained Season 0 models, and every save from 0 to 250 million steps, load by name from Hugging Face. Training helpers cover self-play, opponent pools, curricula and reward shaping. One call evaluates any mix of policies, so cross-play and zero-shot coordination tests come out of the box. MARL-BGs also ships three strong scripted opponents, an LLM API that turns a language model into a team, the BattleClient for playing alongside or against your policies, and the analytical Replay Viewer for stepping through any game tick by tick.

Try it in your browser: the Colab notebook walks through the API, the speed, training a team and the trained Season 0 teams.

Installation

Start here:

You want to Do this
Look around in your browser, with nothing to install Open the walkthrough in Colab: the API, the speed, training a team and the trained Season 0 teams
Play the game The one command below
Use MARL-BattleGrounds in your own project pip, below

You do not need to install JAX first: each route installs the right JAX for your machine.

On Linux, an Apple Silicon Mac or Windows with WSL2, one command installs everything and opens the BattleClient. It picks the right JAX build for your GPU:

curl -L https://github.com/arm-2-lab/MARL-BattleGrounds/archive/refs/tags/v1.0.1.tar.gz | tar xz
cd MARL-BattleGrounds-1.0.1 && sh start.sh

In your own Python project (Python 3.12 to 3.14; 3.14 recommended):

pip install "marl-battlegrounds[all]"           # CPU
pip install "marl-battlegrounds[all,cuda13]"    # NVIDIA GPU, driver 580 or newer

With Docker or Podman, on any system (add --gpus all and the 1.0.1-cuda12 tag for an NVIDIA GPU):

docker run --rm -it -p 127.0.0.1:8765:8766 -p 127.0.0.1:8767:8768 -v "${PWD}:/work" ghcr.io/arm-2-lab/marl-battlegrounds:1.0.1
Route For Tested
One command, sh start.sh Linux (x86_64 and ARM), Apple Silicon Mac, Windows with WSL2 Tested: Linux x86_64 with an NVIDIA RTX 5090; Linux x86_64 and ARM, and an Apple Silicon Mac, on the CPU. Not tested yet; it should work: WSL2
pip Linux, Apple Silicon Mac, WSL2; Python 3.12 to 3.14 Tested: Linux x86_64 and ARM, and an Apple Silicon Mac, each with Python 3.12, 3.13 and 3.14. Not tested yet; it should work: WSL2
Docker Any system with Docker or Podman; NVIDIA GPUs on Linux and Windows Tested: Linux with Docker (CPU, and an NVIDIA GPU with the CUDA 12 image), Linux with Podman, the ARM image on ARM Linux. Not tested yet; it should work: Docker Desktop on a Mac or Windows
Colab, Open in Colab A browser Not tested yet; it should work: Colab's free T4 GPU

"Tested" means we ran exactly these instructions on that system for this release.

Docker runs everything: play, replays, training, evaluation, tournaments, analysis and notebooks. Its limits:

Limit What to do
No GPU in Docker on a Mac Use the CPU image; Docker cannot reach Apple GPUs
study (big multi-run jobs) is not tested in containers and stops when the container stops Start the container with -d; not tested yet
You cannot add packages to a running container Build a two-line Dockerfile: FROM ghcr.io/arm-2-lab/marl-battlegrounds:1.0.1, then RUN uv pip install <package>
Docker Desktop (Mac, Windows) limits the container's memory Raise the memory limit in Docker Desktop's settings for large training runs

Every route, the GPU builds, Windows, clusters and offline use: the install guide.

Quick Start

make, reset and step use JaxMARL's call order (reset(key), step(key, state, actions)), and compile under jit, vmap and lax.scan.

import jax
import marl_battlegrounds as marl_bgs

env = marl_bgs.make("tdm", map_id=0, num_envs=4)  # four games at once
key = jax.random.key(0)
observation, state = env.reset(key)
for _ in range(50):
    key, act_key, step_key, reset_key = jax.random.split(key, 4)
    actions = env.sample_actions(act_key, state)  # legal random actions
    observation, state, reward, done, info = env.step(step_key, state, actions)
    observation, state = env.reset_done(reset_key, state)  # restart finished games

On one RTX 5090 with 1,024 parallel 5v5 games, the simulator alone runs 50,416 game steps per second (random legal actions, compilation excluded), and recurrent MAPPO trains at 29,878 game steps per second in pure self-play.

Train

One command trains recurrent IPPO with the Season 0 learner settings (250 million game steps) against itself, its past copies and the three scripted opponents, and keeps its best checkpoint on the validation maps. What Training Costs gives the time and memory on one GPU.

python -m marl_battlegrounds.experiments.train method=ippo pool=scripted \
  'validation_opponents=[tdm-alpha,tdm-beta,tdm-gamma]' output_dir=runs/my-ippo

Evaluate

import marl_battlegrounds as marl_bgs

results = marl_bgs.evaluate("random", "tdm-alpha", num_episodes=32, num_envs=32)
print(results.head_to_head())  # games, wins, draws, losses, point margin

Your method is always Team A, and paired games swap the spawn ends, so a map's layout cannot favor either side.

Watch And Play

python -m marl_battlegrounds replay path/to/game.marlbg-replay.json
import marl_battlegrounds as marl_bgs

marl_bgs.battle_client(team_b="tdm-alpha")  # play in your browser against a scripted team

Baselines

Method Name Memory Critic Reference Source
Recurrent MAPPO mappo Recurrent Centralized Yu et al., 2022 ppo.py
Feedforward MAPPO ff_mappo None Centralized Yu et al., 2022 ppo.py
Recurrent IPPO ippo Recurrent One per agent de Witt et al., 2020 ppo.py
Feedforward IPPO ff_ippo None One per agent de Witt et al., 2020 ppo.py
QMIX qmix Recurrent QMIX mixer Rashid et al., 2018 qmix.py
PQN-VDN pqn_vdn Recurrent Sum of agent values Gallici et al., 2024; Sunehag et al., 2017 pqn.py

Every save of the six Season 0 models, from 0 to 250 million steps, loads by name from Hugging Face, for example marl_bgs.load_method("MARL-BattleGrounds/tdm-season-0-recurrent-ippo@final") (Released Models).

Your own method can be anything that turns permitted observations into legal actions: a network, a planner, a language model or a mix.

Guides

Guide What It Covers Website
Getting Started Install, your first games, maps, rosters, what agents see start
Game Rules Classes, abilities, scoring, Red Zone, scenarios game
Your Method Plug in your own model or learner, with JAX's jit, vmap and scan your-method
Training The six learners, opponents, curriculum, rewards, checkpoints training
Released Models Load, play, train against or start from the six trained Season 0 models models
Hydra Configs Settings files, overrides, sweeps, studies and the experiment stages configs
Evaluation Fair games, validation and test maps, scenarios, saved results evaluate
Metrics What is measured and how to read it metrics
Tournaments Round robins, Elo ratings and tiers round-robin
Tournament Rules The leaderboard rulebook leaderboard-rules
Replays Save games and watch them in the analytical Replay Viewer record
BattleClient Play against or beside your own agents in the browser battleclient
LLM Agents Run a language model as a team llm
Performance Speed and memory, and how to measure them on your machine performance
Method Sources Where the built-in learners come from method-sources
API Reference Every public name on one page api

Reproduce The Paper

The paper's commands, frozen settings and result tables are in scripts/paper_experiments/.

See Also

JAX-native environments:

  • JaxMARL: Multi-Agent RL Environments and Algorithms in JAX.
  • Gymnax: classic RL environments in JAX.
  • Jumanji: combinatorial and planning environments in JAX.
  • Pgx: board games in JAX.
  • Brax: physics simulation for robot learning in JAX.
  • XLand-MiniGrid: open-ended grid worlds for meta-RL in JAX.
  • Craftax: an open-ended survival game in JAX.

JAX-native algorithms:

  • Mava: JAX implementations of popular MARL algorithms.
  • PureJaxRL: JAX implementation of PPO, and demonstration of end-to-end JAX-based RL training.

Citation

If you use MARL-BattleGrounds in your work, please cite the paper:

@article{hawili2026marlbattlegrounds,
  title   = {{MARL-BattleGrounds}: A Novel {JAX}-native Benchmark for Heterogeneous and Competitive Multi-agent {RL}},
  author  = {Hawili, Ulixes Tariq and Ramamoorthy, Subramanian and Altmann, Yoann and Carlucho, Ignacio},
  journal = {arXiv preprint},
  month   = oct,
  year    = {2026}
}

Contribute

Pull requests are welcome. Start with CONTRIBUTING.md, and open an issue for bugs, ideas, questions and leaderboard submissions.

License

Licensed under the Apache License, Version 2.0.

Metadata

Release files for marl-battlegrounds 1.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for marl-battlegrounds 1.0.1
File Size Uploaded
marl_battlegrounds-1.0.1.tar.gz 4.5 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for marl-battlegrounds 1.0.1
File Interpreter ABI Platform
marl_battlegrounds-1.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 9.4 MB

Release files / marl_battlegrounds-1.0.1.tar.gz

Download URL marl_battlegrounds-1.0.1.tar.gz
Size 4.5 MB
Tags Source
SHA-256 checksum
How to use checksums
e2511113234986ade37c2822a3e901afe911b12efeeb0c2b7fb98cb95e40c6c1
BLAKE2b-256 checksum
How to use checksums
895d0daed71816337fa18d9d746af1109ff94d15f6ab98f85f622b553dd65448
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release files / marl_battlegrounds-1.0.1-py3-none-any.whl

Download URL marl_battlegrounds-1.0.1-py3-none-any.whl
Size 4.9 MB
Tags Python 3
SHA-256 checksum
How to use checksums
12e2c08ef1dcaead0869a224ad6aa0f08f87a8f1e710c7446afff5f3d8b03d4a
BLAKE2b-256 checksum
How to use checksums
6dd6215e618898d23a73d5d50902ba84175a5c28989b85390bdd7e7174bcc6a1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.0.1 This release

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page