Skip to main content

marlbench

Setting up MARL environments can take a surprising amount of time. Each one has its own API, and using more than one environment in the same codebase usually means writing additional code to provide a common API.

marlbench does this work for you. It provides a common interface for several MARL environments, together with environment wrappers and vectorized environments.

This repository contains environment tools only. It does not include MARL algorithms.

How to use

Install it using uv pip install marlbench, or clone the repo.

Full examples are provided in CleanMARL

What is included?

Part Available components
Environments LBF, RWARE, SMAClite, PettingZoo, MaMuJoCo, MAgent2, SMAC, SMACv2
Observation wrappers Transform observations, normalize observations, add agent IDs
Reward wrappers Transform rewards, clip rewards, normalize rewards
Other wrappers Time limits
Vector environments Run multiple environments sequentially or in separate processes

Supported environments

Environment Interface Action space Installation
Level-Based Foraging LBFInterface Discrete pip install lbforaging
Multi-Robot Warehouse RWAREInterface Discrete pip install rware
SMAClite SMACliteInterface Discrete Install it from its GitHub repository
PettingZoo PettingZooInterface Discrete or continuous pip install pettingzoo and install the extra dependencies for the family you use
MaMuJoCo MAmujocoInterface Continuous pip install gymnasium-robotics
MAgent2 MAgent2Interface Discrete pip install magent2
SMAC SMACInterface Discrete Follow the instructions in the SMAC repository
SMACv2 SMACv2Interface Discrete Follow the instructions in the SMACv2 repository

The SMACv2 scenario configuration files are already included in marlbench/configs/smacv2/.

The supported MAgent2 environments are: adversarial_pursuit_v4,battle_v4, battlefield_v5, combined_arms_v6, gather_v5,tiger_deer_v4.

Common API

All interfaces provide the following methods:

obs, info = env.reset(seed=42)
obs, reward, terminated, truncated, info = env.step(actions)

state = env.get_state()
avail_actions = env.get_avail_actions()
agent_mask = env.get_agent_mask() # which agents are active

obs_size = env.get_obs_size()
state_size = env.get_state_size()
action_size = env.get_action_size()

actions = env.sample()
env.close()

Example

The following example runs one LBF episode:

from marlbench.interfaces.lbf import LBFInterface


env = LBFInterface(
    env_name="Foraging-2s-10x10-3p-3f-coop-v3",
    max_episode_steps=150,
    reward_aggr="sum",
    disable_env_checker=True,
)

obs, info = env.reset()
done = False

while not done:
    actions = env.sample()
    obs, reward, terminated, truncated, info = env.step(actions)

    state = env.get_state()
    avail_actions = env.get_avail_actions()
    agent_mask = env.get_agent_mask()

    done = terminated or truncated

env.close()

Other environments can be created in the same way:

from marlbench.interfaces.magent import MAgent2Interface
from marlbench.interfaces.mamujoco import MAmujocoInterface
from marlbench.interfaces.pz import PettingZooInterface
from marlbench.interfaces.rware import RWAREInterface
from marlbench.interfaces.smac import SMACInterface
from marlbench.interfaces.smaclite import SMACliteInterface
from marlbench.interfaces.smacv2 import SMACv2Interface


rware_env = RWAREInterface(
    env_name="rware-tiny-2ag-v2",
    reward_aggr="sum",
)

smaclite_env = SMACliteInterface(env_name="MMM2")

pettingzoo_env = PettingZooInterface(env_name="pursuit_v4",family="sisl")

mamujoco_env = MAmujocoInterface(env_name="Ant-p2x4")

magent_env = MAgent2Interface(env_name="adversarial_pursuit_v4")

smac_env = SMACInterface(env_name="3m")

smacv2_env = SMACv2Interface(env_name="terran_5_vs_5")

Wrappers

The following wrappers are available:

Wrapper Description
TimeLimit Truncates an episode after a fixed number of steps
TransformObservation Applies a function to observations and state, such as clipping
NormalizeObservation Normalizes observations and optionally the global state
NormalizeVecObservation NormalizeObservation for vectorized environments
AddAgentID Adds a one-hot agent ID to each observation
AddAgentIDVec AddAgentID for vectorized environments
TransformReward Applies a function to rewards, such as clipping
NormalizeReward Normalizes shared or individual rewards
NormalizeVecReward NormalizeReward or vectorized environments

Wrappers can be combined:

import numpy as np

from marlbench.interfaces.lbf import LBFInterface
from marlbench.wrappers.obs_wrappers import AddAgentID, NormalizeObservation
from marlbench.wrappers.reward_wrappers import NormalizeReward, TransformReward
from marlbench.wrappers.common import TimeLimit


env = LBFInterface(
    env_name="Foraging-2s-10x10-3p-3f-coop-v3",
    reward_aggr="none",
)

env = TimeLimit(env, max_episode_steps=150)
env = NormalizeObservation(env, normalize_state=True)
env = AddAgentID(env)
env = TransformReward(env, lambda reward: np.clip(reward, -1.0, 1.0))
env = NormalizeReward(env, gamma=0.99)

The order matters. Observations are normalized before the agent IDs are added, so the IDs remain zero or one. Rewards are clipped before they are normalized.

Vector environments

Two vector environment classes are available:

Class Description
SyncVectorEnv Runs multiple environments sequentially
SubprocVectorEnv Runs each environment in a separate process

Both classes receive a list of functions that create environments.

SyncVectorEnv

from marlbench.interfaces.lbf import LBFInterface
from marlbench.vec_envs.sync_vec import SyncVectorEnv
N_ENVS = 4
def make_lbf():
    return LBFInterface(
        env_name="Foraging-2s-10x10-3p-3f-coop-v3",
        max_episode_steps=150,
        reward_aggr="sum")
env_fns = [make_lbf for _ in range(N_ENVS)]
env = SyncVectorEnv(env_fns,auto_reset=False)
observations, infos = env.reset(seed=42)
observations, rewards, dones, truncated, infos = env.step(env.sample())
env.close()

SubprocVectorEnv

SubprocVectorEnv uses the same API:

from marlbench.vec_envs.subproc_vec import SubprocVectorEnv
N_ENVS = 4
def make_lbf():
    return LBFInterface(
        env_name="Foraging-2s-10x10-3p-3f-coop-v3",
        max_episode_steps=150,
        reward_aggr="sum")
if __name__ == "__main__":
    env_fns = [make_lbf for _ in range(N_ENVS)]
    env = SubprocVectorEnv(
        env_fns,
        start_method="spawn",
        auto_reset=False)
    observations, infos = env.reset(seed=42)
    observations, rewards, dones, truncated, infos = env.step(env.sample())
    env.close()

Every returned array has n_envs as its first dimension:

Value Shape
observations (n_envs, n_agents, obs_size)
rewards (n_envs,) shared, (n_envs, n_agents) individual
dones, truncated (n_envs,)
get_state() (n_envs, state_size)
get_avail_actions() (n_envs, n_agents, action_size)
get_agent_mask() (n_envs, n_agents)
get_env_mask() (n_envs,)
infos list of n_envs dicts

How to handle episodes with different lengths

Episodes do not end at the same time. With parallel environments running, one may terminate before the others. When auto_reset=False, an environment that returns done or truncated becomes inactive. It is not stepped again: it keeps returning its final observation with a reward of zero until you reset it. env.get_env_mask() tells you which environments are still running. You can use reset(indices=...) to reset a subset of environment. It accepts an integer, a list of integers, or a boolean mask of shape (n_envs,).

In constrast auto_reset=True automatically resets a finished environment inside step(), and the observation returned by step() then belongs to the new episode. The terminal data is moved into that environment's info dict. It can be accessed using final_obs,final_state, final_avail_actions, final_agent_mask,final_info. dones, truncated and rewards still describe the step that ended the episode, so episode statistics keep working.

Metadata

Release files for marlbench 0.0.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for marlbench 0.0.5
File Size Uploaded
marlbench-0.0.5.tar.gz 23.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for marlbench 0.0.5
File Interpreter ABI Platform
marlbench-0.0.5-py3-none-any.whl Python 3 none any Details

Total release size: 50.1 kB

Release files / marlbench-0.0.5.tar.gz

Download URL marlbench-0.0.5.tar.gz
Size 23.6 kB
Tags Source
SHA-256 checksum
How to use checksums
1b33de73fdd10f15dbaad93a2253be741eab03b3ae6aafabd2eee9fe11c4ad71
BLAKE2b-256 checksum
How to use checksums
1b0f9f4e2572a67ba6c1b8821039a220b97482128d04c2c3ca488c9195eb8c85
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.9.5

Release files / marlbench-0.0.5-py3-none-any.whl

Download URL marlbench-0.0.5-py3-none-any.whl
Size 26.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bd28fb8a5ede5ea9bd47b99731b42ce6b6f0935d4f76243b5975039478ab0ea8
BLAKE2b-256 checksum
How to use checksums
b6271a84dee34d85cf55980c333068e403e33cfe91604155b1f6c314dc210b9a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.9.5

Release history Release notifications | RSS feed

This release

0.0.5 This release

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page