Skip to main content

EvoX Logo

🌟 EvoRL: A GPU-accelerated Framework for Evolutionary Reinforcement Learning 🌟

EvoRL Paper on arXiv

Table of Contents

Introduction

EvoRL is a fully GPU-accelerated framework for Evolutionary Reinforcement Learning, which is implemented by JAX and provides end-to-end GPU-accelerated training pipelines, including following processes:

  • Reinforcement Learning (RL)
  • Evolutionary Computation (EC)
  • Environment Simulation

EvoRL provides a highly efficient and user-friendly platform to develop and evaluate RL, EC and EvoRL algorithms.

Highlight

  • End-to-end training pipelines: The training pipelines for RL, EC and EvoRL are entirely executed on GPUs, eliminating dense communication between CPUs and GPUs in traditional implementations and fully utilizing the parallel computing capabilities of modern GPU architectures.
    • Most algorithms has a Workflow.step() function that is capable of jax.jit and jax.vmap(), supporting parallel training and JIT on full computation graph.
  • Easy integration between EC and RL: Due to modular design, EC components can be easily plug-and-play in workflows and cooperate with RL.
  • Implementation of EvoRL algorithms: Currently, we provide two popular paradigms in Evolutionary Reinforcement Learning: Evolution-guided Reinforcement Learning (ERL): ERL, CEM-RL; and Population-based AutoRL: PBT.
  • Unified Environment API: Support multiple GPU-accelerated RL environment packages (eg: Brax, gymnax, ...). Multiple Env Wrappers are also provided.
  • Object-oriented functional programming model: Classes define the static execution logic and their running states are stored externally.

Update

  • 2025-07-14: Our paper "EvoRL: A GPU-accelerated Framework for Evolutionary Reinforcement Learning" is accepted by ACM TELO.

  • 2025-04-01: Add support for Mujoco Playground Environments.

Documentation

  • For comprehensive guidance, please visit our Documentation, where you'll find detailed installation steps, tutorials, practical examples, and complete API references.

  • EvoRL is also indexed by DeepWiki, providing an AI assistant for beginners. Feel free to ask any question about this repo at https://deepwiki.com/EMI-Group/evorl.

Overview of Key Concepts in EvoRL

  • Workflow defines the training logic of algorithms.
  • Agent defines the behavior of a learning agent, and its optional loss functions.
  • Env provides a unified interface for different environments.
  • SampleBatch is a data structure for continuous trajectories or shuffled transition batch.
  • EC module provide EC components like Evolutionary Algorithms (EAs) and related operators.

Installation

EvoRL is developed on the top of jax. So jax should be installed first, please follow JAX official installation guide. Install the released package from PyPI (available after the first release):

pip install evorl-jax

The distribution name is evorl-jax; the Python import remains import evorl. For the latest development version and the training scripts/configs used below, install from source:

# Install the evorl package from source
git clone https://github.com/EMI-Group/evorl.git
cd evorl
pip install -e .

Aim is included for default experiment logging. WandB, SwanLab, Comet, and Neptune are optional; install their SDKs directly or use EvoRL extras. See Experiment Logging installation.

For developers, see Contributing to EvoRL

Quickstart

Training

EvoRL uses hydra to manage configs and run algorithms. Users can use scripts/train.py or scripts/train_dist.py to run algorithms from CLI.

# hierarchy of folder `configs/`
configs
├── agent
│   ├── ppo.yaml
│   ├── ...
...
├── config.yaml
├── env
│   ├── brax
│   │   ├── ant.yaml
│   │   ├── ...
│   ├── envpool
│   └── gymnax
└── logging.yaml

Specify the agent and env field based on the related config file path (*.yaml) in configs folder. For example: To train the PPO agent with config file in configs/agent/ppo.yaml on the Brax environment Ant with config file in configs/env/brax/ant.yaml, use:

python scripts/train.py agent=ppo env=brax/ant

# Parallel training two seeds on each GPU.
CUDA_VISIBLE_DEVICES=0,5 python scripts/train_dist.py -m hydra/launcher=joblib \
    agent=exp/ppo/brax/ant env=brax/ant seed=114,514

If multiple GPUs are detected, most algorithms will be automatically trained in distributed mode.

For more advanced usage, see our documentation: Training.

Logging

With the default Hydra configuration, a single run stores outputs in outputs/<script>/<timestamp>/, and multi-run mode (-m) uses multirun/<script>/<timestamp>/<overrides>/, where <script> is train or train_dist. LogRecorder writes <experiment-name>.log there; checkpoints use the checkpoints/ subdirectory when checkpoint.enable=true.

By default, the training scripts enable LogRecorder and AimRecorder (recorders: [log, aim]). Aim stores runs locally in the shared aim/.aim repository under the directory where training was launched. View and compare runs from that directory:

aim up --repo aim

To use WandB, install its optional extra (or run pip install wandb) and select it explicitly:

pip install -e ".[wandb]"
wandb login
python scripts/train.py agent=ppo env=brax/ant 'recorders=[log,wandb]'

The supported recorder names are log, aim, wandb, swanlab, comet, and neptune. Multiple installed backends can be selected together, for example 'recorders=[log,aim,wandb]'. See Logging for installation, grouping, and backend behavior.

Example dashboard when using the optional WandB recorder:

Env Rendering

We provide some example visualization scripts for brax and playground environments: visualize_mjx.ipynb.

Algorithms

Currently, EvoRL supports 4 types of algorithms

Type Algorithms
RL A2C, PPO, IMPALA, DQN, DDPG, TD3, SAC, TD7
EA OpenES, VanillaES, ARS, CMA-ES, algorithms from EvoX (PSO, NSGA-II, ...)
Evolution-guided RL ERL-GA, ERL-ES, ERL-EDA, CEMRL, CEMRL-OpenES
Population-based AutoRL PBT family (e.g: PBT-PPO, PBT-SAC, PBT-CSO-PPO)

RL Environments

By default, pip install evorl-jax will automatically install environments on brax. If you want to use other supported environments, please install the additional environment packages. We provide useful extras for different environments.

For example:

# ===== GPU-accelerated Environments =====
# Mujoco playground Envs:
pip install -e ".[mujoco-playground]"
# gymnax Envs:
pip install -e ".[gymnax]"
# Jumanji Envs:
pip install -e ".[jumanji]"
# JaxMARL Envs:
pip install -e ".[jaxmarl]"

# ===== CPU-based Environments =====
# EnvPool Envs:
pip install -e ".[envpool]"
# Gymnasium Envs:
pip install -e ".[gymnasium]"

Current Supported Environments

Environment Library Descriptions
Brax Robotic control
MuJoCo Playground Robotic control
gymnax (experimental) classic control, bsuite, MinAtar
JaxMARL (experimental) Multi-agent Envs
Jumanji (experimental) Game, Combinatorial optimization
EnvPool (experimental) High-performance CPU-based environments
Gymnasium (experimental) Standard CPU-based environments

PRs for other environment libraries are welcomed.

Performance

Test settings:

  • Hardware:
    • 2x Intel Xeon Gold 6132 (56 logical cores in total)
    • 128 GiB RAM
    • 1x Nvidia RTX 3090
  • Task: Swimmer

Bug report & Discussion

To keep our project organized, please use the appropriate GitHub section:

  • Issues – For reporting bugs and PR only. When submitting an issue, please provide clear details to help with troubleshooting.
  • Discussions – For general questions, feature requests, and other topics.

Before posting, kindly check existing issues and discussions to avoid duplicates. Thank you for your contributions!

Acknowledgement

Citing EvoRL

If you use EvoRL in your research and want to cite it in your work, please use:

@article{zheng2025evorl,
  author  = {Zheng, Bowen and Cheng, Ran and Tan, Kay Chen},
  doi     = {10.1145/3750053},
  journal = {ACM Trans. Evol. Learn. Optim.},
  month   = aug,
  title   = {EvoRL: A GPU-accelerated Framework for Evolutionary Reinforcement Learning},
  url     = {https://doi.org/10.1145/3750053},
  year    = {2025}
}

Metadata

Release files for evorl-jax 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for evorl-jax 0.1.0
File Size Uploaded
evorl_jax-0.1.0.tar.gz 220.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for evorl-jax 0.1.0
File Interpreter ABI Platform
evorl_jax-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 489.9 kB

Release files / evorl_jax-0.1.0.tar.gz

Download URL evorl_jax-0.1.0.tar.gz
Size 220.8 kB
Tags Source
SHA-256 checksum
How to use checksums
b5e8df03aa9ca137f8848dd3f1aa7e306c17366adc159c0d31041c31672f6482
BLAKE2b-256 checksum
How to use checksums
85bc38651c86c29033ee5423cc049074f2a5b7dd9bc18dc3d6ed78d1f24d4493
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release files / evorl_jax-0.1.0-py3-none-any.whl

Download URL evorl_jax-0.1.0-py3-none-any.whl
Size 269.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6c1e69b6129cc32fb209b58ed1b79a2820ac46e938cf6ebc30024bc4aaa2ca8c
BLAKE2b-256 checksum
How to use checksums
3f4c76029cba54a12f9abdd519119e6ae52f25d72f501a745efd1c6aa82ea65a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page