Skip to main content

unilab-rl

PyPI CI License

English | 简体中文

Reinforcement learning algorithms and asynchronous runtimes extracted from UniLab, packaged as a standalone, simulator-agnostic library.

Relationship with UniLab

uni_rl is the RL algorithm and async-runtime layer of the UniLab project, split out into its own package. UniLab remains the consumer side: it owns the physics backends, task suites, and training entrypoints, and injects environments into uni_rl through uni_rl.env_contract.EnvFactory. uni_rl never imports unilab / unisim and never constructs environments itself, so any vectorized environment satisfying the contract — including simulators outside UniLab — can drive the algorithms in this package.

If you train with UniLab you already get uni_rl transitively. Install unilab-rl directly when you want to reuse its algorithms and async runtime with your own environment stack.

Naming note: the originally intended distribution name uni-rl is unregistrable on PyPI because it ultranormalizes to the existing unirl project. The distribution is therefore published as unilab-rl; the import namespace remains uni_rl as designed.

Contents

  • On-policy: PPO via rsl_rl (FinalObservationAwarePPO, RslRlVecEnvWrapper), HIM-PPO, and the HORA teacher-policy suite (incl. distillation trainer)
  • Async PPO (APPO): native collector/learner multiprocess implementation
  • Off-policy: FastSAC, FastTD3, and FlashSAC with double-buffer async runners
  • Runtime infrastructure: shared-memory rollout/replay buffers, replay pipelines, data-parallel gradient sync, memory budgeting, tensorboard/wandb training loggers, and a trace recorder

Layout

  • uni_rl.algos.* — the algorithm layer: on-policy (rsl_rl PPO wrappers, him_ppo, hora teacher/distillation suite), async on-policy (appo), off-policy learners (fast_sac, fast_td3, flash_sac), and shared algorithm helpers (common)
  • uni_rl.ipc — runtime infrastructure: async runner, shared-memory rollout/replay buffers, replay pipelines, DP gradient sync, memory budget
  • uni_rl.offpolicy — the generic off-policy double-buffer runner scaffolding
  • uni_rl.logging — tensorboard/wandb training loggers, trace recorder
  • uni_rl.utils — device, seed, nan-guard, observation helpers
  • uni_rl.env_contract — the injected env factory/protocol contract

Installation

pip install unilab-rl
# or, with uv:
uv add unilab-rl

Requires Python 3.10–3.13 and PyTorch ≥ 2.7.

Usage

uni_rl does not construct environments. Inject a picklable env factory (EnvFactory = Callable[[int, Mapping | None], EnvProtocol]) into the runner of your chosen algorithm:

from collections.abc import Mapping

from uni_rl.env_contract import EnvProtocol


def make_env(num_envs: int, cfg: Mapping | None) -> EnvProtocol:
    """Top-level factory (picklable by reference; no closures/lambdas)."""
    ...

The env contract is a minimal numpy-based, autoresetting vectorized-env protocol: dict observations keyed by observation group (obs_groups_spec), step() with final-observation semantics, and reset() returning (obs, info). See the module docstring in src/uni_rl/env_contract.py for the full contract, and the new algorithm recipe section in AGENTS.md for how to plug in a custom algorithm via runtime_resolver without forking.

Design contract

uni_rl does not depend on any simulator or environment library. Algorithm behavior is owned by the algo modules under uni_rl.algos.*; runtime infrastructure (ipc, logging, offpolicy, utils, env_contract) lives at the top level and never depends on the algorithm layer. See UniLab's training entrypoints for reference env integrations.

Development

make sync      # install dependencies (uv)
make test      # pytest
make format    # ruff check --fix + ruff format
uv run mypy src/uni_rl && uv run pyright   # type gates

Citation

If you use unilab-rl in your research, please cite the UniLab paper:

@article{jia2026unilab,
  title   = {UniLab: A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms},
  author  = {Yufei Jia and Zhanxiang Cao and Mingrui Yu and Heng Zhang and Shenyu Chen and Dixuan Jiang and Meng Li and Xiaofan Li and Yiyang Liu and Junzhe Wu and Zheng Li and XiLin Fang and Tingyu Cui and Shengcheng Fu and Haoyang Li and Anqi Wang and Zifan Wang and Dongjie Zhu and Chenyu Cao and Zhenbiao Huang and Ziang Zheng and Jie Lu and Xin Ma and Zhengyang Wei and Xiang Zhao and Tianyue Zhan and Ye He and Yuxiang Chen and Yizhou Jiang and Yue Li and Haizhou Ge and Yuhang Dong and Fan Jia and Ziheng Zhang and Meng Zhang and Xiwa Deng and Zhixing Chen and Hanyang Shao and Chenxin Dong and Yixuan Li and Yizhi Chen and Bokui Chen and Kaifeng Zhang and Hanqing Cui and Yusen Qin and Ruqi Huang and Lei Han and Tiancai Wang and Xiang Li and Yue Gao and Guyue Zhou},
  journal = {arXiv preprint arXiv:2605.30313},
  year    = {2026},
  url     = {https://arxiv.org/abs/2605.30313}
}

License

Apache-2.0, same as UniLab.

Release files for unilab-rl 1.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for unilab-rl 1.1.0
File Size Uploaded
unilab_rl-1.1.0.tar.gz 176.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for unilab-rl 1.1.0
File Interpreter ABI Platform
unilab_rl-1.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 398.8 kB

Release files / unilab_rl-1.1.0.tar.gz

Download URL unilab_rl-1.1.0.tar.gz
Size 176.7 kB
Tags Source
SHA-256 checksum
How to use checksums
1051ef15e0b809fd9a7197458987421fb11052b34ffa49a90a92f76fb5d38e0d
BLAKE2b-256 checksum
How to use checksums
22d293615e9770553ac4c3f67b27203c28509767f994484c2c1237ac1fdce731
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / unilab_rl-1.1.0-py3-none-any.whl

Download URL unilab_rl-1.1.0-py3-none-any.whl
Size 222.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c35e76b4a138a9196f835c5166c1f32c16df82df74a8fcdf356393be53ee5952
BLAKE2b-256 checksum
How to use checksums
0e944300a8b866d275226e78d2d4dbb92d82d857cb328e14ab02cf5bd20b1d90
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

1.4.0

2 release files

1.3.4

2 release files

1.3.3

2 release files

1.3.2

2 release files

1.3.1

2 release files

1.3.0

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.3

2 release files

1.1.1

2 release files

This release

1.1.0 This release

2 release files

1.0.0

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page