Skip to main content

Deterministic, fixed-point, high-throughput Breakout vector environment

Project description

breakout-turbo-env logo
🕹️ Blazing-fast, deterministic Breakout for Reinforcement Learning 🕹️

breakout-turbo-env is a Python environment for running many deterministic Breakout games at once. It is for reinforcement-learning researchers and engineers who need reproducible transitions, fixed observations, and a Gymnasium vector-environment API. Build the native extension, create BreakoutVecEnv, and step it from Python; the repository also includes a playable window, benchmark, and two small training paths.

Fixed-point Rust physics owns game state and parallel stepping. Python exposes the Gymnasium lifecycle, rendering, snapshots, and branching helpers.

Install

Requirements: Python 3.11+, uv, and a Rust toolchain.

git clone https://github.com/tsilva/breakout-turbo-env.git
cd breakout-turbo-env
uv sync --extra dev
uv run maturin develop --release

Use BreakoutVecEnv from the repository root or an environment where the extension has been installed.

import numpy as np
from breakout_turbo_env import BreakoutVecEnv

env = BreakoutVecEnv(num_envs=4096, num_threads=8)
obs, infos = env.reset()
obs, rewards, terminated, truncated, infos = env.step(
    np.zeros(env.num_envs, dtype=np.uint8)
)

done = terminated | truncated
if done.any():
    obs, reset_infos = env.reset(options={"reset_mask": done})

Commands

uv run breakout-turbo-env play                 # open the interactive player
uv run breakout-turbo-env benchmark            # measure the fixed 16-lane policy path
uv run pytest                                  # run Python contract and regression tests
cargo test --lib                               # run Rust library tests
uv run python train.py jerk                    # train a deterministic JERK action tape
uv run python train.py ppo                     # train a PPO policy
uv run python play.py jerk                     # replay the newest JERK policy
uv run python play.py ppo                      # replay the newest PPO policy
make release                                   # validate, tag, and publish a release

For player, benchmark, training, and replay options, append --help to the corresponding command.

Notes

  • The standard observation batch is grayscale uint8, CHW, and defaults to (num_envs, 4, 84, 84). Actions are 0 (noop), 1 (left), and 2 (right).
  • The environment is manual-reset only: after a terminal lane, call reset(options={"reset_mask": mask}) before stepping that lane again. Built-in layouts are full, checker, tunnel, and sparse.
  • render() returns the raw 96×96 RGB game frame; training observations use the processed grayscale stack. The interactive player accepts Left/Right or A/D, Space or R to reset, P to pause, and Escape to quit.
  • Training outputs live in runs/<algorithm>/<timestamp>/. JERK policies use policy.json; PPO policies use policy.npz.
  • make release requires a clean branch synchronized with its upstream. The release workflow builds and audits macOS arm64 and Linux x86_64 wheels before publishing to PyPI.

Architecture

breakout-turbo-env architecture

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

breakout_turbo_env-0.2.0-cp311-abi3-manylinux_2_28_x86_64.whl (288.7 kB view details)

Uploaded CPython 3.11+manylinux: glibc 2.28+ x86-64

breakout_turbo_env-0.2.0-cp311-abi3-macosx_11_0_arm64.whl (258.1 kB view details)

Uploaded CPython 3.11+macOS 11.0+ ARM64

File details

Details for the file breakout_turbo_env-0.2.0-cp311-abi3-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for breakout_turbo_env-0.2.0-cp311-abi3-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 56dd173e9ffa0a9c919fdec6ba2eaf2eb0afba39eab755686ddbf6c587e8df0a
MD5 8166b1cb00d76fe08ed18d563fefbd12
BLAKE2b-256 01018eb9533d62ae7ea059587d0629e535de2908b8cd793f1f19a8a1ac3a35e3

See more details on using hashes here.

Provenance

The following attestation bundles were made for breakout_turbo_env-0.2.0-cp311-abi3-manylinux_2_28_x86_64.whl:

Publisher: release.yml on tsilva/breakout-turbo-env

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file breakout_turbo_env-0.2.0-cp311-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for breakout_turbo_env-0.2.0-cp311-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 e17e95eab55e2ee0d32287d2a6cd4dc2aae4ecdaa86e63783bddc524c0fde520
MD5 59604149fae928242a1a91113f979b07
BLAKE2b-256 f7500c1d51c8602a1dcfef1e4f43cbdd69fa39b93481fd7cc0cc8b2ce97459c4

See more details on using hashes here.

Provenance

The following attestation bundles were made for breakout_turbo_env-0.2.0-cp311-abi3-macosx_11_0_arm64.whl:

Publisher: release.yml on tsilva/breakout-turbo-env

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page