Skip to main content
env-BreakoutAtari2600-turbo-native logo

🕹️ Blazing-fast, deterministic Breakout for Reinforcement Learning 🕹️

env-BreakoutAtari2600-turbo-native is a Python library for reinforcement-learning researchers and engineers who need many reproducible Breakout games behind one Gymnasium vector-environment API. Add it to a uv project from PyPI, create BreakoutVecEnv, and step every lane with one NumPy action batch.

Fixed-point Rust physics owns game state and parallel stepping. Python exposes manual reset, policy-ready observations, native rendering, exact snapshots, and side-effect-free action branching.

Native Breakout gameplay rendered by env-BreakoutAtari2600-turbo-native

Install

Requires Python 3.11+ on Apple-silicon macOS 11+ or x86-64 Linux with glibc 2.28+.

Install uv, then add the library to your project:

uv add env-breakoutatari2600-turbo-native

Choose the corresponding requirement instead when you need an optional tool:

uv add "env-breakoutatari2600-turbo-native[play]"  # interactive Pygame player

To work from source, also install a Rust toolchain, then run:

git clone https://github.com/tsilva/env-BreakoutAtari2600-turbo-native.git
cd env-BreakoutAtari2600-turbo-native
uv sync --frozen --extra dev --extra play
make develop-release

Use

import gymnasium as gym
import numpy as np

env = gym.make_vec(
    "breakout_turbo_env:Breakout-Turbo-v0",
    game="Breakout-Atari2600-v0",
    num_envs=4096,
    num_threads=8,
)
obs, infos = env.reset()
obs, rewards, terminated, truncated, infos = env.step(
    np.zeros(env.num_envs, dtype=np.uint8)
)

done = terminated | truncated
if done.any():
    obs, reset_infos = env.reset(options={"reset_mask": done})

The module-qualified ID imports the package and registers the factory. This ID is vector-only and requires an explicit game; BreakoutVecEnv remains available for direct use.

Train with GradLab

Training recipes and implementations live in GradLab, keeping this repository focused on the environment. Run a published recipe from any directory without installing GradLab or cloning either repository.

For the default high-throughput PPO recipe:

uvx gradlab@0.1.1 train Breakout-Atari2600-v0/ppo

For PPO with learning-rate decay and KL-based update stopping:

uvx gradlab@0.1.1 train Breakout-Atari2600-v0/ppo-stable-updates

Breakout Turbo is ROM-free, so neither command needs a ROM path or registration. GradLab shows live progress, writes a playable final_model.zip below ./runs, and prints the matching version-pinned uvx ... play command when training finishes or is stopped safely. Local runs disable W&B and checkpoint evaluation by default, so they cannot establish acceptance or promotion. These are full-cap research recipes rather than short timed demos.

Turbo Vector API v2

BreakoutVecEnv implements the strict Turbo Vector API v2:

  • metadata["turbo_api_version"] is 2, metadata["transition_transport"] is "numpy", and metadata["render_modes"] advertises rgb_array.
  • Immutable capabilities and signal_schema declarations describe supported features and the dtype, shape, and reset/step availability of every signal.
  • buttons, action_mode, action_preset, action_table, action_meanings, and action_table_hash expose the resolved action semantics without provider-specific probing.
  • state_catalog is an immutable ordered tuple. Callers select reset states with an int32 state_indices array and inspect the read-only active indices with active_state_indices(); state sampling and lane routing remain caller-owned.
  • observation_ownership and observation_buffer_depth declare the exact lifetime of returned observations. Rendering is opt-in: with render_mode="rgb_array", render_lane(index) renders one lane, get_images() renders all lanes, and render() renders lane zero. With the default render_mode=None, the first two methods return None and get_images() returns one None entry per lane.

Interesting live positions can be archived without advancing the game and restored into any lane of the same environment:

capture_mask = np.zeros(env.num_envs, dtype=np.bool_)
capture_mask[0] = True
captured = env.capture_snapshots(capture_mask)

restore_mask = np.zeros(env.num_envs, dtype=np.bool_)
restore_mask[3] = True
starts = [None] * env.num_envs
starts[3] = captured[0]
obs, infos = env.reset(
    options={"reset_mask": restore_mask, "snapshots": starts},
)
env.close()

Importing the package also preserves the Stable Retro-compatible Breakout-Atari2600-v0 vector ID. The complete lifecycle, configuration, snapshot, and branching contract is in the environment documentation.

Stable-Baselines3 users can wrap the already-vectorized environment with the optional, explicitly auto-resetting adapter described in the environment documentation. SB3 remains a separate install and is not part of the core dependency set.

Commands

uv run --frozen --extra play breakout-turbo-env play       # open the player
uv run --frozen --extra play breakout-turbo-env play --uncapped
uv run --frozen breakout-turbo-env benchmark               # benchmark the policy path
uv run --frozen python scripts/compare_stable_retro.py     # run live differential checks
uv run --frozen ruff check .                               # lint Python
uv run --frozen pytest -m "not stable_retro"               # run regular Python tests
cargo test --locked --lib                                  # run Rust tests
make test-stable-retro                                     # require live cartridge parity
make test-semantic-oracle                                  # compare to original Stable Retro authority

Append --help to the player or benchmark command for its options.

Notes

  • Native actions are 0 noop, 1 FIRE, 2 right, and 3 left. The default policy observation is grayscale uint8, CHW, and shaped (num_envs, 4, 84, 84).
  • Rewards are score deltas using Atari row scoring. There is no life-loss or board-clear shaping. The cartridge presents two walls: the first refills after the next paddle return, the second ends at score 864 without another refill, and only losing all five lives terminates the episode.
  • Autoreset is disabled. Reset terminated lanes explicitly with a Boolean reset_mask; unselected lanes remain byte-exact.
  • noop_reset_max=N samples 1..N seeded raw-frame noops for each static reset, matching the conventional Atari reset distribution. FIRE is not issued automatically: use_fire_reset remains unavailable and the policy must start each serve.
  • The canonical Start state targets Stable Retro's native 160×210 Atari Breakout frame, lifecycle, physics, raster, rewards, collision behavior, and public trajectory values. In particular, ball_y uses the Atari RAM convention where zero means the serve is waiting for FIRE. Opt into raw frames with render_mode="rgb_array"; render() then returns lane zero's canonical Stella RGB frame while render_lane(index) selects any lane, separately from policy observations. Original Stable Retro's BGR-labeled RGB565 frame transport is normalized only at this human-facing boundary.
  • Live validation requires a separately obtained lawful ROM. The TurboBench semantic oracle pins original stable-retro==1.0.1 as the authority; the sibling stable-retro-turbo differential remains a secondary regression check. No ROM, save state, or recorded reference frame is distributed by this project.
  • Only Apple-silicon macOS and x86-64 Linux are supported. See support, benchmarking, and release validation for exact boundaries.
  • The project is a 0.x community preview. Public changes are recorded in the changelog. Serialized get_state() snapshots are portable only within the same package version and compatible configuration; live snapshot handles are session-local and intentionally not pickleable.

Architecture

env-BreakoutAtari2600-turbo-native architecture

License

MIT. See third-party notices for Atari, Stable Retro, ROM, and trademark boundaries.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

env_breakoutatari2600_turbo_native-0.5.7.tar.gz (10.3 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

env_breakoutatari2600_turbo_native-0.5.7-cp311-abi3-manylinux_2_28_x86_64.whl (331.8 kB view details)

Uploaded CPython 3.11+manylinux: glibc 2.28+ x86-64

env_breakoutatari2600_turbo_native-0.5.7-cp311-abi3-macosx_11_0_arm64.whl (293.3 kB view details)

Uploaded CPython 3.11+macOS 11.0+ ARM64

File details

Details for the file env_breakoutatari2600_turbo_native-0.5.7.tar.gz.

File metadata

File hashes

Hashes for env_breakoutatari2600_turbo_native-0.5.7.tar.gz
Algorithm Hash digest
SHA256 76fd11b2afc8bc5441b6975ed30575b4f6792fa8c1b3599b20873a9be9ced66e
MD5 f298b81b5b27a15276dda1fa6126ba04
BLAKE2b-256 df5281185f71731191c32307ac9a7f136764a3966b321a91d912a8068edadd7c

See more details on using hashes here.

Provenance

The following attestation bundles were made for env_breakoutatari2600_turbo_native-0.5.7.tar.gz:

Publisher: release.yml on tsilva/env-BreakoutAtari2600-turbo-native

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file env_breakoutatari2600_turbo_native-0.5.7-cp311-abi3-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for env_breakoutatari2600_turbo_native-0.5.7-cp311-abi3-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 9fe9a521ef37f43d6167d795df041bbfac0b7db1c067919061e7b019dbe600aa
MD5 544ddf03e9e0e28afd784857a2d2258d
BLAKE2b-256 bd3685ab680c9c2efc6b0316e3217872d22c91ec8971ad39aa17ba277067db66

See more details on using hashes here.

Provenance

The following attestation bundles were made for env_breakoutatari2600_turbo_native-0.5.7-cp311-abi3-manylinux_2_28_x86_64.whl:

Publisher: release.yml on tsilva/env-BreakoutAtari2600-turbo-native

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file env_breakoutatari2600_turbo_native-0.5.7-cp311-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for env_breakoutatari2600_turbo_native-0.5.7-cp311-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 1dee641f7cf2714d19298cfb0edd5f2aad3e096de1f2a0822241936a0a2fe8b8
MD5 10e57fb1b4f74c8a8bb2867b40a3a299
BLAKE2b-256 5760692172b8a7a936c3ddd9e213f5620d4090fad895ea2034e2264c57474601

See more details on using hashes here.

Provenance

The following attestation bundles were made for env_breakoutatari2600_turbo_native-0.5.7-cp311-abi3-macosx_11_0_arm64.whl:

Publisher: release.yml on tsilva/env-BreakoutAtari2600-turbo-native

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.5.12

3 files

0.5.11

3 files

0.5.10

3 files

0.5.9

3 files

0.5.8

3 files

This release

0.5.7 This release

3 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page