Modular RL (Functional)
A high-performance, researcher-centric Reinforcement Learning library for PyTorch built on the Functional Core, Imperative Shell design pattern. Think of it as RLax for PyTorch—engineered to eliminate the rigidity of deep OOP frameworks and the chaos of monolithic single-file copy-pasting.
💡 The Philosophy: Resolving the RL Paradigm War
Developing Reinforcement Learning algorithms typically forces researchers to choose between three flawed paradigms:
- The Monolith (Single-File Implementations):
- Pros: Everything is in one local scope; logging, file-diffing, and fast prototyping are frictionless.
- Cons: Combinatorial explosion of parameters and massive
if/elseconfiguration trees (e.g., handling frame stacking, continuous actions, recurrent states). Fixing a bug in one DQN implementation does not propagate to others, leading to destructive feature drift. Moving to multi-GPU (DDP) or TPU injections pollutes pure mathematical logic with system-level engineering blocks. Unit testing a loop monolith is practically impossible, forcing reliance on slow, flaky integration tests.
- Standard OOP & Strategy Patterns:
- Pros: Promotes abstraction and modular reuse.
- Cons: Hides design details and heavily relies on internal state mutations (
self._step_count,self._hidden_state), creating subtle off-by-one errors and tracking nightmares. To customize a single feature, you must master deep parent interfaces, inheriting a mountain of hidden knowledge debt. Mixing orthogonal features causes a combinatorial explosion of classes (e.g.,RecurrentContinuousPPO), or God Classes. Strict interfaces strip away critical paper-specific mathematical optimizations to stay general.
- Execution Graphs & DAGs:
- Pros: Maximizes component reuse and layout validation.
- Cons: Deep data-flow graphs aggressively reject standard Python dynamic control flows (like nested dynamic
whileloops inside an MCTS tree search), forcing clunky graph operators. Building a graph system creates immense engineering overhead, resulting in 10–15 node classes for a baseline algorithm, which damages execution speed in PyTorch due to structural dictionary-passing overhead.
🛠️ The Solution: Functional Core, Imperative Shell
This repository maps out a clean compromise. We isolate mathematical and algorithmic actions into a Functional Core composed of stateless, pure, side-effect-free functions. We then assemble these primitives inside an easy-to-read, linear, monolithic loop—the Imperative Shell. This allows our implementations to benefit from the monolithic paradigm, while allowing code reuse, modularity, and testing. Since each function is pure and has a simple interface, and (for the most part) functions don't use other functions internally, it is trivial to read and understand the codebase without getting lost in abstractions or having to dive deep into the codebase.
# --- 1. Initialization (Defining the State) ---
params = init_network()
optimizer_state = init_optimizer()
buffer_state = init_buffer(capacity=10000)
env_state, obs = env.reset()
hidden_state = init_rnn_state()
# --- 2. The Monolithic Loop (The Imperative Shell) ---
for step in range(MAX_STEPS):
# 1. Act (Pure function)
action, next_hidden_state = select_action(params, obs, hidden_state)
# 2. Step Env (Pure-ish function adapter)
next_env_state, next_obs, reward, done = env.step(env_state, action)
# 3. Add to Buffer (Pure state mutation)
transition = (obs, action, reward, hidden_state)
buffer_state = add_to_buffer(buffer_state, transition)
# Update loop states
obs, env_state, hidden_state = next_obs, next_env_state, next_hidden_state
# --- 3. The Functional Update Core ---
if step % UPDATE_FREQ == 0:
batch, rng_key = sample_buffer(buffer_state, rng_key, BATCH_SIZE)
# Calculate Loss & Gradients via Pure Math
loss, grads = calculate_loss(params, batch)
params, optimizer_state = apply_gradients(params, optimizer_state, grads)
# Monolithic layout makes logging and tracking effortless
wandb.log({"loss": loss, "step": step})
⚡ Core Engineering Practices (TODO: improve this section, add more of our actual rules in CONTRIBUTING.md)
-
Explicit Over Implicit We avoid magic configuration objects, automated parameter routing, and hidden global variables. High-level orchestration functions are stateless and clear. You pass tensors explicitly, enabling seamless tracking and debugging.
-
Fail Fast & Guardrails We utilize strong type hints, strict validation checks, and inline assertions at the boundaries of the functional layers. Shape mismatches, device conflicts, and invalid boundaries throw exceptions at their origin point—not deep inside a backend compiled gradient execution.
-
Documentation by Signature Functions are structured so that a researcher can easily interpret the underlying math just by inspecting the name, inputs, and type annotations, removing the need to trace internal source operations across 15 separate tracking scripts.
-
PyTorch Native Performance & Conventions torch.compile Friendly: Pure functions minimize internal state and dictionary unpacks, enabling the compiler to run graph optimizations across the mathematical core.
No CUDA Synchronizations: Absolutely zero internal calls to .item(), .tolist(), or .numpy() inside the computational blocks. This keeps the CPU and GPU timelines decoupled, eliminating latency bubbles.
Device & Dtype Agnostic: We never hardcode strings like device='cuda'. When initializing tensors, we use factory methods like .new_zeros() or torch.zeros_like() relative to the incoming tensor states to protect against distributed data-parallel breaks.
Explicit Dimensions: PyTorch broadcasting can mask bugs. We favor explicit .unsqueeze() calls, allowing code readers to visually audit array matching.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file atomic_rl-0.1.0.tar.gz.
File metadata
- Download URL: atomic_rl-0.1.0.tar.gz
- Upload date:
- Size: 245.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.12.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
44a821d0ba37dd41bdc22fc0d91ed4b1de55390856c74894496a2c663d35e92b
|
|
| MD5 |
510005471fe92639cc174a3d5ea11971
|
|
| BLAKE2b-256 |
67b4e26d16c34cd672a555c148533547a2e7753b827862557cb6f7130d569951
|
Provenance
The following attestation bundles were made for atomic_rl-0.1.0.tar.gz:
Publisher:
publish.yml on epicgamer17/atomic-rl
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
atomic_rl-0.1.0.tar.gz -
Subject digest:
44a821d0ba37dd41bdc22fc0d91ed4b1de55390856c74894496a2c663d35e92b - Sigstore transparency entry: 2373000394
- Sigstore integration time:
-
Permalink:
epicgamer17/atomic-rl@98a8ce35a4689edd058842f1a7f971c2f2bf26f4 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/epicgamer17
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@98a8ce35a4689edd058842f1a7f971c2f2bf26f4 -
Trigger Event:
push
-
Statement type:
File details
Details for the file atomic_rl-0.1.0-py3-none-any.whl.
File metadata
- Download URL: atomic_rl-0.1.0-py3-none-any.whl
- Upload date:
- Size: 141.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.12.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2b4b765b9bf141c2f255eb77ee976f3dac2800fb974917ab01823807ade7e644
|
|
| MD5 |
2a63dd63628f5ff0b47575139b5bc387
|
|
| BLAKE2b-256 |
2c47edeb6dd059377f99d972776497dc014993ef2a83652ab5442a25f77159b7
|
Provenance
The following attestation bundles were made for atomic_rl-0.1.0-py3-none-any.whl:
Publisher:
publish.yml on epicgamer17/atomic-rl
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
atomic_rl-0.1.0-py3-none-any.whl -
Subject digest:
2b4b765b9bf141c2f255eb77ee976f3dac2800fb974917ab01823807ade7e644 - Sigstore transparency entry: 2373000400
- Sigstore integration time:
-
Permalink:
epicgamer17/atomic-rl@98a8ce35a4689edd058842f1a7f971c2f2bf26f4 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/epicgamer17
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@98a8ce35a4689edd058842f1a7f971c2f2bf26f4 -
Trigger Event:
push
-
Statement type: