Overview
Aftab (Persian: آفتاب, meaning "sun" or "sun rays") is a benchmarking framework for evaluating CNN-based encoders in PQN across Atari games. It provides standardized training, evaluation, and reproducibility tools for deep reinforcement learning research.
We have compiled a few videos comparing PQN and Aftab agents. Watch them here.
Encoder Experiments
| IQM HNS |
|---|
|
|
| IQM HNS (Last 50M Frames) |
|
|
Hadamax Experiments
| IQM HNS |
|---|
|
|
| IQM HNS (Last 50M Frames) |
|
|
References:
Q-Values Experiments
| IQM HNS |
|---|
|
|
| IQM HNS (Last 50M Frames) |
|
|
References:
- Stop Regressing
- Deep Exploration via Bootstrapped DQN
- Improving Regression Performance with Distributional Losses
Procgen (Overfitting Prevention) Experiments
Since there are no public benchmarks comparing human-normalized-scores of Procgen environments, we created PNS (Procgen Normalized Score) that is a minimal min-max normalization of scores across seeds.
| IQM PNS |
|---|
|
|
| IQM PNS (Last 50M Frames) |
|
|
Installation
Install via pip:
pip install aftab
Alternatively, you can clone the repository and install in editable mode.
git clone https://github.com/tahashieenavaz/aftab.git aftab_source
pip install -e aftab_source
We highly recommend using Micromamba for creating virtual environments with instructions detailed here.
Training Agents
Currently JAX API is under development and is planned to be finished by the end of 2026. Contributions are highly encouraged.
from aftab import Aftab
from aftab import aftab_environments
seeds = [1, 2, 3, 4]
for environment in aftab_environments:
agent = Aftab(encoder="gamma", frames="pilot")
for seed in seeds:
agent.train(environment=environment, seed=seed)
agent.log()
Custom Encoder Injection
You can define your own encoder as a PyTorch module and pass it to the agent:
import torch
from aftab import Aftab
class CustomImageEncoder(torch.nn.Module):
pass
agent = Aftab(encoder=CustomImageEncoder)
Results
All experimental results are organized by experiment category. Each section contains:
- Tables: numerical results (HNS/PHS and raw scores)
- Charts: IQM normalized scores and training curves
Encoder Experiments
Tables
Charts
Hadamax Experiments
Tables
Charts
Q-Value Experiments
Tables
Charts
Procgen Experiments
Tables
Model Complexity
Base Variants
| Variant | Encoder Parameters | Regression Head Parameters | Total Parameters | Encoder FLOPs | Regression Head FLOPs | Total FLOPs |
|---|---|---|---|---|---|---|
| PQN | 78,304 | 1,686,500 | 1,764,804 | 7.734 | 1.610 | 9.347 |
| Alpha | 174,752 | 1,782,948 | 1,957,700 | 27.541 | 1.610 | 29.151 |
| Beta | 89,008 | 1,782,948 | 1,871,956 | 61.515 | 1.610 | 63.126 |
| Gamma | 117,168 | 1,725,364 | 1,842,532 | 22.901 | 1.610 | 24.512 |
| Delta | 78,552 | 1,850,588 | 1,929,140 | 6.143 | 1.774 | 7.917 |
| Epsilon | 80,112 | 2,179,828 | 2,259,940 | 13.252 | 2.101 | 15.354 |
| Zeta | 77,232 | 2,537,396 | 2,614,628 | 25.362 | 2.462 | 27.824 |
| Eta | 78,400 | 23,739,460 | 23,817,860 | 28.422 | 23.663 | 52.085 |
| Theta | 76,288 | 1,127,428 | 1,203,716 | 9.065 | 1.053 | 10.118 |
Note: The Eta variant has significantly more parameters than other variants, primarily due to the encoder producing a large number of features.
Hadamax Variants
| Variant | Encoder Parameters | Regression Head Parameters | Total Parameters | Encoder FLOPs | Regression Head FLOPs | Total FLOPs |
|---|---|---|---|---|---|---|
| Hadamax | 156,608 | 3,968,516 | 4,125,124 | 159.014 | 3.969 | 162.984 |
| Gamma-Hadamax-Valid | 234,336 | 1,609,220 | 1,843,556 | 122.001 | 1.610 | 123.611 |
| Gamma-Hadamax-Same | 234,336 | 3,280,388 | 3,514,724 | 129.300 | 3.281 | 132.581 |
Hyperparameters
| Hyperparameter | Value |
|---|---|
| Learning rate | $2.5 \times 10^{-4}$ |
| Training environments | 128 |
| Test environments | 8 |
| Optimizer | Rectified Adam |
| Weight decay | 0 |
| $\epsilon$ | $1 \times 10^{-5}$ |
| $\beta_{1}$ | 0.9 |
| $\beta_{2}$ | 0.999 |
| Total Frames | 200,000,000 |
| Loss function | Mean Squared Error |
| Scheduler | Linear Annealing |
| $\epsilon$-greedy exploration | 10% of total frames |
| Discount factor ($\gamma$) | 0.99 |
| GAE ($\lambda$) | 0.65 |
| Epochs | 2 |
| Batch size | 4096 |
Used in encoder and Hadamax experiments.
Statistical Significance
Encoder Experiments
| Wilcoxon Signed Rank Test | Wilcoxon Signed Rank Test (Corrected) |
|---|---|
|
|
|
| Probability of Improvement | |
|
|
|
Hadamax Experiments
| Wilcoxon Signed Rank Test | Wilcoxon Signed Rank Test (Corrected) |
|---|---|
|
|
|
| Probability of Improvement | |
|
|
|
Q-Value Experiments
| Wilcoxon Signed Rank Test | Wilcoxon Signed Rank Test (Corrected) |
|---|---|
|
|
|
| Probability of Improvement | |
|
|
|
Reproducibility
Due to the stochastic nature of deep reinforcement learning, exact reproducibility via fixed datasets is not feasible.
Instead, we provide a set of random seeds used in our experiments.
from aftab import aftab_seeds
print(aftab_seeds)
Full experiment replication:
from aftab import Aftab
from aftab import aftab_environments
from aftab import aftab_seeds
for environment in aftab_environments:
agent = Aftab()
for seed in aftab_seeds:
agent.train(environment=environment, seed=seed)
agent.log()
A comprehensive set of Atari environments is available via EnvPool:
https://envpool.readthedocs.io/en/latest/env/atari.html#available-tasks
Procgen environments use their native RGB observations with shape (3, 64, 64).
Aftab reads each task's EnvPool configuration and only applies supported options.
Atari-only options such as noop, frame_skip, frame_stack,
train_episodic_life, and EnvPool reward clipping are therefore not passed to
Procgen.
A comprehensive set of Procgen environments is available via EnvPool:
https://envpool.readthedocs.io/en/latest/env/procgen.html#available-tasks
Hardware
Nvidia A40 GPUs were used to run all the experiments in this experiment.
| Specification | Details |
|---|---|
| GPU Memory | 48 GB GDDR6 with error-correcting code (ECC) |
| GPU Memory Bandwidth | 696 GB/s |
| Interconnect | NVIDIA NVLink 112.5 GB/s (bidirectional); PCIe Gen4: 64 GB/s |
| NVLink | 2-way low profile (2-slot) |
| Display Ports | 3x DisplayPort 1.4* |
| Max Power Consumption | 300 W |
| Form Factor | 4.4" (H) x 10.5" (L), Dual Slot |
| Thermal | Passive |
| vGPU Software Support | NVIDIA Virtual PC, NVIDIA Virtual Applications, NVIDIA RTX Virtual Workstation, NVIDIA Virtual Compute Server, NVIDIA AI Enterprise |
| vGPU Profiles Supported | See the Virtual GPU Licensing Guide |
| NVENC / NVDEC | 1x / 2x (includes AV1 decode) |
| Secure Boot | Secure and Measured Boot with Hardware Root of Trust (optional) |
| NEBS Ready | Level 3 |
| Power Connector | 8-pin CPU |
Citation
@article{aftab2026drl,
title={Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks},
author={Shieenavaz, Taha and Zareshahraki, Shabnam and Nanni, Loris},
journal={arXiv preprint arXiv:YYMM.NNNNN},
year={2026}
}
Related Works
@misc{2407.04811,
Title = {Simplifying Deep Temporal Difference Learning},
Author = {Matteo Gallici and Mattie Fellows and Benjamin Ellis and Bartomeu Pou and Ivan Masmitja and Jakob Nicolaus Foerster and Mario Martin},
Year = {2024},
Eprint = {arXiv:2407.04811},
}
@misc{2403.03950,
Title = {Stop Regressing: Training Value Functions via Classification for Scalable Deep RL},
Author = {Jesse Farebrother and Jordi Orbay and Quan Vuong and Adrien Ali Taïga and Yevgen Chebotar and Ted Xiao and Alex Irpan and Sergey Levine and Pablo Samuel Castro and Aleksandra Faust and Aviral Kumar and Rishabh Agarwal},
Year = {2024},
Eprint = {arXiv:2403.03950},
}
@misc{1511.06581,
Title = {Dueling Network Architectures for Deep Reinforcement Learning},
Author = {Ziyu Wang and Tom Schaul and Matteo Hessel and Hado van Hasselt and Marc Lanctot and Nando de Freitas},
Year = {2015},
Eprint = {arXiv:1511.06581},
}
@misc{1806.04613,
Title = {Improving Regression Performance with Distributional Losses},
Author = {Ehsan Imani and Martha White},
Year = {2018},
Eprint = {arXiv:1806.04613},
}
@misc{1602.04621,
Title = {Deep Exploration via Bootstrapped DQN},
Author = {Ian Osband and Charles Blundell and Alexander Pritzel and Benjamin Van Roy},
Year = {2016},
Eprint = {arXiv:1602.04621},
}
Useful Links
- Wikipedia: Reinforcement Learning (RL)
- Wikipedia: Deep Reinforcement Learning (DRL)
- Wikipedia: Q-Learning
- Wikipedia: PyTorch
- Wikipedia: Statistical Hypothesis Test
- Wikipedia: Wilcoxon Signed-Rank Test
- PyTorch
License
© 2025 Taha Shieenavaz.
Licensed under CC BY-NC 4.0: https://creativecommons.org/licenses/by-nc/4.0/
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file aftab-1.0.3.tar.gz.
File metadata
- Download URL: aftab-1.0.3.tar.gz
- Upload date:
- Size: 59.5 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2f7ccfda1f166169c5395926bc445445bad7b2de8faf56b437e6c9c670abab69
|
|
| MD5 |
88ba84965b8f4803fe460f5bec41a897
|
|
| BLAKE2b-256 |
5b5ac6ab966d3b5a66c4210b20a80e7f225f4b9ada76147125bff01285c0a195
|
Provenance
The following attestation bundles were made for aftab-1.0.3.tar.gz:
Publisher:
publish.yaml on tahashieenavaz/aftab
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
aftab-1.0.3.tar.gz -
Subject digest:
2f7ccfda1f166169c5395926bc445445bad7b2de8faf56b437e6c9c670abab69 - Sigstore transparency entry: 2424152033
- Sigstore integration time:
-
Permalink:
tahashieenavaz/aftab@fa5fd1dbfa5088e5aa4ce033c049194500c4473a -
Branch / Tag:
refs/tags/v1.0.3 - Owner: https://github.com/tahashieenavaz
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yaml@fa5fd1dbfa5088e5aa4ce033c049194500c4473a -
Trigger Event:
push
-
Statement type:
File details
Details for the file aftab-1.0.3-py3-none-any.whl.
File metadata
- Download URL: aftab-1.0.3-py3-none-any.whl
- Upload date:
- Size: 67.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0c24444a96f2334cd4b736f96febe9fb50ab0d9e0f0e5a2058704731381fff72
|
|
| MD5 |
ceea1635cc92706f4c2ff04e4d76eb66
|
|
| BLAKE2b-256 |
cea2e53a89a05cef903809a75ad253c19ea32be10cde9ef3946f8867917ea4d1
|
Provenance
The following attestation bundles were made for aftab-1.0.3-py3-none-any.whl:
Publisher:
publish.yaml on tahashieenavaz/aftab
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
aftab-1.0.3-py3-none-any.whl -
Subject digest:
0c24444a96f2334cd4b736f96febe9fb50ab0d9e0f0e5a2058704731381fff72 - Sigstore transparency entry: 2424152092
- Sigstore integration time:
-
Permalink:
tahashieenavaz/aftab@fa5fd1dbfa5088e5aa4ce033c049194500c4473a -
Branch / Tag:
refs/tags/v1.0.3 - Owner: https://github.com/tahashieenavaz
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yaml@fa5fd1dbfa5088e5aa4ce033c049194500c4473a -
Trigger Event:
push
-
Statement type: