A simulation framework for stationary and non-stationary multi-armed bandit experiments.
Project description
BanditBench
A clean, object-oriented simulation framework for stationary and non-stationary multi-armed bandit experiments.
This framework was designed with a special focus on comparing different forgetting and exploration mechanisms, such as Sliding-Window UCB and Global Discounted UCB. It includes a variety of dynamic environments to test algorithmic adaptation speeds and robustness against changing reward landscapes.
Why BanditBench?
BanditBench is not intended to replace large general-purpose bandit libraries. Its goal is to provide a compact, readable, and reproducible framework for studying adaptation in non-stationary stochastic bandit settings.
The package focuses on:
- forgetting mechanisms in UCB-style algorithms
- sudden shifts, smooth drifts, and crossing reward landscapes
- pseudo-regret and adaptation delay metrics
- clean experiment scripts for reproducible comparisons
Features
- Modular Agents: Easy-to-extend
Agentbase class. Current implementations include UCB, SW-UCB, and Discounted UCB variants. - Dynamic Environments: Base
Environmentclass enforcing strict regret-tracking capabilities. Supports stationary distributions, piecewise sudden shifts, continuous Brownian drifts, and smooth crossing environments. - Reproducibility: Strict adherence to seeded NumPy Random Generators (
default_rng) for clean, deterministic experimentation across multiple runs. - Research-Ready Metrics: Track instantaneous pseudo-regret, optimal arm selection probabilities, and adaptation delays cleanly across any environment.
Installation
Since this module utilizes standard Python packaging, you can install it easily in your environment.
Option 1: Install from a local checkout
git clone https://github.com/AI-is-fun11/banditbench.git
cd banditbench
pip install -e .
Option 2: Install directly from GitHub after the repository is pushed
pip install git+https://github.com/AI-is-fun11/banditbench.git
Option 2b: Install from PyPI after publication
pip install banditbungee
Option 3: Install test tooling
pip install -e .[dev]
Quick Start Example
Below is a minimal example of how to import the module, spin up a stationary environment, and have an agent interact with it.
import numpy as np
from bandits.agents.ucb import UCB1
from bandits.environments.stationary import StationaryBernoulliEnv
# 1. Initialize a 3-arm stationary environment
env = StationaryBernoulliEnv(means=[0.2, 0.5, 0.8], horizon=1000)
env.reset(seed=42)
# 2. Initialize a standard UCB1 agent
agent = UCB1(n_arms=3, c=2.0)
cumulative_regret = 0.0
# 3. Run the interaction loop
for t in range(env.horizon):
# Agent selects an arm
chosen_arm = agent.select_arm()
# Environment yields a reward
reward = env.step(chosen_arm)
# Agent updates its internal statistics
agent.update(chosen_arm, reward)
# Track pseudo-regret (true best mean - chosen mean)
instant_regret = env.best_mean() - env.current_means()[chosen_arm]
cumulative_regret += instant_regret
print(f"Final Cumulative Pseudo-Regret: {cumulative_regret:.2f}")
Running Experiments
The repository also includes ready-to-run experiment scripts for comparing algorithms in non-stationary settings.
Piecewise-stationary benchmark
python -m bandits.experiments.piecewise_demo
This runs repeated simulations for DiscountedUCB and SlidingWindowUCB, then saves:
- figures to
figures/piecewise_demo/ - summary files to
results/piecewise_demo/
Crossing cosine benchmark
python -m bandits.experiments.cosine_demo
This generates a smooth two-arm crossing environment and saves:
- figures to
figures/cosine_demo/ - summary files to
results/cosine_demo/
Plotting
Plot helpers live in bandits/plots/ and operate on the summarized output returned by the experiment utilities.
Example:
from bandits.agents.sw_ucb import SlidingWindowUCB
from bandits.environments.stationary import StationaryBernoulliEnv
from bandits.experiments.multi_run import run_many
from bandits.metrics.summary import summarize_run
from bandits.plots.piecewise_plots import plot_cumulative_regret
raw = run_many(
agent_factory=lambda: SlidingWindowUCB(n_arms=2, window_size=20),
env_factory=lambda: StationaryBernoulliEnv(means=[0.4, 0.7], horizon=200),
seeds=[0, 1, 2, 3, 4],
)
summary = summarize_run(raw)
plot_cumulative_regret({"SlidingWindowUCB": summary}, out_dir="figures/example")
The main plotting functions include:
plot_cumulative_regretplot_instantaneous_regretplot_optimal_trackingplot_environmentplot_means_pathplot_change_metric_bars
Directory Structure
bandits/agents/: Bandit algorithm implementations. All agents inherit fromAgent.bandits/environments/: Testbeds for both stationary and non-stationary reward distributions. All environments inherit fromEnvironment.bandits/experiments/: Configurable scripts to run large-scale sweeps and comparisons.bandits/metrics/: Calculation of pseudo-regret, adaptation time, and probability of optimal arm selection.bandits/plots/: Visualization tools to compare agent performances seamlessly.tests/: Basic smoke tests.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file banditbungee-0.1.1.tar.gz.
File metadata
- Download URL: banditbungee-0.1.1.tar.gz
- Upload date:
- Size: 18.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ccd864d1f67ef4c7b98138416fdc4936cf01855bf724b2047eaecd9f46abf756
|
|
| MD5 |
12790973c276f30374cf5b33eaa696d8
|
|
| BLAKE2b-256 |
fc192a094059ba78b18b914168006e792407f511e3d581b3e5c7b2563a40ce00
|
Provenance
The following attestation bundles were made for banditbungee-0.1.1.tar.gz:
Publisher:
publish.yml on AI-is-fun11/BanditBench
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
banditbungee-0.1.1.tar.gz -
Subject digest:
ccd864d1f67ef4c7b98138416fdc4936cf01855bf724b2047eaecd9f46abf756 - Sigstore transparency entry: 2033334937
- Sigstore integration time:
-
Permalink:
AI-is-fun11/BanditBench@fe1d8d4166d6364b3e3236527b8e0cae0e0e7c20 -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/AI-is-fun11
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@fe1d8d4166d6364b3e3236527b8e0cae0e0e7c20 -
Trigger Event:
release
-
Statement type:
File details
Details for the file banditbungee-0.1.1-py3-none-any.whl.
File metadata
- Download URL: banditbungee-0.1.1-py3-none-any.whl
- Upload date:
- Size: 30.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8cf18ea0c832abda744eb5c26dc3c70af80a94216e8b1060d4683ec11743f53b
|
|
| MD5 |
efb4ca33a5be87c08f93978e229fc049
|
|
| BLAKE2b-256 |
c5d389469a6bfd9efcd11ba073e6c5ce9714b9d8ff34d5e80c0de8f78c8f169d
|
Provenance
The following attestation bundles were made for banditbungee-0.1.1-py3-none-any.whl:
Publisher:
publish.yml on AI-is-fun11/BanditBench
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
banditbungee-0.1.1-py3-none-any.whl -
Subject digest:
8cf18ea0c832abda744eb5c26dc3c70af80a94216e8b1060d4683ec11743f53b - Sigstore transparency entry: 2033335069
- Sigstore integration time:
-
Permalink:
AI-is-fun11/BanditBench@fe1d8d4166d6364b3e3236527b8e0cae0e0e7c20 -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/AI-is-fun11
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@fe1d8d4166d6364b3e3236527b8e0cae0e0e7c20 -
Trigger Event:
release
-
Statement type: