OR-TS — Odds-Ratio Thompson Sampling
Reference implementation of Odds-Ratio Thompson Sampling, a Thompson sampling policy for batched A/B tests and multi-armed bandits with binary outcomes whose memory is the joint posterior of the treatment contrasts (log odds ratios) rather than each arm's absolute event rate.
S. Kim (2026). Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits. Manuscript. S. Kim and K. Kim (2020). Odds-ratio Thompson sampling to control for time-varying effect. arXiv:2003.01905.
The idea in one paragraph
A platform's event rates move together: a promotion, a layout change, a holiday shifts every arm at once. A per-arm Beta-Bernoulli state remembers each arm's absolute rate and must unlearn all of them after every such shift. OR-TS fits, once per batch, an ordinary reference-coded logistic regression on the batch's counts, with a fresh flat-prior intercept for the batch's common level and the carried posterior of the contrasts as the prior. It then keeps only the contrasts and discards the level. Within a batch the two descriptions are the same thing in different coordinates; across batches they bet on different things staying fixed. On 86 real A/B test series the level moved about twenty-five times as much as the contrast, and in every one of them it moved more.
Left: three arms' observed event rates move almost in parallel because a common level dominates every curve. Right: the same batches in contrast coordinates. That is what OR-TS remembers.
Beta-TS : p_{i,t} = p_{i,t-1} every arm's rate is fixed
Full-TS : (alpha_t, beta_t) = (alpha_{t-1}, beta_{t-1}) same bet, logistic coordinates
OR-TS : beta_t = beta_{t-1}, alpha_t ~ flat only the contrasts are fixed
Five arms, a common shock of sd 0.30 redrawn every batch, mean of five
runs (examples/make_readme_figures.py). The policies that remember the
level keep chasing it; OR-TS never carried it.
Install
pip install -e . # from a clone; numpy and scipy are the only dependencies
pip install -e ".[dev]" # adds pytest
Quick start: Algorithm 1
import numpy as np
from orts import LogisticBandit
bandit = LogisticBandit() # no prior information
# boundary t: the platform hands over the batch's counts {arm: [exposures, events]}
bandit.update({"A": [30000, 300], "B": [30000, 330], "C": [30000, 290]})
# R1 fit the reference-coded logistic model with a fresh flat intercept
# R2 keep the marginal Gaussian of the contrasts (mu, S); discard the intercept
allocation = bandit.win_prop(draw=100_000, rng=np.random.default_rng(0))
# A1 draw contrast vectors, score the reference arm 0, find each draw's winner
# A2 winner shares are the next batch's allocation
# {'A': 0.19, 'B': 0.74, 'C': 0.07}
Repeat update then win_prop at every boundary. The state is the pair
(bandit.mu, bandit.sigma_inv) over bandit.action_list, kept in one
canonical order: arms in the order first seen, the reference arm last (the
first arm of the first batch, or LogisticBandit(reference="control")).
The entries are the contrasts of every arm against the reference, then the
level, which the next update replaces. Batches may name arms in any order
or subset; win_prop(arms) answers in the caller's order; contrasts()
reads the state as {arm: (mean, sd)}; set_reference and drop
re-base or fold the state without losing anything.
What the paper calls it, and where it is in the code
| paper | code |
|---|---|
| Algorithm 1, R1–R2 (fit, marginalize) | LogisticBandit.update(obs) |
| Algorithm 1, A1 (draws) / A2 (allocation) | contrast_draws() / win_prop() |
| Full-TS, the control with Beta-TS's memory | update(obs, odds_ratios_only=False) |
| Beta-TS, the per-arm baseline | TSPar |
| discounted Beta-TS, the matched forgetting baseline | DiscountedTSPar(discount) |
| decay λ (Section 2.3) | update(obs, decay=λ) |
| aggressiveness γ and floors (Section 5.2) | win_prop(aggressive=γ, floor=f) |
| changing arm sets, new reference (Supplement B) | get_par(arms), transform(arms), unseen arms in win_prop(arms) |
| warm start from a Beta-Bernoulli service (Supplement H) | LogisticBandit.from_beta_posteriors({arm: (a, b)}) |
| skipped batches: no events or no non-events (Supplement A) | update returns False and leaves the state |
| stopping and dropping quantities (Section 6.1) | win_prop() and expected_loss() |
| setting λ from measured drift (Supplement H) | implied_decay(excess_sd_beta) |
| diagnostics for the assumption (Sections 4.1, 5.3) | orts.diagnostics |
The two controls
Decay acts on what is carried. update(obs, decay=0.1) scales the
carried contrast precision by 1 - 0.1 before the fit; the effective memory
is roughly 1/decay batches, and the same number on DiscountedTSPar means
the same memory, since tempering a Beta density is count discounting. The
paper's registered simulations say when it pays: where the arm set is fixed
and the contrasts sit still, decay=0 is the setting the data support and
running decay anyway costs regret; where arms are inventory whose relative
appeal drifts, decay is the difference between trailing and leading.
Aggressiveness acts on how strongly the belief drives traffic.
win_prop(aggressive=2.0) raises the winner shares to a power and
renormalizes; floor=0.05 guarantees every arm a share afterwards. Neither
touches the posterior.
Diagnostics: is the assumption holding?
The state-separation assumption is that within a batch the arms share one
level and across batches the contrasts persist. orts.diagnostics computes,
from logged counts alone, what Section 5.3 of the paper says to watch:
from orts import diagnostics as dg
(alpha, var_alpha), contrasts = dg.batch_contrasts(obs_t, reference="A") # one batch
# collect alpha_t, var_alpha_t and contrasts["B"] over batches, then
R = dg.level_contrast_ratio(alphas, alpha_vars, betas, beta_vars) # >>1: level moves, contrast does not
tau = dg.excess_sd(betas, beta_vars) # the contrast's drift beyond noise
lam = bandit.implied_decay(tau) # the decay that drift implies
Plot each batch's contrast against dg.sampling_band(beta_vars); points that
wander outside the band with visible memory mean the contrasts are drifting.
dg.lag1_autocorrelation separates drift (positive) from a constant seen
through noise (near zero). Keep expected events per arm per batch above ten;
below one, the Gaussian state is materially wrong and TSPar is the safer
tool. The lever is the cycle length, not the method.
Stopping and dropping arms
The state computes what a default rule needs: win_prop() gives each arm's
posterior probability of being best, expected_loss() the expected loss of
committing to it now, in log-odds units. A workable default: drop an arm
whose probability stays below 1% for several consecutive batches; stop when
the leader's probability exceeds 95% and its expected loss is below what the
business will forgo. A threshold on absolute-rate posteriors moves when the
level moves; a threshold on the contrast posterior does not. See
examples/ab_testing.py, and remember that checking every batch is a
sequential test.
Migrating a running Beta-Bernoulli service
Same counters in, same probability-matching interface out; three things
change. Counts must be per cycle, not cumulative. The state is a fold over
the batch history, not a cache recomputable from totals, so persist it with
the id of the last batch absorbed. And the incumbent's Beta posteriors can
seed the contrast prior: LogisticBandit.from_beta_posteriors({arm: (a, b)})
inherits its contrast beliefs and, at the first update, discards its level
belief, which is the point.
Examples and tests
python examples/basic_usage.py # Algorithm 1 one boundary at a time
python examples/comparison.py # OR-TS vs Beta-TS vs Full-TS under a common shock
python examples/ab_testing.py # warm start, then the default stopping rule
python examples/make_readme_figures.py # the two README figures (needs matplotlib)
pytest -q
tests/test_paper_features.py checks the paper's claims that are code:
ranking invariance under a level shift, the memory rule, the properness
skip, decay and its Beta-side counterpart, aggressiveness and floors, the
reference transformation, the warm start, the diagnostics.
Layout
orts/ the package
logisticbandit.py LogisticBandit: OR-TS (default) and Full-TS
ts.py TSPar, DiscountedTSPar
diagnostics.py batch contrasts, excess variance, R, implied decay
utils.py the per-batch Laplace fit
examples/ runnable scripts, including the README figure generator
docs/ the README figures
tests/ pytest suite
archive/2020/ the 2020 preprint's synthetic runner and its outputs
logisticbandit.py, ts.py, utils.py deprecated import shims
The registered simulations, dataset analyses and manuscript of the 2026 paper live in a separate research repository; this package is the implementation they run.
Citing
@unpublished{kim2026orts,
author = {Kim, Sulgi},
title = {Odds-Ratio Thompson Sampling: A Specification and Design Guide
for Contrast-Based Multi-Armed Bandits},
year = {2026}
}
@article{kim2020orts,
author = {Kim, Sulgi and Kim, K.},
title = {Odds-ratio Thompson sampling to control for time-varying effect},
journal = {arXiv preprint arXiv:2003.01905},
year = {2020}
}
MIT License.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file orts-2.0.0.tar.gz.
File metadata
- Download URL: orts-2.0.0.tar.gz
- Upload date:
- Size: 28.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
aac884d4b79f790e720367419f734a52049fb23d7c72b32c9437f90a284f09c1
|
|
| MD5 |
b573e3d0613fff35eb2fb7f54397bbff
|
|
| BLAKE2b-256 |
7b4a2d9011fd79ca56e7bc950cfb0796092a947617ad1d83bad5d4503d724bf1
|
Provenance
The following attestation bundles were made for orts-2.0.0.tar.gz:
Publisher:
publish.yml on sulgik/orts
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
orts-2.0.0.tar.gz -
Subject digest:
aac884d4b79f790e720367419f734a52049fb23d7c72b32c9437f90a284f09c1 - Sigstore transparency entry: 2716839644
- Sigstore integration time:
-
Permalink:
sulgik/orts@7a896ccf8792947a1df65e1588c83c4ccb1eae15 -
Branch / Tag:
refs/tags/v2.0.0 - Owner: https://github.com/sulgik
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@7a896ccf8792947a1df65e1588c83c4ccb1eae15 -
Trigger Event:
push
-
Statement type:
File details
Details for the file orts-2.0.0-py3-none-any.whl.
File metadata
- Download URL: orts-2.0.0-py3-none-any.whl
- Upload date:
- Size: 19.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
52273b8c8173e691583d2094f1a989c40d9ccb4059025f61896860ee5eea7298
|
|
| MD5 |
4120118b9c1077a616d3a6ae028a8521
|
|
| BLAKE2b-256 |
0aae1ca75e619db78afe1bbbdabf032001fdfc062725cf553cbccd5a41ad04b9
|
Provenance
The following attestation bundles were made for orts-2.0.0-py3-none-any.whl:
Publisher:
publish.yml on sulgik/orts
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
orts-2.0.0-py3-none-any.whl -
Subject digest:
52273b8c8173e691583d2094f1a989c40d9ccb4059025f61896860ee5eea7298 - Sigstore transparency entry: 2716840648
- Sigstore integration time:
-
Permalink:
sulgik/orts@7a896ccf8792947a1df65e1588c83c4ccb1eae15 -
Branch / Tag:
refs/tags/v2.0.0 - Owner: https://github.com/sulgik
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@7a896ccf8792947a1df65e1588c83c4ccb1eae15 -
Trigger Event:
push
-
Statement type: