factor-qc
A fail-closed quality gate for backtests: one numpy-only engine covering
Deflated Sharpe Ratio, Probability of Backtest Overfitting (CSCV), the
Harvey-Liu multiple-testing haircut and Minimum Track Record Length —
graded P0/P1/P2, and it refuses to judge a backtest that will not declare
how many configurations were tried. Python 3.11+, one dependency
(numpy), Windows / Linux / macOS.
Status: v0.1 — alpha. The statistics are battle-tested inside a production research pipeline and validated against published reference values, but this standalone package is new: expect the CLI to shift before v1.0.
Why this exists
The standard story: you try 200 factor configurations, the best one shows a Sharpe of 1.65, you feel great. The honest story: with 200 trials of pure noise, someone is going to show a Sharpe of 1.65 — the expected maximum of 200 zero-true-SR trials — and it will not be your skill, it will be your selection bias.
Most backtest tooling computes statistics and prints reports. factor-qc
is a gate: it decides, with graded severity, whether a candidate may
pass — and its default answer is no:
- P0 (fatal) — DSR below threshold, PBO above threshold, haircut Sharpe below floor, MinTRL longer than the sample → the candidate must not pass.
- P1 (warning) — weak PSR vs zero, short sample, aggressive trial count → proceed with eyes open.
- P2 (info) — non-normal moments, tiny trial count → recorded, no action.
Philosophy
Honesty is the default; the gate is fail-closed.
The one non-negotiable input is n_trials: the honest count of
configurations you tried. Without it there is no deflation benchmark
(Bailey & López de Prado 2014),
no haircut (Harvey, Liu & Zhu 2016, RFS),
no track-record floor
(Bailey & López de Prado 2018, JPM)
and no overfitting probability
(Bailey, Borwein, López de Prado & Zhu 2017, JCF).
Refuse to declare, and the gate refuses to judge — that asymmetry is the
point. qc check exits non-zero on any P0 failure, so it drops into CI,
pre-commit hooks and research gates as a hard blocker, not a suggestion.
Two design commitments that keep it honest:
- numpy only, no scipy — the standard-normal inverse CDF is Acklam's rational approximation with one Newton refinement; every number in the report is reproducible from the code in this repo, no hidden black box.
- PBO is optional but explicit — without a trials matrix the gate says PBO was not computed, and the DSR cross-trial variance degrades to a conservative single-trial estimate. Absence of evidence is reported as absence, never as evidence.
Quick start
# install from PyPI (once published)
pip install factor-qc
# or run without installing anything:
# PYTHONPATH=src python -m factor_qc --help
python examples/demo.py # try it on reproducible synthetic cases
Your own backtest:
# returns.json = JSON list of per-period returns of the selected candidate
# trials.json = JSON 2D matrix (T x N) of every configuration you tried
qc check --returns returns.json --trials trials.json --n-trials 200
# -> FAIL - P0 blocker(s): dsr: 0.63 vs 0.95; mintrl: 250.8 vs <= 1000; ...
qc check --returns returns.json --n-trials 5 --json # machine-readable
Exit codes: 0 = no P0 failures (P1/P2 may still be failing), 1 = at
least one P0 failure (or missing n_trials), 2 = usage error. Wire it
into CI as a hard gate.
Commands
| Command | What it does |
|---|---|
check |
Run the gate: DSR, PBO (when --trials given), haircut Sharpe, MinTRL as P0; PSR-vs-zero, sample length, trial aggression as P1; moments and trial count as P2. Human-readable or --json output |
version |
Print version |
Flags: --returns (required), --trials (optional), --n-trials
(required unless require_declared_trials is disabled in code),
--periods-per-year (default 252), --n-blocks (CSCV granularity, default
16).
The checks
| Check | Severity | Method | Reference |
|---|---|---|---|
dsr |
P0 | Deflated Sharpe Ratio: P(SR > E[max SR of N trials]) under non-normal moments | Bailey & López de Prado (2014), JPM 40(5) |
pbo |
P0 | Probability of Backtest Overfitting via Combinatorially-Symmetric Cross-Validation (12,870 splits at n_blocks=16) | Bailey, Borwein, López de Prado & Zhu (2017), JCF |
haircut_sharpe |
P0 | Multiple-testing haircut of the Sharpe ratio (Bonferroni/Holm/BHY) | Harvey & Liu (2015) |
mintrl |
P0 | Minimum Track Record Length: observations needed before SR is significant | Bailey & López de Prado (2018), JPM 44(5) |
psr_vs_zero |
P1 | Probabilistic Sharpe vs zero | Bailey & López de Prado (2012) |
sample_length |
P1 | ≥ 252 observations | — |
trial_aggression |
P1 | n_trials ≤ n_obs / 5 | — |
return_moments |
P2 | skew ≈ 0, kurtosis ≈ 3 | — |
trial_count |
P2 | n_trials ≥ 5 | — |
The P0 set mirrors the spirit of Harvey, Liu & Zhu (2016), "…and the Cross-Section of Expected Returns": a factor must survive multiple-testing correction to earn the right to be called a factor. The gate is the machine version of that editorial stance.
Performance note
CSCV enumerates C(n_blocks, n_blocks/2) splits — 12,870 at the default 16.
On large trial matrices (T=1000, N=200) that takes minutes; use
--n-blocks 8 (70 splits) or 10 (252 splits) for interactive speed at
slightly coarser granularity.
Development
python -m pip install -e . pytest
python -m pytest
CI runs the full test suite on Ubuntu, Windows and macOS with Python 3.11 and 3.12. Issues are handled on weekends; pull requests are welcome.
Related work
- Bailey & López de Prado (2014), The Deflated Sharpe Ratio
- Bailey, Borwein, López de Prado & Zhu (2017), The Probability of Backtest Overfitting
- Harvey, Liu & Zhu (2016), …and the Cross-Section of Expected Returns (RFS)
- Harvey & Liu (2021), Lucky Factors (JFE)
- Mobarekeh & López de Prado (2024), Backtest Overfitting in the Machine Learning Era (SSRN 4778909) — why OOS methods still need honest trial accounting
License
MIT
Metadata
Release files for factor-qc 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| factor_qc-0.1.0.tar.gz | 19.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| factor_qc-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 35.0 kB
Release files / factor_qc-0.1.0.tar.gz
| Download URL | factor_qc-0.1.0.tar.gz |
|---|---|
| Size | 19.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ea9d3dabd1f1f09e96fb2c4857fc4b5265e9e8156636a268e2b673166b617f86
|
|
BLAKE2b-256 checksum How to use checksums |
3acb49ee8ea46132ffa9ed0df1a9da010d12b8b8cbe0afce90f4f72b1e3eed0a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.15
|
Release files / factor_qc-0.1.0-py3-none-any.whl
| Download URL | factor_qc-0.1.0-py3-none-any.whl |
|---|---|
| Size | 15.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9e8fd17440f57fa418b6a5d0033ef3d03c32f2d904ece8ec4020848a986f153c
|
|
BLAKE2b-256 checksum How to use checksums |
b4b676c8e27e18fbc69855454b5f0b3ede83878d84541975e6290b0fb9cad90a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.15
|