Skip to main content

factor-qc

A fail-closed quality gate for backtests: one numpy-only engine covering Deflated Sharpe Ratio, Probability of Backtest Overfitting (CSCV), the Harvey-Liu multiple-testing haircut and Minimum Track Record Length 鈥?graded P0/P1/P2, and it refuses to judge a backtest that will not declare how many configurations were tried. Python 3.11+, one dependency (numpy), Windows / Linux / macOS.

Status: v0.1 鈥?alpha. The statistics are battle-tested inside a production research pipeline and validated against published reference values, but this standalone package is new: expect the CLI to shift before v1.0.

Why this exists

The standard story: you try 200 factor configurations, the best one shows a Sharpe of 1.65, you feel great. The honest story: with 200 trials of pure noise, someone is going to show a Sharpe of 1.65 鈥?the expected maximum of 200 zero-true-SR trials 鈥?and it will not be your skill, it will be your selection bias.

Most backtest tooling computes statistics and prints reports. factor-qc is a gate: it decides, with graded severity, whether a candidate may pass 鈥?and its default answer is no:

  • P0 (fatal) 鈥?DSR below threshold, PBO above threshold, haircut Sharpe below floor, MinTRL longer than the sample 鈫?the candidate must not pass.
  • P1 (warning) 鈥?weak PSR vs zero, short sample, aggressive trial count 鈫?proceed with eyes open.
  • P2 (info) 鈥?non-normal moments, tiny trial count 鈫?recorded, no action.

Philosophy

Honesty is the default; the gate is fail-closed.

The one non-negotiable input is n_trials: the honest count of configurations you tried. Without it there is no deflation benchmark (Bailey & L贸pez de Prado 2014), no haircut (Harvey, Liu & Zhu 2016, RFS), no track-record floor (Bailey & L贸pez de Prado 2018, JPM) and no overfitting probability (Bailey, Borwein, L贸pez de Prado & Zhu 2017, JCF). Refuse to declare, and the gate refuses to judge 鈥?that asymmetry is the point. qc check exits non-zero on any P0 failure, so it drops into CI, pre-commit hooks and research gates as a hard blocker, not a suggestion.

Two design commitments that keep it honest:

  1. numpy only, no scipy 鈥?the standard-normal inverse CDF is Acklam's rational approximation with one Newton refinement; every number in the report is reproducible from the code in this repo, no hidden black box.
  2. PBO is optional but explicit 鈥?without a trials matrix the gate says PBO was not computed, and the DSR cross-trial variance degrades to a conservative single-trial estimate. Absence of evidence is reported as absence, never as evidence.

Quick start

# install from PyPI (once published)
pip install factor-qc

# or run without installing anything:
#   PYTHONPATH=src python -m factor_qc --help

python examples/demo.py   # try it on reproducible synthetic cases

Your own backtest:

# returns.json = JSON list of per-period returns of the selected candidate
# trials.json   = JSON 2D matrix (T x N) of every configuration you tried

qc check --returns returns.json --trials trials.json --n-trials 200
# -> FAIL - P0 blocker(s): dsr: 0.63 vs 0.95; mintrl: 250.8 vs <= 1000; ...

qc check --returns returns.json --n-trials 5 --json   # machine-readable

Exit codes: 0 = no P0 failures (P1/P2 may still be failing), 1 = at least one P0 failure (or missing n_trials), 2 = usage error. Wire it into CI as a hard gate.

Commands

Command What it does
check Run the gate: DSR, PBO (when --trials given), haircut Sharpe, MinTRL as P0; PSR-vs-zero, sample length, trial aggression as P1; moments and trial count as P2. Human-readable or --json output
version Print version

Flags: --returns (required), --trials (optional), --n-trials (required unless require_declared_trials is disabled in code), --periods-per-year (default 252), --n-blocks (CSCV granularity, default 16).

The checks

Check Severity Method Reference
dsr P0 Deflated Sharpe Ratio: P(SR > E[max SR of N trials]) under non-normal moments Bailey & L贸pez de Prado (2014), JPM 40(5)
pbo P0 Probability of Backtest Overfitting via Combinatorially-Symmetric Cross-Validation (12,870 splits at n_blocks=16) Bailey, Borwein, L贸pez de Prado & Zhu (2017), JCF
haircut_sharpe P0 Multiple-testing haircut of the Sharpe ratio (Bonferroni/Holm/BHY) Harvey & Liu (2015)
mintrl P0 Minimum Track Record Length: observations needed before SR is significant Bailey & L贸pez de Prado (2018), JPM 44(5)
psr_vs_zero P1 Probabilistic Sharpe vs zero Bailey & L贸pez de Prado (2012)
sample_length P1 鈮?252 observations 鈥?
trial_aggression P1 n_trials 鈮?n_obs / 5 鈥?
return_moments P2 skew 鈮?0, kurtosis 鈮?3 鈥?
trial_count P2 n_trials 鈮?5 鈥?

The P0 set mirrors the spirit of Harvey, Liu & Zhu (2016), "鈥nd the Cross-Section of Expected Returns": a factor must survive multiple-testing correction to earn the right to be called a factor. The gate is the machine version of that editorial stance.

Performance note

CSCV enumerates C(n_blocks, n_blocks/2) splits 鈥?12,870 at the default 16. On large trial matrices (T=1000, N=200) that takes minutes; use --n-blocks 8 (70 splits) or 10 (252 splits) for interactive speed at slightly coarser granularity.

Development

python -m pip install -e . pytest
python -m pytest

CI runs the full test suite on Ubuntu, Windows and macOS with Python 3.11 and 3.12. Issues are handled on weekends; pull requests are welcome.

Related work

Project family

Part of Foolproof Labs — a toolchain against self-deception in quantitative research:

License

MIT

Metadata

Release files for factor-qc 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for factor-qc 0.1.1
File Size Uploaded
factor_qc-0.1.1.tar.gz 19.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for factor-qc 0.1.1
File Interpreter ABI Platform
factor_qc-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 35.9 kB

Release files / factor_qc-0.1.1.tar.gz

Download URL factor_qc-0.1.1.tar.gz
Size 19.8 kB
Tags Source
SHA-256 checksum
How to use checksums
e70503a851c70a76d95b13ad48f66a9ddf7c80a7ae8deffc95aad3405912e5a9
BLAKE2b-256 checksum
How to use checksums
02685def405193c8d5ce374b6bce1737e20cca90f915af6ee07d0db3332980e3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.15

Release files / factor_qc-0.1.1-py3-none-any.whl

Download URL factor_qc-0.1.1-py3-none-any.whl
Size 16.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
fdffd60eaba79869951229a5093e2d56a3c6becbf0fd9a35ca791a22522959fa
BLAKE2b-256 checksum
How to use checksums
264ef8de3c41e8bd5358d21a50c4adf8b6931f847d67d7e464ee03cc69b40d01
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.15

Release history Release notifications | RSS feed

0.1.3

2 release files

0.1.2

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page