Skip to main content

falsify

CI License: MIT Python dependencies

Gate a backtested trading edge before you trust it.

A green backtest is easy to produce and almost always overfits. falsify is a small, pure-stdlib toolkit (no numpy, no pandas, no install-time dependencies) that subjects a candidate strategy to the three ways a good-looking backtest most often fools you — and returns a blunt PASS / WEAK / FAIL.

The point isn't to make the backtest look good. It's to find out whether the edge is real before you risk anything on it.


The three gates

Gate What it catches Method
Multiple testing The best of N tried variants looks good by selection, not skill. Probabilistic & Deflated Sharpe Ratio (Bailey & López de Prado) — deflate by the expected maximum under the null.
Cost floor A gross edge that sits below round-trip cost × turnover is not an edge. Net-of-cost per-trade survival. Most "anomalies" live entirely beneath this line.
Autocorrelation A naive t-stat on overlapping / serially-correlated returns overstates significance. Newey-West (HAC) t-statistic with the automatic Bartlett lag rule.

PASS requires surviving both the multiple-testing haircut and the cost floor.


Install

The PyPI distribution name is falsify-edge; the import name stays falsify. The name falsify on PyPI belongs to an unrelated project, so pip install falsify installs a different package. Until the first falsify-edge release is out, install from source:

pip install git+https://github.com/RAJUSHANIGARAPU/falsify

or, from a clone (editable, for development):

git clone https://github.com/RAJUSHANIGARAPU/falsify
cd falsify
pip install -e .

Python ≥ 3.9. Runtime dependencies: none.

Not to be confused with falsify-quant, a separate AGPL-licensed CLI (numpy/scipy) that scores a strategy file 0–100 across seven checks; this project is an MIT, dependency-free library of three gates you call on return series you already have.


Quickstart

The recommended entry points take raw per-period returns and compute every statistic internally, so you can't mismatch units (see Design below):

from falsify import discovery_verdict_from_returns, newey_west_tstat

# best  = per-trade (or per-period) return series of your chosen strategy
# trials = one return series per variant you tried (best is one of them)
verdict = discovery_verdict_from_returns(
    best, trials,
    gross_edge_per_trade=0.0040,   # mean gross edge per trade
    round_trip_cost=0.0008,        # all-in round-trip cost
    turnover_per_year=12,
)
print(verdict["verdict"])          # "PASS" | "WEAK" | "FAIL"
print(verdict["deflated_sharpe"])  # multiple-testing-adjusted P(true Sharpe > 0)

# Always sanity-check significance under autocorrelation:
nw = newey_west_tstat(best)
print(nw["t_stat"], "vs naive", nw["iid_t"])   # HAC t can be much smaller

Run the worked example (no data needed):

python examples/synthetic_demo.py
candidate                verdict    DSR  net/trade    NW t
------------------------------------------------------------
clean low-noise edge     PASS     1.000    +0.327%  +20.40
noisy 'edge'             WEAK     0.928    +0.592%   +2.61

Both candidates have the same mean return per trade. The clean one passes. The noisy one clears the cost floor, but its deflated Sharpe (0.928) sits between the WEAK threshold (0.90) and the PASS threshold (0.95), so it is held at WEAK and not promoted. It is not outright FAILed; that needs a deflated Sharpe below 0.90 or a cost-floor failure.


Design — why "from returns"

The most common foot-gun in this kind of analysis is a unit mismatch: feeding an annualized Sharpe into a per-period test with n_obs set to a sample/event count silently saturates the deflated Sharpe to ~1.0 — a false PASS. The *_from_returns functions remove the choice entirely: you hand over raw return arrays and the library computes Sharpes, skew, and kurtosis itself, in one consistent frame. The lower-level functions remain available, but they raise on degenerate input (n_obs < 2, non-positive variance) instead of masking it.


API

Function Use
discovery_verdict_from_returns(best, trials, *, gross_edge_per_trade, round_trip_cost, turnover_per_year, …) Start here. Combined gate from raw returns.
deflated_sharpe_from_returns(best, trials) Multiple-testing-adjusted P(Sharpe > 0) from raw returns.
newey_west_tstat(returns, lags=None) HAC t-stat (and the naive iid_t for comparison).
discovery_verdict(best_sr, trial_sharpes, n_obs, …) Lower-level gate (per-period Sharpe inputs).
deflated_sharpe_ratio / probabilistic_sharpe_ratio / expected_max_sharpe Building blocks.
cost_floor_net / min_track_record_length Cost survival and the sample size needed to confirm an edge.

The discipline it encodes

The functions are only half of it; the workflow is the other half:

  1. Screen freely, test rarely. Generating ideas is free and raises no statistical haircut. Running a real test does — so ration tests to ideas with a plausible structural mechanism.
  2. Pre-register. Freeze the rule, the variant grid, the benchmark, and the kill-criteria before pulling data. No goal-post moving after seeing results.
  3. A clean kill is a success. The job is a trustworthy verdict, not a green one.

Field notes

This was extracted from a personal quantitative research platform where it was used in anger as the gate on every candidate strategy. Across ~30 candidates — forecasting, calendar/flow effects, cross-asset signals, carry, event-driven — none survived hostile out-of-sample testing versus simply holding the index after costs. The closest call (a pre-FOMC drift effect) was statistically real in-sample yet still failed out-of-sample. In the course of that work the harness also caught a real unit-mismatch bug in its own significance gate — which is why the current API is built to be misuse-proof. A tool that can falsify its own results is the only kind worth trusting.

Full case study of the platform it came from: https://rajushanigarapu.github.io/autonomous-investor-writeup/


Development

pip install -e ".[dev]"
pytest

Releasing

Publishing to PyPI is automated via GitHub Actions using PyPI Trusted Publishing (OIDC) — no API token is stored in the repo. Every push builds and twine checks the distribution in CI, so main is always release-ready.

To cut a release:

  1. One-time: on PyPI, add a pending Trusted Publisher for project falsify-edge pointing at this repo, workflow publish.yml, and environment pypi.
  2. Bump version in pyproject.toml, commit, and tag (git tag v0.1.1 && git push --tags).
  3. Publish a GitHub Release for that tag — the Publish to PyPI workflow builds and uploads automatically.

License

MIT — see LICENSE.

Metadata

Release files for falsify-edge 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for falsify-edge 0.1.0
File Size Uploaded
falsify_edge-0.1.0.tar.gz 14.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for falsify-edge 0.1.0
File Interpreter ABI Platform
falsify_edge-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 25.2 kB

Release files / falsify_edge-0.1.0.tar.gz

Download URL falsify_edge-0.1.0.tar.gz
Size 14.4 kB
Tags Source
SHA-256 checksum
How to use checksums
42619b5be168bf32db243e9f416cc5b4825d69c4a0c00a429c9044b37ace04b7
BLAKE2b-256 checksum
How to use checksums
63da51f386380fe588a8b05bd67b50f0d1a4d113d91a78b855db5c1478590f07
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / falsify_edge-0.1.0-py3-none-any.whl

Download URL falsify_edge-0.1.0-py3-none-any.whl
Size 10.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4a4b024d64baafa8e84d44b7a6e1d18a9bf8e0e1775ccafe39a5bc25d63d115a
BLAKE2b-256 checksum
How to use checksums
720e2d6901f0b5943c77326488716ad4f5d4bba950b258cbed91f5e87f3342c5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page