Skip to main content

Perception-XAlpha Lite

Quantitative discovery is a multiple-comparisons problem disguised as an optimization problem. This finds fewer factors, on purpose.

A research framework that generates formulaic factors, backtests them on point-in-time data with real costs, and then tries to prove its own findings wrong before believing them. It is deliberately not a trading engine: no broker client, no order path, and CI asserts that mechanically on every commit.

pip install perception-xalpha-lite

Audit a backtest for overfitting

Give it daily returns for the variants you tried, and say how many you actually tried — including the ones you deleted.

import pandas as pd
from xalpha_lite.discovery import pbo, deflated_sharpe_ratio
from xalpha_lite.evidence import white_reality_check

returns = pd.read_csv("returns.csv", index_col=0, parse_dates=True)
sharpes = list(returns.mean() / returns.std(ddof=1))
best = (returns.mean() / returns.std(ddof=1)).idxmax()

print(pbo(returns))                                        # CSCV overfitting probability
print(deflated_sharpe_ratio(returns[best], sharpes, 250))  # against 250 declared trials
print(white_reality_check(returns))                        # family-wide null

On 24 variants of pure random noise, the best has an annualised Sharpe of 1.11 — a number most people would trade. At 24 trials, noise is expected to produce 1.18. PBO comes back 0.64, deflated Sharpe probability 0.46. The verdict is that selection is doing the work.

Four biases, measured on a real equity panel

Four measured biases

flaw reports survives unit
factors chosen with hindsight +2.00 −1.24 bps/day, same panel and cost
limit-locked legs priced as fillable +6.05 +0.38 % forward return of those legs
universe filtered on whole history 391 77 eligible names, first year
overlapping labels scored as independent −5.79 −2.25 t-statistic on pure noise

The first row is the one to sit with. Same data, same cost model, same construction — only the rule for choosing factors differs, and the gap is about 3 bps/day, larger than most published equity-factor results. A pipeline that cannot audit its own selection step cannot tell a discovery from an artifact of choosing.

Reproducible on synthetic data with no signal in it, in ten seconds:

python examples/selection_artifact.py    # IR 4.53 manufactured from pure noise

What is in the box

module what it does
pit disclosure-aware point-in-time alignment; a value appears only after max(notice_date, update_date), and rows without a disclosure date are rejected rather than imputed
dsl allowlisted causal expression language — no eval, no subprocess, no network
discovery bounded synthesis, neutral books, purged walk-forward, counterfactual and placebo controls, PBO and deflated Sharpe
universe point-in-time membership, and limit-locked sessions inferred from the bars themselves
book long-only top-N and dollar-neutral books sharing one cost engine
forward frozen specifications: no overwrite, digest verified on load, one entry per session, scoring only fully elapsed windows
evidence stationary bootstrap, White's Reality Check, Romano–Wolf step-down, BH/BY
decision Top-K pairwise weighting, block replicas, independent probability calibration

Command line: xalpha-lite, xalpha-evidence, xalpha-forward.

Current status, stated plainly

Nothing has graduated. Candidates are generated and fully evaluated; none has cleared the counterfactual, walk-forward and multiple-testing gates together. That is the gates working on a price-and-volume factor library, not the engine failing to run — and unlike most backtests, this one reports the exact count and the reason each candidate died.

Every number in the package and its examples comes from synthetic data it generates itself. It makes no profitability claim and never will.

Limitations

Examples are synthetic. Public financial endpoints may not preserve every restatement vintage. A contemporary security master creates survivorship bias unless replaced by genuine point-in-time membership. A zero-investment research portfolio is not executable in a long-only cash market. Equal overlapping tranches approximate a holding horizon and model neither queue priority nor market impact. Stationary-bootstrap inference assumes weak stationarity, and no resampling procedure repairs contaminated data or an incomplete trial ledger.

MIT. Research and educational use only. No investment advice.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

perception_xalpha_lite-0.5.0.tar.gz (60.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

perception_xalpha_lite-0.5.0-py3-none-any.whl (47.6 kB view details)

Uploaded Python 3

File details

Details for the file perception_xalpha_lite-0.5.0.tar.gz.

File metadata

  • Download URL: perception_xalpha_lite-0.5.0.tar.gz
  • Upload date:
  • Size: 60.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.12

File hashes

Hashes for perception_xalpha_lite-0.5.0.tar.gz
Algorithm Hash digest
SHA256 6d3559853a034f1bf349abb6c90dda71c07c765f8cc5ba38278b4ad07d53c5e1
MD5 7daaa65cad6449df0119345a1bbb4cfe
BLAKE2b-256 966ec3765591448f5ee11b7df884cd6b36278e79a2c9710d193d7f9da76ee220

See more details on using hashes here.

File details

Details for the file perception_xalpha_lite-0.5.0-py3-none-any.whl.

File metadata

File hashes

Hashes for perception_xalpha_lite-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 96573ec54eb469af3952fc762ec6a7ac50cda4941f5ac7e154d21a3eba3c87ea
MD5 2aef2e704f9023fc6e206da313ed48e4
BLAKE2b-256 66148279a9f456fc420ea4df203570000a9be8dbdcbd3c6e8e7a29fda464406b

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.5.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page