Measure the raw mathematical edge of a trading signal — capital-free, with confidence intervals and a reality check.
Project description
Measure the edge before you ever open a $100k account.
Most trading "edges" are artifacts of a small sample. crucible.edge takes a
trade log and tells you, with a confidence interval and a p-value, whether
the edge is real. No account, no position sizing, no equity curve. It's the thing
you run before a backtester.
pip install crucible # core: metrics + stats + simulator (numpy/pandas only)
pip install "crucible[examples]" # + yfinance, to run the demo below on real data
📖 Read the tutorial →, From Trade Log to Verdict: the statistics of a significant edge, every technique worked end to end on a Donchian breakout. (Source:
docs/tutorial.md, rendered with MkDocs Material.) Prefer offline? Download the PDF →
What crucible answers
From a single trade log up to the whole correlated book, one question at a time, each harder than the last, each answered out loud by a named function:
- ✅ Describe the edge: expectancy, profit factor, payoff, SQN, MFE/MAE efficiency →
edge_report - ✅ Quantify sampling noise: a bootstrap CI and p-value on every headline metric →
bootstrap_ci·p_value_positive - ✅ Rule out data-mining luck: sign-permutation p-value (Masters) + a random-entry reality check →
permutation·reality_check - ✅ Rule out drift: a leakage-controlled early-train / late-confirm split →
validation.holdout - ✅ Confirm out-of-sample: Pardo walk-forward with purge / embargo hygiene →
validation.walk_forward - ✅ Price the search itself: how much did selecting the best config overfit? Probability of Backtest Overfitting (CSCV) + deflated Sharpe →
validation.pbo - ✅ Account for correlation: the effective number of independent bets across a correlated book →
breadth.effective_n - ✅ …then all of it, as one gate: run every check above as an ordered, audited gauntlet (REAL → STRONG → DURABLE → GENERAL) that passes only if every gate does →
validation.run_gauntlet
30-second example
import yfinance as yf
from crucible.edge import barrier_trades, edge_report, reality_check
from crucible.strategies import ma_cross
px = yf.download("ES=F", start="2010-01-01") # any OHLC frame works
entries = ma_cross(px, fast=20, slow=50) # your signal: a boolean Series
trades = barrier_trades(px, entries, side="long", # signal -> trade log (in R)
tp=2.0, sl=1.0, timeout=20) # 2R target, 1R stop, 20-bar cap
print(edge_report(trades)) # the full capital-free scorecard
print(reality_check(trades)) # <-- the verdict
======================================================
EDGE REPORT (capital-free)
======================================================
Trades : 214
Win rate : 38.3 %
------------------------------------------------------
Expectancy : +0.081 R [PASS]
Profit factor : 1.34 [PASS]
Payoff ratio : 2.16 [INFO]
SQN-100 : 1.72 [INFO]
------------------------------------------------------
Excursion ratio : 1.28 [PASS]
======================================================
VERDICT (expectancy): +0.081 R 95% CI [-0.031, +0.196]
p(edge>0) = 0.071 -> FRAGILE
point positive, but the CI straddles zero, not distinguishable
from noise at this sample size. Do NOT size it up.
That FRAGILE block is the whole point: a positive expectancy that a backtester
would have shown you as a rising equity curve is, at this sample size,
indistinguishable from noise. crucible says so out loud.
Runnable versions live in
examples/:quickstart.py,validation.py, andbreadth.pyuse synthetic data (no network).real_data_yfinance.pypulls real prices from Yahoo Finance (pip install "crucible[examples]") and runs the full pipeline. Trypython examples/real_data_yfinance.py --ticker QQQ.
What's in the box
-
TradeLog: one documented schema (rin R-multiples, plus optionalmfe/mae/bars_held/prob). Everything speaks it. -
Edge metrics: expectancy, profit factor, payoff ratio, SQN, and the excursion family (MFE/MAE efficiency, E-ratio, time asymmetry, exit efficiency).
-
The honesty layer:
bootstrap_ci,p_value_positive,reality_check(HELD / FRAGILE / FAIL), andrandom_entry_null(did your signal beat coin-flip timing on the same prices?). -
Significance under serial dependence:
block_bootstrap_pvalue/block_bootstrap_ci: the i.i.d. trade bootstrap treats trades as exchangeable, which breaks the time-clustering of a pooled multi-asset book (correlated trades exit the same months). Feed an ordered period-return series (e.g. monthly R) and it resamples contiguous blocks (circular or stationary), so autocorrelation survives in the null, the honest p-value for a clustered book. -
Book-level breadth,
breadth.effective_n: how many independent bets a correlated set of return streams really holds (N_eff, the participation ratio of the correlation eigenvalues), the honest denominator for significance. Still capital-free: correlation structure only, no equity curve.>>> from crucible.breadth import effective_n >>> effective_n(returns).n_eff # 12-market book: 3 correlated blocs + a lone metal 3.8 # ...so it's really ~4 independent bets
See
examples/breadth.pyfor the full factor breakdown. -
A generic barrier simulator,
barrier_trades: OHLC + a boolean entry signal → aTradeLog. No instrument specifics. -
Example signals:
ma_cross,macd_cross. Demos, not endorsed edges.
Does the edge survive out of sample? crucible.validation
The pooled reality check tells you if an edge is real on the whole history.
crucible.validation asks the harder question: does it hold on data the analysis
never touched?
from crucible.validation import holdout, walk_forward, sign_permutation_pvalue
# 1. Early-train / late-confirm, leakage-controlled temporal split
print(holdout(trades, "2019-01-01", embargo_weeks=8)) # verdict is the LATE period
# 2. Sign-permutation p-value (Masters), could the edge come from noise?
print(sign_permutation_pvalue(trades))
# 3. Pardo walk-forward, optimize params in-sample, confirm out-of-sample, stitch
wf = walk_forward(px, ma_cross, param_grid={"fast": [10, 20], "slow": [50, 100]},
is_days=365 * 3, oos_days=365)
print(wf) # per-fold IS->OOS efficiency (WFE)
print(reality_check(wf.stitched)) # the stitched-OOS verdict, the honest one
Also here: sidak_correction and whites_reality_check (max-statistic across every
variant you searched) for when a grid search flatters the best result, plus spa_test,
Hansen's Superior Predictive Ability test, WRC's more powerful successor (studentized, and
it drops clearly-inferior variants so junk can't weaken the test. SPA p ≤ WRC p) shipped
alongside WRC, not replacing it (WRC is the conservative number, SPA the powerful one).
One step further, pbo_cscv + deflated_sharpe (crucible.validation.pbo), which ask
how much the act of selecting the best config overfit: the Probability of Backtest
Overfitting (Bailey/López de Prado CSCV) over a trial matrix, and the Sharpe deflated
for the number of trials and its own skew/kurtosis.
See examples/validation.py.
One verdict for the whole edge: crucible.validation.run_gauntlet
The individual tools above each answer one question. The gauntlet runs them as an ordered set of hard gates and returns a single audited pass/fail, the honest scorecard, capital-free:
from crucible.validation import run_gauntlet, SearchSpaceLog
log = SearchSpaceLog(scope="ES:ma_cross_grid") # record EVERY variant as you try it
for params in grid:
log.record(params, status="tried") # before it runs, so failures still count
...
gauntlet = run_gauntlet(
wf.stitched, # the honest log, stitched out-of-sample
prices=px, # enables REAL's random-timing null
wf=wf, # adds the DURABLE gate
n_variants=log, # the LEDGER -> REAL's Šidák correction (a bare int also works)
)
print(gauntlet.audit_report())
print(gauntlet.passed) # True only if every gate that ran passed
Hand n_variants the ledger, not a number. A count you type in is a
self-attestation, and the honest figure is the one you are least motivated to inflate —
it has to include the variants that errored or scored nothing, because those cost you a
look at the data too. SearchSpaceLog counts them; memory does not.
The gates, REAL (not noise, corrected for the search) → STRONG (real at the
CI lower bound) → DURABLE (holds out-of-sample) → GENERAL (travels across
markets), with two bring-your-own preambles (DECLARE, CLEAN) and a deliberate
handoff (SURVIVE: capital survivability is out of scope). Thresholds live in one
overridable Thresholds. Full write-up in docs/edge_gate.md.
A shareable tearsheet: crucible.report
pip install "crucible[report]"
from crucible.report import tearsheet
tearsheet(trades, "sheet.html", title="SPY 20/50 MA cross")
Writes a self-contained HTML page (plotly.js inlined, renders offline): the
verdict banner, the metric scorecard, the R-multiple distribution, cumulative R,
MFE/MAE excursion, and the bootstrap expectancy distribution behind the CI. Still
capital-free, it charts summed R, never an equity curve. See
examples/tearsheet.py.
Is the ML signal real? crucible.ml
The same honesty question, aimed at a model's scores instead of a trade log:
does a predicted probability actually rank outcomes, or is it noise, leakage, or a
redundant feature wearing a new name? Model-agnostic and capital-free, the core is
numpy/pandas (no sklearn/xgboost). Only the tearsheet needs the report extra.
from crucible.ml import (
information_coefficient, alpha_gate, # predictive power, + a PASS/FAIL gate
quantile_decay, decay_tearsheet, # does a higher score mean a better outcome?
fold_ic, redundancy_droplist, # out-of-fold IC; which features overlap
asof_window, # a point-in-time slice that can't peek ahead
)
# `preds`: a frame with a continuous `score` and the realized `label` (+1 win / -1 or 0 loss)
ic = information_coefficient(preds) # Spearman rank IC of score vs outcome
alpha_gate(ic, min_ic=0.03) # raises AlphaGateError below the bar
decay = quantile_decay(preds) # realized win rate per score quintile
print(decay.monotonic, decay.spread) # a real edge climbs Q1 -> Q5
rep = redundancy_droplist(panel, features, target="fwd")
print(rep.kept, rep.dropped) # keep the highest-|IC| of each redundant cluster
information_coefficient/alpha_gate: Spearman rank IC of a score against its realized label, plus a PASS/FAIL gate to stop an edge-less or leaking model early.quantile_decay/decay_tearsheet: realized win rate per score quantile (a genuine edge rises monotonically Q1→Q5). The tearsheet renders it as self-contained HTML.fold_ic/redundancy_droplist: per-feature out-of-fold IC, and a keep-highest-|IC| drop-list that groups features by |Spearman| / Cramér's V.asof_window/window_before: leakage-safe point-in-time slices, so a live feature is built identically to its training twin.
Still capital-free: it judges signals and features, never an equity curve.
What this is, and isn't
✅ Trade-level edge metrics, excursion efficiency, bootstrap CIs, a random-entry
reality check, all capital-free. Costs are in scope: test on r net of
your commission and slippage. The tutorial provides a simple per-trade cost
model for the netting, or supply your own.
❌ No capital, position sizing, CAGR, drawdown, or
Monte-Carlo-on-equity. If you want an equity curve, hand the TradeLog to
quantstats. crucible stops at the edge.
How this compares to RealTest
RealTest (Marsten Parker / MHP Trading) is a mature, paid, portfolio-level backtester. crucible is not a competitor to it. It's the step before it. If there's overlap worth being honest about, here it is.
Where RealTest is strictly broader. RealTest runs a real capital simulation:
position sizing, commissions, slippage, portfolio-level allocation, and the
equity-curve statistics that follow from it: CAGR/CAR, MaxDD, Sharpe, Sortino,
exposure. crucible has none of that on purpose. It stops at the trade log. If you
want an equity curve, hand the TradeLog to quantstats.
Where the methodologies genuinely differ.
-
Out-of-sample boundary handling. Both tools do OOS/holdout splits and walk-forward with re-optimization. The difference is leakage control at the boundary. crucible's
holdoutrequires a training trade to have entered and exited before the split (a trade whose window straddles the boundary can't leak into the fitted side) and drops anembargo_weeksband at the start of the test period.walk_forwardadds the Pardopurge_days/embargo_dayshygiene per fold. RealTest's documentation describes no equivalent purge/embargo. An open position simply carries across the boundary. -
Walk-forward efficiency. crucible reports a named Pardo WFE per fold (annualized OOS return / annualized IS return) and a mean across folds. RealTest gives you the out-of-sample equity to inspect but does not output a named efficiency ratio.
-
Monte Carlo. Both use bootstrap (random selection with replacement). They resample different things for different questions. RealTest resamples trades or daily account changes to draw percentile bands on the equity curve and drawdown ("how bad could the path get?"). crucible resamples the trade returns to put a confidence interval and a p-value on the edge metric itself (expectancy, profit factor, SQN), "is the edge distinguishable from zero?" crucible does not produce equity/drawdown bands. RealTest does not produce a CI on expectancy.
Reported statistics. Expectancy, profit factor, payoff (RR), and MFE/MAE exist in both. crucible additionally reports SQN natively (RealTest has no built-in SQN statistic. It would take a custom formula) and wraps every headline metric in a bootstrap CI + p-value, which RealTest does not.
The actual reason to reach for crucible: it's open-source (MIT), imports as a Python package, needs no account or capital model, and is free. It answers one narrow question (is this trade-level edge real, or a small-sample artifact?) and answers it out loud, before you ever open a funded account or a full backtester.
Feeding crucible from a RealTest export
The two tools compose cleanly. RealTest runs headless from the command line
(realtest -test script.rts) and, with SaveTradesAs set, writes the trade log
to CSV. RealTest's default trade export carries these columns:
Trade, Strategy, Symbol, Side, DateIn, TimeIn, QtyIn, PriceIn, DateOut, TimeOut,
QtyOut, PriceOut, Reason, Bars, PctGain, Profit, PctMFE, PctMAE, Fraction, Size, Dividends
(The Trades section of your script defines these, so a customized export can
rename or add columns. Treat the list above as the stock layout, not a fixed
schema.) Two things to know before mapping it:
- P&L is in percent and dollars, not R.
PctGainis the return as a percent of position size andProfitis net dollars. Neither is an R-multiple. There is no per-trade risk column in the export, so choosing the 1R denominator is a real modeling decision you own, not a column rename. - MFE/MAE are exported, as
PctMFE/PctMAE(percent runup / drawdown during the trade). Put them in the same R unit as the return.
If you define 1R as a fixed fractional risk R_PCT (e.g. you risked ~1% of the
position per trade), the conversion is one division each:
import pandas as pd
from crucible.edge import TradeLog, edge_report, reality_check
raw = pd.read_csv("trades.csv") # RealTest SaveTradesAs output
R_PCT = 1.0 # your 1R, as a % move (YOUR choice)
raw["r"] = raw["PctGain"].str.rstrip("%").astype(float) / R_PCT
raw["mfe"] = raw["PctMFE"].str.rstrip("%").astype(float) / R_PCT
raw["mae"] = raw["PctMAE"].str.rstrip("%").astype(float) / R_PCT
trades = TradeLog.from_frame(
raw, mapping={"DateIn": "entry_date", "DateOut": "exit_date", "Bars": "bars_held"},
)
print(edge_report(trades))
print(reality_check(trades)) # the verdict RealTest's MC doesn't give
Map your columns onto the TradeLog schema
(r required. mfe, mae, bars_held, prob, entry_date, exit_date
optional) with mapping=. Two honesty notes that matter more than the plumbing:
without a defined per-trade risk, the R denominator is an assumption the whole
report rides on. Pick it deliberately. And if the RealTest run is a
multi-strategy portfolio where several sub-strategies take the same trade, those
rows are duplicated and correlated. Crucible's bootstrap assumes independent
trades, so de-duplicate (or collapse to one book) first or the confidence
interval will look tighter than reality.
(Comparison verified against RealTest's published documentation as of July 2026. RealTest is actively developed. Check its docs for the current feature set.)
Ecosystem
The only input crucible needs is a trade log: the bundled barrier simulator, a
backtester export (see the RealTest section above), or any frame you map onto
TradeLog. If you also need futures/COT data to build a signal from, there is
an optional companion:
- cotdata: an optional data layer. A
local, file-based store for futures prices and CFTC Commitments-of-Traders
positioning: a common wrapper around Norgate, Databento, or yfinance behind
one read API, extensible to other data providers. Point it at a store,
import cotdata, and you have the OHLC frames and COT series you build a signal from. - crucible (this package): the edge layer. Turn that signal into a trade log and find out whether the edge is real.
The flow runs one direction: cotdata (data) → your signal → crucible
(edge). Neither imports the other. crucible works alone with any source of
trades, and together they take you from raw CFTC/price data to a capital-free
verdict with no vendor SDK at runtime.
Releasing
Releases publish to PyPI via GitHub Actions using Trusted Publishing (OIDC,
no API tokens are stored anywhere). Changes are tracked in
CHANGELOG.md.
One-time setup (maintainer, before the first publish):
-
Create a GitHub environment (repo Settings → Environments) named
pypi. (Add a required-reviewer rule on it for a manual approval gate, if you want one.) -
Register a Trusted Publisher from the project's own publishing settings, https://pypi.org/manage/project/crucible/settings/publishing/: owner
mspinola, repocrucible, workflowrelease.yml, environmentpypi.Use that page, not the pending publisher form under https://pypi.org/manage/account/publishing/. The pending form is only for reserving a name that does not exist yet, and it rejects
cruciblewith "this project already exists". The publisher registered against the oldcrucible-quantproject does not carry over either.There is no TestPyPI dry run.
crucibleon TestPyPI belongs to an unrelated project by another author, so nothing can be uploaded there under this name.release.ymlhas no TestPyPI path for that reason.
Cutting a release:
- Bump
versioninpyproject.tomland move theCHANGELOG.mdentry from Unreleased to the new version. - (Optional dry run) Either Actions → Release → Run workflow, which builds and
twine-checks without publishing, or the same thing locally with
python -m build && twine check dist/*. - Tag and push. The tag must match the
pyprojectversion or the run fails:git tag v0.2.0 git push origin v0.2.0 # builds, twine-checks, publishes to PyPI
Version floor. The
crucibleproject on PyPI already carries a 0.1.0 from an unrelated 2011 package, and PyPI never allows a version to be reused. Releases under this name therefore start at 0.2.0. Versions 0.1.0 and below are the historicalcrucible-quantline.
License
MIT
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file crucible-0.3.0.tar.gz.
File metadata
- Download URL: crucible-0.3.0.tar.gz
- Upload date:
- Size: 2.0 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
60526d17cbddcf766810df476830436e9d2131185f26211d42ecfa34503c7f3a
|
|
| MD5 |
00c6ffd5bff1ce11c9cf5df3bb0f7bea
|
|
| BLAKE2b-256 |
8242f5eb7bc65f5687b84679d5867eb0eb7c6bbc175ef100942216d3960a8d09
|
Provenance
The following attestation bundles were made for crucible-0.3.0.tar.gz:
Publisher:
release.yml on mspinola/crucible
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
crucible-0.3.0.tar.gz -
Subject digest:
60526d17cbddcf766810df476830436e9d2131185f26211d42ecfa34503c7f3a - Sigstore transparency entry: 2218595166
- Sigstore integration time:
-
Permalink:
mspinola/crucible@16e49251dcecbb006446a346e18e22333651e454 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/mspinola
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@16e49251dcecbb006446a346e18e22333651e454 -
Trigger Event:
push
-
Statement type:
File details
Details for the file crucible-0.3.0-py3-none-any.whl.
File metadata
- Download URL: crucible-0.3.0-py3-none-any.whl
- Upload date:
- Size: 94.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
beb92513747fd873cfb4a580306f85fe916ae271eee24eba87a5728e2d485940
|
|
| MD5 |
1e98afc9811ecfddae8110a5463f3bd9
|
|
| BLAKE2b-256 |
a16a4dd21d6478a205540d6ea55c0129c6c78bd47d81f83e8067407f4861d754
|
Provenance
The following attestation bundles were made for crucible-0.3.0-py3-none-any.whl:
Publisher:
release.yml on mspinola/crucible
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
crucible-0.3.0-py3-none-any.whl -
Subject digest:
beb92513747fd873cfb4a580306f85fe916ae271eee24eba87a5728e2d485940 - Sigstore transparency entry: 2218595212
- Sigstore integration time:
-
Permalink:
mspinola/crucible@16e49251dcecbb006446a346e18e22333651e454 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/mspinola
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@16e49251dcecbb006446a346e18e22333651e454 -
Trigger Event:
push
-
Statement type: