AnchorTest
Stop trusting a single backtest run.
Most backtests report one run's Sharpe ratio, on one start date, over one universe. That number is a sample
of one. AnchorTest is a small, dependency-light Python library (numpy + pandas, nothing else) that runs
three specific checks that have each caught a real, shipped bug in real trading-strategy research:
- Anchor-average, instead of trusting one lucky start date.
- Compare candidates correctly, with the drawdown sign-convention trap fixed once, in one place.
- Catch a blending/tranching artifact before it becomes your headline result.
Plus a small, general-purpose mutation-testing helper, because a green test suite doesn't tell you it would have caught the bug you didn't think to write a test for.
Why this exists
Three real incidents from one live quant-research project, each of which shipped before it was caught:
1. A single-anchor backtest is a sample of one. A strategy's backtest struck rebalance dates as a fixed stride starting from wherever the data happened to begin. Sweeping every possible phase of a 30-day cycle found Sharpe ranging from 0.35 to 1.46 -- and the ONE phase every earlier test had used happened to land on literally the best of all 30. Every parameter decision made before this was caught had been validated against the luckiest possible number, not a representative one.
2. A sign-convention bug that shipped three false wins. Drawdown is naturally stored as a negative number (a shallower, better drawdown is closer to zero). A comparison written the obvious way --
candidate_dd <= baseline_dd-- silently treats a DEEPER, WORSE drawdown as a pass, because-0.30 <= -0.20is true. Three candidates were reported as "confirmed improvements" on this exact bug before a manual re-check caught it.
3. A blending result that was pure measurement artifact. Averaging several phase-shifted copies of a strategy (offset "tranches") by taking the mean of cycle-i's return across all of them looked like a genuine diversification win: Sharpe up 27%, drawdown several points shallower. The check that caught it: running the IDENTICAL blend on plain buy-and-hold, which cannot have any real phase-dependent skill by construction. Buy-and-hold "improved" by almost the same ratio (Sharpe 0.89 -> 1.20). The blend was smoothing variance across offset windows, not reducing real risk -- and it would have shipped as the project's best result if that one control hadn't been run.
None of these are exotic mistakes. They're the kind of thing that survives code review, survives a reasonable test suite, and looks exactly like a genuine result until someone runs the specific control that exposes it. AnchorTest packages those controls so you run them by default, not by luck.
Install
pip install anchortest # once published -- see below
# or, for now:
pip install -e .
Requires Python 3.9+, numpy, pandas. No scipy, no broker SDK, no options-pricing library -- this works
on any equity-curve-producing backtest, in any asset class, options or not.
Paid audits
I'll run these checks against YOUR backtest and send back a written report: what passed, what didn't, and
why. Quick Check ($199, 3 business days): anchor-average your existing backtest, run it against the
strict bar, check any blending logic for the artifact pattern above. Full Audit ($749, 7 business days):
Quick Check plus a held-out/out-of-sample test designed for your specific strategy, plus a 30-minute call.
This is a review of your validation process, not investment advice and not a claim that any strategy is
profitable. Email grant02339@gmail.com to book, or see the landing page (docs/index.html).
Quickstart
from anchortest import anchor_average, compare, clears_bar
def backtest(params, start_shift=0):
"""Your own backtest. MUST accept start_shift and drop that many rows of your underlying
price data before computing anything -- see examples/random_walk_example.py for a full one."""
... # returns a pandas Series equity curve, indexed by date, starting at 1.0
baseline, _ = anchor_average(backtest, baseline_params, cycle_days=25, n_anchors=25)
candidate, _ = anchor_average(backtest, candidate_params, cycle_days=25, n_anchors=25)
delta = compare(baseline, candidate)
if clears_bar(delta):
print("candidate improves EVERY tracked metric -- worth a closer look")
else:
print("not a promotion:", delta)
Run python examples/random_walk_example.py for a runnable end-to-end demo, including a case where the
single-anchor (shift=0) result actively disagrees with the honest average.
What's in the box
| function | catches |
|---|---|
anchor_average(backtest_fn, params, cycle_days, n_anchors=25) |
trusting one lucky start date |
shifts_for(cycle_days, n_anchors) |
the phase-shift arithmetic itself (deduplicates correctly when n_anchors > cycle_days) |
compare(baseline, candidate) / clears_bar(delta) |
the drawdown sign-convention trap; "5 of 6 metrics improved" being reported as a win |
check_blend_artifact(blend_fn) / assert_no_blend_artifact(result) |
a tranching/blending function that inflates Sharpe with no real skill |
run_mutation_suite(mutations, test_command) / assert_all_caught(results) |
a test suite that would not actually have caught the bug you just fixed |
Each function's docstring explains the real failure mode it generalizes from -- read them; the "why" is not padding.
Philosophy
- Compare against a control before you trust a transform. If you can construct a version of your data
where the "true" answer is known (a phase-invariant control, a random baseline, a null model), run your
method against it before running it on the real thing.
check_blend_artifactis one instance of this principle; it generalizes further than this library currently automates. - The strict bar exists because partial improvement has been about a coin flip. A candidate that
improves mean Sharpe while its minimum-case Sharpe or recent-half Sharpe gets worse is not free money --
in real testing, that exact pattern has predicted a real-world regression as often as an improvement.
clears_bar's default requires every tracked metric to improve, deliberately. - A green test suite is a claim, not a fact, until you've tried to falsify it.
run_mutation_suiteis the smallest possible version of "did this test actually assert anything, or did it just not crash."
What this is not
- Not a backtesting engine. Bring your own
backtest_fn; AnchorTest only tells you how much to trust its output. - Not investment advice, and not a signal that any particular strategy works. It's a set of falsification checks -- what survives them still needs real out-of-sample and (if it trades real markets) real historical-data testing before anyone should trust it with money.
- Not a substitute for testing out-of-sample, on assets or time periods you didn't tune on. That discipline matters at least as much as anything in this library automates; there's no shortcut for it here yet.
Contributing
Issues and PRs welcome, especially: additional artifact-detection controls, a walk-forward / purged-cross-validation helper, and real-world "this caught a bug" case studies to add to the list above.
License
MIT. See LICENSE.
Metadata
Release files for anchortest 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| anchortest-0.1.0.tar.gz | 28.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| anchortest-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 42.3 kB
Release files / anchortest-0.1.0.tar.gz
| Download URL | anchortest-0.1.0.tar.gz |
|---|---|
| Size | 28.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2674cf6c791f270acdca1aeae087fe70c6f035f46d0024628c1580e43108c328
|
|
BLAKE2b-256 checksum How to use checksums |
6dca841ea1d67d1dc504908a2aa1becfbd28be507d85da17e7573645f084f62a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.
Transparency logRelease files / anchortest-0.1.0-py3-none-any.whl
| Download URL | anchortest-0.1.0-py3-none-any.whl |
|---|---|
| Size | 14.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0085b2c04e7c34d4117a36e209120021ed05459f707ff88b69d030ffbac9f9bb
|
|
BLAKE2b-256 checksum How to use checksums |
2872bb28e3d47cbe277744fdf7f6ec2725030905090403f2c0a72d698f34bf81
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.
Transparency log