Skip to main content

AnchorTest

Stop trusting a single backtest run.

Most backtests report one run's Sharpe ratio, on one start date, over one universe. That number is a sample of one. AnchorTest is a small, dependency-light Python library (numpy + pandas, nothing else) that runs three specific checks that have each caught a real, shipped bug in real trading-strategy research:

  1. Anchor-average, instead of trusting one lucky start date.
  2. Compare candidates correctly, with the drawdown sign-convention trap fixed once, in one place.
  3. Catch a blending/tranching artifact before it becomes your headline result.

Plus a small, general-purpose mutation-testing helper, because a green test suite doesn't tell you it would have caught the bug you didn't think to write a test for.

Why this exists

Three real incidents from one live quant-research project, each of which shipped before it was caught:

1. A single-anchor backtest is a sample of one. A strategy's backtest struck rebalance dates as a fixed stride starting from wherever the data happened to begin. Sweeping every possible phase of a 30-day cycle found Sharpe ranging from 0.35 to 1.46 -- and the ONE phase every earlier test had used happened to land on literally the best of all 30. Every parameter decision made before this was caught had been validated against the luckiest possible number, not a representative one.

2. A sign-convention bug that shipped three false wins. Drawdown is naturally stored as a negative number (a shallower, better drawdown is closer to zero). A comparison written the obvious way -- candidate_dd <= baseline_dd -- silently treats a DEEPER, WORSE drawdown as a pass, because -0.30 <= -0.20 is true. Three candidates were reported as "confirmed improvements" on this exact bug before a manual re-check caught it.

3. A blending result that was pure measurement artifact. Averaging several phase-shifted copies of a strategy (offset "tranches") by taking the mean of cycle-i's return across all of them looked like a genuine diversification win: Sharpe up 27%, drawdown several points shallower. The check that caught it: running the IDENTICAL blend on plain buy-and-hold, which cannot have any real phase-dependent skill by construction. Buy-and-hold "improved" by almost the same ratio (Sharpe 0.89 -> 1.20). The blend was smoothing variance across offset windows, not reducing real risk -- and it would have shipped as the project's best result if that one control hadn't been run.

None of these are exotic mistakes. They're the kind of thing that survives code review, survives a reasonable test suite, and looks exactly like a genuine result until someone runs the specific control that exposes it. AnchorTest packages those controls so you run them by default, not by luck.

Install

pip install anchortest    # once published -- see below
# or, for now:
pip install -e .

Requires Python 3.9+, numpy, pandas. No scipy, no broker SDK, no options-pricing library -- this works on any equity-curve-producing backtest, in any asset class, options or not.

Paid audits

I'll run these checks against YOUR backtest and send back a written report: what passed, what didn't, and why. Quick Check ($199, 3 business days): anchor-average your existing backtest, run it against the strict bar, check any blending logic for the artifact pattern above. Full Audit ($749, 7 business days): Quick Check plus a held-out/out-of-sample test designed for your specific strategy, plus a 30-minute call. This is a review of your validation process, not investment advice and not a claim that any strategy is profitable. Email grant02339@gmail.com to book, or see the landing page (docs/index.html).

Quickstart

from anchortest import anchor_average, compare, clears_bar

def backtest(params, start_shift=0):
    """Your own backtest. MUST accept start_shift and drop that many rows of your underlying
    price data before computing anything -- see examples/random_walk_example.py for a full one."""
    ...  # returns a pandas Series equity curve, indexed by date, starting at 1.0

baseline, _ = anchor_average(backtest, baseline_params, cycle_days=25, n_anchors=25)
candidate, _ = anchor_average(backtest, candidate_params, cycle_days=25, n_anchors=25)

delta = compare(baseline, candidate)
if clears_bar(delta):
    print("candidate improves EVERY tracked metric -- worth a closer look")
else:
    print("not a promotion:", delta)

Run python examples/random_walk_example.py for a runnable end-to-end demo, including a case where the single-anchor (shift=0) result actively disagrees with the honest average.

What's in the box

function catches
anchor_average(backtest_fn, params, cycle_days, n_anchors=25) trusting one lucky start date
shifts_for(cycle_days, n_anchors) the phase-shift arithmetic itself (deduplicates correctly when n_anchors > cycle_days)
compare(baseline, candidate) / clears_bar(delta) the drawdown sign-convention trap; "5 of 6 metrics improved" being reported as a win
check_blend_artifact(blend_fn) / assert_no_blend_artifact(result) a tranching/blending function that inflates Sharpe with no real skill
run_mutation_suite(mutations, test_command) / assert_all_caught(results) a test suite that would not actually have caught the bug you just fixed

Each function's docstring explains the real failure mode it generalizes from -- read them; the "why" is not padding.

Philosophy

  • Compare against a control before you trust a transform. If you can construct a version of your data where the "true" answer is known (a phase-invariant control, a random baseline, a null model), run your method against it before running it on the real thing. check_blend_artifact is one instance of this principle; it generalizes further than this library currently automates.
  • The strict bar exists because partial improvement has been about a coin flip. A candidate that improves mean Sharpe while its minimum-case Sharpe or recent-half Sharpe gets worse is not free money -- in real testing, that exact pattern has predicted a real-world regression as often as an improvement. clears_bar's default requires every tracked metric to improve, deliberately.
  • A green test suite is a claim, not a fact, until you've tried to falsify it. run_mutation_suite is the smallest possible version of "did this test actually assert anything, or did it just not crash."

What this is not

  • Not a backtesting engine. Bring your own backtest_fn; AnchorTest only tells you how much to trust its output.
  • Not investment advice, and not a signal that any particular strategy works. It's a set of falsification checks -- what survives them still needs real out-of-sample and (if it trades real markets) real historical-data testing before anyone should trust it with money.
  • Not a substitute for testing out-of-sample, on assets or time periods you didn't tune on. That discipline matters at least as much as anything in this library automates; there's no shortcut for it here yet.

Contributing

Issues and PRs welcome, especially: additional artifact-detection controls, a walk-forward / purged-cross-validation helper, and real-world "this caught a bug" case studies to add to the list above.

License

MIT. See LICENSE.

Metadata

Release files for anchortest 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for anchortest 0.1.0
File Size Uploaded
anchortest-0.1.0.tar.gz 28.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for anchortest 0.1.0
File Interpreter ABI Platform
anchortest-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 42.3 kB

Release files / anchortest-0.1.0.tar.gz

Download URL anchortest-0.1.0.tar.gz
Size 28.3 kB
Tags Source
SHA-256 checksum
How to use checksums
2674cf6c791f270acdca1aeae087fe70c6f035f46d0024628c1580e43108c328
BLAKE2b-256 checksum
How to use checksums
6dca841ea1d67d1dc504908a2aa1becfbd28be507d85da17e7573645f084f62a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release files / anchortest-0.1.0-py3-none-any.whl

Download URL anchortest-0.1.0-py3-none-any.whl
Size 14.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0085b2c04e7c34d4117a36e209120021ed05459f707ff88b69d030ffbac9f9bb
BLAKE2b-256 checksum
How to use checksums
2872bb28e3d47cbe277744fdf7f6ec2725030905090403f2c0a72d698f34bf81
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page