Skip to main content

topstep-backtest

An event-driven backtesting framework for developing futures strategies that can profitably pass the Topstep Trading Combine.

Two things make it different. Backtest/live parity: a strategy is written once against structural protocols that both the deterministic SimBroker and the live AsyncTopstepClient from topstep-sdk satisfy. A first-class prop-firm rule engine: the two-state trailing Maximum Loss Limit, optional Daily Loss Limit, consistency target, position caps and session flatten are enforced in real time — including intrabar forced liquidation with adverse slippage — not scored after the fact.

Money is exact Decimal on the tick grid, lots are FIFO, reruns are bit-for-bit reproducible. SimBroker runs the full order lifecycle (market/limit/stop/trailing, signed-tick OCO brackets, gateway-parity APIError rejections) and resolves fill-vs-breach intrabar along one pessimistic price path. Every indicator is TA-Lib, driven bar by bar, so no formula is re-implemented here to drift. Module map in AGENTS.md.

class SmaCross(SymbolStrategy):
    def __init__(self, contract_id: str) -> None:
        super().__init__(contract_id)
        self.fast = self.use(Sma(20))  # TA-Lib SMA — causal by construction
        self.slow = self.use(Sma(50))  # == TalibIndicator("SMA", timeperiod=50)
        self.cross = self.use(Cross(self.fast, self.slow))  # compares them; not TA-Lib

    async def on_bar(self, bar: Bar) -> None:  # gated until every use()d indicator is ready
        if self.cross.up and self.position.flat:
            await self.buy(2, stop_loss_ticks=40, take_profit_ticks=80)  # signed-tick OCO

Every named indicator is a typed alias for a TA-Lib function, not a reimplementation: Sma(20) is TalibIndicator("SMA", timeperiod=20), and the generic form reaches 152 of TA-Lib's 161 functions directly. Nothing in this repo implements an indicator formula, so there is no second implementation to drift from the reference one. Cross is the deliberate exception — TA-Lib has no crossover primitive, so it is a framework helper that compares two TA-Lib outputs rather than computing anything.

Install

pip install topstep-backtest
pip install "topstep-backtest[data]"   # adds pandas + pyarrow: DataFrame and Parquet input

Requires Python 3.12+. TA-Lib is a core dependency and ships wheels for common platforms; on others you will need the TA-Lib C library first.

from datetime import date
from decimal import Decimal

from topstep_backtest import AccountSize, Backtest, SymbolStrategy
from topstep_backtest.core.instruments import spec_for_symbol
from topstep_backtest.data.synthetic import synthetic_bars
from topstep_backtest.indicators import Cross, Sma

MNQ = "CON.F.US.MNQ.U26"


class SmaCross(SymbolStrategy):
    def __init__(self, contract_id: str) -> None:
        super().__init__(contract_id)
        self.fast = self.use(Sma(10))  # TA-Lib SMA
        self.slow = self.use(Sma(30))  # TA-Lib SMA
        self.cross = self.use(Cross(self.fast, self.slow))  # framework helper, not TA-Lib

    async def on_bar(self, bar) -> None:
        if self.cross.up and self.position.flat:
            await self.buy(1, stop_loss_ticks=40, take_profit_ticks=80)
        elif self.cross.down and self.position.is_long:
            await self.close()


bars = synthetic_bars(
    contract_id=MNQ,
    spec=spec_for_symbol("MNQ"),
    start_day=date(2026, 5, 4),
    days=5,
    seed=7,
    start_price=Decimal("23000.00"),
    bars_per_day=120,
    vol_ticks=12,
)
print(Backtest(bars, SmaCross(MNQ), account=AccountSize.S50K).run())

For your own data, data.wrangler.bars_from_dataframe (pandas) and bars_from_records (no pandas) turn candles into validated Bar streams. Both make you declare stamp="open" or stamp="close" — what your timestamps mean is the difference between a causal backtest and an off-by-one-bar look-ahead.

What you get back

print(report) renders the verdict, the day-by-day MLL trail, and four statistics blocks. Every figure states its basis, because most admit two honest answers and mixing them silently is how a report lies:

  • Trade statistics — expectancy, payoff ratio, win rate, profit factor, longest losing streak, and breakeven_cost_per_half_turn (the extra cost per half-turn that would zero the run). Gross-basis, alongside net P&L. Under 200 closes the report prints a PROVISIONAL banner: nothing is suppressed, but a measured edge that thin is not distinguishable from sampling noise, and the report says so rather than letting you read it as a finding.
  • Drawdown, three ways — static from the initial balance, eod_trailing (Topstep's actual MLL mechanic) and intraday_trailing (Apex-style, ratchets on unrealized highs). These are different numbers on the same path and a strategy can survive one while violating another. Plus duration, time-to-recovery, avg_eod_trailing (the mean episode depth, so you can see whether the worst one was typical), and min_floor_headroom — how close the account ever came to termination, as against where it merely ended.
  • Round trips — flat-to-flat excursions with true R-multiples: net P&L over the dollars actually risked at entry, taken from the bracket stop. Net-basis, deliberately opposite to the gross trade statistics, because a round trip is a complete decision. Plus the dollar extremes and holding times: R says how a trade went against its own plan, worst_trade says whether the account could absorb it.
  • Daily P&L distribution — worst day, p05/p25/median/p75/p95, and stdev in dollars, for reasoning about a daily loss limit against the day you should size for, not just the one you drew.
  • exposure — the fraction of bars that actually held a position, which is what tells you how to read every figure above. The same drawdown at 5% and at 95% exposure are not the same risk. Alongside equity_peak (what a trailing floor anchors to) and the run's window.

The same report renders as an interactive HTML tearsheet — one self-contained file with no server and no network, so it opens offline and archives next to the run:

report.to_html("tearsheet.html")  # name the file
report.show()  # or write a temp file and open a browser

It draws the candlestick tape with every fill marked, the equity curve against the trailing MLL floor, daily P&L and the R-multiple distribution, plus every statistic the text render prints, basis labels included — the two renders share their formatting helpers, so they cannot disagree. For a sheet from every run without naming a file each time, Backtest(...).run_with_tearsheet("runs") writes tearsheet-<UTC stamp>.html and returns both the report and the path.

Then stop trusting one sample:

from topstep_backtest.metrics import monte_carlo
from topstep_backtest.rules.params import AccountSize, combine_params

mc = monte_carlo(report.result, params=combine_params(AccountSize.S50K), paths=3000, seed=7)
print(
    f"P(pass) {mc.pass_probability:.1%}  |  died: "
    f"MLL {mc.mll_breach_probability:.1%} · "
    f"consistency {mc.consistency_blocked_probability:.1%} · "
    f"too slow {mc.target_not_reached_probability:.1%}"
)

monte_carlo block-bootstraps the run's own trading days — in contiguous blocks, never i.i.d., because day-to-day clustering is exactly what a trailing drawdown bets against — and replays thousands of synthetic Combines through the real rule kernel, carrying each day's intraday equity excursion so an intraday breach is reproduced rather than missed. What matters is not the pass probability but the autopsy, because the three failure modes imply three different fixes:

mode what to do about it
mll_breach too much risk per day — resize
consistency_blocked money made in too few days — throttle the outsized day; the edge is fine
target_not_reached the edge is too slow for the window — nothing risk-side helps
uv run python examples/run_montecarlo.py   # the same edge at two sizes, failing two ways

Status and limitations

Pre-alpha (0.2.0). The engine core is well covered — exact-Decimal money on the tick grid, FIFO lot accounting, a structurally enforced no-look-ahead firewall, byte-identical reruns — but read these before trusting a number:

  • The rule and fee constants are NOT calibrated against a live account. They are researched, source-cited config (docs/topstep-rules.md §9 — the "trading day" definition was calibrated 2026-08-03 and the engine already matched; the rest is still unchecked). Treat a PASSED/FAILED verdict as a diagnostic, not an answer, and distrust any result landing within a tick or a fee of a limit. This applies with more force to the Monte-Carlo pass probability: a figure printed to one decimal from unverified constants is precise, not accurate.
  • Analytics can only resample the tape you gave them. Walk-forward, PBO, deflated Sharpe and the EV-per-attempt model all ship (metrics/walkforward.py, metrics/overfitting.py, metrics/economics.py), but none of them escapes your sample: the Monte-Carlo resamples a strategy's own observed days, so it cannot invent a market regime your tape never contained and will understate tail risk on a short or single-regime sample.
  • No exchange holiday calendar ships with this package. Bars on market holidays and past early-close halts are not detected, flagged or filtered anywhere — filter them upstream. (A built-in calendar was removed in 0.1.0: it disagreed with CME on several dates a year and silently discarded tradable sessions, which is worse than not having one.) Weekends and the 17:00–18:00 ET maintenance halt are modelled.
  • Multi-year runs need stitching, which now ships. data.continuous.stitch_continuous back-adjusts several expiries into one continuous series labelled with the bare ticker ("MNQ"). A raw splice does not corrupt P&L — the 16:10 flatten means no position spans a roll seam — but it does corrupt indicator state, producing spurious crossovers and ATR spikes at every roll. Additive adjustment is exactly P&L-neutral here and is pinned by an end-to-end test; it does distort logic keyed to absolute price levels.
  • Tier-0 bar fills only. Market orders fill at the next bar's open and default to zero slippage, so a strategy whose edge is thinner than roughly 8–10 ticks per round turn is inside the model's error bars. Event windows (08:30 ET releases and the like) are not honestly modelled at bar resolution.
  • Untested end to end: multi-symbol runs, dll_enabled=True, and non-quarter-tick products (CL, GC).

What is proven is structural parity — pyright-strict protocol conformance plus a place() signature-diff test against the live SDK. The behavioural gate (an intent-sequence test against a recording live broker) is not built yet.

Development

From a checkout (the sibling topstep-sdk repo is expected at ../topstep-sdk; set UV_NO_SOURCES=1 to resolve it from PyPI instead):

uv sync --extra dev
uv run pytest && uv run ruff check . && uv run ruff format --check . && uv run pyright
uv run python scripts/gen_api_surface.py --check   # AGENTS.md §3 must not be stale
uv run python examples/run_combine.py              # end-to-end combine verdict
uv run python examples/run_montecarlo.py           # outcome distribution + autopsy

Those checks are the gate: CI runs all of them on Python 3.12/3.13/3.14, then installs the built wheel into a clean venv and smoke-tests it. ruff format --check is part of it — ruff check passing is not sufficient, and the two fail independently.

Documentation

There is a full documentation site — a browsable version of everything below, plus a quickstart, a page on reading the report, and an API reference generated from the live docstrings:

https://tarricsookdeo.github.io/topstep-backtest/

It is rebuilt from main on every push. The site is public; the source repository is not, so there are no "Edit this page" links and there is no public issue tracker. To read it offline, or to preview a change before pushing it:

uv run --extra docs mkdocs serve      # http://127.0.0.1:8000

Everything the site renders also ships as plain markdown inside the sdist — pip download --no-binary :all: topstep-backtest and unpack it, or read the copies in your environment. There is no public issue tracker.

  • docs/TUTORIAL_EMA_CROSSOVER.md — start here: one strategy end to end, raw candles to Combine verdict.
  • website/results.md — reading the report: what every figure means, what basis it is on, which ones mislead alone, and what this framework deliberately does not report.
  • AGENTS.md — dense reference for using the framework: strategy dialect, module map, the invariants you must not break, and §6, the end-to-end workflow (run → check the wiring → read the run minding each metric's basis → Monte-Carlo → act on the autopsy).
  • docs/INDICATORS.md — the TA-Lib indicator surface, wrapper by wrapper.
  • docs/topstep-rules.md — the rulebook being enforced, with sources, confidence levels, and a verify-before-trusting checklist.
  • docs/DESIGN.md — architecture contract, for modifying the framework.
  • docs/ROADMAP.md — what is built, partial, and not started. Next: live adapter + calibration → per-year / per-regime breakdowns → L1/L2/MBO fill tiers; funded-account (XFA) modeling is deliberately parked.
  • examples/ — runnable: run_real_data.py (your CSV/Parquet → verdict), run_combine.py (synthetic end to end), run_tearsheet.py (the same run as one HTML file), run_montecarlo.py (outcome distribution + autopsy), run_windows.py (a long tape replayed as consecutive independent Combine attempts), ema_cross.py, sma_cross.py, talib_macd.py, hand_wired.py (what the facade assembles).

Stack

Python 3.12+ · topstep-sdk · msgspec · TA-Lib · uv · ruff · pyright (strict) · pytest + hypothesis. Canonical timezone ET (America/New_York); the internal hot path is int-ns UTC; all money is exact Decimal on the tick grid — indicator values cross from float64 into Decimal and are deliberately not tick-snapped, because an indicator level is not a tradeable price.

License

MIT — see LICENSE.

Unofficial. Not affiliated with, endorsed by, or sponsored by Topstep, LLC or ProjectX Trading, LLC. "Topstep" is a trademark of its respective owner and is used here only to identify the evaluation program this tool models. Rule numbers researched July 2026 — re-verify against Topstep's help center before trusting a pass verdict (see the checklist in docs/topstep-rules.md §9).

Not financial advice. This software simulates a trading evaluation and can be wrong. You are solely responsible for any capital you risk.

Metadata

Release files for topstep-backtest 0.2.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for topstep-backtest 0.2.3
File Size Uploaded
topstep_backtest-0.2.3.tar.gz 493.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for topstep-backtest 0.2.3
File Interpreter ABI Platform
topstep_backtest-0.2.3-py3-none-any.whl Python 3 none any Details

Total release size: 725.5 kB

Release files / topstep_backtest-0.2.3.tar.gz

Download URL topstep_backtest-0.2.3.tar.gz
Size 493.0 kB
Tags Source
SHA-256 checksum
How to use checksums
c0db0a0feec7900da8d75e832734c321bf3b82d90e72d8fdbb7f644f2739bde1
BLAKE2b-256 checksum
How to use checksums
2a5a869c4eb42318fc3590c6bc2d1e8a0a5267b9b7c84a4b3af83273a41809b7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 16, 2026.

Transparency log

Release files / topstep_backtest-0.2.3-py3-none-any.whl

Download URL topstep_backtest-0.2.3-py3-none-any.whl
Size 232.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b88f633d795929a05c6072d3e80d848d7500cae9e97e56d314dcfc14b627a23a
BLAKE2b-256 checksum
How to use checksums
e5efee349c892aa189fb2c34cc54ee8a4babf3c0ca0a5715bc34e5bd4da56ce5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 16, 2026.

Transparency log

Release history Release notifications | RSS feed

0.4.0

2 release files

0.3.0

2 release files

This release

0.2.3 This release

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page