topstep-backtest
An event-driven backtesting framework for developing futures strategies that can profitably pass the Topstep Trading Combine.
Two things make it different. Backtest/live parity: a strategy is written once against
structural protocols that both the deterministic SimBroker and the live
AsyncTopstepClient from topstep-sdk satisfy. A first-class prop-firm rule engine:
the two-state trailing Maximum Loss Limit, optional Daily Loss Limit, consistency target,
position caps and session flatten are enforced in real time — including intrabar forced
liquidation with adverse slippage — not scored after the fact.
Money is exact Decimal on the tick grid, lots are FIFO, reruns are bit-for-bit
reproducible. SimBroker runs the full order lifecycle (market/limit/stop/trailing,
signed-tick OCO brackets, gateway-parity APIError rejections) and resolves fill-vs-breach
intrabar along one pessimistic price path. Every indicator is TA-Lib, driven bar by bar, so
no formula is re-implemented here to drift. Module map in AGENTS.md.
class SmaCross(SymbolStrategy):
def __init__(self, contract_id: str) -> None:
super().__init__(contract_id)
self.fast = self.use(Sma(20)) # TA-Lib SMA — causal by construction
self.slow = self.use(Sma(50)) # == TalibIndicator("SMA", timeperiod=50)
self.cross = self.use(Cross(self.fast, self.slow)) # compares them; not TA-Lib
async def on_bar(self, bar: Bar) -> None: # gated until every use()d indicator is ready
if self.cross.up and self.position.flat:
await self.buy(2, stop_loss_ticks=40, take_profit_ticks=80) # signed-tick OCO
Every named indicator is a typed alias for a TA-Lib function, not a
reimplementation: Sma(20) is TalibIndicator("SMA", timeperiod=20), and the generic
form reaches 152 of TA-Lib's 161 functions directly. Nothing in this repo implements an
indicator formula, so there is no second implementation to drift from the reference one.
Cross is the deliberate exception — TA-Lib has no crossover primitive, so it is a
framework helper that compares two TA-Lib outputs rather than computing anything.
Install
pip install topstep-backtest
pip install "topstep-backtest[data]" # adds pandas, for DataFrame input
Requires Python 3.12+. TA-Lib is a core dependency and ships wheels for common platforms; on others you will need the TA-Lib C library first.
from datetime import date
from decimal import Decimal
from topstep_backtest import AccountSize, Backtest, SymbolStrategy
from topstep_backtest.core.instruments import spec_for_symbol
from topstep_backtest.data.synthetic import synthetic_bars
from topstep_backtest.indicators import Cross, Sma
MNQ = "CON.F.US.MNQ.U26"
class SmaCross(SymbolStrategy):
def __init__(self, contract_id: str) -> None:
super().__init__(contract_id)
self.fast = self.use(Sma(10)) # TA-Lib SMA
self.slow = self.use(Sma(30)) # TA-Lib SMA
self.cross = self.use(Cross(self.fast, self.slow)) # framework helper, not TA-Lib
async def on_bar(self, bar) -> None:
if self.cross.up and self.position.flat:
await self.buy(1, stop_loss_ticks=40, take_profit_ticks=80)
elif self.cross.down and self.position.is_long:
await self.close()
bars = synthetic_bars(
contract_id=MNQ,
spec=spec_for_symbol("MNQ"),
start_day=date(2026, 5, 4),
days=5,
seed=7,
start_price=Decimal("23000.00"),
bars_per_day=120,
vol_ticks=12,
)
print(Backtest(bars, SmaCross(MNQ), account=AccountSize.S50K).run())
For your own data, data.wrangler.bars_from_dataframe (pandas) and bars_from_records (no
pandas) turn candles into validated Bar streams. Both make you declare stamp="open" or
stamp="close" — what your timestamps mean is the difference between a causal backtest and
an off-by-one-bar look-ahead.
What you get back
print(report) renders the verdict, the day-by-day MLL trail, and four statistics blocks.
Every figure states its basis, because most admit two honest answers and mixing them
silently is how a report lies:
- Trade statistics — expectancy, payoff ratio, win rate, profit factor, longest losing
streak, and
breakeven_cost_per_half_turn(the extra cost per half-turn that would zero the run). Gross-basis, alongside net P&L. Under 200 closes the report prints aPROVISIONALbanner: nothing is suppressed, but a measured edge that thin is not distinguishable from sampling noise, and the report says so rather than letting you read it as a finding. - Drawdown, three ways —
staticfrom the initial balance,eod_trailing(Topstep's actual MLL mechanic) andintraday_trailing(Apex-style, ratchets on unrealized highs). These are different numbers on the same path and a strategy can survive one while violating another. Plus duration, time-to-recovery,avg_eod_trailing(the mean episode depth, so you can see whether the worst one was typical), andmin_floor_headroom— how close the account ever came to termination, as against where it merely ended. - Round trips — flat-to-flat excursions with true R-multiples: net P&L over the
dollars actually risked at entry, taken from the bracket stop. Net-basis, deliberately
opposite to the gross trade statistics, because a round trip is a complete decision. Plus
the dollar extremes and holding times: R says how a trade went against its own plan,
worst_tradesays whether the account could absorb it. - Daily P&L distribution — worst day, p05/p25/median/p75/p95, and
stdevin dollars, for reasoning about a daily loss limit against the day you should size for, not just the one you drew. exposure— the fraction of bars that actually held a position, which is what tells you how to read every figure above. The same drawdown at 5% and at 95% exposure are not the same risk. Alongsideequity_peak(what a trailing floor anchors to) and the run's window.
Then stop trusting one sample:
from topstep_backtest.metrics import monte_carlo
from topstep_backtest.rules.params import AccountSize, combine_params
mc = monte_carlo(report.result, params=combine_params(AccountSize.S50K), paths=3000, seed=7)
print(
f"P(pass) {mc.pass_probability:.1%} | died: "
f"MLL {mc.mll_breach_probability:.1%} · "
f"consistency {mc.consistency_blocked_probability:.1%} · "
f"too slow {mc.target_not_reached_probability:.1%}"
)
monte_carlo block-bootstraps the run's own trading days — in contiguous blocks, never
i.i.d., because day-to-day clustering is exactly what a trailing drawdown bets against — and
replays thousands of synthetic Combines through the real rule kernel, carrying each day's
intraday equity excursion so an intraday breach is reproduced rather than missed. What
matters is not the pass probability but the autopsy, because the three failure modes
imply three different fixes:
| mode | what to do about it |
|---|---|
mll_breach |
too much risk per day — resize |
consistency_blocked |
money made in too few days — throttle the outsized day; the edge is fine |
target_not_reached |
the edge is too slow for the window — nothing risk-side helps |
uv run python examples/run_montecarlo.py # the same edge at two sizes, failing two ways
Status and limitations
Pre-alpha (0.1.0). The engine core is well covered — exact-Decimal money on the tick grid, FIFO lot accounting, a structurally enforced no-look-ahead firewall, byte-identical reruns — but read these before trusting a number:
- The rule and fee constants are NOT calibrated against a live account. They are
researched, source-cited config (
docs/topstep-rules.md§9 — the "trading day" definition was calibrated 2026-08-03 and the engine already matched; the rest is still unchecked). Treat aPASSED/FAILEDverdict as a diagnostic, not an answer, and distrust any result landing within a tick or a fee of a limit. This applies with more force to the Monte-Carlo pass probability: a figure printed to one decimal from unverified constants is precise, not accurate. - Analytics are single-run plus resampling. There is no walk-forward, no PBO/DSR overfitting guard, and no EV-per-attempt model. The Monte-Carlo resamples a strategy's own observed days, so it cannot invent a market regime your tape never contained and will understate tail risk on a short or single-regime sample.
- No exchange holiday calendar ships with this package. Bars on market holidays and past early-close halts are not detected, flagged or filtered anywhere — filter them upstream. (A built-in calendar was removed in 0.1.0: it disagreed with CME on several dates a year and silently discarded tradable sessions, which is worse than not having one.) Weekends and the 17:00–18:00 ET maintenance halt are modelled.
- Multi-year runs need stitching, which now ships.
data.continuous.stitch_continuousback-adjusts several expiries into one continuous series labelled with the bare ticker ("MNQ"). A raw splice does not corrupt P&L — the 16:10 flatten means no position spans a roll seam — but it does corrupt indicator state, producing spurious crossovers and ATR spikes at every roll. Additive adjustment is exactly P&L-neutral here and is pinned by an end-to-end test; it does distort logic keyed to absolute price levels. - Tier-0 bar fills only. Market orders fill at the next bar's open and default to zero slippage, so a strategy whose edge is thinner than roughly 8–10 ticks per round turn is inside the model's error bars. Event windows (08:30 ET releases and the like) are not honestly modelled at bar resolution.
- Untested end to end: multi-symbol runs,
dll_enabled=True, and non-quarter-tick products (CL, GC).
What is proven is structural parity — pyright-strict protocol conformance plus a
place() signature-diff test against the live SDK. The behavioural gate (an
intent-sequence test against a recording live broker) is not built yet.
Development
From a checkout (the sibling topstep-sdk repo is expected at ../topstep-sdk; set
UV_NO_SOURCES=1 to resolve it from PyPI instead):
uv sync --extra dev
uv run pytest && uv run ruff check . && uv run ruff format --check . && uv run pyright
uv run python scripts/gen_api_surface.py --check # AGENTS.md §3 must not be stale
uv run python examples/run_combine.py # end-to-end combine verdict
uv run python examples/run_montecarlo.py # outcome distribution + autopsy
Those checks are the gate: CI runs all of them on Python 3.12/3.13/3.14, then installs the
built wheel into a clean venv and smoke-tests it. ruff format --check is part of it —
ruff check passing is not sufficient, and the two fail independently.
Documentation
There is a full documentation site — a browsable version of everything below, plus a quickstart, a page on reading the report, and an API reference generated from the live docstrings. The source repository is private, so it is not hosted publicly; build and read it locally from a checkout:
uv run --extra docs mkdocs serve # http://127.0.0.1:8000
Everything the site renders also ships as plain markdown inside the sdist — pip download --no-binary :all: topstep-backtest and unpack it, or read the copies in your
environment. There is no public issue tracker.
docs/TUTORIAL_EMA_CROSSOVER.md— start here: one strategy end to end, raw candles to Combine verdict.website/results.md— reading the report: what every figure means, what basis it is on, which ones mislead alone, and what this framework deliberately does not report.AGENTS.md— dense reference for using the framework: strategy dialect, module map, the invariants you must not break, and §6, the end-to-end workflow (run → check the wiring → read the run minding each metric's basis → Monte-Carlo → act on the autopsy).docs/INDICATORS.md— the TA-Lib indicator surface, wrapper by wrapper.docs/topstep-rules.md— the rulebook being enforced, with sources, confidence levels, and a verify-before-trusting checklist.docs/DESIGN.md— architecture contract, for modifying the framework.docs/ROADMAP.md— what is built, partial, and not started. Next: live adapter + calibration → remaining analytics (EV per attempt, PBO/DSR, walk-forward) → L1/L2/MBO fill tiers; funded-account (XFA) modeling is deliberately parked.examples/— runnable:run_real_data.py(your CSV/Parquet → verdict),run_combine.py(synthetic end to end),run_montecarlo.py(outcome distribution + autopsy),run_windows.py(a long tape replayed as consecutive independent Combine attempts),ema_cross.py,sma_cross.py,talib_macd.py,hand_wired.py(what the facade assembles).
Stack
Python 3.12+ · topstep-sdk · msgspec · TA-Lib · uv · ruff · pyright (strict) ·
pytest + hypothesis. Canonical timezone ET (America/New_York); the internal hot
path is int-ns UTC; all money is exact Decimal on the tick grid — indicator values cross
from float64 into Decimal and are deliberately not tick-snapped, because an indicator
level is not a tradeable price.
License
MIT — see LICENSE.
Unofficial. Not affiliated with, endorsed by, or sponsored by Topstep, LLC or ProjectX Trading, LLC. "Topstep" is a trademark of its respective owner and is used here only to identify the evaluation program this tool models. Rule numbers researched July 2026 — re-verify against Topstep's help center before trusting a pass verdict (see the checklist in
docs/topstep-rules.md§9).Not financial advice. This software simulates a trading evaluation and can be wrong. You are solely responsible for any capital you risk.
Metadata
Release files for topstep-backtest 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| topstep_backtest-0.2.0.tar.gz | 489.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| topstep_backtest-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 721.7 kB
Release files / topstep_backtest-0.2.0.tar.gz
| Download URL | topstep_backtest-0.2.0.tar.gz |
|---|---|
| Size | 489.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7d66005b5edbc2f8b9c8f791bdd0301dbe33ef1bb7aabd6a0f7cb5f368a2f5a7
|
|
BLAKE2b-256 checksum How to use checksums |
5ee2d4b5e3a66add63a043ea01043ba5c5168eb6dd2bdca573258a3e3564c2ed
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 16, 2026.
Transparency logRelease files / topstep_backtest-0.2.0-py3-none-any.whl
| Download URL | topstep_backtest-0.2.0-py3-none-any.whl |
|---|---|
| Size | 232.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e1b033a7ce8fdb816ce6c529c8a6e08283cdc7c103257c49a3998657d51ca2d9
|
|
BLAKE2b-256 checksum How to use checksums |
38f3d3f451bd200b06b125c929c34f36e547a433df2624a1c0ef59db4f62f6f4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 16, 2026.
Transparency log