market-data-normalizer (mdnorm)
Normalize heterogeneous market-data feeds — CSV tick dumps, exchange WebSocket JSON, and FIX — into a single, exchange-agnostic event schema, so downstream research and execution code never has to care where a tick came from.
Zero runtime dependencies. Pure Python (3.10+). Decimal prices, integer
nanosecond timestamps.
Why
Every venue spells the same thing differently: BTCUSDT vs XBT/USD,
millisecond epochs vs FIX UTCTimestamp, is_buyer_maker booleans vs side
codes. Research notebooks and backtesters end up littered with per-venue
parsing branches. mdnorm pushes that mess to the edge and hands the rest of
your stack one clean type.
Install
pip install market-data-normalizer
The distribution is named market-data-normalizer; the import name is
mdnorm:
import mdnorm
Pure Python, no runtime dependencies, Python 3.10+.
Quick start
from mdnorm import from_csv_row, from_ws_json, from_fix
# CSV row (ISO-8601 timestamp)
from_csv_row(
{"symbol": "btc/usd", "ts": "2026-01-02T00:00:00Z",
"price": "42000.5", "size": "0.25", "side": "buy"},
venue="coinbase",
)
# Exchange WebSocket trade message
from_ws_json({"s": "BTCUSDT", "p": "42000.5", "q": "0.25",
"T": 1767312000000, "m": False}, venue="binance")
# FIX execution report (SOH-delimited in the wild; "|" here for readability)
from_fix("55=BTC/USD|31=42000.5|32=0.25|54=1|60=20260102-00:00:00",
venue="lmax", sep="|")
All three calls above produce the same MarketEvent.
Quotes (bid/ask)
from mdnorm import from_ws_quote
q = from_ws_quote(
{"s": "BTCUSDT", "b": "41999.5", "B": "1.2",
"a": "42000.5", "A": "0.8", "T": 1767312000000},
venue="binance",
)
q.mid_price # Decimal("42000.0")
q.spread # Decimal("1.0")
from_csv_quote does the same for CSV rows with bid/ask columns.
OHLCV bars
from mdnorm import time_bars
bars = time_bars(events, interval_ns=60_000_000_000) # 1-minute bars
bars[0].open, bars[0].high, bars[0].low, bars[0].close, bars[0].volume, bars[0].vwap
time_bars reduces a stream of trade events into fixed-interval OHLCV Bars
(with VWAP and trade count), sorting out-of-order input and skipping quotes.
resample_bars(bars, interval_ns) downsamples bars to a coarser interval
(e.g. 1-minute → 5-minute) with correct OHLC aggregation and volume-weighted
VWAP.
fill_gaps(bars) returns a gapless series, inserting flat zero-volume bars
(OHLC = previous close) for any interval with no trades — a continuous grid for
backtests and feature pipelines.
Event-driven bars
Time bars are not the only clock. Sample by activity instead:
from decimal import Decimal
from mdnorm import count_bars, volume_bars, dollar_bars
count_bars(events, every=500) # tick bars
volume_bars(events, min_volume=Decimal("100")) # volume bars
dollar_bars(events, min_notional=Decimal("1e6")) # dollar bars
Trading sessions
Filter a feed down to the hours that matter, with daylight saving handled for you:
from mdnorm import US_EQUITY_RTH, filter_session, group_by_session_date
rth = filter_session(events, US_EQUITY_RTH) # 09:30-16:00 New York
by_day = group_by_session_date(events, US_EQUITY_RTH)
Overnight windows (a session that opens at 18:00 and closes at 17:00 the
next day) are supported, and session_date keeps a whole night in one
bucket. From the command line:
$ mdnorm bars trades.csv --interval 5m --session 09:30-16:00 --tz America/New_York -o rth.csv
Corporate actions and contract rolls
A raw price series is not continuous. A 4-for-1 split divides the printed price by four overnight, a cash dividend drops it by the amount paid, and a futures roll steps it by the spread between the two contracts. None of them are market moves, but all of them look like returns:
from decimal import Decimal
from mdnorm import adjust_bars, split, dividend, roll, iso_to_ns
actions = [
split(iso_to_ns("2026-06-06T00:00:00Z"), Decimal("4")),
dividend(iso_to_ns("2026-05-09T00:00:00Z"), Decimal("0.25")),
]
clean = adjust_bars(bars, actions)
Back-adjustment leaves the most recent segment at the prices that actually printed and restates everything before each event, so the joins are seamless:
raw closes 500 502 498 504 │ 126 125.5 127 126.5
raw returns +0.4% -0.8% +1.2% │ -75.0% -0.4% +1.2% -0.4%
^ the split, not a crash
adj closes 125 125.5 124.5 126 │ 126 125.5 127 126.5
adj returns +0.4% -0.8% +1.2% │ +0.0% -0.4% +1.2% -0.4%
Splits scale volume as well as price. Dividends take their reference price
from the last print before the ex-date unless you pass one. Rolls support
both conventions — AdjustMethod.RATIO (default, preserves returns) and
AdjustMethod.DIFFERENCE (preserves price differences, the usual choice for
futures). Factors are composed as exact rationals, so a 1-for-2 followed by a
1-for-3 restates 600 to exactly 100 rather than 99.999...96.
Actions can come from a file, and the CLI wires it up:
$ mdnorm bars trades.csv --interval 1d --actions actions.csv -o adjusted.csv
$ mdnorm bars tape.jsonl --infer-sides --every-imbalance 500 -o imbalance.csv
$ mdnorm book deltas.csv --symbol BTC-USD -o quotes.jsonl
$ mdnorm nbbo quotes.jsonl --max-age 2s -o top.jsonl
$ mdnorm tca fills.csv --market tape.jsonl --decision-price 100
ts,kind,value,ref_price
2026-06-06T00:00:00Z,split,4,
2026-05-09T00:00:00Z,dividend,0.25,190.50
2026-03-14T00:00:00Z,roll,5312.50,5290.25
Who crossed the spread
Most trade tapes give you a price and a size but not the aggressor. That one
missing field is what separates a price series from an order-flow series, and
signed volume, order imbalance and imbalance bars are all defined in terms of
it. mdnorm.micro infers it, using the three rules the literature settled on:
from mdnorm import SideRule, infer_sides, trade_imbalance, mean_effective_spread
classified = infer_sides(events) # Lee-Ready by default
print(trade_imbalance(classified)) # -1 selling .. +1 buying
print(mean_effective_spread(classified)) # 2 * |price - mid|
SideRule.TICK compares each trade with the previous different price and
needs trades only. SideRule.QUOTE compares the trade with the prevailing
mid. SideRule.LEE_READY — the default — uses the quote rule and falls back
to the tick rule at the mid. A side reported by the venue always wins;
inference only fills gaps, and trades it cannot resolve stay None rather
than being guessed at. Published accuracy of these rules is roughly 75-85% on
liquid names, so treat an inferred side as an estimate.
roll_spread estimates the effective spread from trade prices alone, via the
serial covariance that bid-ask bounce induces. It needs no quotes, which makes
it a useful cross-check on the rest — and it returns None rather than zero
when the covariance comes out non-negative and the estimator is undefined.
Imbalance bars
Once trades carry a side, the sampling clock can follow order flow instead of time or volume:
from mdnorm import Pipeline
bars = Pipeline().infer_sides().imbalance_bars(Decimal("500")).run(events)
A bar runs until buyers have outbought sellers, or the reverse, by the
threshold. Balanced two-sided periods produce one long bar; a sustained
one-sided push produces several short ones. by="tick" measures the imbalance
in trade count rather than size. From the command line:
$ mdnorm bars tape.jsonl --infer-sides --every-imbalance 500 -o imbalance.csv
Rebuilding the order book
Exchanges do not send you a book. They send a snapshot and then a stream of deltas, and the book only exists if you apply every one of them, in order:
from mdnorm import BookDelta, OrderBook, Side, replay_book
book = OrderBook("BTC-USD", "binance")
book.apply_snapshot(ts, bids=[(D("100"), D("2"))], asks=[(D("101"), D("3"))], seq=10)
quotes = list(replay_book(book, deltas)) # one quote per change in the top
print(book.best_bid, book.spread, book.imbalance(levels=5))
Two failure modes make a reconstructed book silently untrue, and this implementation refuses to hide either.
A sequence gap means a message was missed, and no later update repairs the
damage — the book is simply wrong from then on, in a way that looks completely
normal. OrderBook raises SequenceGapError the moment a number is skipped,
naming how many updates went missing, because the correct response is to
resynchronise from a snapshot rather than carry on. Duplicated or replayed
messages are rejected the same way. Feeds without sequence numbers work fine;
pass strict_sequence=False to opt out entirely.
A crossed book — best bid at or above best ask — is not a market state but
a symptom: a dropped delete, a stale snapshot, two venues merged by mistake.
It is exposed as is_crossed, and the spread goes negative rather than being
quietly clamped to zero.
to_quote() turns the top of the book into an ordinary MarketEvent, so a
reconstructed book feeds straight into session filtering, trade classification
and effective spreads with nothing in between. From the command line:
$ mdnorm book deltas.csv --symbol BTC-USD --venue binance -o quotes.jsonl
One instrument, several venues
When something trades in more than one place, "the price" is a question. The consolidated top of book is the answer, and it is where three problems live that a maximum over venues will not warn you about:
from mdnorm import consolidate
top = consolidate(quotes, max_age_ns=2_000_000_000) # 2s staleness cutoff
A venue that goes quiet keeps voting. When a feed disconnects, its last
quote stays in the consolidation forever — and a stale price is very often the
best price, so the dead venue ends up setting the top of book. This is the
failure that produces a consolidated feed which looks excellent and is
fiction. max_age_ns retires a venue that has not spoken recently;
stale_venues() names them.
A consolidated book can appear crossed. A bid on one venue above the offer
on another looks like free money and is almost always clock skew between two
feeds timestamped by different machines. is_crossed reports it and
crossed_updates counts it, because the useful response is to check the
clocks rather than to trade the spread.
Ties need a rule. Equal best prices are broken by size, then by venue name, so the same input always produces the same output.
Which venue actually sets the price is a measurement in its own right, and
leadership counts it. The pieces compose: an order book becomes a quote,
quotes from several venues consolidate into one, and the result feeds trade
classification and effective spreads unchanged.
$ mdnorm nbbo quotes.jsonl --symbol BTC-USD --max-age 2s -o top.jsonl
Did I execute well?
Once the tape is clean the next question is about you rather than the market, and every standard benchmark has a way of quietly flattering the person running it:
from mdnorm import Fill, Side, evaluate
report = evaluate(my_fills, market_trades, decision_price=D("100"))
print(report.slippage_vs_vwap_bps, report.participation_rate)
Your own trades are in the benchmark. A VWAP over the public tape includes
the prints you just made, so you end up partly benchmarking yourself against
yourself — and the bigger your share of volume, the more the benchmark bends
toward your own average price. evaluate removes your fills from the tape
before computing anything; exclude_fills does it on its own if you want the
benchmark separately. In the library's own test suite, leaving them in turns a
100 VWAP into 109 and a losing execution into a winning one.
Participation decides whether the number means anything. Beating VWAP by two basis points on 0.1% of volume is a result; the same number on 30% of volume mostly measures your own impact. The summary always reports the two together, and the CLI says so out loud above 10%.
Sign conventions are stated, not assumed. Positive basis points always mean better than the benchmark — paying below it on a buy, selling above it on a sell. Mixed-side fills are refused rather than netted, because one number covering both directions has no meaning.
By default the window runs from your first fill to your last. That is right
for a worked order and wrong for a single fill — the only print in the window
is then your own — so start_ns and end_ns let you score against an
interval you chose instead.
TWAP skips intervals that never traded instead of carrying the last price forward, for the same reason nothing else here invents data.
$ mdnorm tca fills.csv --market tape.jsonl --decision-price 100
Several instruments, one time grid
Research wants a matrix — one row per timestamp, one column per instrument — and building it from independent tick streams is where look-ahead bias gets in, because every mistake here makes the backtest better rather than raising:
from mdnorm import Field, align
rows = align({"BTC": btc_events, "ETH": eth_events},
interval_ns=60_000_000_000, # a one-minute grid
max_age_ns=5 * 60_000_000_000) # nothing older than 5 minutes
rows[0].values # {"BTC": Decimal("60000"), "ETH": Decimal("3000")}
rows[0].ages_ns # how old each value was at that grid point
rows[0].complete # False if any column had nothing to show
The join only looks backwards. A value is visible at a grid point only if it was observed at or before it. "Nearest observation" is the expensive default in this area: on a one-minute grid it lets a print from 09:30:20 be read at 09:30:00, and twenty seconds of hindsight is enough to make a mediocre signal look tradeable.
A bar labelled 09:30 is not knowable at 09:30. It contains everything that
traded until 09:31, so joining bars on their label imports an interval of the
future. AsOfSeries.from_bars timestamps each bar at its end, and
align_bars therefore gives you the last closed bar per column — one
interval further back than the naive join, and the version you could have
traded.
Forward-filling has no natural end. A halted or delisted stream otherwise
contributes its last price forever, and a frozen price correlates with nothing,
which reads as diversification. With max_age_ns a quiet column becomes
None; the age is still reported, so row.stale (had data, too old) and
row.missing (never had data) stay distinguishable.
A feed you get late was not available on time. AsOfSeries.delayed(250ms)
shifts observation times forward by the delivery delay, so alignment reflects
when you could have acted rather than when the source stamped it.
Nothing interpolates or smooths. align_on takes timestamps you supply, for
one row per print of a reference instrument, per signal, or per fill.
$ mdnorm align BTC=btc.csv ETH=eth.jsonl --interval 1m --max-age 5m -o matrix.csv
Features that cannot see the future
With the matrix built, the next step is turning prices into returns, z-scores, volatility and correlations. This is the second place look-ahead gets in, and it gets in just as quietly:
from mdnorm import ReturnMethod, column, returns, rolling_zscore, realized_volatility
px = column(rows, "BTC")
r = returns(px, method=ReturnMethod.LOG)
z = rolling_zscore(px, window=60) # trailing, never full-sample
vol = realized_volatility(r, window=60) # per period until you annualise it
A full-sample z-score is look-ahead. Subtracting the mean and dividing by
the standard deviation of the whole series gives every observation knowledge
of the distribution it sits in — including the part that had not happened yet.
It is one line of code and it is everywhere. Every statistic here is trailing:
the value at i comes from values[i-window+1 : i+1] and nothing else. There
is a test that pins this as a property — change the tail of the input and every
earlier output must be byte-identical — and a second test that shows the
full-sample form failing it.
A partial window is not a result. Until the window fills you get None,
not a twenty-period statistic computed from three observations. A gap inside a
window propagates for the same reason: stepping over the hole would compute a
twenty-period number from nineteen and label it twenty.
A frozen series has no z-score and no correlation. Zero dispersion makes
both undefined, so both return None rather than 0. Reading that zero as a
correlation is how a dead feed becomes an apparent diversifier.
There is no default annualisation factor. √252 is right for daily bars on a
252-day calendar and wrong for almost everything else. realized_volatility
returns per-period volatility unless you pass a factor, and periods_per_year
makes you state the calendar rather than assume one — the same minute bars are
525,600 periods a year on a continuous venue and 98,280 on a cash equity
session.
$ mdnorm features matrix.csv --returns log --zscore 60 --vol 60 \
--interval 1m --sessions-per-year 365 --session-length 24h -o feats.csv
Data quality
from mdnorm.quality import find_issues, clean
find_issues(events) # list of QualityIssue (outlier / gap / out_of_order / non_positive)
cleaned, issues = clean(events) # drop bad ticks & invalid rows, keep a report
clean removes price outliers and non-positive price/size records and returns
the surviving events plus everything it flagged.
Serialization
from mdnorm import to_records
to_records(events) # list of flat dicts (Decimals as strings)
to_records(bars, as_float=True) # numeric output for DataFrames
to_records (and event_to_dict / bar_to_dict) flatten events and bars into
plain, JSON-serialisable dicts — drop straight into pandas.DataFrame, a
csv.DictWriter, or json.dumps.
Consolidating streams
from mdnorm import merge_streams, dedupe
timeline = dedupe(merge_streams(binance_events, coinbase_events))
merge_streams interleaves multiple venue feeds into one timestamp-ordered
timeline; dedupe drops exact duplicate events left behind by reconnects and
replays.
CSV files
from mdnorm import read_csv_trades, write_records_csv
events = read_csv_trades("trades.csv", venue="coinbase") # file -> events
write_records_csv(bars, "bars.csv", as_float=True) # events/bars -> file
read_csv_trades parses a whole CSV of trades into normalized events;
write_records_csv writes events or bars back out. Standard library only.
NDJSON / JSON Lines
from mdnorm import write_jsonl, read_jsonl_events
write_jsonl(events, "events.jsonl") # one JSON object per line
events2 = read_jsonl_events("events.jsonl") # lossless round-trip
# large files: stream lazily, .gz handled transparently
for e in iter_jsonl_events("dump.jsonl.gz"):
...
Pipelines
Declare a processing chain once, reuse it everywhere:
from decimal import Decimal
from mdnorm import Pipeline
pipe = (
Pipeline()
.dedupe()
.clean(max_return=Decimal("0.1"))
.time_bars(60_000_000_000) # 1-minute bars
.fill_gaps()
)
bars = pipe.run(events)
print(pipe.last_issues) # quality report from clean()
Command line
The common conversions ship as a zero-dependency CLI:
$ mdnorm bars trades.csv --venue binance --interval 1m -o bars.csv
$ mdnorm quality trades.csv --max-gap 5m
$ mdnorm convert trades.csv -o trades.jsonl
$ mdnorm bars trades.csv --interval 1d --actions actions.csv -o adjusted.csv
$ mdnorm align BTC=btc.csv ETH=eth.csv --interval 1m --max-age 5m -o matrix.csv
$ mdnorm features matrix.csv --returns log --zscore 60 --vol 60 -o feats.csv
Also available as python -m mdnorm.
The unified schema
@dataclass(frozen=True, slots=True)
class MarketEvent:
symbol: str # canonical "BASE-QUOTE", e.g. "BTC-USD"
venue: str # source venue
event_type: EventType # TRADE | QUOTE
ts_ns: int # nanoseconds since Unix epoch (UTC)
price: Decimal | None
size: Decimal | None
side: Side | None # BUY | SELL
# ... plus bid/ask fields for quotes
Design notes
- Money is
Decimal. Prices and sizes never touch binary floats, so42000.10stays42000.10. - Time is integer nanoseconds, UTC. One comparable integer regardless of whether the source gave seconds, milliseconds, or a FIX timestamp string.
- Symbols are canonicalized to
BASE-QUOTE, with venue aliases resolved (XBT→BTC) and quote currencies detected longest-match-first soUSDTwins overUSD. - Normalizers are pure functions — one raw record in, one
MarketEventout — which keeps them trivial to unit-test and compose into any streaming or batch pipeline.
Architecture
raw feed ──► normalizer ─────────────► MarketEvent ──► your pipeline
(CSV / (from_csv_row / (unified, (research,
WS JSON / from_ws_json / immutable) backtest,
FIX) from_fix) execution)
│
├── symbols.canonical_symbol() BTCUSDT → BTC-USDT
├── timeutil.*_to_ns() any time → ns UTC
├── adjust.adjust_events() splits/divs/rolls
├── micro.infer_sides() who crossed the spread
├── book.OrderBook() deltas → live book → quotes
├── consolidate() many venues → one best bid/offer
├── evaluate() your fills vs the market
├── align() N instruments → one time grid
└── returns() / rolling_*() features, trailing windows only
Tests
pip install pytest
pytest -q
The suite includes a cross-venue equivalence test proving CSV, WebSocket and FIX representations of one trade collapse to an identical event.
License
MIT © HarvestGroup360 (AMII LTD). See LICENSE.
Maintained by HarvestGroup360 as part of our open quantitative-infrastructure tooling.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file market_data_normalizer-1.10.0.tar.gz.
File metadata
- Download URL: market_data_normalizer-1.10.0.tar.gz
- Upload date:
- Size: 95.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8f1f85ccd0562d8004e352535f9f2ee1bf56520d7d2ec25a5940e64aa2537600
|
|
| MD5 |
da140c21535dc085de8e3c9649a264ab
|
|
| BLAKE2b-256 |
08bd297b49140058ce5dc695fc10dd24f89d55cdbc69538702dc9a9ef4167bd9
|
Provenance
The following attestation bundles were made for market_data_normalizer-1.10.0.tar.gz:
Publisher:
publish.yml on Harvestgroup360/market-data-normalizer
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
market_data_normalizer-1.10.0.tar.gz -
Subject digest:
8f1f85ccd0562d8004e352535f9f2ee1bf56520d7d2ec25a5940e64aa2537600 - Sigstore transparency entry: 2500980656
- Sigstore integration time:
-
Permalink:
Harvestgroup360/market-data-normalizer@1126a7f3112ed016e568b9f47adb44f7dcae0b66 -
Branch / Tag:
refs/tags/v1.10.0 - Owner: https://github.com/Harvestgroup360
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@1126a7f3112ed016e568b9f47adb44f7dcae0b66 -
Trigger Event:
push
-
Statement type:
File details
Details for the file market_data_normalizer-1.10.0-py3-none-any.whl.
File metadata
- Download URL: market_data_normalizer-1.10.0-py3-none-any.whl
- Upload date:
- Size: 74.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
41e77c78c0c974c221b0b756a0e1d67035de14ab7d5a5147f8dc947bc8c0a79a
|
|
| MD5 |
5841b5872cb7f3dd0265e1625fbccea9
|
|
| BLAKE2b-256 |
253135f99b35426a05ca97ac48f7aab08b0678cff7de928ec993f78fa0cb7053
|
Provenance
The following attestation bundles were made for market_data_normalizer-1.10.0-py3-none-any.whl:
Publisher:
publish.yml on Harvestgroup360/market-data-normalizer
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
market_data_normalizer-1.10.0-py3-none-any.whl -
Subject digest:
41e77c78c0c974c221b0b756a0e1d67035de14ab7d5a5147f8dc947bc8c0a79a - Sigstore transparency entry: 2500980668
- Sigstore integration time:
-
Permalink:
Harvestgroup360/market-data-normalizer@1126a7f3112ed016e568b9f47adb44f7dcae0b66 -
Branch / Tag:
refs/tags/v1.10.0 - Owner: https://github.com/Harvestgroup360
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@1126a7f3112ed016e568b9f47adb44f7dcae0b66 -
Trigger Event:
push
-
Statement type: