cotdata
A local, file-based data layer for futures prices and CFTC Commitments of Traders (COT) positioning.
cotdata separates fetching data (a "producer" that talks to vendors) from using it (any number of "consumers" that just read Parquet through a small, stable API). Point every tool at one synced store, and none of them ever call a vendor SDK at runtime — so the same data feeds your research, backtests, and dashboards identically, on any OS.
- One store, many readers. Consumers
import cotdataand read; they never touch a vendor SDK. Swapping a data vendor is a producer-only change. - Free COT, optional paid prices. CFTC Commitments of Traders data (1986–present) downloads free from cftc.gov on any OS. Futures prices/specs come from Norgate (paid, Windows) and are optional.
- Cross-platform reads. Produce on Windows (for Norgate); read anywhere (Mac/Linux/Windows), offline.
- Predecessor stitching.
get_cot()transparently stitches migrated CFTC codes (e.g. the Russell 2000) and rescales tick-size changes (e.g. Lumber) into one continuous series. - Atomic writes. Read the store safely even while the producer is downloading and writing.
- New-data signal. Every run writes a structured
status.jsonso downstream tools can poll one file to detect fresh data.
Data sources at a glance
| Data | Source | Cost | Runs on |
|---|---|---|---|
| CFTC COT — legacy / disaggregated / TFF | cftc.gov | Free | any OS |
| Futures prices + contract specs | Norgate Data | Paid subscription | Windows (producer only) |
| Futures prices, back-adjusted | Databento GLBX.MDP3 | Paid per query | any OS |
| Futures prices, research-grade fallback | Yahoo Finance | Free | any OS |
| Reading the store (any of the above) | — | Free | any OS |
Contents
- Quickstart · How it works · Reading data · Producing data · Windows setup · Scheduling on Windows · Scheduling on Linux · Syncing the store · Operations · Concepts & design · COT vintage tracking · Reference: schemas · Reference: COT formats · Diagnostics · Development · Contributing · License
Quickstart
The fastest zero-cost path uses free CFTC COT data — no account, any OS:
pip install cotdata
export COTDATA_STORE=~/cotdata_store # where the shared store lives
cotdata-update --cot-legacy # free CFTC download (first run pulls history; cached after)
python -c "import cotdata; print(cotdata.get_cot('ES').tail())"
That downloads the CFTC Legacy COT history and reads the S&P 500 (ES) positioning back out:
Open_Interest_All Comm_Positions_Long_All Comm_Positions_Short_All NonComm_Positions_Long_All NonComm_Positions_Short_All
Report_Date_as_MM_DD_YYYY
2026-06-23 1980254 1444102 1531232 251385 286833
2026-06-30 1967167 1422155 1509889 249934 287526
2026-07-07 1969636 1435736 1502199 244103 286994
Futures prices additionally require a Norgate subscription on Windows — see Producing data.
How it works
The store is the API boundary — not Python imports. Producers write Parquet + manifest.json; consumers only read. Nobody touches a vendor SDK at app runtime, so swapping a vendor is a producer-only change.
PRODUCER — runs where each source is reachable
Norgate export (Windows) CFTC COT download (any OS)
│ │
└──────────────┬───────────────┘
▼ write parquet + manifest
┌────────────────────────────────────────────────────────────┐
│ CANONICAL STORE ($COTDATA_STORE) │
│ prices/ cot_legacy/ cot_disagg/ cot_tff/ │
│ metadata/ manifest.json status.json │
└────────────────────────────────────────────────────────────┘
│ read (offline, any OS)
┌──────────────┴───────────────┐
▼ ▼
your signal research your backtest / dashboards
both just: import cotdata · store synced via rsync / Dropbox / S3
The store layout:
prices/{symbol}_{adjustment}.parquet— Open/High/Low/Close/Volume/Open Interest, tz-naiveDateindex.adjustment∈ {backadj,unadj} on disk;propadjis a third view derived on read (not stored). Close = exchange settlement.cot_legacy/{symbol}_{code}.parquet— weekly CFTC Legacy positioning.cot_disagg/{symbol}_{code}.parquet— weekly CFTC Disaggregated positioning.cot_tff/{symbol}_{code}.parquet— weekly CFTC Traders in Financial Futures positioning.metadata/contract_specs.parquet— Norgate contract specifications (tick size, point value, margin).manifest.json— per-tablelast_date,n_rows,source,updated_at,schema_version.status.json— machine-readable new-data signal for downstream tools (see Operations).vintage/— optional as-published (vintage) capture: retained raw CFTC downloads plus change-only observations and field-level revisions. Purely additive; the tables above are unchanged whether or not it is enabled. See COT vintage tracking.
Reading data (consumer)
Set COTDATA_STORE to the synced store directory, then:
import cotdata
# Prices — pick the adjustment that matches your use:
signals = cotdata.get_prices("ES", adjustment="backadj") # signals + stops (gap-free rolls)
sizing = cotdata.get_prices("ES", adjustment="unadj") # position sizing (true dollar prices)
milk = cotdata.get_prices("DC", adjustment="propadj") # ratio-adjusted: strictly positive, %-return preserving
# COT — three CFTC report families:
legacy = cotdata.get_cot("ES", report="legacy") # Commercial / Non-Commercial
disagg = cotdata.get_cot("ES", report="disagg") # Managed Money, Swap Dealers, ... (commodities)
tff = cotdata.get_cot("ES", report="tff") # Leveraged Funds, Asset Managers, ... (financials)
A price frame (get_prices("ES", adjustment="backadj").tail(3)):
Open High Low Close Volume Open Interest
Date
2026-07-10 7587.25 7628.75 7552.75 7620.25 1078031.0 1966297.0
2026-07-13 7607.00 7615.25 7547.25 7563.00 1274520.0 1945908.0
2026-07-14 7557.00 7613.75 7531.50 7591.25 1139735.0 0.0
Predecessor stitching & scaling: get_cot() doesn't just read a file — it stitches historical CFTC codes for contracts that migrated exchanges (e.g. the Russell 2000) and rescales data for contracts that changed tick sizes (e.g. Lumber), so downstream models see one clean, continuous asset.
Producing data (producer)
Run on the machine that can reach the source. Norgate prices require Windows, CFTC COT runs anywhere, and a server without Norgate can build prices from Databento instead (see Cross-platform prices without Norgate).
COTDATA_STORE=/store cotdata-update --prices # Norgate prices, ALL registry symbols (Windows)
COTDATA_STORE=/store cotdata-update --prices --symbols ES NQ # ...or a subset
COTDATA_STORE=/store cotdata-update --metadata # Norgate contract specs (Windows)
COTDATA_STORE=/store cotdata-update --cot-legacy # CFTC Legacy (any OS)
COTDATA_STORE=/store cotdata-update --cot-disagg # CFTC Disaggregated (any OS)
COTDATA_STORE=/store cotdata-update --cot-tff # CFTC Traders in Financial Futures (any OS)
COTDATA_STORE=/store cotdata-update --cot-all # all three CFTC COT reports
COTDATA_STORE=/store cotdata-vintage fetch # optional: capture as-published COT (any OS)
--prices with no --symbols updates every symbol in the registry; add --symbols to scope it. Each run prints a per-symbol line with the date advance (e.g. ES: … [2026-07-13 -> 2026-07-14]) and a summary footer (OK/failed counts, rows written, elapsed, newest date). A run exits non-zero if a fetch hard-fails (Norgate/CFTC unreachable), so a scheduler can retry — see Scheduling on Windows.
Installation for the producer
pip install "cotdata[norgate]" # adds the norgatedata dependency (Windows)
The norgatedata package talks locally to the Norgate Data Updater application — there are no API keys. You just need the Updater installed, authenticated, and running.
Cross-platform prices without Norgate (databento)
A server that cannot run Norgate (for example the public dashboard host) can build the price store from Databento instead. One provider owns each symbol end to end, so this is a full replacement, not a blend. Install the extras and set the environment:
pip install "cotdata[databento,yahoo]" # databento producer + the Yahoo fallback
export COTDATA_STORE=/path/to/store # the store the dashboard reads
export COTDATA_PRICE_SOURCE=databento # deployment default, so softs/MSCI fall to Yahoo
export DATABENTO_API_KEY=db-... # for the paid ingest step only
# optional: export COTDATA_DATABENTO_RAW=/path/to/raw # defaults to $COTDATA_STORE/_raw/databento
Then build the store in order (ingest before build):
cotdata-update --ingest-databento # Stage 1 (PAID): raw .n.0/.n.1 ohlcv-1d + statistics -> raw store
cotdata-update --build-databento # Stage 2 (FREE): additive back-adjustment -> $COTDATA_STORE/prices
cotdata-update --prices-yahoo # softs, lumber, MSCI proxies (resolve to Yahoo on this deployment)
cotdata-update --cot-all # CFTC COT, the dashboard needs it too
cotdata-update --check # coverage, newest dates, staleness
- Two stages, one paid. Stage 1 is the only step that hits the API. It writes an append-only raw store and resumes from the last fetched date, so re-runs pull only new days. Stage 2 reads that raw store with no API cost, so the back-adjustment can be iterated offline. The raw store is producer-internal, so keep it out of any sync to consumers.
- History starts 2010-06-06 (the GLBX floor), shallower than Norgate. Markets not on CME Globex (ICE softs, lumber, MSCI intl) fall back to Yahoo.
- First-run check. A healthy symbol prints
built unadj+backadj (N bars, K rolls). If it printsno rolls detected, back-adjustment is a no-op for that symbol, so investigate before trusting it. - Validate against Norgate (optional gate) with
scripts/validate_databento_vs_norgate.pyif you have both stores. - Schedule the two price commands nightly and
--cot-allweekly — see Scheduling on Linux.
Producer halves: one host, one job
cotdata has two producers by design: the CFTC downloader (free, any OS) and the price
producer (Norgate needs Windows). Two entry points scope a host to one of them:
cotdata-cot --cot-all # CFTC half, any OS
cotdata-prices --prices --metadata --require-final # price half, Windows for Norgate
cotdata-update ... # both, for a single-machine deployment
Each scoped entry point refuses the other half's flags, so a price box cannot quietly
become a second COT producer racing the first. --check and --reconcile are read-only
and work from either.
Each half also owns its own manifest (manifests/cot.json, manifests/prices.json).
The legacy top-level manifest.json held both halves in ONE file, which was unsafe two
ways: the update is a read-modify-write, so two producers lose each other's entries, and
a file-level sync between two stores resolves it last-writer-wins and silently discards
one side. The per-half files are disjoint, so both problems go away.
Nothing writes manifest.json any more. Migrate a store once:
cotdata-update --migrate-manifests
Idempotent, and it never touches data. Until you run it, a domain missing from the
per-half files is still read from the aggregate with a warning. Delete manifest.json
once every consumer of that store is on this version. See ADR-0007.
Scheduling on Windows (Task Scheduler)
Full setup, including wrapper scripts, the three-task layout (daily prices, daily COT catch-up, Friday release-window poller), --require-final event-driven pricing, restart-on-failure retry settings, and Norgate/Task-Scheduler troubleshooting (notably: NDU needs an interactive session), is in docs/WINDOWS_SCHEDULING.md. Start with the Windows Setup Guide first if Python/the venv/COTDATA_STORE aren't configured yet.
The short version: prices fire once daily near the Norgate Continuous Futures Final (~8:55pm ET), COT gets a daily morning catch-up plus a tight Friday-afternoon poll around its ~3:30pm ET release, and every task uses restart-on-failure so idempotent, cheap re-runs absorb both transient errors and "not published yet."
Scheduling on Linux (cron)
Full setup, including wrapper scripts, the crontab entries (nightly prices, daily COT catch-up, Friday release-window poller), flock overlap protection, and troubleshooting (cron's bare environment, timezone conversion, DATABENTO_API_KEY not being picked up), is in docs/LINUX_SCHEDULING.md.
The short version: a databento server schedules the same way as the Windows/Norgate producer — prices nightly, COT gets a daily morning catch-up plus a tight Friday-afternoon poll around its ~3:30pm ET release, all idempotent and safe to over-run.
Syncing the store between machines
Norgate needs Windows, so a research Mac or a Linux dashboard is usually a read-only replica of a store produced elsewhere. Prefer one producer writing everything and a strictly one-directional sync.
Two directories must be excluded for size: _cache/ (cotdata's cache of downloaded
CFTC source zips) and _raw/ (the paid databento raw store) are producer-internal
and together are ~70% of the bytes. The legacy manifest.json should be excluded too.
Anything a consumer put in the store by hand is a correctness issue rather than a saving. No producer creates it, so a mirroring sync deletes it. Exclude it, but the real fix is to keep it out of the store: the store belongs to its producer.
The same rule bites hardest on vintage/ if you enable it, because that data cannot be
re-fetched: capture it on the producer so it syncs outward, or keep it outside the
mirrored store with COTDATA_VINTAGE_ROOT. Its provenance index is deliberately named
snapshots.json, since the usual manifest.json exclusion matches by name at any depth
and would otherwise strip it in transit, delivering raw archives with no index.
Consumer cloud sync (Dropbox, Google Drive) is a poor fit here: conflict copies land
inside the store, and on-demand placeholder files break read_parquet on the machine
doing research.
Full guidance, exclusion table and example scripts: docs/SYNCING.md.
Operations
Read-only and maintenance commands, all cross-platform (they work off the store, no network):
cotdata-update --check # store status: row counts, newest data, staleness
cotdata-update --reconcile # prune stale manifest entries (see below)
--check reports per-domain row counts, newest data date, last write, and any entries lagging behind their peers (a partial-run signal):
domain entries rows newest data last write (UTC) behind
prices 84 829,096 2026-07-14 2026-07-15T10:15:24Z 1d
cot_legacy 44 70,201 2026-07-07 2026-07-14T04:26:55Z 8d
...
✓ all entries current (none lag behind their domain's newest).
status.json — new-data signal for downstream tools
Every producer run writes $COTDATA_STORE/status.json (atomically, beside the data), so tools that trigger on fresh data poll one small structured file instead of scanning the store:
{
"generated_at": "2026-07-15T10:15:24Z",
"schema_version": 2,
"newest_data": { "prices": "2026-07-14", "cot_legacy": "2026-07-07", "cot_disagg": "2026-07-07", "cot_tff": "2026-07-07" },
"domains": { "prices": { "newest_data": "2026-07-14", "last_write": "2026-07-15T10:15:24Z", "entries": 84, "rows": 829096, "lagging": 0 }, "...": {} },
"last_run": { "kinds": ["prices"], "ok": ["ES", "..."], "symbols_failed": [], "rows": 1658000, "seconds": 88, "at": "2026-07-15T10:15:24Z" }
}
Polling contract:
- To detect new data, compare
newest_data.<domain>(e.g.newest_data.prices,newest_data.cot_legacy) against your last-seen value. It advances only when genuinely new daily data arrives — a no-op run leaves it unchanged. - To detect that a run happened at all (new data or not), use
generated_at. last_runcarries the most recent run's outcome (which domains, per-symbol failures) for alerting.
Prices and each COT report are separate domains, so a price-triggered tool and a COT-triggered tool each watch their own key.
--reconcile — manifest hygiene
COT tables are stored per code as {symbol}_{code} (e.g. RTY_23977A), so a symbol's current and predecessor (hist_codes) contracts are both attributable to it. --reconcile drops manifest entries whose parquet file is missing — bare-code ghosts and retired domains left by older naming schemes — so --check and status.json show only real, consistently-named entries. It never touches data (only removes bookkeeping for files that don't exist).
Concepts & design
Back-adjusted vs unadjusted prices
Futures contracts expire, forcing traders to "roll" into the next contract, which usually trades at a slightly different price. Simply stitching contracts together creates artificial price gaps, so cotdata stores two series and derives a third:
backadj(signals & stops). Gap-free arithmetic (additive) rolls shift historical prices to align with the new contract, preserving absolute daily point moves. Use this for indicators, signals, and stop-losses to avoid false triggers on rollover gaps.unadj(position sizing). Back-adjustment shifts historical prices (sometimes negative), so you can't use it for dollar values. Useunadj(raw, real-life prices) for that day to compute true dollar risk and contract counts.propadj(proportional / ratio adjustment — strictly positive). Derived on read fromunadj+backadj; preserves daily percentage returns and never goes non-positive. Use it for low-priced, long-history contracts where additive back-adjustment accumulates roll gaps below zero and breaks price-based stops and R-multiples. See Class III Milk (DC) below.
Why propadj exists — Class III Milk (DC)
Norgate publishes continuous futures in only two forms: unadjusted and additive back-adjusted (_CCB) — there is no native ratio-adjusted series. Additive adjustment subtracts each roll's calendar spread from all prior history, and for a low-priced, seasonal, ~29-year contract like DC (Class III Milk, ~$15–20/cwt) those gaps accumulate past zero: 46.7% of DC_backadj closes are ≤ 0 (range −9.83 to 23.09). A price-based stop, an R-multiple, or a percentage return is meaningless on a non-positive series, so CMR cannot use DC's backadj at all — even though DC is the flagship new-asset-class (Dairy) held-out generalization market.
propadj salvages it. Because the additive series B and unadjusted series U differ by an offset O = B − U that steps only at rolls, each roll's calendar spread is recoverable (s = O[r−1] − O[r]) and convertible to a multiplicative roll ratio k = (U[r−1] + s)/U[r−1]. Scaling each historical segment by the cumulative product of k (most-recent segment anchored to actual prices) yields a series that is strictly positive over the full 1997–2026 history (DC range 4.68–25.01), preserves within-segment percentage returns exactly, and is sign-identical to backadj on every day including rolls. It is a pure function of two already-stored series, so it needs no producer re-run — get_prices("DC", adjustment="propadj") works today. Recommendation: CMR reads DC (and any similarly low-priced contract) with adjustment="propadj". Restricting DC to its positive-price era (2011→present, ~15y) or dropping it were the fallbacks; neither is needed.
Providers & authentication
Which vendor prices a symbol is a deployment choice, not a fixed fact (the same ES is Norgate for local research and databento on a public-dash server). It is resolved when a producer runs, from three inputs: the deployment default COTDATA_PRICE_SOURCE (norgate if unset), per-symbol capability (the norgate / databento / yahoo mappings in the registry, null where a vendor has no series), and an optional per-symbol price_source override. A symbol uses its override if set, otherwise the default when that vendor can serve it, otherwise a Yahoo fallback where a ticker exists. Each producer writes only the symbols that resolve to it, so a symbol is never blended across vendors. One provider owns each symbol end to end.
- Norgate Data (paid, Windows). No Python API key. The
norgatedatapackage talks locally to the Norgate Data Updater app, which must be installed, authenticated, and running on Windows. This is the default for local research (cotdata-update --prices). - Databento (paid, cross-platform). A two-stage producer for a server that cannot run Norgate. Stage 1 (
cotdata-update --ingest-databento) pulls raw.n.0/.n.1ohlcv-1dandstatisticsinto an append-only raw store ($COTDATA_DATABENTO_RAW, else_raw/databentounder the store). This is the only paid step, and it is resumable, so re-runs fetch only new dates. Stage 2 (cotdata-update --build-databento) derives the back-adjusted prices from the raw store with no API cost, so the build logic can be iterated offline. SetDATABENTO_API_KEYandCOTDATA_PRICE_SOURCE=databento. History starts 2010-06-06, and markets not on CME Globex (ICE softs, lumber, MSCI intl) fall back to Yahoo. The raw store is producer-internal, so exclude it from any consumer sync. - Yahoo Finance (free, research-grade).
cotdata-update --prices-yahooprices the markets that resolve to yfinance on this deployment: the MSCI ETF proxies always, plus the softs and lumber on a databento server. Expect gaps and silent revisions. This is not a production replacement for the paid feeds. Requires the[yahoo]extra.
The symbol registry
The supported futures contracts are defined in a YAML registry, so adding a market needs no code:
- Add a market: edit
src/cotdata/registry.yamlunder its asset class. The registry handles metadata likeis_equityand predecessorhist_codes. - Centralize it: set
COTDATA_REGISTRYto a sharedregistry.yaml(e.g. inside$COTDATA_STORE) so producer and consumers use identical asset definitions without agit pull.
Atomic store
The store uses atomic writes (write-temp-then-rename). Consumers can safely query via get_prices / get_cot even while cotdata-update is actively downloading and writing.
COT vintage tracking (as-published history)
CFTC revises COT data after publication — most consequentially through trader reclassification, which moves positions between categories retroactively. Because downstream signals are rolling z-scores and percentiles against years of history, a restatement silently rewrites the baseline every historical reading was computed against. There is precedent: in July 2008 the Commission revised reports back to July 3, 2007.
CFTC serves current state only. There is no vintage archive and no as-published endpoint, so vintage data can only be accumulated going forward — every uncaptured week is a permanent blind spot in the part of the series most likely to have been revised.
This is opt-in and purely additive: if you never run it, the store behaves exactly as
before. Enabling it adds a vintage/ subtree.
cotdata-vintage fetch # capture current + prior year + weekly static (daily)
cotdata-vintage fetch --all # every year 1986-present (see below)
cotdata-vintage ingest --pending # parse retained raw -> observations + revisions
cotdata-vintage diff --since 2026-01-01 # field-level revisions, with revision depth
cotdata-vintage asof --as-of 2026-07-24T18:00:00 --report-date 2026-07-21
cotdata-vintage flow --market 088691 # weekly flow decomposition (see below)
cotdata-schedule sync # CFTC Special Announcements
cotdata-schedule published # true publication dates from retained weekly statics
cotdata-schedule backfill # resolve release_date + its provenance
How it works:
-
Immutable landing zone. Every fetch is recorded (including 304s) and raw bytes are retained permanently under
vintage/raw/, written atomically and never rewritten. A byte-identical regeneration is deduped — a changed download is not itself a revision. -
All three reports. Legacy, Disaggregated and TFF canonicalise into one long schema. Disagg and TFF also populate per-category spreading, per-category trader counts and CR4/CR8 concentration, none of which the Legacy file carries, and they are where Managed Money and Leveraged Funds live. A suppressed trader count (CFTC writes
.) canonicalises to null rather than to a string. -
Change-only observations. A row is written only when its value hash differs from the latest for its natural key
(report_date, market_code, report_type, combined, category), so storage grows with actual revisions rather than with time. -
Field-level revisions carry
age_days(revision depth): whether revisions stay in recent weeks or reach back into the calibration window determines how much the rest of a system has to care. -
Point-in-time reads.
asof(t)returns each key's latest value observed at or beforet, reconstructing what was actually knowable then. -
Release dates with provenance.
report_dateis stored exactly as reported (never normalized to Tuesday), andrelease_dateis resolved throughpublished > observed > announced > scheduled > derived, with the source recorded — a release date without provenance is worse than none, since indexing onreport_dateembeds a lookahead (three days normally, weeks during a backlog).publishedis the weekly static's HTTPLast-Modified, a true publication timestamp; it is forward-only (that file holds one week and is overwritten), so weeks predating capture fall back down the chain.announcedcomes from the republication tables CFTC posts on the Special Announcements page after a disruption, and is what puts the Oct–Dec 2025 appropriations-lapse backlog on its real dates: 36,296 stored rows thatderivedotherwise places up to 47 days early, before the lapse that stopped them being published at all. -
Flow decomposition.
cotdata-vintage flowlabels each week's ΔLong versus ΔShort asnew_longs/short_covering/new_shorts/long_liquidation, by dominant leg, with no parameters to tune. A rally driven by short covering has a finite fuel supply and one driven by fresh longs does not, and they are indistinguishable on a price chart. It also emitsdays_elapsed, because COT was fortnightly until 1992-10-13 and holidays shift the rest, so a "weekly" change is not always weekly.
Run capture on the producer, not a replica, and schedule it daily: nearly every request returns 304, so a daily run is close to free while catching holiday-shifted and backlog releases with no schedule logic.
The frozen-year tripwire. CFTC regenerates a rolling two-year window (current plus the
immediately-prior year) and nothing older. The prior year is therefore re-served every week
but byte-identical, which is the one place a content check on closed data comes free.
It is in the default fetch set for that reason: it costs one roughly 7 MB transfer per week
and zero bytes on disk, and it is the only automated retroactive-restatement detector here.
Anything other than unchanged bytes (deduped) on it raises an alert, which ingest
re-raises as a non-zero exit and a REVISIONS_<date>.txt marker file. --no-prior-year
turns it off. --all extends the same check to every year, but since nothing older is ever
re-served it mostly confirms 304s; monthly or quarterly is the right cadence for that. Full
design notes, including the measured CFTC caching behaviour, are in
docs/design/cot_vintage.md.
Replica warning. The vintage tree must not be written on a machine whose store is mirrored (
robocopy /MIR,rsync --delete) from a producer: the mirror deletes destination-only files and the data is irreplaceable. Capture on the producer, or setCOTDATA_VINTAGE_ROOTto a path outside the mirrored store. See docs/SYNCING.md.
Local development
uv venv # create .venv
uv pip install -e . # install cotdata + deps
export COTDATA_STORE=/path/to/synced/store # the shared store
uv run pytest # run the tests
On the Windows producer, install the Norgate extra with uv pip install -e ".[norgate]" (tested on Python 3.10, within Norgate's supported versions). Use uv run <cmd>, or activate with source .venv/bin/activate (Mac/Linux) / .venv\Scripts\activate (Windows).
Reference: Data schemas
The canonical store uses standard Parquet files. Loaded with pd.read_parquet(), they conform to the following schemas.
Price Data (prices/{symbol}_{adjustment}.parquet)
Primary price history (Norgate Data), indexed by tz-naive Date. The pipeline downloads both the back-adjusted (backadj) series for signals/stops and the unadjusted (unadj) series for true transaction-cost modeling.
Reading reconstructed volume: the reconstruction columns below are internal storage. Consumers should not read Volume_Reconstructed directly — call get_prices(symbol, volume="reconstructed") and the Volume column is served as reconstructed-with-per-row-raw-fallback, plus a Volume_Source column for audit. The default volume="front" returns the front-month series unchanged (byte-identical to the pre-v2 API).
Schema versioning: schema_version in manifest.json records the on-disk data version (v2 = reconstructed volume promoted). Consumers key cache invalidation on cotdata.schema_version() and can guard with cotdata.require_schema(min_version).
| Column | Type | Description |
|---|---|---|
Date |
DatetimeIndex | Trading day (tz-naive, normalized to midnight). |
Open |
float | Opening price. |
High |
float | High price. |
Low |
float | Low price. |
Close |
float | Settlement Close price. |
Volume |
float | Continuous contract trading volume (front-month only). |
Open Interest |
float | Continuous contract open interest. |
Volume_Reconstructed |
float | True market volume (sum of First and Second contract). Differs from raw Volume by symbol — typically higher for products whose rolls spread volume across contracts, but roughly equal or lower for symbols with a near-empty back month (e.g. crypto). Not a drop-in replacement. |
Volume_Source |
string | reconstructed if First+Second available, raw fallback if not. |
FirstVolume / SecondVolume |
float | Trading volume of the specific first and second expiring contracts. |
FirstContract / SecondContract |
string | Contract names for the first and second expirations (e.g., ES-2024H). |
Delivery Month |
float | Expiration month of the active contract (e.g. 202609). Used to detect contract rolls. |
Contract Specifications (metadata/contract_specs.parquet)
Contract metadata (Norgate Data), used for exact point-value risk sizing and transaction cost models.
| Column | Type | Description |
|---|---|---|
Symbol |
string | Internal ticker symbol (e.g., ES). |
Norgate_Symbol |
string | Raw Norgate symbol used to query the API (e.g., &ES_CCB). |
Name |
string | Full name of the contract. |
Exchange |
string | Name of the listing exchange. |
Group |
string | Norgate asset classification group. |
Contract Size |
float | Size multiplier (e.g., $50 for ES). Also called Point Value. |
Tick Size |
float | Minimum price fluctuation (e.g., 0.25 for ES). |
Tick Value |
float | Dollar value of one tick (Tick Size * Contract Size). |
Point Value |
float | Same as Contract Size. |
Currency |
string | Base currency of the contract. |
Margin |
float | Initial margin requirement (if provided by Norgate). |
COT Legacy Data (cot_legacy/{symbol}_{code}.parquet)
Legacy positioning data (CFTC Legacy Futures Report). History starts in 1986. Indexed by tz-naive Report_Date_as_MM_DD_YYYY.
[!NOTE] Legacy Reports: broken down by exchange, with futures-only and combined futures-and-options variants. Legacy classifies reportable open interest into non-commercial and commercial traders. The
cotdatapipeline strictly downloads the Futures-only reports (https://www.cftc.gov/files/dea/history/dea_fut_xls_{YEAR}.zip).
[!NOTE] Column Subset: The raw CFTC
.xlsfiles contain well over 100 columns; the pipeline keeps the focused 15-column subset below to keep files small. To include more, add the exact CFTC column name toTARGET_COLSinsrc/cotdata/providers/cftc.py.
| Column | Type | Description |
|---|---|---|
Report_Date_as_MM_DD_YYYY |
DatetimeIndex | Reporting date (typically Tuesday). |
Market_and_Exchange_Names |
string | Name of the contract and exchange. |
CFTC_Contract_Market_Code |
string | 6-digit CFTC contract code. |
Open_Interest_All |
float | Total open interest for the contract. |
Comm_Positions_Long_All |
float | Commercial Long positions. |
Comm_Positions_Short_All |
float | Commercial Short positions. |
NonComm_Positions_Long_All |
float | Non-Commercial (Large Speculator) Long positions. |
NonComm_Positions_Short_All |
float | Non-Commercial (Large Speculator) Short positions. |
NonRept_Positions_Long_All |
float | Non-Reportable (Small Speculator) Long positions. |
NonRept_Positions_Short_All |
float | Non-Reportable (Small Speculator) Short positions. |
Traders_Tot_All |
float | Total number of reportable traders. |
Traders_Comm_Long_All |
float | Number of Commercial Long traders. |
Traders_Comm_Short_All |
float | Number of Commercial Short traders. |
Traders_NonComm_Long_All |
float | Number of Non-Commercial Long traders. |
Traders_NonComm_Short_All |
float | Number of Non-Commercial Short traders. |
COT Disaggregated Data (cot_disagg/{symbol}_{code}.parquet)
Entity-specific positioning and trader counts (CFTC Disaggregated Futures-Only Report). History starts in 2006. Indexed by tz-naive Report_Date_as_MM_DD_YYYY.
[!NOTE] Lossless Image: Unlike the filtered Legacy schema, the Disaggregated parquets are a lossless image of the source CFTC
txtfiles — all granular entity groups (Money Manager, Swap Dealer, Producer/Merchant, Other Reportable) and theirTraders_*counts. Required for computing Position Size and Clustering metrics.
COT Traders in Financial Futures (TFF) Data (cot_tff/{symbol}_{code}.parquet)
Entity-specific positioning and trader counts for financial markets (CFTC TFF Futures-Only Report). History starts in 2006. Indexed by tz-naive Report_Date_as_MM_DD_YYYY.
[!NOTE] Financials Counterpart: TFF is the exact counterpart to Disaggregated, used for financial markets (Equities, FX, Rates), which have no Disaggregated report.
[!NOTE] Lossless Image: Like Disaggregated, TFF parquets are a lossless image of the source CFTC
txtfiles — the financial entity groups (Dealer,Asset_Mgr,Lev_Money,Other_Rept) and theirTraders_*counts.
Reference: COT formats explained
The CFTC publishes positioning data in three formats; cotdata manages all three for complete coverage and the deepest history.
- Legacy (1986–Present) — all markets. Divides traders into Commercial (hedgers) and Non-Commercial (large speculators). The only format with pre-2006 data, so it's essential for long-term backtesting.
- Disaggregated / DIS (2006–Present) — physical commodities only (Agriculture, Energy, Metals). Splits traders into Producer/Merchant, Swap Dealers, Managed Money, and Other Reportables — a clearer view of "smart money" (Managed Money) in commodities.
- Traders in Financial Futures / TFF (2006–Present) — financial markets only (Equities, Rates, Currencies). Splits traders into Dealer/Intermediary, Asset Manager, Leveraged Funds, and Other Reportables — the definitive source for speculative flow (Leveraged Funds) in financials.
Diagnostics
Verify your Norgate subscription and configuration with the included smoke test, on the Windows producer:
python tests/test_adjustment.py
It checks: (1) Local communication — Python can reach the Norgate Data Updater; (2) Subscription access — your subscription includes the required CME futures package; (3) Roll-gap validation — proves whether the Updater is returning back-adjusted (gap-free) vs unadjusted continuous contracts, by hunting for calendar-spread gaps at roll dates. Gap-free data is vital for accurate stop-loss modeling.
Ecosystem
cotdata is the data layer of a small, unbundled toolchain — it stops at "clean data behind a stable API" on purpose. What you do with that data is a separate, swappable step:
- cotdata (this package) — the data layer. One synced store of futures prices and CFTC COT positioning; many readers, no vendor SDK at read time.
- crucible — the edge layer. Feed a signal built on cotdata frames into crucible and it tells you — with a confidence interval and a p-value — whether the trade-level edge is real, before you open a funded account.
The flow runs one direction: cotdata (data) → your signal → crucible
(edge). Neither imports the other, so cotdata stays useful on its own for any
COT/futures research — crucible is just the most common thing to point at it next.
Development
Want to contribute or work on cotdata locally? See CONTRIBUTING.md for:
- Virtual environment setup with
uvor standardpip - Running the test suite
- Platform-specific notes (Norgate is Windows-only; CFTC parsing runs anywhere)
- Code style guidelines
Contributing
Issues and pull requests are welcome. Please see CONTRIBUTING.md for setup, tests, and conventions. When filing a bug, include your OS — Norgate features require Windows, while store reads and CFTC COT run anywhere.
License
Released under the MIT License — see LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file cotdata-0.3.0.tar.gz.
File metadata
- Download URL: cotdata-0.3.0.tar.gz
- Upload date:
- Size: 201.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
193a30ed5af86b5cae6043c8891fafc3d6c0590dc32eaec11fff164e0a2c9e9b
|
|
| MD5 |
227559195fc8fa2b8978479cf2598d12
|
|
| BLAKE2b-256 |
45db99b872689c79b5e01f449b35ac3fff2dc9e0e8613c66783986dcf71dc547
|
Provenance
The following attestation bundles were made for cotdata-0.3.0.tar.gz:
Publisher:
release.yml on mspinola/cotdata
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
cotdata-0.3.0.tar.gz -
Subject digest:
193a30ed5af86b5cae6043c8891fafc3d6c0590dc32eaec11fff164e0a2c9e9b - Sigstore transparency entry: 2329430655
- Sigstore integration time:
-
Permalink:
mspinola/cotdata@452e6c9cde108807b56788a0ac0cab887c5c3eab -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/mspinola
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@452e6c9cde108807b56788a0ac0cab887c5c3eab -
Trigger Event:
push
-
Statement type:
File details
Details for the file cotdata-0.3.0-py3-none-any.whl.
File metadata
- Download URL: cotdata-0.3.0-py3-none-any.whl
- Upload date:
- Size: 127.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7324c420173ef195806a3c9c6e74032ee23744337e935cda685f0365a8355030
|
|
| MD5 |
82ff6652f2910412db8b282f465fd8b6
|
|
| BLAKE2b-256 |
4d69eff298c2610bf7fa8bd177e91993cae83c89b91357ed18053a24c7e977c0
|
Provenance
The following attestation bundles were made for cotdata-0.3.0-py3-none-any.whl:
Publisher:
release.yml on mspinola/cotdata
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
cotdata-0.3.0-py3-none-any.whl -
Subject digest:
7324c420173ef195806a3c9c6e74032ee23744337e935cda685f0365a8355030 - Sigstore transparency entry: 2329430686
- Sigstore integration time:
-
Permalink:
mspinola/cotdata@452e6c9cde108807b56788a0ac0cab887c5c3eab -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/mspinola
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@452e6c9cde108807b56788a0ac0cab887c5c3eab -
Trigger Event:
push
-
Statement type: