Skip to main content

histdata-fetcher

Download free historical Forex data from HistData.com straight into a pandas DataFrame — 1-minute OHLC bars or raw bid/ask ticks, for any of the ~66 instruments the site publishes.

HistData.com has no official API. This library drives the same request flow the download page itself uses, fetches the per-period zip files concurrently, unpacks and parses them, and hands back one tidy, sorted DataFrame.

from histdata_fetcher import fetch_data

result = fetch_data("EUR/USD", "2024-01-01", "2024-03-31", "1min")

print(result.data.head())
print(f"{len(result.data):,} bars -> {result.output_path}")
             datetime     open     high      low    close  volume
0 2024-01-01 17:00:00  1.10441  1.10448  1.10441  1.10448     0.0
1 2024-01-01 17:01:00  1.10450  1.10453  1.10444  1.10444     0.0

Install

pip install histdata-fetcher

Requires Python 3.9+. Pulls in pandas, requests, and pyarrow (for the default Parquet output).

Usage

Fetch data

from histdata_fetcher import fetch_data

result = fetch_data(
    pair="EURUSD",           # "EUR/USD", "eur-usd" etc. all work
    start_date="2024-01-01", # str, datetime.date, or datetime.datetime
    end_date="2024-06-30",   # inclusive
    timeframe="1min",        # "1min" (M1 bars) or "tick" (raw bid/ask)
    output_format="parquet", # "parquet", "csv", or None to skip writing
    output_path=None,        # defaults to ./<PAIR>_<tf>_<start>_<end>.parquet
    max_workers=8,           # zip files downloaded concurrently
)

fetch_data returns a FetchResult:

Attribute Meaning
.data the combined pandas.DataFrame, sorted by datetime
.output_path Path written, or None if nothing was written
.fetched_periods period labels that downloaded, e.g. ["2024", "2025-01"]
.failed_periods FailedPeriod records (label, start, end, reason)
.ok True when .data is non-empty

Periods the site has no data for are reported, not raised — a gap in the middle of a long range will not abort the whole pull:

result = fetch_data("XAUUSD", "2005-01-01", "2024-12-31", "1min", output_format=None)
for f in result.failed_periods:
    print(f"{f.period_label}: {f.reason}")

To work purely in memory, pass output_format=None.

List available instruments

from histdata_fetcher import get_available_pairs

catalog = get_available_pairs("1min", resolve_end_date=False)  # one HTTP request
print(len(catalog))                       # 66
print(catalog["EURUSD"].start_date)       # 2000-05-01

resolve_end_date=True (the default) additionally resolves each pair's most recent published period, which costs one request per pair. Use resolve_end_date=False when you only need the pair list and start dates.

Data notes

These are properties of HistData's data, not of this client — worth knowing before you build on it.

Timestamps are EST without DST. Per HistData's FAQ, every timestamp is Eastern Standard Time (UTC−5) year-round, with no daylight-savings shift. This library leaves them tz-naive, exactly as published. Localize them yourself if you need UTC:

df["datetime"] = df["datetime"].dt.tz_localize("Etc/GMT+5").dt.tz_convert("UTC")

Volume is always 0. HistData does not publish volume for forex/CFD data. The column is kept so the schema matches the source files.

Ticks share timestamps, and are not de-duplicated. Tick timestamps are at best millisecond-resolution, so genuinely distinct quotes routinely land on the same timestamp. Worse, the resolution is not stable over time — measured on EURUSD, about 4% of rows in June 2026 share a timestamp, rising to ~50% in July and August 2026, where HistData publishes whole-second timestamps (milliseconds always 000). There is no unique key, so tick rows are returned exactly as published, in published order; treating datetime as unique will silently throw away real market data. 1-minute bars are de-duplicated on datetime, since one bar per minute is a true unique key.

File granularity differs by timeframe. 1-minute data is served as one zip per year for elapsed years and one per month for the current year; tick data is monthly only. The client works this out for you — a request that spans both simply produces a mix, visible in .fetched_periods.

Schemas

timeframe="1min"

column dtype
datetime datetime64 (EST, tz-naive)
open / high / low / close float64
volume float64 (always 0)

timeframe="tick"

column dtype
datetime datetime64 (EST, tz-naive, millisecond resolution)
bid / ask float64
volume float64 (always 0)

Sizing your requests

Tick data is large: one month of EURUSD ticks is roughly 1.4 million rows (~9 MB compressed). Pulling several years of ticks in one call will hold all of it in memory before writing. For big historical pulls, loop a year at a time and write each to its own file.

Errors

Exception Raised when
PairNotAvailableError the pair isn't offered for that timeframe
PeriodUnavailableError a single period failed (caught internally; surfaces via .failed_periods)
HistDataError base class for the above; also raised if the site layout can no longer be parsed
ValueError bad arguments — unknown timeframe, start > end, range entirely before the pair's first data

A start_date earlier than the pair's first published month is clamped forward, and an end_date in the future is clamped to today; both log a warning.

Logging

import logging
logging.basicConfig(level=logging.INFO)
logging.getLogger("histdata_fetcher").setLevel(logging.DEBUG)

Development

git clone https://github.com/Njenjo/histdata-fetcher
cd histdata-fetcher
pip install -e ".[dev]"
pytest

pytest runs the offline suite against an in-process fake of the site — no network needed. The live end-to-end tests are opt-in:

pytest -m network

Stability

This library scrapes an HTML download flow rather than a documented API, so a redesign of histdata.com can break it. Parsing failures raise HistDataError with a clear message rather than returning silently wrong data. If the pair list or download form stops parsing, please open an issue.

Legal

The data belongs to HistData.com and is provided under their terms of use — free for personal and educational use, with redistribution restrictions. This library is an unaffiliated client that automates the public download flow; you are responsible for using it within those terms. Please keep max_workers modest and don't hammer the site.

License

MIT — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

histdata_fetcher-0.1.0.tar.gz (21.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

histdata_fetcher-0.1.0-py3-none-any.whl (15.9 kB view details)

Uploaded Python 3

File details

Details for the file histdata_fetcher-0.1.0.tar.gz.

File metadata

  • Download URL: histdata_fetcher-0.1.0.tar.gz
  • Upload date:
  • Size: 21.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.11.0

File hashes

Hashes for histdata_fetcher-0.1.0.tar.gz
Algorithm Hash digest
SHA256 cbbcfc5eebfb0cec7d6263f4f1acb6b53880ea9669cb895534ac4660cd22b96d
MD5 c7e4e850c8c63e623fb3d130765590b3
BLAKE2b-256 59e852aad5de52b5be26da5f8d7e265e25bf3d7ca8e76cf521caa0f8ac4ac0af

See more details on using hashes here.

File details

Details for the file histdata_fetcher-0.1.0-py3-none-any.whl.

File metadata

File hashes

Hashes for histdata_fetcher-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ba5f6b342aae08737e60e1972c9a5d9ae0152b8c815bedab646d722027a5b43c
MD5 9032886d5f32c43ef3f01936ae0aafa2
BLAKE2b-256 bd82880dca35169cf60d5e87478145af302b121d0021166351ca7864c3f1565d

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page