histdata-fetcher
Download free historical Forex data from HistData.com straight into a pandas DataFrame — 1-minute OHLC bars or raw bid/ask ticks, for any of the ~66 instruments the site publishes.
HistData.com has no official API. This library drives the same request flow the download page itself uses, fetches the per-period zip files concurrently, unpacks and parses them, and hands back one tidy, sorted DataFrame.
from histdata_fetcher import fetch_data
result = fetch_data("EUR/USD", "2024-01-01", "2024-03-31", "1min")
print(result.data.head())
print(f"{len(result.data):,} bars -> {result.output_path}")
datetime open high low close volume
0 2024-01-01 17:00:00 1.10441 1.10448 1.10441 1.10448 0.0
1 2024-01-01 17:01:00 1.10450 1.10453 1.10444 1.10444 0.0
Install
pip install histdata-fetcher
Requires Python 3.9+. Pulls in pandas, requests, and pyarrow (for the default Parquet output).
Usage
Fetch data
from histdata_fetcher import fetch_data
result = fetch_data(
pair="EURUSD", # "EUR/USD", "eur-usd" etc. all work
start_date="2024-01-01", # str, datetime.date, or datetime.datetime
end_date="2024-06-30", # inclusive
timeframe="1min", # "1min" (M1 bars) or "tick" (raw bid/ask)
output_format="parquet", # "parquet", "csv", or None to skip writing
output_path=None, # defaults to ./<PAIR>_<tf>_<start>_<end>.parquet
max_workers=8, # zip files downloaded concurrently
)
fetch_data returns a FetchResult:
| Attribute | Meaning |
|---|---|
.data |
the combined pandas.DataFrame, sorted by datetime |
.output_path |
Path written, or None if nothing was written |
.fetched_periods |
period labels that downloaded, e.g. ["2024", "2025-01"] |
.failed_periods |
FailedPeriod records (label, start, end, reason) |
.ok |
True when .data is non-empty |
Periods the site has no data for are reported, not raised — a gap in the middle of a long range will not abort the whole pull:
result = fetch_data("XAUUSD", "2005-01-01", "2024-12-31", "1min", output_format=None)
for f in result.failed_periods:
print(f"{f.period_label}: {f.reason}")
To work purely in memory, pass output_format=None.
List available instruments
from histdata_fetcher import get_available_pairs
catalog = get_available_pairs("1min", resolve_end_date=False) # one HTTP request
print(len(catalog)) # 66
print(catalog["EURUSD"].start_date) # 2000-05-01
resolve_end_date=True (the default) additionally resolves each pair's most recent published period, which costs one request per pair. Use resolve_end_date=False when you only need the pair list and start dates.
Data notes
These are properties of HistData's data, not of this client — worth knowing before you build on it.
Timestamps are EST without DST. Per HistData's FAQ, every timestamp is Eastern Standard Time (UTC−5) year-round, with no daylight-savings shift. This library leaves them tz-naive, exactly as published. Localize them yourself if you need UTC:
df["datetime"] = df["datetime"].dt.tz_localize("Etc/GMT+5").dt.tz_convert("UTC")
Volume is always 0. HistData does not publish volume for forex/CFD data. The column is kept so the schema matches the source files.
Ticks share timestamps, and are not de-duplicated. Tick timestamps are at best millisecond-resolution, so genuinely distinct quotes routinely land on the same timestamp. Worse, the resolution is not stable over time — measured on EURUSD, about 4% of rows in June 2026 share a timestamp, rising to ~50% in July and August 2026, where HistData publishes whole-second timestamps (milliseconds always 000). There is no unique key, so tick rows are returned exactly as published, in published order; treating datetime as unique will silently throw away real market data. 1-minute bars are de-duplicated on datetime, since one bar per minute is a true unique key.
File granularity differs by timeframe. 1-minute data is served as one zip per year for elapsed years and one per month for the current year; tick data is monthly only. The client works this out for you — a request that spans both simply produces a mix, visible in .fetched_periods.
Schemas
timeframe="1min"
| column | dtype |
|---|---|
datetime |
datetime64 (EST, tz-naive) |
open / high / low / close |
float64 |
volume |
float64 (always 0) |
timeframe="tick"
| column | dtype |
|---|---|
datetime |
datetime64 (EST, tz-naive, millisecond resolution) |
bid / ask |
float64 |
volume |
float64 (always 0) |
Sizing your requests
Tick data is large: one month of EURUSD ticks is roughly 1.4 million rows (~9 MB compressed). Pulling several years of ticks in one call will hold all of it in memory before writing. For big historical pulls, loop a year at a time and write each to its own file.
Errors
| Exception | Raised when |
|---|---|
PairNotAvailableError |
the pair isn't offered for that timeframe |
PeriodUnavailableError |
a single period failed (caught internally; surfaces via .failed_periods) |
HistDataError |
base class for the above; also raised if the site layout can no longer be parsed |
ValueError |
bad arguments — unknown timeframe, start > end, range entirely before the pair's first data |
A start_date earlier than the pair's first published month is clamped forward, and an end_date in the future is clamped to today; both log a warning.
Logging
import logging
logging.basicConfig(level=logging.INFO)
logging.getLogger("histdata_fetcher").setLevel(logging.DEBUG)
Development
git clone https://github.com/Njenjo/histdata-fetcher
cd histdata-fetcher
pip install -e ".[dev]"
pytest
pytest runs the offline suite against an in-process fake of the site — no network needed. The live end-to-end tests are opt-in:
pytest -m network
Stability
This library scrapes an HTML download flow rather than a documented API, so a redesign of histdata.com can break it. Parsing failures raise HistDataError with a clear message rather than returning silently wrong data. If the pair list or download form stops parsing, please open an issue.
Legal
The data belongs to HistData.com and is provided under their terms of use — free for personal and educational use, with redistribution restrictions. This library is an unaffiliated client that automates the public download flow; you are responsible for using it within those terms. Please keep max_workers modest and don't hammer the site.
License
MIT — see LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file histdata_fetcher-0.1.0.tar.gz.
File metadata
- Download URL: histdata_fetcher-0.1.0.tar.gz
- Upload date:
- Size: 21.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cbbcfc5eebfb0cec7d6263f4f1acb6b53880ea9669cb895534ac4660cd22b96d
|
|
| MD5 |
c7e4e850c8c63e623fb3d130765590b3
|
|
| BLAKE2b-256 |
59e852aad5de52b5be26da5f8d7e265e25bf3d7ca8e76cf521caa0f8ac4ac0af
|
File details
Details for the file histdata_fetcher-0.1.0-py3-none-any.whl.
File metadata
- Download URL: histdata_fetcher-0.1.0-py3-none-any.whl
- Upload date:
- Size: 15.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ba5f6b342aae08737e60e1972c9a5d9ae0152b8c815bedab646d722027a5b43c
|
|
| MD5 |
9032886d5f32c43ef3f01936ae0aafa2
|
|
| BLAKE2b-256 |
bd82880dca35169cf60d5e87478145af302b121d0021166351ca7864c3f1565d
|