bhavkit
Download, clean, and store historical NSE bhavcopy data in a local DuckDB database.
bhavkit is a typed Python CLI and library that turns the National Stock
Exchange's raw bhavcopy archives into a queryable, validated, analytics-ready
dataset. It handles the fiddly parts of working with NSE data — archive
downloading, format normalization, validation, deduplication, and gap tracking —
so you can spend your time analyzing markets instead of wrestling with files.
Features
- Four NSE datasets — CM (equity bhavcopy), IDX (index values), FO (futures & options), DELIV (deliverable positions) — plus the equity master (symbol metadata), normalized into a versioned schema.
- Reliable ingestion — asynchronous downloads with retries, exponential
backoff + jitter, polite rate limiting, a resumable on-disk cache, and
idempotent
INSERT OR REPLACEupserts keyed onsha256. Re-running a range is a no-op. - Built-in data-quality tooling — monthly coverage reports (Markdown + JSON), gap detection, anomaly flags (extreme moves, invalid OHLC, delivery mismatches), equity-master drift, and cross-dataset coverage.
- Query everything — a read-only interactive SQL shell with template queries, plus parquet/CSV export with symbol/date filters.
- Offline mode — run the entire pipeline against a local mirror of
pre-downloaded files (
source = "local"), no NSE dependency. - Configuration that configures itself —
bhavkit initcreates the database and a defaultbhavkit.tomlwith zero manual setup; tune via a TOML file,BHAV_*env vars, CLI flags, or thebhavkit configcommand.
Built with
Python ≥ 3.11 · Typer · DuckDB · Polars · httpx · pydantic · Rich
Table of contents
- Requirements
- Installation
- Quickstart
- Usage
- Configuration
- Data model
- Library usage
- Reliability & design notes
- Development
- Documentation
- License
Requirements
- Python 3.11+
- ~1 GB disk for a multi-year dataset (a decade of daily CM data is roughly a few hundred MB compressed in DuckDB)
Installation
As a CLI
# global CLI via uv — from a checkout, a wheel, or a git URL
uv tool install . # local checkout / wheel
uv tool install "bhavkit @ git+https://github.com/<you>/bhavkit" # from GitHub
# ephemeral run, no install
uvx --from . bhavkit --help
# or a classic pip install (Python 3.11+)
pip install .
pip install git+https://github.com/<you>/bhavkit
This puts a bhavkit executable on your path:
bhavkit --help
bhavkit init
bhavkit update --start 2024-01-01 --end 2024-01-31
From a GitHub clone
git clone https://github.com/<you>/bhavkit.git
cd bhavkit
uv sync # or: pip install .
uv run bhavkit --help
As a Python library
bhavkit is a regular package — add it to any project and import bhavkit:
uv add --editable ../bhavkit # local checkout
uv add "bhavkit @ git+https://github.com/<you>/bhavkit" # from GitHub
# (once published): uv add bhavkit
Quickstart
# 1. create the database, schema, and a default config — nothing to configure
bhavkit init
# 2. sanity-check the NSE archive endpoints
bhavkit probe
# 3. pull a month of data across all datasets
bhavkit update --start 2024-01-01 --end 2024-01-31 --datasets cm,idx,fo,deliv
# 4. refresh the equity master, then run a coverage / quality report
bhavkit master refresh
bhavkit report --month 2024-01
# 5. query the result
bhavkit query --template top_gainers
Your data now lives in bhavkit.duckdb (data/ holds the file cache and
report output). Wait for the market close (~15:45 IST) to see the day's data,
then bhavkit update for an incremental refresh.
Usage
| Command | What it does |
|---|---|
bhavkit init |
Create the DB, schema, migration records, cache dirs, and a default bhavkit.toml. |
bhavkit probe |
Test the configured NSE base URLs and report reachability. |
bhavkit update |
Download + ingest a date range (default: last 7 days). |
bhavkit backfill --start <date> |
Same engine as update, sized for large historical ranges; resumable via the cache. |
bhavkit master refresh |
Refresh the equity master from EQUITY_L.csv. |
bhavkit status |
Bar/master/gap counts + a coverage-by-month table from ingest_log. |
bhavkit report |
Data-quality report (report_<product>.md + .json); refreshes data_gaps. |
bhavkit query |
Read-only interactive SQL shell; one-shot with --sql / --template. |
bhavkit export |
Export a table to parquet or CSV with symbol/date filters. |
bhavkit holidays |
Show the advisory NSE holiday calendar. |
bhavkit config |
Inspect/edit the config file: show, set <key> <value>, unset, path, keys. |
Every command accepts --config <path>, --db-path <path>, and
--data-dir <path> to override configuration. A full reference with every flag
is in docs/cli.md.
Updating. update / backfill share one engine:
--start, --end date range (YYYY-MM-DD)
--datasets comma-separated: cm, idx, fo, deliv (default: cm)
--source nse | local (default: nse)
--local-dir mirror folder when --source local
--nse-base-url override the archive base URL
--force re-ingest even if already ingested
Querying. bhavkit query is a read-only shell (write statements are
blocked). In the shell: \q \h \d \t \r <name>, or run any read query
ending in ;. Built-in templates: latest, table_counts, top_gainers,
top_losers, volume_leaders, delivery_spike, fo_most_traded,
fo_open_interest, index_history, delivery_ratio.
bhavkit query --template top_gainers # most recent day, top 10 by % change
bhavkit query --sql "SELECT * FROM bhav_daily WHERE symbol='SBIN' LIMIT 5"
Exporting.
bhavkit export --table bhav_daily --symbols SBIN,RELIANCE --format csv --out sb.csv
bhavkit export --table fo_daily --start 2024-01-01 --end 2024-01-05 --format parquet
bhavkit export --table index_daily --indexes "Nifty 50" --format csv
Tables: bhav_daily, index_daily, fo_daily, deliverable_daily,
equity_master. Output defaults to data/exports/ with a descriptive filename.
Configuration
Precedence, highest wins: CLI flags > BHAV_* env vars > TOML file >
defaults.
bhavkit init writes a default bhavkit.toml automatically, and TOML is read
from bhavkit.toml in the project root (falling back to
~/.config/bhavkit/config.toml). There is nothing to copy by hand — edit the
file or use bhavkit config set.
bhavkit config show # resolved settings (defaults + file + env)
bhavkit config set concurrency 8
bhavkit config set source local # switch to an offline mirror
bhavkit config unset source # revert to default
| Key | Default | Description |
|---|---|---|
db_path |
bhavkit.duckdb |
DuckDB database file. |
data_dir |
data |
Cache, report, and export output root. |
source |
nse |
nse (live archive) or local (offline mirror). |
nse_base_url |
https://nsearchives.nseindia.com |
NSE archive root. |
local_dir |
— | Mirror root when source = "local". |
concurrency |
4 |
Max parallel downloads (1–32). |
retries |
5 |
Retries for 408/429/5xx, beyond the first attempt. |
backoff_base |
1.0 |
Exponential backoff base (seconds) + jitter. |
rate_limit_sleep |
0.35 |
Minimum interval between requests. |
timeout |
30.0 |
Per-request timeout (seconds). |
user_agent |
browsersish Mozilla/5.0 … bhavkit/0.1 |
HTTP User-Agent. |
verbose |
false |
Debug logging. |
Env vars use the dot-flattened form: BHAV_CONCURRENCY=8,
BHAV_NSE_BASE_URL=https://…, BHAV_SOURCE=local. A fully commented sample
lives at bhavkit.toml.example.
Sources. nse resolves against nsearchives.nseindia.com (alternate
mirrors are probed by bhavkit probe):
cm:content/historical/EQUITIES/<YYYY>/<MON>/cmDDMONYYYYbhav.csv.zipidx:content/indices/ind_close_all_DDMMYYYY.csvfo:content/historical/DERIVATIVES/<YYYY>/<MON>/foDDMONYYYYbhav.csv.zipdeliv:products/content/sec_bhavdata_full_DDMMYYYY.csv
local reads an offline mirror laid out as
<local_dir>/<product>/<YYYY>/<MMM>/<filename> (same filenames as NSE), which
lets you run the full pipeline against pre-downloaded files.
NSE can rate-limit bulk scraping — keep
rate_limit_sleep >= 0.3and prefer incrementalupdatefor daily refreshes.
Data model
The schema is versioned via schema_migrations; bhavkit init migrates older
databases in place.
| Table | Primary key | Notes |
|---|---|---|
bhav_daily |
(symbol, series, date) |
OHLC, prev_close, last, turnover (INR), traded qty, trades, delivery qty/%. |
index_daily |
(index_name, date) |
Daily index OHLC, volume, turnover. |
fo_daily |
(symbol, instrument, expiry_date, option_type, strike_price, date) |
F&O; settle_price, contracts, value, open_interest, change_in_oi. |
deliverable_daily |
(symbol, series, date) |
Traded vs delivered quantity, delivery-to-traded ratio. |
equity_master |
(symbol, series) |
Name, ISIN, industry, status from EQUITY_L.csv. |
ingest_log |
auto | Audit: product, date, url, sha256, rows, status, error, duration. |
qc_issues |
auto | Per-file validation findings (bad dates, empties, OHLC violations…). |
data_gaps |
(product, date, gap_type) |
Derived from the latest report. |
Notes: current NSE bhavcopies omit AVG_PRICE (stored NULL); F&O zero OHLC
is legitimate for illiquid strikes; both the classic (DELIV/LACS) and
current (ISIN/TIMESTAMP) CM formats are normalized to the same schema with
turnover always in rupees.
Library usage
Fine-grained control does not require the CLI — everything is a plain Python API:
from pathlib import Path
from bhavkit.config import load_config
from bhavkit.db import Database
from bhavkit.query import run_sql, execute_template
cfg = load_config()
db = Database(cfg.db_path)
db.bootstrap() # idempotent schema setup
df = run_sql(db, "SELECT symbol, close FROM bhav_daily WHERE date = '2024-01-05' ORDER BY close DESC LIMIT 5")
print(df)
runners = execute_template(db, "top_gainers") # returns a polars DataFrame
db.close()
See the executable workbook for the complete walkthrough — bootstrap,
the ingest pipeline (read_bhav_file → clean_bars → upsert_bars), template
queries, reports, parquet export, and driving the CLI as a module:
uv sync
uv run --with jupyter --with ipykernel \
jupyter nbconvert --to notebook --execute --inplace workbook/bhavkit.ipynb
Records a single-writer DuckDB database: only one process may hold the file open at a time — close any open connection before shelling out to the CLI.
Reliability & design notes
- Resumable by design. Files are cached on disk keyed by their URL; ingest
is idempotent and skipped when the
sha256is unchanged. An interruptedbackfillresumes on the next run instead of restarting. - Non-trading days. A weekday that returns
404from NSE is logged asnot_a_trading_day— the server is the authority, not the advisory calendar (which is used only for reporting). Zero-error, resumable backfills make reporting's gap table accurate. - Polite by default. Rate-limited requests, exponential backoff with jitter
on
408/429/5xx, and cleanup of partial files on failure. - Validation with context. Extreme single-day moves are usually corporate
actions (split/bonus), not bad data — the DQ report flags them rather than
hiding them. ETFs trade as
EQbut aren't in the equity master, so they surface as "master drift" for you to judge.
Development
uv sync --extra dev
uv run pytest # test suite (81 tests)
uv run ruff check . # lint
uv run mypy src # typecheck
The implementation tree lives under src/bhavkit/ — cli.py (commands),
config.py (pydantic + precedence), db.py + schema.sql (DuckDB, versioned),
sources/ (nse/local), download.py (async pool + cache), ingest.py
(per-product pipelines), clean.py (validation), report.py (DQ reports),
query.py (SQL shell + templates), export.py, metadata.py, calendar.py.
Documentation
- CLI reference — every command and flag:
docs/cli.md - Library walkthrough —
workbook/bhavkit.ipynb - Roadmap & design notes —
plan.md
License
Released under the MIT License.
Release files for bhavkit 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| bhavkit-0.1.1.tar.gz | 103.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| bhavkit-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 145.4 kB
Release files / bhavkit-0.1.1.tar.gz
| Download URL | bhavkit-0.1.1.tar.gz |
|---|---|
| Size | 103.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d4edb943704f08ec978d5b15edf4dee2eb8a394d157345cd345b88db29295636
|
|
BLAKE2b-256 checksum How to use checksums |
f2a539710b698be3d04d037837bff79392159c37d77a2fa18b689b24519d140d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.
Transparency logRelease files / bhavkit-0.1.1-py3-none-any.whl
| Download URL | bhavkit-0.1.1-py3-none-any.whl |
|---|---|
| Size | 42.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8725bffdd6cae4e977993d71ee8e42515124e9664ab8cb8ec85982ea47c8195f
|
|
BLAKE2b-256 checksum How to use checksums |
0eb0794a47f2612ba3a1b4f3f00f972bf1f89c64bd559e861cc7d85ab7a2245a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.
Transparency log