Skip to main content

pyAlexS3

OpenAlex S3 → DuckDB loader powered by rich progress bars.

Reads OpenAlex NDJSON dumps directly from S3 via DuckDB's httpfs extension — no downloading required.

Features

  • 🚀 Direct S3 reads via DuckDB httpfs — no local downloads
  • 🦆 Zero-setup DuckDB loading via read_json_auto(...)
  • 🎯 Filter by date range (YYYY-MM-DD) and by part numbers
  • 🔁 Resume from a specific date and part after a failure
  • 🔎 Optional SQL-style WHERE predicate
  • 📊 Optional rich progress bar showing batch progress

Installation

pip install pyalexs3

or with uv:

uv add pyalexs3

Python 3.10+ is required.

Quick Start

from pyalexs3.core import OpenAlexS3Processor

p = OpenAlexS3Processor(n_workers=4)

for file_batch, rel in p.lazy_load(
    obj_type="works",
    start_date="2025-01-01",
    end_date="2025-03-01",
    columns=["id", "title", "publication_year"],
):
    df = rel.df()
    print(df.head())

Filter with WHERE clause

for file_batch, rel in p.lazy_load(
    obj_type="works",
    start_date="2025-01-01",
    end_date="2025-03-01",
    columns=["id", "title", "publication_year"],
    where_clause="title IS NOT NULL AND language='en'",
):
    df = rel.df()

Resume After Failure

If your pipeline fails midway, resume from a specific date and part number:

for file_batch, rel in p.lazy_load(
    obj_type="works",
    start_date="2025-01-01",
    end_date="2025-03-01",
    resume_from="2025-01-15/5",  # skip everything before 2025-01-15 part 5
):
    df = rel.df()

Load Specific Parts Only

for file_batch, rel in p.lazy_load(
    obj_type="works",
    start_date="2025-01-01",
    end_date="2025-01-01",
    parts=[0, 1, 2],  # only load part_000.gz, part_001.gz, part_002.gz
):
    df = rel.df()

Show Progress

p = OpenAlexS3Processor(n_workers=4, show_progress=True)

for file_batch, rel in p.lazy_load(obj_type="works"):
    df = rel.df()

Track Which Files Were Processed

Each lazy_load iteration yields both the file batch and the relation:

for file_batch, rel in p.lazy_load(obj_type="works"):
    print(f"Processing: {file_batch}")  # list of S3 keys in this batch
    df = rel.df()

API

OpenAlexS3Processor(n_workers=4, **kwargs)

Parameter Type Default Description
n_workers int 4 DuckDB thread count
show_progress bool False Show rich progress bar
pragma_show_progress bool False Enable DuckDB internal progress bar

lazy_load(...) -> Generator[tuple[list[str], DuckDBPyRelation], None, None]

Parameter Type Default Description
obj_type str required OpenAlex object type e.g. works, authors
data_type str parquet OpenAlex S3 bucket type e.g. parquet, jsonl
columns list[str] | None None Columns to select. None = all
limit int | None None Max records per batch
start_date str | None 2016-06-24 Start of date range YYYY-mm-dd (inclusive)
end_date str | None today End of date range YYYY-mm-dd (inclusive)
parts list[int] | None None Specific part numbers to load. None = all
where_clause str | None None SQL filter. Do not include WHERE keyword
resume_from str | None None Resume from YYYY-mm-dd/<part> e.g. 2025-01-15/5
batch_size int 10 Number of S3 files per batch

Yields tuple[list[str], duckdb.DuckDBPyRelation]:

  • list[str] — S3 keys in this batch (useful for progress tracking)
  • DuckDBPyRelation — query the batch with .df(), .arrow(), .fetchall()

Supported Object Types

works, authors, sources, institutions, topics, keywords, publishers, funders, concepts

Behavior & Notes

  • No downloads — data is read directly from S3 via DuckDB httpfs. No temp files, no cleanup needed.
  • DuckDB — installs and loads httpfs automatically on init. Sets PRAGMA threads to n_workers.
  • Object cache — PRAGMA enable_object_cache=true is set by default for repeated queries on the same files.
  • S3 auth — OpenAlex S3 is public. No credentials needed.

Testing

Dev dependencies include pytest.

uv sync --extra dev
uv run pytest -q

Tests mock the S3 client directly using unittest.mock to test the file listing and filtering logic without hitting real S3.

Development

  • Source layout: src/pyalexs3/
  • Typed package marker: src/pyalexs3/py.typed

License

MIT © EurekAI

Citation

If you are using this for research purposes please use this BibTeX for citation:

@misc{pyalexs32025,
    author = {Adityam Ghosh},
    title = {pyalexs3},
    howpublished = {\url{https://github.com/EurekAI-Org/pyalexs3}},
    year = {2025},
    note = {[Accessed 09-10-2025]},
}

Metadata

Release files for pyalexs3 0.1.9

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyalexs3 0.1.9
File Size Uploaded
pyalexs3-0.1.9.tar.gz 10.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyalexs3 0.1.9
File Interpreter ABI Platform
pyalexs3-0.1.9-py3-none-any.whl Python 3 none any Details

Total release size: 19.9 kB

Release files / pyalexs3-0.1.9.tar.gz

Download URL pyalexs3-0.1.9.tar.gz
Size 10.0 kB
Tags Source
SHA-256 checksum
How to use checksums
a716d95a11b6e2c0101ae489396fd27563a1ab42de4ad9054dc4ea452cb89a94
BLAKE2b-256 checksum
How to use checksums
05534f68fb2758c5f49ec56dce6f26f70a0c4aa41399b1827a2335cb5f6eab46
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / pyalexs3-0.1.9-py3-none-any.whl

Download URL pyalexs3-0.1.9-py3-none-any.whl
Size 9.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
52798f2328e27e1ed7025a7d718b8a1af285df2b7914cbb43a1de3ee7ac7a040
BLAKE2b-256 checksum
How to use checksums
0cd9179f917e2164a9b488f7bb0dc2a3010b8a21fd4e42e9fc77aafdf67fc339
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.1.9 This release

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page