Skip to main content

recount3

The Clear BSD License with Extra Clause Current library version Supported Python versions CodeCov GitHub project

recount3 is a typed Python library and command-line tool for the recount3 data repository, a uniformly processed collection of RNA-seq studies spanning tens of thousands of human and mouse samples from SRA, GTEx, and TCGA. It discovers, downloads, and assembles recount3 resources into analysis-ready objects. These resources include gene, exon, and junction count matrices, sample metadata, genome annotations, and BigWig coverage files.

The package provides two interfaces:

  • A Python library that assembles count matrices, sample metadata, and genomic coordinates into BiocPy SummarizedExperiment and RangedSummarizedExperiment objects, with recount3-compatible scaling and normalization utilities (approximate read counts, AUC- and mapped-reads-based scaling, and TPM).

  • A command-line tool (recount3) that implements a discover -> manifest -> materialize workflow for scripts and pipelines. It emits JSONL/TSV manifests and materializes resources to a directory or a .zip archive with parallel downloads.

Contents

Installation

The core package supports Python 3.10 through 3.15 and depends on NumPy, pandas, and SciPy. Five optional extras enable additional features:

python3 -m pip install recount3                # core
python3 -m pip install "recount3[biocpy]"      # + SummarizedExperiment builders
python3 -m pip install "recount3[bigwig]"      # + BigWig coverage access
python3 -m pip install "recount3[parquet]"     # + .parquet output
python3 -m pip install "recount3[anndata]"     # + .h5ad output
python3 -m pip install "recount3[pybiocfilecache]"  # shared R/Python cache
python3 -m pip install "recount3[all]"         # every optional feature
  • biocpy (biocframe, genomicranges, summarizedexperiment) is required for create_rse and every helper that returns or operates on a BiocPy object.

  • bigwig (pyBigWig) is required only for BigWig coverage access.

  • parquet (pyarrow) is required only to write .parquet output. pandas accepts either pyarrow or fastparquet; an existing fastparquet installation is used as-is.

  • anndata (anndata, delayedarray) is required only to write .h5ad output, and implies biocpy.

  • pybiocfilecache enables opt-in R BiocFileCache interoperability. The default URL-hash cache is ~3x faster and needs no extra. Select sharing with RECOUNT3_CACHE_BACKEND=pybiocfilecache; it defaults to R’s recount3 cache directory (normally ~/.cache/R/recount3 on Linux/WSL). Explicit cache directories still take precedence. See “Cache and configuration” in the tutorial for concurrency, refresh semantics, maintenance, and measured overhead.

all installs every optional feature listed above. The dev and docs extras hold the test and documentation toolchains and are installed separately.

On Windows, substitute py -3 -m pip install .... Upgrade an existing installation with python3 -m pip install --upgrade recount3.

Note: The optional extras have platform constraints. The bigwig extra can be difficult or impossible to install on Windows and macOS. The biocpy extra can be difficult or impossible to install on Windows, and anndata inherits that constraint because it requires biocpy. On Python 3.15, anndata (and therefore all) cannot be installed yet: its h5py dependency publishes no Python 3.15 wheels, and building it from source requires the HDF5 C library. The parquet and pybiocfilecache extras install cleanly on every supported platform. The core package and the command-line workflow do not depend on any extra and work on all supported platforms.

The three-layer API

recount3 exposes the same workflow at three levels of abstraction, so a single project can be assembled in one call while multi-project or custom workflows retain full control:

  • High level. create_rse() builds one project into a RangedSummarizedExperiment and performs discovery, downloading, metadata merging, and range assembly in a single call.

  • Mid level. R3ResourceBundle is a filterable container of resources for combining multiple projects, selecting subsets, and stacking matrices.

  • Low level. R3Resource represents a single file and manages its URL, cache entry, and parser.

See the Tutorial for a complete walkthrough.

Quickstart

Python API

Assemble a project into a RangedSummarizedExperiment (requires the recount3[biocpy] extra):

>>> import recount3 as r3
>>>
>>> rse = r3.create_rse(
...     project="SRP009615",
...     organism="human",
...     annotation_label="gencode_v26",
... )
>>> rse.shape
(63856, 12)

For multi-project or custom workflows, use the bundle layer to filter resources and stack matrices directly:

>>> import recount3 as r3
>>>
>>> bundle = r3.R3ResourceBundle.discover(
...     organism="human",
...     data_source="sra",
...     project="SRP009615",
... )
>>> print(f"Found {len(bundle.resources)} resources.")
Found 10 resources.
>>>
>>> gene_counts = bundle.filter(
...     resource_type="count_files_gene_or_exon",
...     genomic_unit="gene",
...     annotation_extension="G026",
... ).stack_count_matrices(compat="feature")
>>> gene_counts.shape
(63856, 12)

Command-line tool

Discover resources, write a JSONL manifest, and download in parallel:

# Search for gene-level count files and write a manifest.
recount3 search gene-exon \
    organism=human data_source=sra genomic_unit=gene project=SRP009615 \
    --format=jsonl > manifest.jsonl

# Materialize all resources from the manifest (8 parallel jobs).
recount3 download --from=manifest.jsonl --dest=./downloads --jobs=8

Because both subcommands operate on JSONL via standard streams, search and download compose into a single pipeline without an intermediate file:

recount3 search annotations \
    organism=human genomic_unit=gene annotation_extension=G026 \
    --format=jsonl | \
recount3 download --from=- --dest=./annotations

The bundle subcommands assemble analysis-ready outputs without a Python session. Supported outputs are a stacked count matrix (TSV, gzip-compressed TSV, or Parquet) and a pickled SummarizedExperiment or RangedSummarizedExperiment:

recount3 bundle rse --from=manifest.jsonl --genomic-unit=gene --out=rse.pkl

Note: Read the full documentation on Pages for the complete API reference, the CLI guide, and worked examples.

Data mirrors

recount3 publishes the same relative file layout on several interchangeable public mirrors. recount3 targets this layout rather than any single host, so selecting a different mirror requires only a change to the base URL (the RECOUNT3_URL environment variable, the --base-url CLI flag, or the base_url field of recount3.config.Config):

Mirror

Base URL

Duffel load balancer (default)

http://duffel.rail.bio/recount3/

AWS Open Data

https://recount-opendata.s3.amazonaws.com/recount3/release/

JHU IDIES (Dataverse)

https://data.idies.jhu.edu/recount3/data/

Dependencies

Core (installed automatically):

numpy>=2.0
pandas>=2.2
scipy>=1.13

Optional: BiocPy integration (recount3[biocpy]):

biocframe>=0.7
genomicranges>=0.8
summarizedexperiment>=0.7.1

Optional: BigWig support (recount3[bigwig]):

pybigwig>=0.3.18

Optional: Parquet output (recount3[parquet]):

pyarrow>=23.0.1

Optional: AnnData output (recount3[anndata], implies biocpy):

anndata>=0.11
delayedarray>=0.5

Optional: shared R/Python cache (recount3[pybiocfilecache]):

pybiocfilecache>=0.7.0,<0.8

Questions, Feature Requests, and Bug Reports

Please submit questions, feature requests, and bug reports on Issues.

License

This package is distributed under The Clear BSD License with Extra Clause.

Citation

When using the recount3 Python package in published work, please cite:

  • Alexander Alsalihi, Robert M Flight, and Hunter N.B. Moseley. “The recount3 Python package for programmatic access to uniformly processed RNA-seq data.” bioRxiv (2026). doi: 10.64898/2026.06.17.732943

If your work uses the recount3 data resource, please also cite the original recount3 publication:

  • Wilks, C., Zheng, S.C., Chen, F.Y. et al. “recount3: summaries and queries for large-scale RNA-seq expression and splicing.” Genome Biology 22, 323 (2021). doi: 10.1186/s13059-021-02533-6

Metadata

Release files for recount3 1.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for recount3 1.3.0
File Size Uploaded
recount3-1.3.0.tar.gz 236.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for recount3 1.3.0
File Interpreter ABI Platform
recount3-1.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 371.6 kB

Release files / recount3-1.3.0.tar.gz

Download URL recount3-1.3.0.tar.gz
Size 236.9 kB
Tags Source
SHA-256 checksum
How to use checksums
ad35da48e0fd2c98d3176312641290d8e91144bf812613ca901034c6118ac833
BLAKE2b-256 checksum
How to use checksums
4f30241105b090af995d0aabbd97e3c8d41426d0edc9c6ab1b551f6f623d1c2e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / recount3-1.3.0-py3-none-any.whl

Download URL recount3-1.3.0-py3-none-any.whl
Size 134.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f83029cf7a2670c46347786d82d89df73b6008751d05de9daf4adac3e247277e
BLAKE2b-256 checksum
How to use checksums
c472752465f0779e13668d0f56a4a7d1e1c8b59a6df771a6a33abae09267584e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

1.3.0 This release

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page