Skip to main content

Tests Coverage Status PyPI

PyEnsembl

PyEnsembl is a Python interface to Ensembl reference genome metadata such as exons and transcripts. PyEnsembl downloads GTF and FASTA files from the Ensembl FTP server and loads them into a local database. PyEnsembl can also work with custom reference data specified using user-supplied GTF and FASTA files.

Example Usage

from pyensembl import EnsemblRelease

# release 77 uses human reference genome GRCh38
data = EnsemblRelease(77)

# will return ['HLA-A']
gene_names = data.gene_names_at_locus(contig=6, position=29945884)

# get all exons associated with HLA-A
exon_ids  = data.exon_ids_of_gene_name('HLA-A')

Installation

PyEnsembl requires Python 3.9 or later. You can install PyEnsembl using pip:

pip install pyensembl

This should also install any required packages such as datacache.

Before using PyEnsembl, run the following command to download and install Ensembl data:

pyensembl install --release <list of Ensembl release numbers> --species <species-name>

For example, pyensembl install --release 75 76 --species human will download and install all human reference data from Ensembl releases 75 and 76.

To install the newest supported Ensembl release for a reference assembly:

pyensembl install --reference-name GRCh37

Reference names are case-insensitive. This selects human release 75 for GRCh37; the species is inferred from the reference. You can also specify --release to select older releases for that assembly. Conflicting --species or --release selections are rejected before downloading data. Deletion commands still require an explicit --release.

Alternatively, you can create the EnsemblRelease object from inside a Python process and call ensembl_object.download() followed by ensembl_object.index().

Annotation coverage

PyEnsembl uses Ensembl's complete chr_patch_hapl_scaff GTF for human GRCh38 from release 82, mouse GRCm38 releases 82–102, and zebrafish GRCz11 from release 92. These files include additional genes on assembly patches and haplotypes. Other assemblies and earlier releases use the standard GTF filename.

Patch and haplotype contig names are preserved, for example CHR_HG2263_PATCH. Gene-name searches can return additional genes on these contigs; use stable gene IDs or a contig filter when selecting a particular locus.

After upgrading from versions before 2.10.17, rerun installation for each affected release you use, for example pyensembl install --release 97 --species human. The complete GTF creates a separate index, so an older index cannot hide the additional genes. Existing source files and indexes are retained, and unchanged FASTA files are reused. Custom mirrors must provide the complete GTF filename; to use a deliberately restricted annotation, supply its GTF as custom data.

Development Setup

For development, install PyEnsembl in editable mode with development dependencies:

git clone https://github.com/openvax/pyensembl.git
cd pyensembl
pip install -e .[dev]

This installs the package in development mode along with tools for testing, linting, and building:

  • pytest for running tests
  • ruff for code linting
  • pytest-cov for coverage reporting
  • build for package building

Run lint and tests with:

./lint.sh
./test.sh

Most tests need Ensembl data installed first; .github/workflows/tests.yml lists the releases CI installs.

Cache Location

By default, PyEnsembl uses the platform-specific Cache folder and caches the files into the pyensembl sub-directory. You can override this default by setting the environment key PYENSEMBL_CACHE_DIR as your preferred location for caching:

export PYENSEMBL_CACHE_DIR=/custom/cache/dir

or

import os

os.environ['PYENSEMBL_CACHE_DIR'] = '/custom/cache/dir'
# ... PyEnsembl API usage

Usage tips

List installed genomes

To see the genomes for which PyEnsembl has already downloaded and indexed metadata you can run:

pyensembl list

Or equivalently do this in Python:

from pyensembl.shell import collect_all_installed_ensembl_releases
collect_all_installed_ensembl_releases()

Reference DNA sequences (optional)

Reference DNA enables intronic, intergenic, and flanking sequence queries. It is opt-in: a normal installation does not download a whole genome. Human DNA needs roughly 1 GB compressed and several GB of disk space after decompression.

from pyensembl import EnsemblRelease

release = EnsemblRelease(81, download_genome_fasta=True)
release.download_genome_fasta()  # DNA only; reuses an existing compatible cache
release.index_genome_fasta()
with release:
    bases = release.sequence("7", 117_480_000, 117_480_100)

sequence(contig, start, end) returns plus-strand, one-based inclusive bases, including for loci annotated on the minus strand. Contig names must match the FASTA exactly ("1" and "chr1" are distinct). Coordinates must be integers with 1 <= start <= end <= contig length. Missing contigs and invalid ranges raise ValueError. Unconfigured or uninstalled DNA raises MissingGenomeFastaError, a ValueError subclass whose message includes the install command. Reads never download missing remote data implicitly. The default result is uppercase; mask="raw" preserves soft-masked lowercase.

The lazy .fasta reader also supports fasta[contig][start-1:end].seq for consumers such as Varcode. .fasta is None when DNA is unconfigured or not installed. Readers already handed out stay usable after clear_cache(); close() closes them. genome_fasta_path reports the existing uncompressed file, or None. download() and index() include DNA when configured, alongside annotation, transcript, and protein files. Attached DNA does not affect equality: genes and transcripts from the same release compare equal with or without it.

# Install all data, including DNA.
pyensembl install --release 81 --with-genome-fasta

# Install DNA alone; annotation/transcript/protein data are not needed.
pyensembl install --release 81 --only-genome-fasta

# An optional subset and soft masking.
pyensembl install --release 81 --only-genome-fasta \
    --genome-fasta-type primary_assembly --masked soft

The default downloaded DNA is unmasked toplevel, covering the patch and haplotype contigs included in Ensembl annotations. primary_assembly includes chromosomes and unplaced/unlocalized sequences, but excludes patches and haplotypes. It is unavailable for some older releases and species. Mask choices are none, soft (lowercase repeats), and hard (repeats replaced with N). The corresponding Python options are genome_fasta_type and genome_fasta_mask. See Ensembl's DNA file definitions.

Attach a local FASTA

release = EnsemblRelease(81, genome_fasta_path="/data/my_reference.fa.gz")
bases = release.sequence("7", 117_480_000, 117_480_100)

# Custom annotations can also attach DNA:
from pyensembl import Genome
custom = Genome("custom", "my_annotations", genome_fasta_path_or_url="/data/reference.fa")
pyensembl install --release 81 --genome-fasta-path /data/my_reference.fa
# Add --only-genome-fasta to skip annotation installation.

Local FASTAs take precedence over canonical downloads. Plain FASTAs are read in place; gzip/BGZF files are decompressed into an uncompressed cache copy on first use. Indexes always live in PyEnsembl's cache, so read-only source directories work and user-owned files/indexes remain untouched. Full indexing warns about annotation contigs missing from a local FASTA. Matching contig names alone do not verify assembly identity. Supply the same local path when constructing subsequent Python objects; CLI metadata does not silently change constructors. Custom Genome sources also accept a FASTA URL.

Shared DNA cache and cleanup

Canonical Ensembl downloads use a readable hierarchy under pyensembl/dna_cache/:

<species>/<provider>/<reference>-<assembly-accession>/<coverage>/<masking>/fasta/<file-key>/

For example, compatible human releases 81 and 82 share a directory shaped like:

pyensembl/dna_cache/
  homo_sapiens/ftp.ensembl.org/GRCh38-GCA_000001405.18/
    toplevel/unmasked/fasta/<file-key>/
      sequence.fa
      sequence.fa.fai
      object.json
      index.json

Coverage is toplevel or primary_assembly; masking is unmasked, softmasked, or hardmasked. fasta describes the stored encoding: the cached sequence is uncompressed, even when downloaded from a .fa.gz file. The familiar reference name and versioned assembly accession distinguish assembly patches.

The 16-character file key distinguishes upstream file revisions within those semantic directories. It is a prefix of the SHA-256 of the complete identity metadata, including the Ensembl Unix checksum and compressed size. object.json retains the full readable identity; conflicting identities are rejected before reuse or overwrite. The hash covers metadata, not the FASTA bytes, and Ensembl's Unix checksums are not cryptographic hashes.

Compatible releases reuse both the FASTA and FAI through per-release JSON references. Local FASTAs and custom mirrors remain separate. If upstream identity metadata is incomplete, the assembly directory ends in -unverified and the file key uses the full source URL, keeping releases isolated.

A new release installation fetches small metadata files before deciding whether DNA can be reused. Registered releases work offline. Each release retains its references, including different installed masking/flavor choices. Downloads retry transient HTTP failures and are checked against the upstream size.

Reads take no locks and write nothing, so a fully installed and indexed cache can be read-only for other users. A download or index build locks only the DNA file it writes, so other species and releases stay usable. Registering references, removing releases, and pruning briefly lock the whole cache. Files follow your umask, as do lock files on Python 3.10+ (use umask 002 or default ACLs for a group-shared cache). dna_cache itself may be a symlink, e.g. to a larger disk. Sharing requires release caches beside dna_cache; this holds on Linux, macOS, and with PYENSEMBL_CACHE_DIR. Windows' default cache layout differs, so there DNA stays per release unless PYENSEMBL_CACHE_DIR is set.

pyensembl list --check-genome-fasta  # inspect source, presence, and index
pyensembl delete-all-files --release 81  # remove this release's references
pyensembl prune --orphan-genome-fastas --dry-run
pyensembl prune --orphan-genome-fastas

list includes DNA-only installations and shows the most recently installed DNA choice for each release. Inspection does not download or rebuild files. Deleting a release preserves shared DNA still referenced by any release. Pruning removes only cache-owned objects with no release references and skips objects being downloaded or indexed; malformed reference metadata aborts pruning, while list reports it for the affected release. delete-index-files preserves shared DNA indexes to keep other releases usable; call index_genome_fasta(overwrite=True) to rebuild one. Local source files are never pruned. Python callers can use prune_genome_fastas(dry_run=True) to inspect (path, bytes) candidates.

List supported species

To see every species PyEnsembl knows about, with its assemblies and supported Ensembl release ranges:

pyensembl available

Load genome in Python

Here's an example Python snippet that loads fly genome data from Ensembl release v100:

from pyensembl import EnsemblRelease
data = EnsemblRelease(release=100, species='drosophila_melanogaster')

Data structures

Gene

gene = data.gene_by_id(gene_id='FBgn0011747')

Transcript

transcript = gene.transcripts[0]

Protein information

transcript.protein_id
transcript.protein_sequence

Non-Ensembl Data

PyEnsembl also allows arbitrary genomes via the specification of local file paths or remote URLs to both Ensembl and non-Ensembl GTF and FASTA files. (Warning: GTF formats can vary, and handling of non-Ensembl data is still very much in development.)

For example:

from pyensembl import Genome
data = Genome(
    reference_name='GRCh38',
    annotation_name='my_genome_features',
    # annotation_version=None,
    gtf_path_or_url='/My/local/gtf/path_to_my_genome_features.gtf', # Path or URL of GTF file
    # transcript_fasta_paths_or_urls=None, # List of paths or URLs of FASTA files containing transcript sequences
    # protein_fasta_paths_or_urls=None, # List of paths or URLs of FASTA files containing protein sequences
    # cache_directory_path=None, # Where to place downloaded and cached files for this genome
)
# parse GTF and construct database of genomic features
data.index()
gene_names = data.gene_names_at_locus(contig=6, position=29945884)

API

The EnsemblRelease object has methods to let you access all possible combinations of the annotation features gene_name, gene_id, transcript_name, transcript_id, exon_id as well as the location of these genomic elements (contig, start position, end position, strand).

Genes

genes(contig=None, strand=None, biotype=None)
Returns a list of Gene objects, optionally restricted to a particular contig, strand, or gene_biotype.
genes_at_locus(contig, position, end=None, strand=None)
Returns a list of Gene objects overlapping a particular position on a contig, optionally extend into a range with the end parameter and restrict to forward or backward strand by passing strand='+' or strand='-'.
gene_by_id(gene_id)
Return a Gene object for given Ensembl gene ID (e.g. "ENSG00000068793").
gene_names(contig=None, strand=None)
Returns all gene names in the annotation database, optionally restricted to a particular contig or strand.
genes_by_name(gene_name)
Get all the unique genes with the given name (there might be multiple due to copies in the genome), return a list containing a Gene object for each distinct ID.
gene_by_protein_id(protein_id)
Find Gene associated with the given Ensembl protein ID (e.g. "ENSP00000350283")
gene_names_at_locus(contig, position, end=None, strand=None)
Names of genes overlapping with the given locus, optionally restricted by strand. (returns a list to account for overlapping genes)
gene_name_of_gene_id(gene_id)
Returns name of gene with given gene ID.
gene_name_of_transcript_id(transcript_id)
Returns name of gene associated with given transcript ID.
gene_name_of_transcript_name(transcript_name)
Returns name of gene associated with given transcript name.
gene_name_of_exon_id(exon_id)
Returns name of gene associated with given exon ID.
gene_ids(contig=None, strand=None, biotype=None)
Return all gene IDs in the annotation database, optionally restricted by chromosome name, strand, or gene_biotype.
gene_ids_of_gene_name(gene_name)
Returns all Ensembl gene IDs with the given name.
nearest_gene(contig, position, end=None, strand=None)
Returns (distance, Gene) for the gene whose locus is nearest to the position (or position..end interval) on the given contig — even when no gene overlaps. Returns (inf, None) when no candidates exist.
merged_gene_intervals(contig, strand=None)
Returns the union of all gene loci on the contig as a sorted list of non-overlapping (start, end) tuples. Adjacent intervals (end+1 == next start) are merged into one.

Transcripts

transcripts(contig=None, strand=None, biotype=None)
Returns a list of Transcript objects for all transcript entries in the Ensembl database, optionally restricted to a particular contig, strand, or transcript_biotype.
transcript_by_id(transcript_id)
Construct a Transcript object for given Ensembl transcript ID (e.g. "ENST00000369985")
transcripts_by_name(transcript_name)
Returns a list of Transcript objects for every transcript matching the given name.
transcript_names(contig=None, strand=None)
Returns all transcript names in the annotation database.
transcript_ids(contig=None, strand=None, biotype=None)
Returns all transcript IDs in the annotation database.
transcript_ids_of_gene_id(gene_id)
Return IDs of all transcripts associated with given gene ID.
transcript_ids_of_gene_name(gene_name)
Return IDs of all transcripts associated with given gene name.
transcript_ids_of_transcript_name(transcript_name)
Find all Ensembl transcript IDs with the given name.
transcript_ids_of_exon_id(exon_id)
Return IDs of all transcripts associated with given exon ID.
nearest_transcript(contig, position, end=None, strand=None)
Returns (distance, Transcript) to the closest transcript on the contig. Returns (inf, None) when no candidates exist.

Exons

exon_ids(contig=None, strand=None)
Returns a list of exon IDs in the annotation database, optionally restricted by the given chromosome and strand.
exon_by_id(exon_id)
Construct an Exon object for given Ensembl exon ID (e.g. "ENSE00001209410")
exon_ids_of_gene_id(gene_id)
Returns a list of exon IDs associated with a given gene ID.
exon_ids_of_gene_name(gene_name)
Returns a list of exon IDs associated with a given gene name.
exon_ids_of_transcript_id(transcript_id)
Returns a list of exon IDs associated with a given transcript ID.
exon_ids_of_transcript_name(transcript_name)
Returns a list of exon IDs associated with a given transcript name.

Release files for pyensembl 2.11.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyensembl 2.11.0
File Size Uploaded
pyensembl-2.11.0.tar.gz 129.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyensembl 2.11.0
File Interpreter ABI Platform
pyensembl-2.11.0-py3-none-any.whl Python 3 none any Details

Total release size: 215.7 kB

Release files / pyensembl-2.11.0.tar.gz

Download URL pyensembl-2.11.0.tar.gz
Size 129.7 kB
Tags Source
SHA-256 checksum
How to use checksums
584669210cf505abc5d561c09b1afb09ba0d24bfb7aae0efd3ffd57142fbebeb
BLAKE2b-256 checksum
How to use checksums
49aec7fbaeb94cfcbb108de5665b73500d8690051eb0024d831123c70160310c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.6

Release files / pyensembl-2.11.0-py3-none-any.whl

Download URL pyensembl-2.11.0-py3-none-any.whl
Size 86.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
870ddd73426a7ff7ca75ea972d166b9c7b458123f6e9a4790ceab158797e44ad
BLAKE2b-256 checksum
How to use checksums
138775815f8da68e297754aa835b9d6f18f34914276c85eaef9869a48ef14af1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.6

Release history Release notifications | RSS feed

2.13.0

2 release files

2.12.0

2 release files

This release

2.11.0 This release

2 release files

2.10.9

2 release files

2.10.8

2 release files

2.10.7

2 release files

2.10.6

2 release files

2.10.5

2 release files

2.10.1

2 release files

2.10.0

2 release files

2.9.8

2 release files

2.9.7

2 release files

2.9.6

2 release files

2.9.5

2 release files

2.9.4

2 release files

2.9.3

2 release files

2.9.2

2 release files

2.9.1

2 release files

2.9.0

2 release files

2.8.0

2 release files

2.7.0

2 release files

2.6.13

2 release files

2.6.7

2 release files

2.6.6

2 release files

2.6.5

2 release files

2.6.4

2 release files

2.6.2

2 release files

2.6.1

2 release files

2.6.0

2 release files

2.3.13

2 release files

2.3.12

2 release files

2.3.11

2 release files

2.3.10

2 release files

2.3.9

2 release files

2.3.8

2 release files

2.3.7

2 release files

2.3.6

2 release files

2.3.4

2 release files

2.3.3

2 release files

2.3.2

2 release files

2.3.1

2 release files

2.3.0

1 release file

2.2.9

1 release file

2.2.8

1 release file

2.2.7

1 release file

2.2.6

1 release file

2.2.5

1 release file

2.2.4

1 release file

2.2.3

1 release file

2.2.2

1 release file

2.2.1

1 release file

2.2.0

1 release file

2.1.0

1 release file

2.0.2

1 release file

2.0.1

1 release file

2.0.0

1 release file

1.9.4

1 release file

1.9.3

1 release file

1.9.2

1 release file

1.9.1

1 release file

1.9.0

1 release file

1.8.8

1 release file

1.8.7

1 release file

1.8.6

1 release file

1.8.5

1 release file

1.8.4

1 release file

1.8.3

1 release file

1.8.2

1 release file

1.8.1

1 release file

1.8.0

1 release file

1.7.5

1 release file

1.7.4

1 release file

1.7.3

1 release file

1.7.2

1 release file

1.7.1

1 release file

1.7.0

1 release file

1.6.0

1 release file

1.5.2

1 release file

1.5.0

1 release file

1.4.0

1 release file

1.3.0

1 release file

1.2.6

1 release file

1.2.4

1 release file

1.2.3

1 release file

1.2.2

1 release file

1.2.1

1 release file

1.1.0

1 release file

1.0.3

1 release file

1.0.2

1 release file

1.0.1

1 release file

1.0.0

1 release file

0.9.7

1 release file

0.9.6

1 release file

0.9.5

1 release file

0.9.4

1 release file

0.9.3

1 release file

0.9.1

1 release file

0.9.0

1 release file

0.8.14

1 release file

0.8.13

1 release file

0.8.12

1 release file

0.8.11

1 release file

0.8.10

1 release file

0.8.9

1 release file

0.8.8

1 release file

0.8.7

1 release file

0.8.5

1 release file

0.8.4

1 release file

0.8.3

1 release file

0.8.2

1 release file

0.8.1

1 release file

0.7.0

1 release file

0.6.11

1 release file

0.6.10

1 release file

0.6.9

1 release file

0.6.8

1 release file

0.6.7

1 release file

0.6.5

1 release file

0.6.4

1 release file

0.6.3

1 release file

0.6.2

1 release file

0.6.1

1 release file

0.6.0

1 release file

0.5.13

1 release file

0.5.12

1 release file

0.5.11

1 release file

0.5.10

1 release file

0.5.9

1 release file

0.5.8

1 release file

0.5.7

1 release file

0.5.4

1 release file

0.5.3

1 release file

0.5.2

1 release file

0.5.1

1 release file

0.5.0

1 release file

0.4.0

1 release file

0.3.3

1 release file

0.3

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page