Skip to main content

Tests Coverage Status PyPI

PyEnsembl

PyEnsembl is a Python interface to Ensembl reference genome metadata such as exons and transcripts. PyEnsembl downloads GTF and FASTA files from the Ensembl FTP server and loads them into a local database. PyEnsembl can also work with custom reference data specified using user-supplied GTF and FASTA files.

Example Usage

from pyensembl import EnsemblRelease

# release 77 uses human reference genome GRCh38
data = EnsemblRelease(77)

# will return ['HLA-A']
gene_names = data.gene_names_at_locus(contig=6, position=29945884)

# get all exons associated with HLA-A
exon_ids  = data.exon_ids_of_gene_name('HLA-A')

Installation

PyEnsembl requires Python 3.9 or later. You can install PyEnsembl using pip:

pip install pyensembl

This should also install any required packages such as datacache.

Before using PyEnsembl, run the following command to download and install Ensembl data:

pyensembl install --release <list of Ensembl release numbers> --species <species-name>

For example, pyensembl install --release 75 76 --species human will download and install all human reference data from Ensembl releases 75 and 76.

To install the newest supported Ensembl release for a reference assembly:

pyensembl install --reference-name GRCh37

Reference names are case-insensitive. This selects human release 75 for GRCh37; the species is inferred from the reference. You can also specify --release to select older releases for that assembly. Conflicting --species or --release selections are rejected before downloading data. Deletion commands still require an explicit --release.

Alternatively, you can create the EnsemblRelease object from inside a Python process and call ensembl_object.download() followed by ensembl_object.index().

Whole-genome reference DNA, for intronic and intergenic sequence, is optional and not installed by default; see Reference DNA.

Annotation coverage

PyEnsembl uses Ensembl's complete chr_patch_hapl_scaff GTF for human GRCh38 from release 82, mouse GRCm38 releases 82–102, and zebrafish GRCz11 from release 92. These files include additional genes on assembly patches and haplotypes. Other assemblies and earlier releases use the standard GTF filename.

Patch and haplotype contig names are preserved, for example CHR_HG2263_PATCH. Gene-name searches can return additional genes on these contigs; use stable gene IDs or a contig filter when selecting a particular locus.

After upgrading from versions before 2.10.17, rerun installation for each affected release you use, for example pyensembl install --release 97 --species human. The complete GTF creates a separate index, so an older index cannot hide the additional genes. Existing source files and indexes are retained, and unchanged FASTA files are reused. Custom mirrors must provide the complete GTF filename; to use a deliberately restricted annotation, supply its GTF as custom data.

Development Setup

For development, install PyEnsembl in editable mode with development dependencies:

git clone https://github.com/openvax/pyensembl.git
cd pyensembl
pip install -e .[dev]

This installs the package in development mode along with tools for testing, linting, and building:

  • pytest for running tests
  • ruff for code linting
  • pytest-cov for coverage reporting
  • build for package building

Run lint and tests with:

./lint.sh
./test.sh

Most tests need Ensembl data installed first; .github/workflows/tests.yml lists the releases CI installs.

Species assembly ranges are checked against Ensembl's archive. After raising MAX_ENSEMBL_RELEASE, recheck every assembly boundary on the live FTP servers:

PYENSEMBL_NETWORK_TESTS=1 ./test.sh tests/test_species_assemblies.py

Timed benchmarks are opt-in because wall-clock limits depend on the machine and its load. Run them on an otherwise idle machine:

PYENSEMBL_BENCHMARKS=1 ./test.sh tests/test_timings.py -s

Cache Location

By default, PyEnsembl uses the platform-specific Cache folder and caches the files into the pyensembl sub-directory. You can override this default by setting the environment key PYENSEMBL_CACHE_DIR as your preferred location for caching:

export PYENSEMBL_CACHE_DIR=/custom/cache/dir

or

import os

os.environ['PYENSEMBL_CACHE_DIR'] = '/custom/cache/dir'
# ... PyEnsembl API usage

Usage tips

List installed genomes

To see which genomes are in the local cache, whether each is ready to use, and any reference DNA:

pyensembl list
Species  Assembly  Release  Annotation   Reference DNA      Location
human    GRCh38    81       indexed      toplevel, indexed  ~/Library/Caches/pyensembl/GRCh38/ensembl81
human    GRCh38    82       not indexed  -                  ~/Library/Caches/pyensembl/GRCh38/ensembl82
custom   GRCm38    mine1    indexed      -                  ~/Library/Caches/pyensembl/GRCm38/mine1

Annotation is indexed when everything is downloaded and indexed, so queries need no network access or setup. not indexed means the files are downloaded but the first query would spend minutes indexing them, and incomplete means some downloads are missing. Run pyensembl install for that release to finish (add --species for non-human genomes; custom genomes need their original install options). Custom genomes are listed on Linux and macOS, or wherever PYENSEMBL_CACHE_DIR is set.

install prints progress on stderr, one line per step; add --verbose (-v) to see every download and database step.

To get the installed Ensembl releases in Python:

from pyensembl.shell import collect_all_installed_ensembl_releases
collect_all_installed_ensembl_releases()

List supported species

To see every species PyEnsembl knows about, with its assemblies and supported Ensembl release ranges:

pyensembl available

Load genome in Python

Here's an example Python snippet that loads fly genome data from Ensembl release v100:

from pyensembl import EnsemblRelease
data = EnsemblRelease(release=100, species='drosophila_melanogaster')

Data structures

Gene

gene = data.gene_by_id(gene_id='FBgn0011747')

Transcript

transcript = gene.transcripts[0]

Protein information

transcript.protein_id
transcript.protein_sequence

Non-Ensembl Data

PyEnsembl also allows arbitrary genomes via the specification of local file paths or remote URLs to both Ensembl and non-Ensembl GTF and FASTA files. (Warning: GTF formats can vary, and handling of non-Ensembl data is still very much in development.)

For example:

from pyensembl import Genome
data = Genome(
    reference_name='GRCh38',
    annotation_name='my_genome_features',
    # annotation_version=None,
    gtf_path_or_url='/My/local/gtf/path_to_my_genome_features.gtf', # Path or URL of GTF file
    # transcript_fasta_paths_or_urls=None, # List of paths or URLs of FASTA files containing transcript sequences
    # protein_fasta_paths_or_urls=None, # List of paths or URLs of FASTA files containing protein sequences
    # cache_directory_path=None, # Where to place downloaded and cached files for this genome
)
# parse GTF and construct database of genomic features
data.index()
gene_names = data.gene_names_at_locus(contig=6, position=29945884)

Reference DNA (optional)

Reference DNA lets you read any genomic interval, including introns, intergenic regions, and flanking sequence. It is opt-in: a normal installation does not download a whole genome. Human DNA is about 1 GB to download and takes several GB of disk once decompressed.

Quick start

pyensembl install --release 81 --with-genome-fasta  # annotation and DNA
pyensembl install --release 81 --only-genome-fasta  # just the DNA
from pyensembl import EnsemblRelease

release = EnsemblRelease(81, genome_fasta=True)
release.download_genome_fasta()  # does nothing if the DNA is already installed
with release:
    bases = release.sequence("7", 117_480_000, 117_480_100)
    tp53 = release.genes_by_name("TP53")[0]  # annotated on the minus strand
    tp53_dna = release.sequence(tp53.contig, tp53.start, tp53.end, strand=tp53.strand)

genome_fasta=True only chooses the DNA. Nothing is downloaded until you call download_genome_fasta() or download(), or run pyensembl install. Python objects use reference DNA only when constructed with genome_fasta: a plain EnsemblRelease(81) does not pick up DNA installed by the CLI, and its error message names the call that does.

Reading sequences

sequence(contig, start, end, mask="upper", *, strand="+"):

  • Coordinates are one-based and inclusive, like the rest of PyEnsembl, with 1 <= start <= end <= contig length.
  • Bases are read from the plus strand; strand="-" returns the reverse complement, so a gene or transcript reads 5' to 3'.
  • Contigs can be named as in the FASTA or as PyEnsembl reports them (gene.contig). If the names differ only by a chr prefix, the error suggests the right one.
  • Results are uppercase; mask="raw" keeps soft-masked repeats in lowercase.
  • Absent contigs and invalid ranges raise ValueError. Missing DNA raises MissingGenomeFastaError, a ValueError whose message explains how to install or enable it. Reads never download anything.

Related attributes and methods:

  • fasta is a pyfaidx reader with zero-based, half-open slices (fasta[contig][start - 1:end].seq), as used by Varcode. It needs the FASTA's own contig names, and is None when DNA is not configured or not installed.
  • genome_fasta_path is the uncompressed FASTA on disk, or None.
  • download() and index() include DNA when it is configured. index_genome_fasta() builds the DNA index before the first query needs it.
  • close() closes the reader. Readers already handed out stay usable after clear_cache().
  • Attached DNA does not affect equality: genes and transcripts from the same release compare equal with or without it.

Choosing DNA

Python (EnsemblRelease) CLI (pyensembl install)
Ensembl's DNA genome_fasta=True --with-genome-fasta or --only-genome-fasta
A local FASTA genome_fasta="/data/ref.fa.gz" --genome-fasta-path /data/ref.fa.gz
Coverage genome_fasta_type="primary_assembly" --genome-fasta-type primary_assembly
Masking genome_fasta_mask="soft" --masked soft

The default is unmasked toplevel DNA, which covers the patch and haplotype contigs in Ensembl annotations. primary_assembly has the chromosomes and unplaced/unlocalized sequences but no patches or haplotypes, and some older releases and species don't provide it. Masking is none, soft (repeats in lowercase), or hard (repeats replaced with N). See Ensembl's DNA file definitions.

Local FASTA files

release = EnsemblRelease(81, genome_fasta="/data/my_reference.fa.gz")

# Custom annotations can attach DNA too, from a path or URL:
from pyensembl import Genome
custom = Genome("custom", "my_annotations", genome_fasta_path_or_url="/data/reference.fa")

Plain FASTA files are read in place; gzip and BGZF files are decompressed into PyEnsembl's cache on first use. Indexes always live in the cache, so read-only source directories work and your own files and indexes are never modified. index() warns about annotation contigs that are missing from a local FASTA, but matching contig names don't prove that the assembly matches.

Managing disk space

pyensembl list --check-genome-fasta      # DNA for each release, verifying indexes
pyensembl delete-all-files --release 81  # release 81's files and DNA references
pyensembl prune --dry-run                # shared DNA that no installed release uses
pyensembl prune

Compatible releases share one copy of Ensembl DNA, so deleting a release keeps DNA that other releases still use, and prune removes DNA that no release references. It skips DNA that is being downloaded or indexed, never touches local FASTA files, and deletes nothing if any release's DNA metadata is malformed (list shows which one). list includes DNA-only installs and shows each release's most recently installed DNA. delete-index-files keeps shared DNA indexes because other releases may use them; rebuild one with index_genome_fasta(overwrite=True). In Python, prune_genome_fastas(dry_run=True) returns (path, bytes) candidates.

How the shared DNA cache works

Ensembl DNA is stored once per upstream file under pyensembl/dna_cache/, in <species>/<provider>/<reference>-<assembly accession>/<coverage>/<masking>/fasta/<file key>/:

pyensembl/dna_cache/
  homo_sapiens/ftp.ensembl.org/GRCh38-GCA_000001405.18/
    toplevel/unmasked/fasta/<file key>/
      sequence.fa        uncompressed, even when downloaded as .fa.gz
      sequence.fa.fai
      object.json        full identity of the upstream file
      index.json

Installing a release first reads Ensembl's small README and CHECKSUMS files to see whether another release already downloaded the same file. The versioned assembly accession distinguishes assembly patches. The 16-character file key is a prefix of the SHA-256 of the file's identity (assembly, Ensembl's Unix checksum, and compressed size), which separates upstream revisions of the same file, and conflicting identities are never reused. These are metadata checks: Ensembl's Unix checksums are not cryptographic hashes. If the metadata is incomplete, the assembly directory ends in -unverified and each release keeps its own copy. Local FASTA files and custom mirrors are never shared. Downloads retry transient HTTP failures and are checked against the upstream size, and installed releases work offline.

Reads take no locks and write nothing, so a fully installed and indexed cache can be read-only for other users. A download or index build locks only the file it writes; registering, deleting, and pruning releases briefly lock the whole cache. Files follow your umask, as do lock files on Python 3.10+, so use umask 002 or default ACLs for a group-shared cache. dna_cache itself may be a symlink, e.g. to a larger disk.

Sharing needs release caches next to dna_cache. That is the case on Linux and macOS, and whenever PYENSEMBL_CACHE_DIR is set. Windows' default cache layout is different, so there each release keeps its own DNA unless PYENSEMBL_CACHE_DIR is set.

Upgrading from 2.11.0

2.11.0 2.12.0 and later
EnsemblRelease(81, download_genome_fasta=True) EnsemblRelease(81, genome_fasta=True)
EnsemblRelease(81, genome_fasta_path="/data/ref.fa") EnsemblRelease(81, genome_fasta="/data/ref.fa")

The old keywords still work but emit a DeprecationWarning. Objects pickled or serialized by 2.11.0 still load. CLI flags are unchanged.

API

The EnsemblRelease object has methods to let you access all possible combinations of the annotation features gene_name, gene_id, transcript_name, transcript_id, exon_id as well as the location of these genomic elements (contig, start position, end position, strand).

Genes

genes(contig=None, strand=None, biotype=None)
Returns a list of Gene objects, optionally restricted to a particular contig, strand, or gene_biotype.
genes_at_locus(contig, position, end=None, strand=None)
Returns a list of Gene objects overlapping a particular position on a contig, optionally extend into a range with the end parameter and restrict to forward or backward strand by passing strand='+' or strand='-'.
gene_by_id(gene_id)
Return a Gene object for given Ensembl gene ID (e.g. "ENSG00000068793").
gene_names(contig=None, strand=None)
Returns all gene names in the annotation database, optionally restricted to a particular contig or strand.
genes_by_name(gene_name)
Get all the unique genes with the given name (there might be multiple due to copies in the genome), return a list containing a Gene object for each distinct ID.
gene_by_protein_id(protein_id)
Find Gene associated with the given Ensembl protein ID (e.g. "ENSP00000350283")
gene_names_at_locus(contig, position, end=None, strand=None)
Names of genes overlapping with the given locus, optionally restricted by strand. (returns a list to account for overlapping genes)
gene_name_of_gene_id(gene_id)
Returns name of gene with given gene ID.
gene_name_of_transcript_id(transcript_id)
Returns name of gene associated with given transcript ID.
gene_name_of_transcript_name(transcript_name)
Returns name of gene associated with given transcript name.
gene_name_of_exon_id(exon_id)
Returns name of gene associated with given exon ID.
gene_ids(contig=None, strand=None, biotype=None)
Return all gene IDs in the annotation database, optionally restricted by chromosome name, strand, or gene_biotype.
gene_ids_of_gene_name(gene_name)
Returns all Ensembl gene IDs with the given name.
nearest_gene(contig, position, end=None, strand=None)
Returns (distance, Gene) for the gene whose locus is nearest to the position (or position..end interval) on the given contig — even when no gene overlaps. Returns (inf, None) when no candidates exist.
merged_gene_intervals(contig, strand=None)
Returns the union of all gene loci on the contig as a sorted list of non-overlapping (start, end) tuples. Adjacent intervals (end+1 == next start) are merged into one.

Transcripts

transcripts(contig=None, strand=None, biotype=None)
Returns a list of Transcript objects for all transcript entries in the Ensembl database, optionally restricted to a particular contig, strand, or transcript_biotype.
transcript_by_id(transcript_id)
Construct a Transcript object for given Ensembl transcript ID (e.g. "ENST00000369985")
transcripts_by_name(transcript_name)
Returns a list of Transcript objects for every transcript matching the given name.
transcript_names(contig=None, strand=None)
Returns all transcript names in the annotation database.
transcript_ids(contig=None, strand=None, biotype=None)
Returns all transcript IDs in the annotation database.
transcript_ids_of_gene_id(gene_id)
Return IDs of all transcripts associated with given gene ID.
transcript_ids_of_gene_name(gene_name)
Return IDs of all transcripts associated with given gene name.
transcript_ids_of_transcript_name(transcript_name)
Find all Ensembl transcript IDs with the given name.
transcript_ids_of_exon_id(exon_id)
Return IDs of all transcripts associated with given exon ID.
nearest_transcript(contig, position, end=None, strand=None)
Returns (distance, Transcript) to the closest transcript on the contig. Returns (inf, None) when no candidates exist.

Exons

exon_ids(contig=None, strand=None)
Returns a list of exon IDs in the annotation database, optionally restricted by the given chromosome and strand.
exon_by_id(exon_id)
Construct an Exon object for given Ensembl exon ID (e.g. "ENSE00001209410")
exon_ids_of_gene_id(gene_id)
Returns a list of exon IDs associated with a given gene ID.
exon_ids_of_gene_name(gene_name)
Returns a list of exon IDs associated with a given gene name.
exon_ids_of_transcript_id(transcript_id)
Returns a list of exon IDs associated with a given transcript ID.
exon_ids_of_transcript_name(transcript_name)
Returns a list of exon IDs associated with a given transcript name.

Reference DNA

These need reference DNA; see Reference DNA.

sequence(contig, start, end, mask="upper", strand="+")
Returns the bases from start to end (one-based, inclusive) on the plus strand, or their reverse complement with strand="-". mask="raw" keeps soft-masked lowercase.
download_genome_fasta(overwrite=False)
Downloads the configured reference DNA without annotation data; does nothing if it is already installed.
index_genome_fasta(overwrite=False)
Builds the DNA index now rather than on the first query.
fasta
A pyfaidx reader with zero-based, half-open slices, or None when DNA is not configured or not installed.
genome_fasta_path
Path of the installed, uncompressed FASTA, or None.

Release files for pyensembl 2.14.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyensembl 2.14.0
File Size Uploaded
pyensembl-2.14.0.tar.gz 141.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyensembl 2.14.0
File Interpreter ABI Platform
pyensembl-2.14.0-py3-none-any.whl Python 3 none any Details

Total release size: 232.7 kB

Release files / pyensembl-2.14.0.tar.gz

Download URL pyensembl-2.14.0.tar.gz
Size 141.4 kB
Tags Source
SHA-256 checksum
How to use checksums
c4ca56eaae3e9b7ccc72c7e5972efc28724a986b9291a8343438618e9545dfd2
BLAKE2b-256 checksum
How to use checksums
9378523716a9b282f2642c0ff37498e0ddd9c4c4cdf57b79ed69fdf96ad325b0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.6

Release files / pyensembl-2.14.0-py3-none-any.whl

Download URL pyensembl-2.14.0-py3-none-any.whl
Size 91.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e17cefb4cfa7fca8957ab9939a5f215cc7ca1ab0ec266887d6e417757678b492
BLAKE2b-256 checksum
How to use checksums
8a729a2d99f2a88485fefbecf3a9de7ead12c16100b676525e4c6f2daf1dc448
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.6

Release history Release notifications | RSS feed

This release

2.14.0 This release

2 release files

2.13.3

2 release files

2.13.2

2 release files

2.13.1

2 release files

2.13.0

2 release files

2.12.0

2 release files

2.11.0

2 release files

2.10.9

2 release files

2.10.8

2 release files

2.10.7

2 release files

2.10.6

2 release files

2.10.5

2 release files

2.10.1

2 release files

2.10.0

2 release files

2.9.8

2 release files

2.9.7

2 release files

2.9.6

2 release files

2.9.5

2 release files

2.9.4

2 release files

2.9.3

2 release files

2.9.2

2 release files

2.9.1

2 release files

2.9.0

2 release files

2.8.0

2 release files

2.7.0

2 release files

2.6.13

2 release files

2.6.7

2 release files

2.6.6

2 release files

2.6.5

2 release files

2.6.4

2 release files

2.6.2

2 release files

2.6.1

2 release files

2.6.0

2 release files

2.3.13

2 release files

2.3.12

2 release files

2.3.11

2 release files

2.3.10

2 release files

2.3.9

2 release files

2.3.8

2 release files

2.3.7

2 release files

2.3.6

2 release files

2.3.4

2 release files

2.3.3

2 release files

2.3.2

2 release files

2.3.1

2 release files

2.3.0

1 release file

2.2.9

1 release file

2.2.8

1 release file

2.2.7

1 release file

2.2.6

1 release file

2.2.5

1 release file

2.2.4

1 release file

2.2.3

1 release file

2.2.2

1 release file

2.2.1

1 release file

2.2.0

1 release file

2.1.0

1 release file

2.0.2

1 release file

2.0.1

1 release file

2.0.0

1 release file

1.9.4

1 release file

1.9.3

1 release file

1.9.2

1 release file

1.9.1

1 release file

1.9.0

1 release file

1.8.8

1 release file

1.8.7

1 release file

1.8.6

1 release file

1.8.5

1 release file

1.8.4

1 release file

1.8.3

1 release file

1.8.2

1 release file

1.8.1

1 release file

1.8.0

1 release file

1.7.5

1 release file

1.7.4

1 release file

1.7.3

1 release file

1.7.2

1 release file

1.7.1

1 release file

1.7.0

1 release file

1.6.0

1 release file

1.5.2

1 release file

1.5.0

1 release file

1.4.0

1 release file

1.3.0

1 release file

1.2.6

1 release file

1.2.4

1 release file

1.2.3

1 release file

1.2.2

1 release file

1.2.1

1 release file

1.1.0

1 release file

1.0.3

1 release file

1.0.2

1 release file

1.0.1

1 release file

1.0.0

1 release file

0.9.7

1 release file

0.9.6

1 release file

0.9.5

1 release file

0.9.4

1 release file

0.9.3

1 release file

0.9.1

1 release file

0.9.0

1 release file

0.8.14

1 release file

0.8.13

1 release file

0.8.12

1 release file

0.8.11

1 release file

0.8.10

1 release file

0.8.9

1 release file

0.8.8

1 release file

0.8.7

1 release file

0.8.5

1 release file

0.8.4

1 release file

0.8.3

1 release file

0.8.2

1 release file

0.8.1

1 release file

0.7.0

1 release file

0.6.11

1 release file

0.6.10

1 release file

0.6.9

1 release file

0.6.8

1 release file

0.6.7

1 release file

0.6.5

1 release file

0.6.4

1 release file

0.6.3

1 release file

0.6.2

1 release file

0.6.1

1 release file

0.6.0

1 release file

0.5.13

1 release file

0.5.12

1 release file

0.5.11

1 release file

0.5.10

1 release file

0.5.9

1 release file

0.5.8

1 release file

0.5.7

1 release file

0.5.4

1 release file

0.5.3

1 release file

0.5.2

1 release file

0.5.1

1 release file

0.5.0

1 release file

0.4.0

1 release file

0.3.3

1 release file

0.3

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page