ORFeus: alternative ORF predictor
Project description
ORFeus is an ORF prediction tool designed to detect alternative translation events, including programmed ribosomal frameshifts, stop codon readthrough, and short upstream or downstream ORFs. It requires aligned ribosome profiling (ribo-seq) reads, reference transcript annotations, and a reference genome. ORFeus can be run on both bacterial and eukaryotic data.
Note that high-resolution (even nucleotide-resolution) ribosome profiling data is ideal. The higher-resolution the data and the deeper the sequencing, the better predictions ORFeus will make.
Quick Start
orfeusbuild forward.wig reverse.wig genome.fa annotations.gtf
orfeusrun data.txt.gz parameters_h1.npy parameters_h0.npy
Overview
ORFeus requires the following input files. (See input files for details on preparing these to work with ORFeus, particularly the aligned ribo-seq reads file.)
- Aligned ribo-seq read counts (format: wiggle or bedgraph)
- Transcript annotations (format: gtf or gff)
- Genome (format: fasta)
Reasons to use ORFeus:
- You are interested in finding alternative translation events in an annotated species. (Run ORFeus and look at the top predictions.)
- You want to find changes in translation across multiple conditions or timepoints. (Run ORFeus separately on ribo-seq data from each condition and compare predictions.)
Reasons not to use ORFeus:
- You need a de novo ORF caller for a novel genome. (ORFeus requires input transcript annotations.)
- You don't have high-resolution ribo-seq data. (ORFeus infers translation based on ribosome profiling reads.)
ORFeus can predict the following types of canonical and alternative ORFs:
Dependencies
We recommend installing Anaconda, which is a Python distribution that comes with all of these packages.
Installation
Install from source
git clone https://github.com/morichardson/ORFeus
Input files
Transcript annotations
The annotations file must be in GTF or GFF
format.
At minimum, the annotations file must contain the following columns
(placeholder columns are not used by ORFeus and can be populated with any value,
though the standard is .):
seqname: name of the chromosome (note this must match exactly theseqnamein the genome sequence file)source: placeholderfeature: feature type (note onlyfive_prime_utr,exon, andthree_prime_utrfeatures will be kept)start: first position of the feature, 1-indexedend: last position of the feature, 1-indexedscore: placeholderstrand: + (forward) or - (reverse)frame: placeholderattribute: semicolon-separated list with additional information (note only the info below will be kept)transcript_idtranscript_nametranscript_biotype(note only"protein_coding"features will be kept)
Below is an example transcript from a GTF file that meets the minimum requirements. All placeholder fields have been populated with a period.
V . five_prime_utr 546794 546816 . + . transcript_id "YER178W_mRNA"; transcript_name "PDA1"; transcript_biotype "protein_coding";
V . exon 546817 548079 . + . transcript_id "YER178W_mRNA"; transcript_name "PDA1"; transcript_biotype "protein_coding";
V . three_prime_utr 548080 548208 . + . transcript_id "YER178W_mRNA"; transcript_name "PDA1"; transcript_biotype "protein_coding";
Genome sequence
The genome sequence file must be in
FASTA format. There should
be one sequence entry for each unique seqname (chromosome) in the annotations
file.
Below is an example FASTA file excerpt for the chromosome of the above
example transcript. Note that the seqname matches the seqname column entries
in the annotations example.
>V dna:chromosome chromosome:R64-1-1:V:1:576874:1 REF
CGTCTCCTCCAAGCCCTGTTGTCTCTTACCCGGATGTTCAACCAAAAGCTACTTACTACC
TTTATTTTATGTTTACTTTTTATAGATTGTCTTTTTATCCTACTCTTTCCCACTTGTCTC
TCGCTACTGCCGTGCAACAAACACTAAATCAAAACAGTGAAATACTACTACATCAAAACG
CATATTCCCTAGAAAAAAAAATTTCTTACAATATACTATACTACACAATACATAATCACT
...
Aligned ribo-seq read counts
The final aligned ribo-seq read counts must be in either WIG format or BedGraph format. The raw reads must be aligned and then the read ends should be offset to correspond to the P-site of the ribosome. The count of read ends at each position of the genome should be stored, with one file for the forward strand and one for the reverse strand.
Align raw reads to genome
Align raw ribo-seq reads to the genome. Filter out reads mapping to annotated ncRNA sequences. You should decide whether uniqely-mapping or multi-mapping is appropriate for your data set.
Uniquely-mapping reads:
- filters out reads that map to repetitive regions (e.g. regions with repeated sequences may appear as gaps in the read density, even though they may actually be translated)
- filters out reads that map to similar or related sequences (e.g. insertion sequences that have multiple copies in the genome will have no reads, even though they may actually be translated)
Multi-mapping reads:
- generates confounding signals from mis-mapped multi-mapping reads (e.g. reads that were actually generated from one transcript also map to another transcript, adding noise)
- complicates interpretation of predictions (e.g. predictions of alternative events may be due to reads from that transcript or another transcript)
In some cases, you may want to run ORFeus twice: once on the uniquely-mapped reads and once on the multi-mapped reads. This will allow you to compare the predictions and determine which events may be artifacts of read mapping. Any predictions that differ between the two runs should be examined more closely, since they might arise from mapping artifacts.
Offset aligned reads to P-site
Before passing the data to ORFeus, you need to offset the read ends so they align to a position within the P-site of the ribosome. This lets ORFeus infer the exact codon being translated for each read.
You can determine the offset for each read length and export the resulting read counts using existing software packages like Shoelaces or using your own custom scripts.
Usage
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file orfeus-1.0.tar.gz.
File metadata
- Download URL: orfeus-1.0.tar.gz
- Upload date:
- Size: 62.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/4.0.2 CPython/3.7.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b2586a68373911ca4c8ffa24978dd784d8a5afd421be55d888c186b669bd2d54
|
|
| MD5 |
5e9714108822b3b072ab9009a06575c7
|
|
| BLAKE2b-256 |
41fbe7a05fbbff43688e75bf7a272663fd50a683851817f8658022dd18a0bb7b
|
File details
Details for the file orfeus-1.0-py3-none-any.whl.
File metadata
- Download URL: orfeus-1.0-py3-none-any.whl
- Upload date:
- Size: 79.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/4.0.2 CPython/3.7.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8d5f549005372f826337e524e7ad1ccb8c8303826e3dd75a90dcf782772c6374
|
|
| MD5 |
3e909fa81f899b6443722289f2168359
|
|
| BLAKE2b-256 |
f7e7db131be50e55930b8ac8012463276a9b641b82c34eb8148056fa99786cff
|