pyfastx is a python module for fast randomaccess to FASTA sequences from flat text fileor even from gzip compressed file.
Project description
a robust python module for fast random access to FASTA sequences
About
The pyfastx is a lightweight Python C extension that enables you to randomly access FASTA sequences in flat text file, even in gzip compressed file. This module uses kseq.h written by Heng Li to parse FASTA file and zran.c written by Paul McCarthy in indexed_gzip project to index gzipped file.
Installation
Make sure you have both pip and at least version 3.5 of Python before starting.
You can install pyfastx via the Python Package Index (PyPI)
pip install pyfastx
Read FASTA file
The fastest way to parse flat or gzipped FASTA file without building index.
>>> import pyfastx
>>> for name, seq in pyfastx.Fasta('test/data/test.fa.gz', build_index=False):
>>> print(name, seq)
Read flat or gzipped FASTA file and build index, support for random access to FASTA.
>>> import pyfastx
>>> fa = pyfastx.Fasta('test/data/test.fa.gz')
>>> fa
<Fasta> test/data/test.fa.gz contains 211 seqs
Note: Building index may take some times. The time required to build index depends on the size of FASTA file. If index built, you can randomly access to any sequences in FASTA file.
Get FASTA information
>>> # get sequence counts in FASTA
>>> len(fa)
211
>>> # get total sequnce length of FASTA
>>> fa.size
86262
>>> # get GC content of DNA sequence of FASTA
>>> fa.gc_content
43.529014587402344
>>> # get composition of nucleotides in FASTA
>>> fa.composition
{'A': 24534, 'C': 18694, 'G': 18855, 'T': 24179, 'N': 0}
Get sequence from FASTA
>>> # get sequence like dictionary
>>> s1 = fa['JZ822577.1']
>>> s1
<Sequence> JZ822577.1 with length of 333
>>> # get sequence like list
>>> s2 = fa[2]
>>> s2
<Sequence> JZ822579.1 with length of 176
>>> # get last sequence
>>> s3 = fa[-1]
>>> s3
<Sequence> JZ840318.1 with length of 134
>>> # check name weather in FASTA file
>>> 'JZ822577.1' in fa
True
Get sequence information
>>> s = fa[-1]
>>> s
<Sequence> JZ840318.1 with length of 134
>>> # get sequence name
>>> s.name
'JZ840318.1'
>>> # get sequence string
>>> s.seq
'ACTGGAGGTTCTTCTTCCTGTGGAAAGTAACTTGTTTTGCCTTCACCTGCCTGTTCTTCACATCAACCTTGTTCCCACACAAAACAATGGGAATGTTCTCACACACCCTGCAGAGATCACGATGCCATGTTGGT'
>>> # get sequence length
>>> len(s)
134
>>> # get GC content if dna sequence
>>> s.gc_content
46.26865768432617
>>> # get nucleotide composition if dna sequence
>>> s.composition
{'A': 31, 'C': 37, 'G': 25, 'T': 41, 'N': 0}
Sequence slice
Sequence object can be sliced like a python string
>>> # get a sub seq from sequence
>>> ss = seq[10:30]
>>> ss
<Sequence> JZ840318.1 from 11 to 30
>>> ss.name
'JZ840318.1:11-30'
>>> ss.seq
'CTTCTTCCTGTGGAAAGTAA'
>>> ss = s[-10:]
>>> ss
<Sequence> JZ840318.1 from 125 to 134
>>> ss.name
'JZ840318.1:125-134'
>>> ss.seq
'CCATGTTGGT'
Note: Slicing start and end coordinates are 0-based. Currently, pyfastx does not support an optional third step or stride argument. For example ss[::-1]
Reverse and complement sequence
>>> # get sliced sequence
>>> fa[0][10:20].seq
'GTCAATTTCC'
>>> # get reverse of sliced sequence
>>> fa[0][10:20].reverse
'CCTTTAACTG'
>>> # get complement of sliced sequence
>>> fa[0][10:20].complement
'CAGTTAAAGG'
>>> # get reversed complement sequence, corresponding to sequence in antisense strand
>>> fa[0][10:20].antisense
'GGAAATTGAC'
Get subsequences
Subseuqneces can be retrieved from FASTA file by using a list of [start, end] coordinates
>>> # get subsequence with start and end position
>>> interval = (1, 10)
>>> fa.get_seq('JZ822577.1', interval)
'CTCTAGAGAT'
>>> # get subsequences with a list of start and end position
>>> intervals = [(1, 10), (50, 60)]
>>> fa.get_seq('JZ822577.1', intervals)
'CTCTAGAGATTTTAGTTTGAC'
>>> # get subsequences with reverse strand
>>> fa.get_seq('JZ822577.1', (1, 10), strand='-')
'ATCTCTAGAG'
Get identifiers
Get all identifiers of sequence as a list-like object.
>>> ids = fa.keys()
>>> ids
<Identifier> contains 211 identifiers
>>> # get count of sequence
>>> len(ids)
211
>>> # get identifier by index
>>> ids[0]
'JZ822577.1'
>>> # check identifier where in fasta
>>> 'JZ822577.1' in ids
True
>>> # iter identifiers
>>> for name in ids:
>>> print(name)
>>> # convert to a list
>>> list(ids)
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Hashes for pyfastx-0.2.0-cp37-cp37m-win_amd64.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 3ea499f226db33b4ed2ecdf8482fe7cb5ad3ea7951065de22aef417d513eb4d0 |
|
MD5 | 087806f3d5e5ec988cb0f877a6252e01 |
|
BLAKE2b-256 | b8800f9e50fc7b3aaf57e25c0593c842dbd810116858f69e8dfd2c181f93cf9a |
Hashes for pyfastx-0.2.0-cp37-cp37m-win32.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | cd06da2121c075b66b2da1691da1e5520929789ba353ff5c7013cfcba456947f |
|
MD5 | 33c28795dd052f9f0ac4c7d2fd32a988 |
|
BLAKE2b-256 | 6812968cc6f3a07446bf58fcbd7c946a1f972dd5f0971db18fbb529fffd76f18 |
Hashes for pyfastx-0.2.0-cp37-cp37m-manylinux1_x86_64.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | f7fd574ff29a13ee54b2ffe09789f45ccf66d4f9effadef8afebc17f6ac4fb29 |
|
MD5 | 01b72d9dc5b2bb5fd8a6a39922f99f5c |
|
BLAKE2b-256 | ff9b98b3211740a3dafb3519a83b301110fec6022ffa838fe3d05de37721662e |
Hashes for pyfastx-0.2.0-cp37-cp37m-manylinux1_i686.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | ba0ca9749ba9df35b4ba5b263401b89b932f4a36942fcbde3c8da49804bfb306 |
|
MD5 | 7aa95e8827180aff8b3b6316e2db40f0 |
|
BLAKE2b-256 | 73ea2b3360e8b4ba4e19645a84a875d0cf5f8ffae0e0c8ce3a814cf2e45e40c4 |
Hashes for pyfastx-0.2.0-cp37-cp37m-macosx_10_6_intel.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | b1818028f286ec92b27643de13a4f011c5f10b6f5b99022350fab9a1bc0fcf2f |
|
MD5 | afae2a3251a598309a16bb8466ab527d |
|
BLAKE2b-256 | b082e4658d5b2b7b892356012468fd08e56cad7158785ff38b5b1a2e85e86f7d |
Hashes for pyfastx-0.2.0-cp36-cp36m-win_amd64.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 76a8f8628b724fa7e75950e6adb80805c01e8ea4c7569480daf4f2e29b06393c |
|
MD5 | 3c3e5ba64884b9368d148844d1ba9236 |
|
BLAKE2b-256 | 75a0358126652b6dbf28b9ab29092dccb209bc29739390a03214d8b9f2107fed |
Hashes for pyfastx-0.2.0-cp36-cp36m-win32.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 57fe3a60837c147048bf98521bf61b111623efe50b502e3908645859f23de56e |
|
MD5 | 82d45b76cadb28b307285863cb209b5f |
|
BLAKE2b-256 | 16357f842a11efdb58f65c0ad71b6e10c29dc1ef850e8c857fac3bd837715ddb |
Hashes for pyfastx-0.2.0-cp36-cp36m-manylinux1_x86_64.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 3fdbeacdbed450e96282d1a488bdff50a1aea9a55ebeb3046729e14ffc4d85d7 |
|
MD5 | 9eb0c3bbecd74f4b66db34db2340d21b |
|
BLAKE2b-256 | 8274796e3555fb1189346282bebbb1a1cffd4e8f2fd1c1aa259be922b0448617 |
Hashes for pyfastx-0.2.0-cp36-cp36m-manylinux1_i686.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | cee97da119d52b50b38869ad009411d187ca18d88024cbbb225bccaeabcd6341 |
|
MD5 | 3b2c9ce0b06da4a0a820060f4e7cacac |
|
BLAKE2b-256 | d03d7a3196eddfb1a6e1b17144fc8d74dc18a64d8772a0ccf19db71f84fd8822 |
Hashes for pyfastx-0.2.0-cp36-cp36m-macosx_10_6_intel.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 71e077aeef571c23fb9bad5ee84f600921d396be37b25ccaff14d82528f65e6e |
|
MD5 | 491312b6cd94ea80cbefcba1e7e1fe8a |
|
BLAKE2b-256 | d3e64cbe7469576eb890eca9482f0cc1e306095240f76f6927d1109b1f81293c |
Hashes for pyfastx-0.2.0-cp35-cp35m-win_amd64.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 4def3f7adf2a64145707ca18e53bb7cf6c62eb2f61c39fa2aea0ce91005f5f86 |
|
MD5 | d1b0f8f3f8c02ca8a0aac90c524b754c |
|
BLAKE2b-256 | 89d61d3c7f7bb6e6c0d433dc5d87b633adb8a361ca58e812bf0813836464093f |
Hashes for pyfastx-0.2.0-cp35-cp35m-win32.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 33ff4c45049617b17ebbb036d770d7f1fc5a480b6d090cfe3e5450f1aabb98d5 |
|
MD5 | 6aea8cba498c9cfeb1f1d51d25193f44 |
|
BLAKE2b-256 | 8c052f7b6d2d579989baf19938b7875ca7040468a2c388a219fb4b11611e2ecc |
Hashes for pyfastx-0.2.0-cp35-cp35m-manylinux1_x86_64.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 49b36aca5736c0e9dfc3a4783f3a97d6a6bc1731e0fc25c4b5f39d9756686c66 |
|
MD5 | 2c26c90927b1338ab9bde68bcf42f3a7 |
|
BLAKE2b-256 | 797cd838b408e6dd1341160fe5132169da95a76813496518b08f7d5dc18fe046 |
Hashes for pyfastx-0.2.0-cp35-cp35m-manylinux1_i686.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 61ff99d344af5ed76db28c5fb3e57d39b83ef2ff9431efaf794731ca35ea1f02 |
|
MD5 | 9f223d995fcfd9352b3e2ef16f6f0785 |
|
BLAKE2b-256 | a8b1dacb7f90bcbab532b711441fca59d84a61ca80c8438dfa986a10740217b8 |
Hashes for pyfastx-0.2.0-cp35-cp35m-macosx_10_6_intel.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 59925cabf7cf85b5aff098626027d9cc326d3f83cc04d2a524f45b7c5df0e6d6 |
|
MD5 | 3722a09b3d40f91c4928e6addb0d678f |
|
BLAKE2b-256 | e0fa77502caf329c6824a513676fbf9cd54b1c5679f0da13ca16d394729f7c7e |