biopython.convert

Interconvert various file formats supported by biopython. Supports querying records with JMESPath.

These details have not been verified by PyPI

Project links

Homepage

Project description

BioPython-Convert

Interconvert various file formats supported by BioPython.

Supports querying records with JMESPath.

Installation

pip install biopython-convert

or:

conda install biopython-convert

or:

git clone https://github.com/brinkmanlab/BioPython-Convert.git
cd BioPython-Convert
./setup.py install

Use

biopython.convert [-s] [-v] [-i] [-q JMESPath] input_file input_type output_file output_type
    -s Split records into seperate files
    -q JMESPath to select records. Must return list of SeqIO records or mappings. Root is list of input SeqIO records.
    -i Print out details of records during conversion
    -v Print version and exit

Supported formats: abi, abi-trim, ace, cif-atom, cif-seqres, clustal, embl, fasta, fasta-2line, fastq-sanger, fastq, fastq-solexa, fastq-illumina, genbank, gb, ig, imgt, nexus, pdb-seqres, pdb-atom, phd, phylip, pir, seqxml, sff, sff-trim, stockholm, swiss, tab, qual, uniprot-xml, gff3, txt, json, yaml

JMESPath

The root node for a query is a list of SeqRecord objects. The query can return a list with a subset of these or a mapping, keying to the constructor parameters of a SeqRecord object.

If the formats are txt, json, or yaml, then the JMESPath resulting object will simply be dumped in those formats.

A web based tool is available to experiment with constructing queries in real time on your data. Simply convert your dataset to JSON and load it into the JMESPath playground to begin composing your query. It supports loading JSON files directly rather than trying to copy/paste the data.

split() and let() functions are available in addition to the JMESPath standard functions

extract(Seq, SeqFeature) is also made available to allow access to the SeqFeature.extract() function within the query

Examples:

Append a new record:

[@, [{'seq': 'AAAA', 'name': 'my_new_record'}]] | []

Filter out any plasmids:

[?!(features[?type=='source'].qualifiers.plasmid)]

Keep only the first record:

[0]

Output taxonomy of each record (txt output):

[*].annotations.taxonomy

Output json object containing id and molecule type:

[*].{id: id, type: annotations.molecule_type}

Convert dataset to PTT format using text output:

[0].[join(' - 1..', [description, to_string(length(seq))]), join(' ', [to_string(length(features[?type=='CDS' && qualifiers.translation])), 'proteins']), join(`"\t"`, ['Location', 'Strand', 'Length', 'PID', 'Gene', 'Synonym', 'Code', 'COG', 'Product']), (features[?type=='CDS' && qualifiers.translation].[join('..', [to_string(sum([location.start, `1`])), to_string(location.end)]), [location.strand][?@==`1`] && '+' || '-', length(qualifiers.translation[0]), (qualifiers.db_xref[?starts_with(@, 'GI')].split(':', @)[1])[0] || '-', qualifiers.gene[0] || '-', qualifiers.locus_tag[0] || '-', '-', '-', qualifiers.product[0] ] | [*].join(`"\t"`, [*].to_string(@)) )] | []

        Convert dataset to faa format using fasta output::

                        [0].let({org: (annotations.organism || annotations.source)}, &(features[?type=='CDS' && qualifiers.translation].{id:
                        join('|', [
                                (qualifiers.db_xref[?starts_with(@, 'GI')].['gi', split(':', @)[1]]),
                                (qualifiers.protein_id[*].['ref', @]),
                                (qualifiers.locus_tag[*].['locus', @]),
                                join('', [':', [location][?strand==`-1`] && 'c' || '', to_string(sum([location.start, `1`])), '..', to_string(location.end)])
                        ][][]),
                        seq: qualifiers.translation[0],
                        description: (org && join('', [qualifiers.product[0], ' [', org, ']']) || qualifiers.product[0])}))

See CONTRIBUTING.rst for information on contributing to this repo.

Project details

These details have not been verified by PyPI

Project links

Homepage

Release history Release notifications | RSS feed

1.3.3

Jun 23, 2022

This version

1.3.2

Sep 15, 2021

1.3.1

Sep 13, 2021

1.2.0

Feb 25, 2021

1.1.0

Feb 4, 2021

1.0.4

Sep 8, 2020

1.0.3

Sep 25, 2019

1.0.2

Sep 5, 2019

1.0.0

Aug 9, 2019

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

biopython.convert-1.3.2.tar.gz (30.1 MB view details)

Uploaded Sep 15, 2021 Source

File details

Details for the file biopython.convert-1.3.2.tar.gz.

File metadata

Download URL: biopython.convert-1.3.2.tar.gz
Upload date: Sep 15, 2021
Size: 30.1 MB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.4.2 importlib_metadata/4.8.1 pkginfo/1.7.1 requests/2.26.0 requests-toolbelt/0.9.1 tqdm/4.62.2 CPython/3.9.7

File hashes

Hashes for biopython.convert-1.3.2.tar.gz
Algorithm	Hash digest
SHA256	`c88f96b672d0e4c53ffbecb6e8864f632a5bd0e01a8c65cded8d58ad9edc0352`
MD5	`4eeb998b8bd58bac1f297f7039a1fad8`
BLAKE2b-256	`7615b0512068d417375042ae3401d97c220b5ce44db4d8c326c19e99c3dfdcb8`

See more details on using hashes here.

biopython.convert 1.3.2

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

BioPython-Convert

Installation

Use

JMESPath

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

File details

File metadata

File hashes