Skip to main content

Inputing a VCF file, it returns the genomic sequence at the specified length (31 by default).

Project description

vcf2seq

Aim

Similar to seqtailor [PMID:31045209] : reads a VCF file, outputs a genomic sequence (default length: 31)

Unlike seqtailor, all sequences will have the same length. Moreover, it is possible to have an absence character (by default the dot . ) for indels.

  • When a insertion is larger than --size parameter, only first --size nucleotides are outputed.
  • Sequence headers are formated as "_".

VCF format specifications: https://github.com/samtools/hts-specs/blob/master/VCFv4.4.pdf

Installation

pip install vcf2seq

usage

usage: vcf2seq.py [-h] -g genome [-s SIZE] [-t {alt,ref,both}] [-b BLANK] [-a ADD_COLUMNS [ADD_COLUMNS ...]] [-o OUTPUT] [-v] vcf


positional arguments:
  vcf                   vcf file (mandatory)

options:
  -h, --help            show this help message and exit
  -g genome, --genome genome
                        genome as fasta file (mandatory)
  -s SIZE, --size SIZE  size of the output sequence (default: 31)
  -t {alt,ref,both}, --type {alt,ref,both}
                        alt, ref, or both output? (default: alt)
  -b BLANK, --blank BLANK
                        Missing nucleotide character, default is dot (.)
  -a ADD_COLUMNS [ADD_COLUMNS ...], --add-columns ADD_COLUMNS [ADD_COLUMNS ...]
                        Add one or more columns to header (ex: '-a 3 AA' will add columns 3 and 27). The first column is '1' (or 'A')
  -o OUTPUT, --output OUTPUT
                        Output file (default: <input_file>-vcf2seq.fa/tsv)
  -f {fa,tsv}, --output-format {fa,tsv}
                        Output file format (default: fa)
  -v, --version         show program's version number and exit

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vcf2seq-0.6.5a0.tar.gz (19.8 kB view details)

Uploaded Source

Built Distribution

vcf2seq-0.6.5a0-py3-none-any.whl (20.7 kB view details)

Uploaded Python 3

File details

Details for the file vcf2seq-0.6.5a0.tar.gz.

File metadata

  • Download URL: vcf2seq-0.6.5a0.tar.gz
  • Upload date:
  • Size: 19.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.0.0 CPython/3.11.5

File hashes

Hashes for vcf2seq-0.6.5a0.tar.gz
Algorithm Hash digest
SHA256 2e4c6bd5844ae5b7f9d03a31538c411c048e410a733bdd1c6f9195dbff1fa010
MD5 9feb04930daa6cdb88eb486bf37fc997
BLAKE2b-256 b9e9fc9140546d24127cde47d59f5ff0279626cc1ed02d3dc3240e6577db1336

See more details on using hashes here.

File details

Details for the file vcf2seq-0.6.5a0-py3-none-any.whl.

File metadata

  • Download URL: vcf2seq-0.6.5a0-py3-none-any.whl
  • Upload date:
  • Size: 20.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.0.0 CPython/3.11.5

File hashes

Hashes for vcf2seq-0.6.5a0-py3-none-any.whl
Algorithm Hash digest
SHA256 fca9c005bfdb9aa14f2f4d1362d40fe1fb71797b55098a3651f35a8a4a05098c
MD5 bf92ba355c8c24abe0b6ee8da489f4e3
BLAKE2b-256 89247cad4afa6da412181466ab00da82e8331135b1a7d6c3c1643ff3b3b96323

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page