Skip to main content

A pure-Python library for reading and writing PEFF (PSI Extended FASTA Format) files.

Project description

pefftacular

PyPI Python Package License Python

Python library for reading and writing PEFF (PSI Extended FASTA Format) files. PEFF is a superset of FASTA used in proteomics that carries rich per-entry annotations — PTMs, variants, processed forms, and more — encoded directly in the sequence header.

Install

pip install pefftacular

Dev install:

just install

Quick start

read_peff — load everything into memory at once:

from pefftacular import read_peff

header, entries = read_peff("proteins.peff")

for entry in entries:
    print(entry.db_unique_id, entry.pname, len(entry.sequence))

PeffReader — iterate lazily without loading the full file:

from pefftacular import PeffReader

with PeffReader("proteins.peff") as reader:
    file_header = reader.header
    for entry in reader:
        process(entry)

Data model

read_peff and PeffReader yield SequenceEntry objects with these fields:

Field Type Description
prefix str Database prefix (e.g. sp, tr)
db_unique_id str Accession (e.g. P12345)
sequence str Amino acid sequence
pname str | None Protein name (\\PName=)
gname str | None Gene name (\\GName=)
ncbi_tax_id int | None NCBI taxonomy ID (\\NcbiTaxId=)
length int | None Sequence length (\\Length=)
sv int | None Sequence version (\\SV=)
ev int | None Entry version (\\EV=)
pe int | None Protein existence level (\\PE=)
variant_simple tuple[VariantSimple, ...] Simple sequence variants
variant_complex tuple[VariantComplex, ...] Multi-residue variants (start, end, new sequence, optional tag)
mod_res_unimod tuple[ModResUnimod, ...] UniMod modification sites
mod_res_psi tuple[ModResPsi, ...] PSI-MOD modification sites
mod_res tuple[ModRes, ...] Other named modification sites
processed tuple[Processed, ...] Processed sequence forms
custom_values dict[str, tuple[CustomKeyValue, ...]] Header-declared custom keys, parsed by their CustomKeyDef
extra dict[str, str] Non-standard keys with no CustomKeyDef

Annotations

Variants:

from pefftacular import read_peff

_, entries = read_peff("proteins.peff")
entry = entries[0]

for v in entry.variant_simple:
    print(v.position, v.new_amino_acid, v.tag)
    # e.g. 42, "K", "rs12345"

Modifications (UniMod):

for mod in entry.mod_res_unimod:
    print(mod.position, mod.accession, mod.name)
    # e.g. 17, "21", "Phospho"

Modifications (PSI-MOD):

for mod in entry.mod_res_psi:
    print(mod.position, mod.accession, mod.name)
    # e.g. 17, "MOD:00696", "phosphorylated residue"

Processed forms:

for proc in entry.processed:
    print(proc.start_pos, proc.end_pos, proc.accession, proc.name)
    # e.g. 1, 24, "PRO_0000012345", "Signal peptide"

Custom keys (declared via # CustomKeyDef= in the header):

When the database header declares a custom key, entry values for that key are parsed using its RegExp / FieldNames / FieldTypes and exposed as typed fields on entry.custom_values. The original item text is preserved in raw for lossless round-trips.

Header excerpt:

# CustomKeyDef=(KeyName=SecondaryStructure|Description="..."|ConceptCURIE=BAO:0000014|RegExp="([0-9]+)\|([0-9]+)\|([A-Za-z]+:[0-9]+)?\|(.+)"|FieldNames=StartPosition,EndPosition,CURIE,Description|FieldTypes=integer,integer,string,string)

Entry usage:

>cu:P00001 \SecondaryStructure=(10|20|ncithesaurus:C47937|Helix)

Access:

ss = entry.custom_values["SecondaryStructure"]
ss[0].fields["StartPosition"]    # 10 (int)
ss[0].fields["Description"]      # "Helix"

Supported FieldTypes are XSD basic types (string, integer, decimal, boolean, date, time) plus enumeration(a|b|c). Coercion failures and enumeration mismatches emit UserWarning and fall back to the raw string. If no RegExp is declared, the value is split on | and zipped with FieldNames.

Other non-standard keys (no CustomKeyDef registered) still land in entry.extra as raw strings:

value = entry.extra.get("MyCustomKey")

Writing

Build a header and entries, then write:

from pefftacular import DatabaseHeader, FileHeader, SequenceEntry, write_peff

db_header = DatabaseHeader(
    prefix="sp",
    db_name="SwissProt",
    db_version="2024_01",
    number_of_entries=1,
)

file_header = FileHeader(
    peff_version="1.0",
    databases=(db_header,),
)

entry = SequenceEntry(
    prefix="sp",
    db_unique_id="P12345",
    sequence="MKTIIALSYIFCLVFA",
    pname="Example protein",
    gname="EXMP",
)

write_peff(file_header, [entry], "output.peff")

dest can be a file path string, a pathlib.Path, or a text-mode file object.

Error handling

Every exception derives from PeffError (a ValueError subclass), so you can catch any failure with one clause. Parse errors carry structured, actionable detail — .line, .context, and a .hint — and attach the offending text and the hint as exception notes, so they also show up in tracebacks:

from pefftacular import PeffError, PeffParseError, read_peff

try:
    header, entries = read_peff("malformed.peff")
except PeffParseError as e:
    print(e.line)     # 1-based line number where it failed
    print(e.context)  # the exact offending text
    print(e.hint)     # a short suggestion for how to fix it
except PeffError:
    ...               # any other pefftacular failure

Write errors raise PeffWriteError (also a PeffError), with a .hint:

from pefftacular import PeffWriteError

try:
    write_peff(file_header, entries, "/read-only/output.peff")
except PeffWriteError as e:
    print(e, e.hint)

Spec-violation warnings

Reading is permissive: the data is always returned, but anything that violates a PEFF MUST rule (out-of-range positions, missing required fields, NumberOfEntries mismatches, un-coercible custom values, …) is reported through the PeffWarning category. Promote them to errors when you want strict parsing:

import warnings
from pefftacular import PeffWarning, read_peff

warnings.simplefilter("error", PeffWarning)
header, entries = read_peff("suspect.peff")  # now raises on any spec violation

Logging

The library follows the standard logging convention (it attaches a NullHandler and never configures logging itself). Enable a behavioral trace — useful when scripting or debugging with an AI coding agent:

import logging
logging.basicConfig(level=logging.DEBUG)
logging.getLogger("pefftacular").setLevel(logging.DEBUG)

Milestones (entries read/written) log at INFO; file open, header parse, and entry counts log at DEBUG, under the pefftacular.parser / pefftacular.writer loggers.

Development

Contributor and AI-agent guidance lives in AGENTS.md. The one command to run before committing is just check (formatting, lint, types, and tests — the same gate CI enforces); just fix auto-applies formatting.

just install      # install dependencies
just check        # format-check + lint + type-check + test (pre-commit gate)
just fix          # auto-fix lint + formatting
just test         # run tests
just test-file tests/test_errors.py   # run a single test file
just cov          # run tests with coverage
just build        # build the package
just clean        # remove cache files

Run just with no arguments to list every recipe.

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pefftacular-0.4.0.tar.gz (760.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pefftacular-0.4.0-py3-none-any.whl (24.1 kB view details)

Uploaded Python 3

File details

Details for the file pefftacular-0.4.0.tar.gz.

File metadata

  • Download URL: pefftacular-0.4.0.tar.gz
  • Upload date:
  • Size: 760.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for pefftacular-0.4.0.tar.gz
Algorithm Hash digest
SHA256 fea4c2ee230c7b2596c3c1a0c841a1bc279d7cb59080242786588e0b78a55a1c
MD5 d3ac4efaa919d95c9b4e8ba130a310cc
BLAKE2b-256 be949000a41ba87d38667bbcdce8c44278cb1f00aa40ee7c8cf312d0ea33eae4

See more details on using hashes here.

File details

Details for the file pefftacular-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: pefftacular-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 24.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for pefftacular-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 be18a91aa6d9e8fa3cef6acc8c2c7832ed2321e90ddd5a0cd75424e7500165c4
MD5 df6238845a042da2e8a82bdfb3db577e
BLAKE2b-256 f6baaff4de294ce9a5508537c0a0662c241850cda114890169148d14011fd84b

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page