fastatacular
A small, dependency-free library for reading and writing FASTA sequence files in Python. It's built for proteomics and genomics pipelines that need fast, predictable FASTA parsing without pulling in a bioinformatics megapackage.
It understands UniProt-style description keys (OS=, OX=, GN=, PE=, SV=) and pipe-delimited identifiers (sp|P12345|EX_HUMAN, gi|12345|ref|NP_000001.1|) out of the box, so you get structured fields instead of a header string to parse yourself.
Highlights
- Zero dependencies — pure Python, nothing else to install.
- Two ways to read —
read_fastafor the whole file at once,FastaReaderto stream entries lazily without loading everything into memory. - Compressed input —
.gz,.bz2and.xzfiles are read directly (detected from the file's magic bytes). - UniProt headers parsed for you — accession, organism, gene name, protein existence, and sequence version come back as typed fields, not a string you have to split yourself.
- Round-trip safe — entries produced by
read_fastawrite back out byte-for-byte compatible headers. - Actionable parse errors —
FastaParseErrorreports the offending line number and surrounding context. - Shares its API shape with pefftacular, the PEFF (PSI Extended FASTA) sibling library, so switching formats doesn't mean relearning the interface.
Install
pip install fastatacular
Dev install:
just install
Quick start
read_fasta — load everything into memory at once:
from fastatacular import read_fasta
entries = read_fasta("proteins.fasta")
for entry in entries:
print(entry.identifier, len(entry.sequence))
FastaReader — iterate lazily without loading the full file:
from fastatacular import FastaReader
with FastaReader("proteins.fasta") as reader:
for entry in reader:
process(entry)
Compressed files are read transparently: gzip, bzip2 and xz, detected from the magic
bytes, not the file name. Pipes, FIFOs and /dev/stdin work too:
entries = read_fasta("uniprot_sprot.fasta.gz")
A PEFF file also reads as plain FASTA: its # file header lines are skipped and each
entry keeps its identifier, sequence and raw description. Use
pefftacular to parse the PEFF annotations.
Data model
Each entry is a SequenceEntry:
| Field | Type | Description |
|---|---|---|
identifier |
str |
Token immediately after > (e.g. `sp |
sequence |
str |
Concatenated sequence with whitespace stripped |
prefix |
str | None |
Database prefix (sp, tr, gi, ...) when the id is pipe-delimited |
accession |
str | None |
Second pipe field (e.g. P12345 in sp|P12345|EX_HUMAN) |
entry_name |
str | None |
Third pipe field on UniProt ids (e.g. EX_HUMAN) |
description |
str | None |
Free text after the identifier |
pname |
str | None |
Protein name (description text, minus KEY=value pairs) |
gname |
str | None |
Gene name (GN=) |
os_name |
str | None |
Organism name (OS=) |
ncbi_tax_id |
int | None |
NCBI taxonomy ID (OX=) |
pe |
int | None |
Protein existence level (PE=) |
sv |
int | None |
Sequence version (SV=) |
extra |
dict[str, str] |
Any other KEY=value pairs found in the header |
raw_header |
str |
The original header line (without leading >) |
UniProt-style headers
from fastatacular import read_fasta
[entry] = read_fasta("one.fasta")
# >sp|P12345|EX_HUMAN Example protein OS=Homo sapiens OX=9606 GN=EXMP PE=1 SV=2
entry.prefix # "sp"
entry.accession # "P12345"
entry.entry_name # "EX_HUMAN"
entry.pname # "Example protein"
entry.os_name # "Homo sapiens"
entry.ncbi_tax_id # 9606
entry.gname # "EXMP"
entry.pe # 1
entry.sv # 2
Non-standard KEY=value pairs are captured in entry.extra. Headers with no KEY=value tokens leave description and pname populated and extra empty.
Writing
Construct entries and write them out:
from fastatacular import SequenceEntry, write_fasta
entries = [
SequenceEntry(
identifier="sp|P12345|EX_HUMAN",
sequence="MKTIIALSYIFCLVFA",
pname="Example protein",
os_name="Homo sapiens",
ncbi_tax_id=9606,
gname="EXMP",
pe=1,
sv=2,
),
]
write_fasta(entries, "output.fasta")
dest accepts a path string, a pathlib.Path, or a text-mode file object.
Sequence lines wrap at 60 characters by default. Override with line_width= (pass 0 to disable wrapping):
write_fasta(entries, "output.fasta", line_width=80)
write_fasta(entries, "single-line.fasta", line_width=0)
If raw_header is set on an entry (as it is on every entry produced by read_fasta) and still matches the entry's structured fields, the writer round-trips it verbatim. If you changed a field (for example dataclasses.replace(entry, gname="XYZ")), or raw_header is empty, the header is rebuilt from the structured fields, so your edit is written.
Decoy databases
Build target-decoy databases for FDR estimation with reverse, pseudo_reverse,
shuffle, debruijn (repeat-preserving, Moosa et al. 2020) or markov (order-2
chains shipped for human, mouse, yeast and E. coli, trained on UniProt, CC BY 4.0).
Every method takes keep_residues (e.g. "KR"), keep_nterm and keep_cterm, and the
same seed always gives the same decoys. Pure Python, no extra dependencies.
from fastatacular import is_decoy, make_decoys, read_fasta, write_decoy_fasta
write_decoy_fasta("human.fasta", "human_td.fasta", method="pseudo_reverse", seed=1)
decoys = make_decoys(read_fasta("human.fasta"), method="markov", model="human", keep_residues="KR", seed=1)
next(decoys).identifier # "DECOY_sp|..."; is_decoy(entry) checks the prefix
See docs/decoys.md for every option, the Markov model data and per-method quality numbers on the human proteome.
Random access by identifier or accession
FastaIndex(path) reads the file once and keeps only each entry's key and byte range;
index[key] then reads and parses just that entry.
from fastatacular import FastaIndex
index = FastaIndex("human.fasta")
entry = index["sp|P31946|1433B_HUMAN"] # SequenceEntry, read from disk on demand
"sp|P31946|1433B_HUMAN" in index, len(index) # no file access
index.write_fai() # human.fasta.fai, samtools-compatible
index = FastaIndex.from_fai("human.fasta") # later: load the .fai instead of scanning
by_acc = FastaIndex("human.fasta", key="accession")
by_acc["P31946"]
- By default keys are full identifiers (the first header word, the
.fainame, as in samtools). This works on target-decoy databases, wheresp|P1|XandDECOY_sp|P1|XorReverse_sp|P1|Xshare an accession. key="accession"keys entries by accession (P31946forsp|P31946|1433B_HUMAN, or the whole identifier when it has no|). It needs unique accessions.- A repeated key raises
FastaErrornaming the first one. Real databases do repeat identifiers (IP2 exports, merged databases):FastaIndex(path, duplicates="first")keeps the first entry for each key and skips the rest, assamtools faidxdoes, and logs one warning with the number skipped.from_faitakes the same option. - A missing key raises
FastaKeyError, which is also aKeyError. - The file must be a regular, uncompressed file. A FIFO, pipe or directory raises
FastaError: save the input to a file first, or read it once withFastaReader. gzip (including bgzip), bzip2 and xz raiseFastaError: decompress first (gunzip -k human.fasta.gz). bgzip/.gziis not supported. .faicaveats: the name column is the first header word (mapped to the accession on load withkey="accession"). Like samtools,write_fai()needs every sequence line of an entry but the last to have the same length, and no comment or blank lines inside a sequence; otherwise it raisesFastaError(rewrite the file withwrite_fastafirst).from_faichecks each entry's header against the.fainame but does not re-count residues, so rebuild the.faiwhenever the FASTA changes. Entries with no sequence raise (samtools skips them). samtools and pysam refuse a FASTA file that starts with a byte-order mark, so the.faiof such a file works only withFastaIndex.
Tables with pandas or polars
to_records(source) returns one plain dict per entry, so any data-frame library can
take the result directly. fastatacular does not ship or require pandas or polars;
install whichever you use. (The test suite runs these examples only when the library is
installed.)
import pandas as pd
import polars as pl
from fastatacular import to_records
records = to_records("human.fasta") # a path, an open text handle, or entries
df = pd.DataFrame(records)
human = pl.DataFrame(records).filter(pl.col("ncbi_tax_id") == 9606)
FastaReader.to_records() and SequenceEntry.to_record() give the same dicts. Every
record has these keys, in this order (fastatacular.RECORD_KEYS):
| key | type | value |
|---|---|---|
identifier |
str | text after > up to the first whitespace |
prefix, accession, entry_name |
str or None | sp, P12345, NAME from sp|P12345|NAME |
pname |
str or None | protein name (description before the first KEY=) |
gname, os_name |
str or None | GN=, OS= |
ncbi_tax_id, pe, sv |
int or None | OX=, PE=, SV= |
description |
str or None | everything after the identifier |
extra |
str or None | other KEY=value pairs, space-separated |
raw_header |
str | the header line without > |
length |
int | sequence length |
sequence |
str | the residues |
The same fields in pefftacular records have other names where each package follows its own model. Rename these to put both in one frame:
| field | fastatacular | pefftacular |
|---|---|---|
| database prefix | prefix |
prefix |
| accession | accession |
db_unique_id |
| entry name | entry_name |
id |
| protein name, gene | pname, gname |
pname, gname |
| organism name | os_name |
tax_name |
| taxon id, PE, SV | ncbi_tax_id, pe, sv |
ncbi_tax_id, pe, sv |
| other keys | extra, KEY=value pairs |
extra, \Key=value pairs |
| length, residues | length, sequence |
length, sequence |
Error handling
Parse errors raise FastaParseError:
from fastatacular import FastaParseError, read_fasta
try:
entries = read_fasta("malformed.fasta")
except FastaParseError as e:
print(e.line) # offending line number
print(e.context) # surrounding line content
Input that is not UTF-8, or a corrupt compressed file, also raises FastaParseError
(chained to the underlying UnicodeDecodeError or OSError).
Write errors raise FastaWriteError, whose index names the bad entry. Every entry is
validated before anything is written, so a failed write_fasta leaves no partial file.
Both errors subclass FastaError (a ValueError), so except FastaError catches either.
SequenceEntry is frozen but not hashable (its extra field is a dict), so key sets and
dicts by entry.identifier, not by the entry. A FastaReader is single-pass: iterate it
once, or open a new one to read the file again.
Development
just install # install dependencies
just test # run tests
just test-v # run tests (verbose)
just cov # run tests with coverage
just lint # ruff lint
just format # ruff format
just check # lint + type check + test
just build # build the package
just clean # remove cache files
Citation
If you use fastatacular in research, please cite the archived software release. Machine-readable citation metadata is available in CITATION.cff; GitHub's Cite this repository menu can render it as APA or BibTeX. DOI: 10.5281/zenodo.22926358.
License
Release files for fastatacular 1.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| fastatacular-1.1.0.tar.gz | 193.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| fastatacular-1.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 294.7 kB
Release files / fastatacular-1.1.0.tar.gz
| Download URL | fastatacular-1.1.0.tar.gz |
|---|---|
| Size | 193.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bf968169dd597c2f903887067f9fab1024659de0195186670b3df70ab359965f
|
|
BLAKE2b-256 checksum How to use checksums |
4656b0338f619c1e0f6d9ebe87ac35d2bd00b319c48d3048a11a50b06f392444
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency logRelease files / fastatacular-1.1.0-py3-none-any.whl
| Download URL | fastatacular-1.1.0-py3-none-any.whl |
|---|---|
| Size | 101.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
41f6082234631e895f6bbbfab706cdd1f7d1ee2eb41f72ef5d7c48cf90da87cc
|
|
BLAKE2b-256 checksum How to use checksums |
4830a7c07f732127921737f394d62df539f4fa294d8ece2bb7d85d254026cb52
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency log