fastatacular
A small, dependency-free library for reading and writing FASTA sequence files in Python. It's built for proteomics and genomics pipelines that need fast, predictable FASTA parsing without pulling in a bioinformatics megapackage.
It understands UniProt-style description keys (OS=, OX=, GN=, PE=, SV=) and pipe-delimited identifiers (sp|P12345|EX_HUMAN, gi|12345|ref|NP_000001.1|) out of the box, so you get structured fields instead of a header string to parse yourself.
Highlights
- Zero dependencies — pure Python, nothing else to install.
- Two ways to read —
read_fastafor the whole file at once,FastaReaderto stream entries lazily without loading everything into memory. - UniProt headers parsed for you — accession, organism, gene name, protein existence, and sequence version come back as typed fields, not a string you have to split yourself.
- Round-trip safe — entries produced by
read_fastawrite back out byte-for-byte compatible headers. - Actionable parse errors —
FastaParseErrorreports the offending line number and surrounding context. - Shares its API shape with pefftacular, the PEFF (PSI Extended FASTA) sibling library, so switching formats doesn't mean relearning the interface.
Install
pip install fastatacular
Dev install:
just install
Quick start
read_fasta — load everything into memory at once:
from fastatacular import read_fasta
entries = read_fasta("proteins.fasta")
for entry in entries:
print(entry.identifier, len(entry.sequence))
FastaReader — iterate lazily without loading the full file:
from fastatacular import FastaReader
with FastaReader("proteins.fasta") as reader:
for entry in reader:
process(entry)
Data model
Each entry is a SequenceEntry:
| Field | Type | Description |
|---|---|---|
identifier |
str |
Token immediately after > (e.g. `sp |
sequence |
str |
Concatenated sequence with whitespace stripped |
prefix |
str | None |
Database prefix (sp, tr, gi, ...) when the id is pipe-delimited |
accession |
str | None |
Second pipe field (e.g. P12345 in sp|P12345|EX_HUMAN) |
entry_name |
str | None |
Third pipe field on UniProt ids (e.g. EX_HUMAN) |
description |
str | None |
Free text after the identifier |
pname |
str | None |
Protein name (description text, minus KEY=value pairs) |
gname |
str | None |
Gene name (GN=) |
os_name |
str | None |
Organism name (OS=) |
ncbi_tax_id |
int | None |
NCBI taxonomy ID (OX=) |
pe |
int | None |
Protein existence level (PE=) |
sv |
int | None |
Sequence version (SV=) |
extra |
dict[str, str] |
Any other KEY=value pairs found in the header |
raw_header |
str |
The original header line (without leading >) |
UniProt-style headers
from fastatacular import read_fasta
[entry] = read_fasta("one.fasta")
# >sp|P12345|EX_HUMAN Example protein OS=Homo sapiens OX=9606 GN=EXMP PE=1 SV=2
entry.prefix # "sp"
entry.accession # "P12345"
entry.entry_name # "EX_HUMAN"
entry.pname # "Example protein"
entry.os_name # "Homo sapiens"
entry.ncbi_tax_id # 9606
entry.gname # "EXMP"
entry.pe # 1
entry.sv # 2
Non-standard KEY=value pairs are captured in entry.extra. Headers with no KEY=value tokens leave description and pname populated and extra empty.
Writing
Construct entries and write them out:
from fastatacular import SequenceEntry, write_fasta
entries = [
SequenceEntry(
identifier="sp|P12345|EX_HUMAN",
sequence="MKTIIALSYIFCLVFA",
pname="Example protein",
os_name="Homo sapiens",
ncbi_tax_id=9606,
gname="EXMP",
pe=1,
sv=2,
),
]
write_fasta(entries, "output.fasta")
dest accepts a path string, a pathlib.Path, or a text-mode file object.
Sequence lines wrap at 60 characters by default. Override with line_width= (pass 0 to disable wrapping):
write_fasta(entries, "output.fasta", line_width=80)
write_fasta(entries, "single-line.fasta", line_width=0)
If raw_header is set on an entry (as it is on every entry produced by read_fasta), the writer round-trips it verbatim. Otherwise the header is rebuilt from the structured fields.
Error handling
Parse errors raise FastaParseError:
from fastatacular import FastaParseError, read_fasta
try:
entries = read_fasta("malformed.fasta")
except FastaParseError as e:
print(e.line) # offending line number
print(e.context) # surrounding line content
Write errors raise FastaWriteError.
Development
just install # install dependencies
just test # run tests
just test-v # run tests (verbose)
just cov # run tests with coverage
just lint # ruff lint
just format # ruff format
just check # lint + type check + test
just build # build the package
just clean # remove cache files
Citation
If you use fastatacular in research, please cite the archived software release. Machine-readable citation metadata is available in CITATION.cff; GitHub's Cite this repository menu can render it as APA or BibTeX. DOI: 10.5281/zenodo.22926358.
License
Release files for fastatacular 0.1.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| fastatacular-0.1.3.tar.gz | 43.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| fastatacular-0.1.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 53.5 kB
Release files / fastatacular-0.1.3.tar.gz
| Download URL | fastatacular-0.1.3.tar.gz |
|---|---|
| Size | 43.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b51940f25f26d9a0b72936759ad8158e2e57b90999edede24f2b00b895716b2d
|
|
BLAKE2b-256 checksum How to use checksums |
9c7b00ee30422a4d80dfd5780f59599c0889066491c5c03297145c8b6044c228
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency logRelease files / fastatacular-0.1.3-py3-none-any.whl
| Download URL | fastatacular-0.1.3-py3-none-any.whl |
|---|---|
| Size | 10.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6114cd80b2647cda8ead1a1554ef44f02af1b12f4c81bcab24bd52bac2ec881e
|
|
BLAKE2b-256 checksum How to use checksums |
8afac7d1ff760f4b054068cda35a9a9299fd28c59b9cc2ccacf4e6025cc45197
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency log