fastatacular
Pure-Python library for reading and writing FASTA sequence files, with optional parsing of UniProt-style description keys (OS=, OX=, GN=, PE=, SV=) and pipe-delimited identifiers (sp|P12345|EX_HUMAN, gi|12345|ref|NP_000001.1|).
It's the plain-FASTA companion to pefftacular and ships with the same read_* / *Reader / write_* shape.
Install
pip install fastatacular
Dev install:
just install
Quick start
read_fasta — load everything into memory at once:
from fastatacular import read_fasta
entries = read_fasta("proteins.fasta")
for entry in entries:
print(entry.identifier, len(entry.sequence))
FastaReader — iterate lazily without loading the full file:
from fastatacular import FastaReader
with FastaReader("proteins.fasta") as reader:
for entry in reader:
process(entry)
Data model
Each entry is a SequenceEntry:
| Field | Type | Description |
|---|---|---|
identifier |
str |
Token immediately after > (e.g. `sp |
sequence |
str |
Concatenated sequence with whitespace stripped |
prefix |
str | None |
Database prefix (sp, tr, gi, ...) when the id is pipe-delimited |
accession |
str | None |
First pipe field (e.g. P12345) |
entry_name |
str | None |
Third pipe field on UniProt ids (e.g. EX_HUMAN) |
description |
str | None |
Free text after the identifier |
pname |
str | None |
Protein name (description text, minus KEY=value pairs) |
gname |
str | None |
Gene name (GN=) |
os_name |
str | None |
Organism name (OS=) |
ncbi_tax_id |
int | None |
NCBI taxonomy ID (OX=) |
pe |
int | None |
Protein existence level (PE=) |
sv |
int | None |
Sequence version (SV=) |
extra |
dict[str, str] |
Any other KEY=value pairs found in the header |
raw_header |
str |
The original header line (without leading >) |
UniProt-style headers
from fastatacular import read_fasta
[entry] = read_fasta("one.fasta")
# >sp|P12345|EX_HUMAN Example protein OS=Homo sapiens OX=9606 GN=EXMP PE=1 SV=2
entry.prefix # "sp"
entry.accession # "P12345"
entry.entry_name # "EX_HUMAN"
entry.pname # "Example protein"
entry.os_name # "Homo sapiens"
entry.ncbi_tax_id # 9606
entry.gname # "EXMP"
entry.pe # 1
entry.sv # 2
Non-standard KEY=value pairs are captured in entry.extra. Headers with no KEY=value tokens leave description and pname populated and extra empty.
Writing
Construct entries and write them out:
from fastatacular import SequenceEntry, write_fasta
entries = [
SequenceEntry(
identifier="sp|P12345|EX_HUMAN",
sequence="MKTIIALSYIFCLVFA",
pname="Example protein",
os_name="Homo sapiens",
ncbi_tax_id=9606,
gname="EXMP",
pe=1,
sv=2,
),
]
write_fasta(entries, "output.fasta")
dest accepts a path string, a pathlib.Path, or a text-mode file object.
Sequence lines wrap at 60 characters by default. Override with line_width= (pass 0 to disable wrapping):
write_fasta(entries, "output.fasta", line_width=80)
write_fasta(entries, "single-line.fasta", line_width=0)
If raw_header is set on an entry (as it is on every entry produced by read_fasta), the writer round-trips it verbatim. Otherwise the header is rebuilt from the structured fields.
Error handling
Parse errors raise FastaParseError:
from fastatacular import FastaParseError, read_fasta
try:
entries = read_fasta("malformed.fasta")
except FastaParseError as e:
print(e.line) # offending line number
print(e.context) # surrounding line content
Write errors raise FastaWriteError.
Development
just install # install dependencies
just test # run tests
just test-v # run tests (verbose)
just cov # run tests with coverage
just lint # ruff lint
just format # ruff format
just check # lint + type check + test
just build # build the package
just clean # remove cache files
License
Release files for fastatacular 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| fastatacular-0.1.1.tar.gz | 30.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| fastatacular-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 39.9 kB
Release files / fastatacular-0.1.1.tar.gz
| Download URL | fastatacular-0.1.1.tar.gz |
|---|---|
| Size | 30.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cf47aab3d2a84280deb4b33a86fcef8d917a9a6482f5def5ee18897e3b1bc4cd
|
|
BLAKE2b-256 checksum How to use checksums |
5d2b5025d92d5ea394492d97d121b6cb7bc8bd762bacaa5c3e50cb9013626db8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency logRelease files / fastatacular-0.1.1-py3-none-any.whl
| Download URL | fastatacular-0.1.1-py3-none-any.whl |
|---|---|
| Size | 9.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
42f61054f074c775bbc2503dd7c5ae13dbc344fe4ad08d675ebfa1538bf3d2b1
|
|
BLAKE2b-256 checksum How to use checksums |
bc3bb2a18074fb8f912f2f3ac2a8fc1fb0b49546c36ad6fb251e23b5080f4c16
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency log