TaxonWeave
TaxonWeave is a Python framework for provenance-aware reconciliation of taxonomic concepts and molecular sequence records.
Taxonomic names change through time, but molecular records deposited in public databases retain the names and annotations associated with their original submissions. As a result, sequence data relevant to a currently accepted species may be distributed across accepted names, historical synonyms, and other nomenclatural combinations.
TaxonWeave connects taxonomic information from the World Register of Marine Species (WoRMS) with nucleotide records from GenBank to make those relationships explicit and reproducible.
TaxonWeave is currently under active development.
Current version: 0.1.0
What TaxonWeave does
Given a scientific name, TaxonWeave:
- resolves the supplied name through WoRMS;
- identifies the currently accepted taxonomic concept and AphiaID;
- retrieves nomenclatural synonyms associated with that concept;
- searches GenBank using the accepted name and WoRMS synonyms;
- tracks which taxonomic search name retrieved each GenBank record;
- deduplicates records retrieved through multiple names;
- preserves the taxonomic name originally represented on each GenBank record;
- reconciles deposited names against the current WoRMS concept;
- extracts molecular, geographic, collection, and specimen-associated metadata;
- conservatively normalizes commonly used molecular-marker annotations; and
- produces a structured summary that can be exported as CSV files.
- Supports multi-taxon batch reconciliation and comparative auditing of molecular-data and metadata coverage.
The objective is not to replace taxonomic judgment. TaxonWeave provides a reproducible evidence-reconciliation layer connecting changing taxonomic concepts with molecular records.
Search provenance and deposited names
A central design principle of TaxonWeave is the distinction between search provenance and original taxonomic assertion.
For each GenBank record, TaxonWeave distinguishes:
found_via_name— the taxonomic name used in a GenBank search that retrieved the record;deposited_name— the organism name represented on the GenBank record itself;accepted_name— the currently accepted taxonomic concept resolved through WoRMS.
These fields should not be treated as interchangeable.
For example, a historical synonym may retrieve a GenBank record even when that record is currently represented in GenBank under the accepted species name. TaxonWeave preserves both pieces of information rather than interpreting the search term as the deposited identification.
This distinction allows synonym-mediated retrieval to be documented without rewriting the original database assertion.
Taxonomic reconciliation
TaxonWeave currently classifies deposited GenBank names relative to the WoRMS concept as:
ACCEPTED_NAMEWORMS_SYNONYMUNRESOLVED_NAMEMISSING_NAME
Associated reconciliation states distinguish records that match the accepted name, records that can be reconciled nomenclaturally through WoRMS, and records requiring further investigation.
A WoRMS synonym is treated as a nomenclatural assertion, not as independent biological evidence that two sampled organisms are conspecific.
Molecular-marker normalization
GenBank records contain heterogeneous gene and product annotations. TaxonWeave preserves the raw annotations while also applying conservative normalization rules for commonly encountered markers.
Examples include:
- COI, COII, and COIII
- CytB
- ATP6 and ATP8
- ND1–ND6 and ND4L
- 12S and 16S mitochondrial rRNA
- 18S and 28S nuclear rRNA
- histone H3
Normalization status is explicitly recorded as:
NORMALIZEDRECOGNIZEDAMBIGUOUSUNRESOLVED
TaxonWeave deliberately avoids making unsupported inferences. For example, a generic annotation such as large subunit ribosomal RNA is retained as ambiguous rather than automatically being interpreted as 28S.
Installation
TaxonWeave currently requires Python 3.9 or later.
Clone the repository:
git clone https://github.com/parasiteguy/taxonweave.git
cd taxonweave
Create and activate a virtual environment:
python -m venv .venv
source .venv/bin/activate
Install TaxonWeave in editable mode:
pip install -e .
Command-line usage
A species can be queried directly from the command line:
taxonweave query "Ficopomatus enigmaticus"
TaxonWeave will resolve the taxonomic concept through WoRMS, search GenBank across the accepted name and associated synonyms, reconcile the resulting records, and print a summary report.
An obsolete or historical name can also be supplied:
taxonweave query "Mercierella enigmatica"
TaxonWeave first resolves the supplied name through WoRMS before constructing the molecular search.
Batch queries
TaxonWeave can reconcile multiple taxonomic concepts in a single batch workflow. The batch interface applies the same single-taxon reconciliation engine to each submitted name and produces both individual TaxonWeave reports and a combined comparative summary.
Create a CSV file containing a scientific_name column:
scientific_name
Alitta succinea
Ficopomatus enigmaticus
Boccardia proboscidea
Run the batch query:
taxonweave batch species.csv --output results
TaxonWeave creates a separate report directory for each successfully resolved taxon:
results/
├── batch_summary.csv
├── Alitta_succinea/
│ ├── taxonomy.csv
│ ├── synonyms.csv
│ ├── genbank_searches.csv
│ ├── sequences.csv
│ └── summary.csv
├── Ficopomatus_enigmaticus/
│ └── ...
└── Boccardia_proboscidea/
└── ...
The batch_summary.csv file contains one row per submitted taxon and combines the standard TaxonWeave summary statistics with comparative measures of taxonomic reconciliation and GenBank metadata completeness.
Derived percentage fields include:
accepted_name_pct— percentage of retrieved GenBank records deposited under the current accepted WoRMS namesynonym_name_pct— percentage deposited under a current WoRMS synonymunresolved_name_pct— percentage whose deposited taxon name could not be reconciled with the accepted name or WoRMS synonym setgeographic_metadata_pct— percentage containing geographic metadatacoordinate_metadata_pct— percentage containing coordinatescollection_date_pct— percentage containing collection-date metadatavoucher_metadata_pct— percentage containing aspecimen_voucherqualifier
These percentages use the number of unique GenBank records retrieved for each taxon as the denominator.
A failure for one submitted taxon does not terminate the batch. Failed queries are recorded in batch_summary.csv with batch_status and batch_error, while TaxonWeave continues processing the remaining taxa.
Exporting results
Results can be exported as CSV files:
taxonweave query "Ficopomatus enigmaticus" --export
A specific output directory can be supplied:
taxonweave query "Ficopomatus enigmaticus" \
--export \
--output ficopomatus_results
Python usage
TaxonWeave can also be used directly from Python:
from taxonweave import query_species
report = query_species("Ficopomatus enigmaticus")
report.summary()
The report can then be exported:
report.export("ficopomatus_results")
Exported files
A TaxonWeave report currently contains five CSV files:
| File | Contents |
|---|---|
taxonomy.csv |
Accepted WoRMS taxonomic concept and classification |
synonyms.csv |
WoRMS nomenclatural synonyms associated with the concept |
genbank_searches.csv |
GenBank search names and retrieval counts |
sequences.csv |
Reconciled GenBank records and associated molecular/specimen metadata |
summary.csv |
Summary statistics for the complete query |
The sequences.csv file preserves both raw database annotations and TaxonWeave-derived reconciliation fields so that downstream analyses can distinguish source data from interpreted or normalized information.
Example: Ficopomatus enigmaticus
The invasive serpulid polychaete Ficopomatus enigmaticus provides a useful example of the TaxonWeave workflow.
WoRMS recognizes historical names including Mercierella enigmatica and Phycopomatus enigmaticus. Searching these names independently can retrieve molecular records associated with the same contemporary taxonomic concept.
TaxonWeave resolves the nomenclatural relationships, records which search names recover each sequence, deduplicates overlapping GenBank results, and preserves the taxonomic identification represented on each individual record.
This makes it possible to distinguish the history of the name used to discover a record from the taxonomic assertion associated with the record itself.
Current scope
TaxonWeave v0.1.0 currently focuses on:
WoRMS → taxonomic concept and nomenclatural history
GenBank → molecular records and associated metadata
Occurrence-data integration is not part of the core v0.1.0 workflow.
TaxonWeave currently relies on live external database services. Results may therefore change as WoRMS and GenBank records are added, revised, or reannotated.
Taxonomic reconciliation should be interpreted as evidence organization rather than automated species delimitation or taxonomic revision.
Development
Run the automated test suite with:
pytest -v
The test suite currently covers package integrity, taxonomic reconciliation logic, molecular-marker normalization, and command-line argument parsing.
Live database queries are conceptually distinct from the deterministic unit tests because external API content can change independently of TaxonWeave.
Roadmap
Planned development includes:
- expanded provenance reporting;
- additional reconciliation diagnostics;
- improved handling of specimen and voucher metadata;
- stable command-line workflows;
- programmatic report access for downstream biodiversity-informatics analyses; and
- a web interface built on the same TaxonWeave scientific engine.
Citation
TaxonWeave is under active development. Formal citation information will be added with the first archived software release.
License
TaxonWeave is released under the MIT License.
Copyright (c) 2026 Andrew Davinack
Metadata
Release files for TaxonWeave 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| taxonweave-0.2.1.tar.gz | 28.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| taxonweave-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 53.8 kB
Release files / taxonweave-0.2.1.tar.gz
| Download URL | taxonweave-0.2.1.tar.gz |
|---|---|
| Size | 28.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
97f7d7b9f1b5fdc716e0e9f979c38920ffa7d858fe76076ca576a137ea1b164d
|
|
BLAKE2b-256 checksum How to use checksums |
af09224cacb444c9358aeedd4d1c034a346a7ba7f4a574851b425b1e77fca8ce
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / taxonweave-0.2.1-py3-none-any.whl
| Download URL | taxonweave-0.2.1-py3-none-any.whl |
|---|---|
| Size | 25.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
50013d66514808444962d0ddf8bfb16fc267ab74fb4921473677b68962adb642
|
|
BLAKE2b-256 checksum How to use checksums |
d9de16ac860f812cc916388c93ce31df4b8738c25d934d724e405fd46dda1517
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency log