Skip to main content

TaxonWeave

TaxonWeave is a Python framework for provenance-aware reconciliation of taxonomic concepts and molecular sequence records.

Taxonomic names change through time, but molecular records deposited in public databases retain the names and annotations associated with their original submissions. As a result, sequence data relevant to a currently accepted species may be distributed across accepted names, historical synonyms, and other nomenclatural combinations.

TaxonWeave connects taxonomic information from the World Register of Marine Species (WoRMS) with nucleotide records from GenBank to make those relationships explicit and reproducible.

TaxonWeave is currently under active development.

Current version: 0.1.0

What TaxonWeave does

Given a scientific name, TaxonWeave:

  1. resolves the supplied name through WoRMS;
  2. identifies the currently accepted taxonomic concept and AphiaID;
  3. retrieves nomenclatural synonyms associated with that concept;
  4. searches GenBank using the accepted name and WoRMS synonyms;
  5. tracks which taxonomic search name retrieved each GenBank record;
  6. deduplicates records retrieved through multiple names;
  7. preserves the taxonomic name originally represented on each GenBank record;
  8. reconciles deposited names against the current WoRMS concept;
  9. extracts molecular, geographic, collection, and specimen-associated metadata;
  10. conservatively normalizes commonly used molecular-marker annotations; and
  11. produces a structured summary that can be exported as CSV files.
  12. Supports multi-taxon batch reconciliation and comparative auditing of molecular-data and metadata coverage.

The objective is not to replace taxonomic judgment. TaxonWeave provides a reproducible evidence-reconciliation layer connecting changing taxonomic concepts with molecular records.

Search provenance and deposited names

A central design principle of TaxonWeave is the distinction between search provenance and original taxonomic assertion.

For each GenBank record, TaxonWeave distinguishes:

  • found_via_name — the taxonomic name used in a GenBank search that retrieved the record;
  • deposited_name — the organism name represented on the GenBank record itself;
  • accepted_name — the currently accepted taxonomic concept resolved through WoRMS.

These fields should not be treated as interchangeable.

For example, a historical synonym may retrieve a GenBank record even when that record is currently represented in GenBank under the accepted species name. TaxonWeave preserves both pieces of information rather than interpreting the search term as the deposited identification.

This distinction allows synonym-mediated retrieval to be documented without rewriting the original database assertion.

Taxonomic reconciliation

TaxonWeave currently classifies deposited GenBank names relative to the WoRMS concept as:

  • ACCEPTED_NAME
  • WORMS_SYNONYM
  • UNRESOLVED_NAME
  • MISSING_NAME

Associated reconciliation states distinguish records that match the accepted name, records that can be reconciled nomenclaturally through WoRMS, and records requiring further investigation.

A WoRMS synonym is treated as a nomenclatural assertion, not as independent biological evidence that two sampled organisms are conspecific.

Molecular-marker normalization

GenBank records contain heterogeneous gene and product annotations. TaxonWeave preserves the raw annotations while also applying conservative normalization rules for commonly encountered markers.

Examples include:

  • COI, COII, and COIII
  • CytB
  • ATP6 and ATP8
  • ND1–ND6 and ND4L
  • 12S and 16S mitochondrial rRNA
  • 18S and 28S nuclear rRNA
  • histone H3

Normalization status is explicitly recorded as:

  • NORMALIZED
  • RECOGNIZED
  • AMBIGUOUS
  • UNRESOLVED

TaxonWeave deliberately avoids making unsupported inferences. For example, a generic annotation such as large subunit ribosomal RNA is retained as ambiguous rather than automatically being interpreted as 28S.

Installation

TaxonWeave currently requires Python 3.9 or later.

Clone the repository:

git clone https://github.com/parasiteguy/taxonweave.git
cd taxonweave

Create and activate a virtual environment:

python -m venv .venv
source .venv/bin/activate

Install TaxonWeave in editable mode:

pip install -e .

Command-line usage

A species can be queried directly from the command line:

taxonweave query "Ficopomatus enigmaticus"

TaxonWeave will resolve the taxonomic concept through WoRMS, search GenBank across the accepted name and associated synonyms, reconcile the resulting records, and print a summary report.

An obsolete or historical name can also be supplied:

taxonweave query "Mercierella enigmatica"

TaxonWeave first resolves the supplied name through WoRMS before constructing the molecular search.

Batch queries

TaxonWeave can reconcile multiple taxonomic concepts in a single batch workflow. The batch interface applies the same single-taxon reconciliation engine to each submitted name and produces both individual TaxonWeave reports and a combined comparative summary.

Create a CSV file containing a scientific_name column:

scientific_name
Alitta succinea
Ficopomatus enigmaticus
Boccardia proboscidea

Run the batch query:

taxonweave batch species.csv --output results

TaxonWeave creates a separate report directory for each successfully resolved taxon:

results/
├── batch_summary.csv
├── Alitta_succinea/
│   ├── taxonomy.csv
│   ├── synonyms.csv
│   ├── genbank_searches.csv
│   ├── sequences.csv
│   └── summary.csv
├── Ficopomatus_enigmaticus/
│   └── ...
└── Boccardia_proboscidea/
    └── ...

The batch_summary.csv file contains one row per submitted taxon and combines the standard TaxonWeave summary statistics with comparative measures of taxonomic reconciliation and GenBank metadata completeness.

Derived percentage fields include:

  • accepted_name_pct — percentage of retrieved GenBank records deposited under the current accepted WoRMS name
  • synonym_name_pct — percentage deposited under a current WoRMS synonym
  • unresolved_name_pct — percentage whose deposited taxon name could not be reconciled with the accepted name or WoRMS synonym set
  • geographic_metadata_pct — percentage containing geographic metadata
  • coordinate_metadata_pct — percentage containing coordinates
  • collection_date_pct — percentage containing collection-date metadata
  • voucher_metadata_pct — percentage containing a specimen_voucher qualifier

These percentages use the number of unique GenBank records retrieved for each taxon as the denominator.

A failure for one submitted taxon does not terminate the batch. Failed queries are recorded in batch_summary.csv with batch_status and batch_error, while TaxonWeave continues processing the remaining taxa.

Exporting results

Results can be exported as CSV files:

taxonweave query "Ficopomatus enigmaticus" --export

A specific output directory can be supplied:

taxonweave query "Ficopomatus enigmaticus" \
    --export \
    --output ficopomatus_results

Python usage

TaxonWeave can also be used directly from Python:

from taxonweave import query_species

report = query_species("Ficopomatus enigmaticus")

report.summary()

The report can then be exported:

report.export("ficopomatus_results")

Exported files

A TaxonWeave report currently contains five CSV files:

File Contents
taxonomy.csv Accepted WoRMS taxonomic concept and classification
synonyms.csv WoRMS nomenclatural synonyms associated with the concept
genbank_searches.csv GenBank search names and retrieval counts
sequences.csv Reconciled GenBank records and associated molecular/specimen metadata
summary.csv Summary statistics for the complete query

The sequences.csv file preserves both raw database annotations and TaxonWeave-derived reconciliation fields so that downstream analyses can distinguish source data from interpreted or normalized information.

Example: Ficopomatus enigmaticus

The invasive serpulid polychaete Ficopomatus enigmaticus provides a useful example of the TaxonWeave workflow.

WoRMS recognizes historical names including Mercierella enigmatica and Phycopomatus enigmaticus. Searching these names independently can retrieve molecular records associated with the same contemporary taxonomic concept.

TaxonWeave resolves the nomenclatural relationships, records which search names recover each sequence, deduplicates overlapping GenBank results, and preserves the taxonomic identification represented on each individual record.

This makes it possible to distinguish the history of the name used to discover a record from the taxonomic assertion associated with the record itself.

Current scope

TaxonWeave v0.1.0 currently focuses on:

WoRMS → taxonomic concept and nomenclatural history

GenBank → molecular records and associated metadata

Occurrence-data integration is not part of the core v0.1.0 workflow.

TaxonWeave currently relies on live external database services. Results may therefore change as WoRMS and GenBank records are added, revised, or reannotated.

Taxonomic reconciliation should be interpreted as evidence organization rather than automated species delimitation or taxonomic revision.

Development

Run the automated test suite with:

pytest -v

The test suite currently covers package integrity, taxonomic reconciliation logic, molecular-marker normalization, and command-line argument parsing.

Live database queries are conceptually distinct from the deterministic unit tests because external API content can change independently of TaxonWeave.

Roadmap

Planned development includes:

  • expanded provenance reporting;
  • additional reconciliation diagnostics;
  • improved handling of specimen and voucher metadata;
  • stable command-line workflows;
  • programmatic report access for downstream biodiversity-informatics analyses; and
  • a web interface built on the same TaxonWeave scientific engine.

Citation

TaxonWeave is under active development. Formal citation information will be added with the first archived software release.

License

TaxonWeave is released under the MIT License.

Copyright (c) 2026 Andrew Davinack

Metadata

Release files for TaxonWeave 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for TaxonWeave 0.2.1
File Size Uploaded
taxonweave-0.2.1.tar.gz 28.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for TaxonWeave 0.2.1
File Interpreter ABI Platform
taxonweave-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 53.8 kB

Release files / taxonweave-0.2.1.tar.gz

Download URL taxonweave-0.2.1.tar.gz
Size 28.1 kB
Tags Source
SHA-256 checksum
How to use checksums
97f7d7b9f1b5fdc716e0e9f979c38920ffa7d858fe76076ca576a137ea1b164d
BLAKE2b-256 checksum
How to use checksums
af09224cacb444c9358aeedd4d1c034a346a7ba7f4a574851b425b1e77fca8ce
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / taxonweave-0.2.1-py3-none-any.whl

Download URL taxonweave-0.2.1-py3-none-any.whl
Size 25.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
50013d66514808444962d0ddf8bfb16fc267ab74fb4921473677b68962adb642
BLAKE2b-256 checksum
How to use checksums
d9de16ac860f812cc916388c93ce31df4b8738c25d934d724e405fd46dda1517
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.2

2 release files

This release

0.2.1 This release

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page