Skip to main content

Classify sequencing reads by non-human taxonomic content (kraken2), including the allele-based non-human fraction of reads supporting a variant.

Project description

nonhuman-screen

PyPI Python versions CI License: MIT

Classify sequencing reads by non-human taxonomic content using kraken2, and reduce the per-read verdicts to a non-human fraction (NHF) plus per-domain breakdowns (bacterial / archaeal / fungal / protist / viral / UniVec-Core).

The headline capability is the allele-based NHF: given a BAM/CRAM and a variant, compute the non-human fraction of the reads that support that variant's ALT allele — in one call, no boilerplate:

from nonhuman_screen import classify_variant_alt_reads

v = classify_variant_alt_reads(
    "sample.bam", "kraken2_db", "chr1", 12345, "A", "T", ref_fasta="ref.fa",
)
v.nonhuman_fraction   # 0.0–1.0 over ALT-supporting reads
v.fractions.bacterial # per-domain breakdown
v.supporting_reads    # denominator

nonhuman-screen is sample-agnostic: it classifies whatever reads you hand it and never asks whose they are. The same call works on a proband, a parent, a tumour, or a plain FASTQ — so it can back higher-level analyses (e.g. flagging inherited variant calls that are actually parental contamination) without knowing anything about pedigrees.

Extracted from kmer-denovo-filter, where this methodology was originally developed, so it can be reused independently.

How it works (in one paragraph)

Reads are written to a temporary FASTQ and classified by kraken2's LCA algorithm. Each read's assigned taxid is mapped to a domain using the database's nodes.dmp taxonomy. A read counts as non-human only if it is classified and lies outside the human lineage, outside the human clade, and outside UniVec-Core. Two conservative guards reduce false positives: a human-homology guard (any k-mer voting for human, taxid 9606, drops the read from every non-human numerator — important for integrating viruses such as HBV/HPV/ERVs) and UniVec-Core exclusion (synthetic vector/adapter sequences, taxid 81077, are tracked separately and never counted as contamination). See docs/methodology.md.

Install

pip install nonhuman-screen           # core engine (standard library only)
pip install 'nonhuman-screen[bam]'    # + BAM/CRAM & allele-based NHF (pysam)

The bam extra is required for the BAM/allele helpers and for the nonhuman-screen CLI (both classify modes import pysam). The core install covers only Kraken2Runner.classify_sequences on in-memory reads.

You also need the kraken2 binary on PATH and a kraken2 database — see docs/database.md.

Usage

Allele-based NHF for many variants (batched — one kraken2 call)

from nonhuman_screen import classify_variants_alt_reads

variants = [("chr1", 12345, "A", "T"), ("chr2", 999, "G", "GAC")]  # pos 0-based
for v in classify_variants_alt_reads("sample.bam", "kraken2_db", variants,
                                      ref_fasta="ref.fa"):
    print(v.variant_key, v.nonhuman_fraction, v.supporting_reads)

The reads supporting all variants' ALT alleles are gathered and classified in a single kraken2 invocation (its database load dominates runtime), then split back per variant. You never re-implement fetch/allele-filter/classify yourself.

Classify an arbitrary set of reads

from nonhuman_screen import Kraken2Runner

result = Kraken2Runner("kraken2_db").classify_sequences(
    {"read1": "ACGT...", "read2": "TTGCA..."}
)
result.nonhuman_fraction        # 0.0–1.0
result.fractions()              # TaxonomicFractions(nonhuman=…, bacterial=…, …)
result.taxonomy_available       # False if nodes.dmp was missing (fractions unreliable)
result.nonhuman_read_names      # set of read names

Command line

# Per-variant allele NHF table (headline)
nonhuman-screen classify \
    --bam sample.bam --kraken2-db kraken2_db --ref-fasta ref.fa \
    --variants candidates.vcf.gz --out-prefix sample_contam
# -> sample_contam.variant_nhf.tsv, sample_contam.summary.json

# Whole-BAM summary
nonhuman-screen classify --bam sample.bam --kraken2-db kraken2_db

See docs/cli.md.

Coordinate convention

All pos arguments in the Python API are 0-based (pysam reference coordinates); variant keys are "{chrom}:{pos}:{ref}:{alt}". The CLI reads a VCF and handles the conversion for you.

API summary

Symbol Needs [bam] Purpose
Kraken2Runner.classify_sequences(seqs) no Classify {name: seq}ClassificationResult
ClassificationResult no Per-domain read-name sets, counts, fractions(), nonhuman_fraction, taxonomy_available
TaxonomicFractions no Per-domain fractions; from_result / over_reads
read_supports_alt(read, variant_pos, ref, alt) no Does an aligned read carry the ALT allele?
parse_kmer_votes(kmer_string) no Parse kraken2 per-read k-mer detail into taxid votes
VariantNHF no Result of the allele-based helpers: variant_key, nonhuman_fraction, supporting_reads, fractions, to_dict() (pos 0-based)
classify_variant_alt_reads(...) yes Allele-based NHF for one variant → VariantNHF
classify_variants_alt_reads(...) yes Batched allele-based NHF for many variants
classify_reads_from_bam(...) yes Classify reads by name and/or locus
reads_supporting_alt(...) yes Read names supporting an ALT allele

Docker

The included Dockerfile installs the pinned kraken2 (v2.17.1) and the package with the [bam] extra:

docker build -t nonhuman-screen .

docker run --rm -v "$PWD:/data" nonhuman-screen \
    classify --bam /data/sample.bam --kraken2-db /data/kraken2_db \
             --variants /data/calls.vcf.gz --out-prefix /data/contam

The image entrypoint is nonhuman-screen; mount your BAM and kraken2 database in. You still supply the kraken2 database yourself (see docs/database.md).

Testing

pip install -e '.[bam,dev]'
pytest

Unit tests need nothing extra. The integration and functional tests additionally need the kraken2 binary and self-skip without it. The functional tests run the CLI end-to-end against committed positive/negative Streptococcus controls (a strep-injected BAM vs. clean human BAMs) — their rationale, how the fixtures are built, and what they assert are documented in docs/testing.md.

Releasing

Releases are published to PyPI automatically by GitHub Actions using Trusted Publishing (OIDC — no API tokens), via .github/workflows/release.yml:

  • Continuous builds — every push to main publishes <major>.<minor>.<run-number> (the <major>.<minor> base is read from pyproject.toml), so the latest code is always installable from PyPI.
  • Tagged releases — pushing a vX.Y.Z tag publishes exactly that version:
    git tag v0.2.0
    git push origin v0.2.0
    
  • Bumping the minor/major — edit version in pyproject.toml on main; continuous builds then track the new base.

To switch to deliberate, semver-only releases, remove main from the workflow's push.branches trigger (leaving only tags).

License

MIT — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

nonhuman_screen-0.1.6.tar.gz (41.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

nonhuman_screen-0.1.6-py3-none-any.whl (25.4 kB view details)

Uploaded Python 3

File details

Details for the file nonhuman_screen-0.1.6.tar.gz.

File metadata

  • Download URL: nonhuman_screen-0.1.6.tar.gz
  • Upload date:
  • Size: 41.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for nonhuman_screen-0.1.6.tar.gz
Algorithm Hash digest
SHA256 5e283ebe2f01fa5c15d1e9213624775a9f9ae2f7f0b7254aeef9981e8a84ff2d
MD5 e9c854a7c80f02bd2f553c4fdc05ea61
BLAKE2b-256 d1075f72705b992095b9c5fa687261d8a578701ce2fbef74e966a6f663d84c54

See more details on using hashes here.

Provenance

The following attestation bundles were made for nonhuman_screen-0.1.6.tar.gz:

Publisher: release.yml on jlanej/nonhuman-screen

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file nonhuman_screen-0.1.6-py3-none-any.whl.

File metadata

  • Download URL: nonhuman_screen-0.1.6-py3-none-any.whl
  • Upload date:
  • Size: 25.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for nonhuman_screen-0.1.6-py3-none-any.whl
Algorithm Hash digest
SHA256 9f40277f14e305c9f3b83bac8f0dcf247439c50468f3e6abdbf31d22c16542b5
MD5 2508c82f8082c2e17a079d32f7ca4a74
BLAKE2b-256 ce8a717d306fb973f8fabb3f2987761e452f8e4af8227708eda2eae23d25a066

See more details on using hashes here.

Provenance

The following attestation bundles were made for nonhuman_screen-0.1.6-py3-none-any.whl:

Publisher: release.yml on jlanej/nonhuman-screen

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page