Skip to main content

Classify sequencing reads by non-human taxonomic content (kraken2), including the allele-based non-human fraction of reads supporting a variant.

Project description

nonhuman-screen

Classify sequencing reads by non-human taxonomic content using kraken2, and reduce the per-read verdicts to a non-human fraction (NHF) plus per-domain breakdowns (bacterial / archaeal / fungal / protist / viral / UniVec-Core).

The headline capability is the allele-based NHF: given a BAM/CRAM and a variant, compute the non-human fraction of the reads that support that variant's ALT allele — in one call, no boilerplate:

from nonhuman_screen import classify_variant_alt_reads

v = classify_variant_alt_reads(
    "sample.bam", "kraken2_db", "chr1", 12345, "A", "T", ref_fasta="ref.fa",
)
v.nonhuman_fraction   # 0.0–1.0 over ALT-supporting reads
v.fractions.bacterial # per-domain breakdown
v.supporting_reads    # denominator

nonhuman-screen is sample-agnostic: it classifies whatever reads you hand it and never asks whose they are. The same call works on a proband, a parent, a tumour, or a plain FASTQ — so it can back higher-level analyses (e.g. flagging inherited variant calls that are actually parental contamination) without knowing anything about pedigrees.

Extracted from kmer-denovo-filter, where this methodology was originally developed, so it can be reused independently.

How it works (in one paragraph)

Reads are written to a temporary FASTQ and classified by kraken2's LCA algorithm. Each read's assigned taxid is mapped to a domain using the database's nodes.dmp taxonomy. A read counts as non-human only if it is classified and lies outside the human lineage, outside the human clade, and outside UniVec-Core. Two conservative guards reduce false positives: a human-homology guard (any k-mer voting for human, taxid 9606, drops the read from every non-human numerator — important for integrating viruses such as HBV/HPV/ERVs) and UniVec-Core exclusion (synthetic vector/adapter sequences, taxid 81077, are tracked separately and never counted as contamination). See docs/methodology.md.

Install

pip install nonhuman-screen           # core engine (standard library only)
pip install 'nonhuman-screen[bam]'    # + BAM/CRAM & allele-based NHF (pysam)

The bam extra is required for the BAM/allele helpers and for the nonhuman-screen CLI (both classify modes import pysam). The core install covers only Kraken2Runner.classify_sequences on in-memory reads.

You also need the kraken2 binary on PATH and a kraken2 database — see docs/database.md.

Usage

Allele-based NHF for many variants (batched — one kraken2 call)

from nonhuman_screen import classify_variants_alt_reads

variants = [("chr1", 12345, "A", "T"), ("chr2", 999, "G", "GAC")]  # pos 0-based
for v in classify_variants_alt_reads("sample.bam", "kraken2_db", variants,
                                      ref_fasta="ref.fa"):
    print(v.variant_key, v.nonhuman_fraction, v.supporting_reads)

The reads supporting all variants' ALT alleles are gathered and classified in a single kraken2 invocation (its database load dominates runtime), then split back per variant. You never re-implement fetch/allele-filter/classify yourself.

Classify an arbitrary set of reads

from nonhuman_screen import Kraken2Runner

result = Kraken2Runner("kraken2_db").classify_sequences(
    {"read1": "ACGT...", "read2": "TTGCA..."}
)
result.nonhuman_fraction        # 0.0–1.0
result.fractions()              # TaxonomicFractions(nonhuman=…, bacterial=…, …)
result.taxonomy_available       # False if nodes.dmp was missing (fractions unreliable)
result.nonhuman_read_names      # set of read names

Command line

# Per-variant allele NHF table (headline)
nonhuman-screen classify \
    --bam sample.bam --kraken2-db kraken2_db --ref-fasta ref.fa \
    --variants candidates.vcf.gz --out-prefix sample_contam
# -> sample_contam.variant_nhf.tsv, sample_contam.summary.json

# Whole-BAM summary
nonhuman-screen classify --bam sample.bam --kraken2-db kraken2_db

See docs/cli.md.

Coordinate convention

All pos arguments in the Python API are 0-based (pysam reference coordinates); variant keys are "{chrom}:{pos}:{ref}:{alt}". The CLI reads a VCF and handles the conversion for you.

API summary

Symbol Needs [bam] Purpose
Kraken2Runner.classify_sequences(seqs) no Classify {name: seq}ClassificationResult
ClassificationResult no Per-domain read-name sets, counts, fractions(), nonhuman_fraction, taxonomy_available
TaxonomicFractions no Per-domain fractions; from_result / over_reads
read_supports_alt(read, variant_pos, ref, alt) no Does an aligned read carry the ALT allele?
parse_kmer_votes(kmer_string) no Parse kraken2 per-read k-mer detail into taxid votes
VariantNHF no Result of the allele-based helpers: variant_key, nonhuman_fraction, supporting_reads, fractions, to_dict() (pos 0-based)
classify_variant_alt_reads(...) yes Allele-based NHF for one variant → VariantNHF
classify_variants_alt_reads(...) yes Batched allele-based NHF for many variants
classify_reads_from_bam(...) yes Classify reads by name and/or locus
reads_supporting_alt(...) yes Read names supporting an ALT allele

Docker

The included Dockerfile installs the pinned kraken2 (v2.17.1) and the package with the [bam] extra:

docker build -t nonhuman-screen .

docker run --rm -v "$PWD:/data" nonhuman-screen \
    classify --bam /data/sample.bam --kraken2-db /data/kraken2_db \
             --variants /data/calls.vcf.gz --out-prefix /data/contam

The image entrypoint is nonhuman-screen; mount your BAM and kraken2 database in. You still supply the kraken2 database yourself (see docs/database.md).

License

MIT — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

nonhuman_screen-0.1.0.tar.gz (38.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

nonhuman_screen-0.1.0-py3-none-any.whl (24.7 kB view details)

Uploaded Python 3

File details

Details for the file nonhuman_screen-0.1.0.tar.gz.

File metadata

  • Download URL: nonhuman_screen-0.1.0.tar.gz
  • Upload date:
  • Size: 38.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for nonhuman_screen-0.1.0.tar.gz
Algorithm Hash digest
SHA256 8ef1eb8a9e82ece7ff5b22122039af6971a777e06f32859b2192e86a7966439c
MD5 f4d0d26d231b62198f0eef72742a68e5
BLAKE2b-256 424c41006e9aebd2af1e7328ec6760c2c35a5a207c29f492cfa3f8f749f67c68

See more details on using hashes here.

Provenance

The following attestation bundles were made for nonhuman_screen-0.1.0.tar.gz:

Publisher: release.yml on jlanej/nonhuman-screen

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file nonhuman_screen-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: nonhuman_screen-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 24.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for nonhuman_screen-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 d32283d46585cf1ef6328c40b94d7a52fecaab84232bc7e846f0fa3f3fb5ffbb
MD5 7ccf722bf0dff2c495f08b10a7b0b721
BLAKE2b-256 236d9ff91554a48f85521b02d256dc7917a296cb68b5a9bfef99c9f3bf99f0ed

See more details on using hashes here.

Provenance

The following attestation bundles were made for nonhuman_screen-0.1.0-py3-none-any.whl:

Publisher: release.yml on jlanej/nonhuman-screen

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page