Skip to main content

Epidemiological linkage inference from temporal and genetic data with an E/P/I infectiousness model.

Project description

EpiLink

EpiLink scores how compatible a pair of samples is with recent transmission scenarios using sampling-time differences and consensus genetic distance.

It is useful when you have:

  • a sampling-time difference in days
  • a consensus genetic distance in mutations
  • a question like "is this pair more compatible with direct transmission or a recent shared ancestor?"

EpiLink returns per-scenario compatibility scores and can also sum scores across a user-defined target subset such as ["ad(0)", "ca(0,0)"].

Installation

Clone the repository first if you are starting from GitHub:

git clone https://github.com/ydnkka/epilink.git
cd epilink

The repository environment is the easiest way to get everything needed for the package, examples, and simulation helpers:

conda env create -f environment.yml
conda activate epilink

If you prefer pip:

python -m pip install -e .
python -m pip install networkx pandas

EpiLink requires Python 3.10 or newer.

Scenario labels

  • ad(0): direct ancestor-descendant transmission
  • ad(1): ancestor-descendant transmission with one hidden intermediate
  • ca(0,0): a recent shared common ancestor with one branch to each sampled case
  • ca(m_i,m_j): a common-ancestor scenario with m_i and m_j hidden generations on each branch

maximum_depth controls how many of these latent scenarios are generated.

Which method to use

  • score_pair(...): one observed pair, plus a full per-scenario breakdown
  • score_target(...): only the target score, for scalar or array inputs
  • pairwise_model(...): a cached scorer for repeatedly evaluating the same target subset

Each individual scenario compatibility lies in [0, 1]. If target contains multiple scenarios, target_compatibility is the sum across that subset, so it can be greater than 1.

Quick start

from epilink import EpiLink, InfectiousnessToTransmission

profile = InfectiousnessToTransmission(rng_seed=2026)

model = EpiLink(
    transmission_profile=profile,
    maximum_depth=2,
    mc_samples=20000,
    target=["ad(0)", "ca(0,0)"],
    mutation_process="stochastic",
)

result = model.score_pair(
    sample_time_difference=3.0,
    genetic_distance=2.0,
)

print(result["target_labels"])
print(result["target_compatibility"])
print(result["scenario_scores"]["ad(0)"]["compatibility"])

More examples

Score only a target subset

Use score_target when you only care about the combined score:

score = model.score_target(
    sample_time_difference=3.0,
    genetic_distance=2.0,
    target=["ad(0)", "ad(1)", "ca(0,0)"],
)

print(score)

Use Scenario objects instead of strings

from epilink import Scenario

score = model.score_target(
    sample_time_difference=3.0,
    genetic_distance=2.0,
    target=[
        Scenario(kind="ad", intermediates=0),
        Scenario(kind="ca", branch_to_i=0, branch_to_j=0),
    ],
)

Score many pairs at once

score_target and pairwise_model broadcast NumPy inputs, so you can score a whole grid or batch efficiently:

import numpy as np

pairwise = model.pairwise_model(target=["ad(0)", "ca(0,0)"])

time_differences = np.array([[0.0], [2.0], [4.0]])
genetic_distances = np.array([[0.0, 1.0, 2.0, 3.0]])

scores = pairwise(time_differences, genetic_distances)
print(scores.shape)  # (3, 4)

Build a toy simulated pair table

The simulation helpers are useful for generating synthetic examples and benchmarking downstream workflows:

import networkx as nx

from epilink import (
    build_pairwise_case_table,
    simulate_epidemic_dates,
    simulate_genomic_sequences,
)

tree = nx.DiGraph(
    [
        ("case-0", "case-1"),
        ("case-0", "case-2"),
    ]
)

dated_tree = simulate_epidemic_dates(profile, tree, fraction_sampled=1.0)
simulated = simulate_genomic_sequences(profile, dated_tree, genome_length=500)
pair_table = build_pairwise_case_table(simulated["packed"], dated_tree)

print(pair_table.head())

Mutation models

  • mutation_process="deterministic" compares the observation with expected mutation counts
  • mutation_process="stochastic" compares the observation with Poisson mutation-count draws

The stochastic option is usually the better choice when you want mutation-count variability to be part of the score.

Background and model characteristics

Citation

If you use EpiLink in research, please cite the software metadata in CITATION.cff. The underlying infectiousness model is:

  1. Hart WS, Maini PK, Thompson RN. High infectiousness immediately before COVID-19 symptom onset highlights the importance of continued contact tracing. eLife. 2021;10:e65534. http://dx.doi.org/10.7554/eLife.65534

License

MIT. See LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

epilink-0.1.3.tar.gz (1.7 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

epilink-0.1.3-py3-none-any.whl (20.3 kB view details)

Uploaded Python 3

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page