admixture-cache
Precomputed-P supervised-ADMIXTURE projection cache. Build the slow training pass once per panel × K × clusters_yaml combo; project new targets in ~2 seconds.
Why this exists
Supervised ADMIXTURE training on a real-world panel takes hours to days per restart (K=21 regional cache: ~12-14 hr × 5 restarts; K=4 ancestral_cluster: ~5-7 hr × 5 restarts). For consumer pipelines serving many users, re-running this training per target is wasteful — the P matrix is determined almost entirely by the panel, not the target.
admixture-cache splits the supervised-ADMIXTURE workflow into:
- Panel cache build (operator, slow, one-time per panel update): stock ADMIXTURE × N restarts → cache best-LL P matrix + multimodality SD + manifest.
- Per-target projection (consumer, fast, every run): align target.bed to cached panel variants + axes (plink2), load dosages, solve for Q via scipy SLSQP under the standard binomial admixture likelihood.
The projection math matches stock ADMIXTURE Q values to within ~1e-5 absolute on representative workloads (15K × 850K matrix at K=4).
Install
pip install admixture-cache
Python 3.11 through 3.14 are supported. End-to-end paths require ADMIXTURE (for build) and plink2 (for project / verify) on PATH. Pure-library use without those binaries is fine — only the build/projection orchestrators shell out.
Quickstart — library
from pathlib import Path
from admixture_cache import build_panel_cache, project_target
# One-time, slow (~hours per restart per cache)
manifest = build_panel_cache(
panel_bed=Path("panel.bed"),
panel_pop_file=Path("panel.pop"),
clusters_yaml=Path("clusters.yaml"),
k=21,
cache_dir=Path("data/regional_k21_cache/"),
admixture_runner=my_tool_runner, # see ToolRunner Protocol below
track="regional",
panel_id="aadr_v66_ho",
panel_version="v66.0",
admixture_version="1.4.0",
seeds=[1, 2, 3, 4, 5],
sd_threshold=0.02,
)
# Per-target, fast (~2 seconds end-to-end)
result = project_target(
target_bed=Path("target.bed"),
cache_dir=Path("data/regional_k21_cache/"),
plink2_runner=my_plink2_runner,
work_dir=Path("scratch/projection/"),
)
print(result.target_q) # K-vector
print(result.cluster_order) # K names
print(result.panel_stability_max_sd) # cached panel restart_sd
Quickstart — CLI
Installing the package registers the admixture-cache console script with four subcommands:
# 1. Build a panel cache (slow, one-time).
admixture-cache build \
--panel-bed panel.bed \
--panel-pop panel.pop \
--clusters-yaml clusters.yaml \
--k 21 \
--cache-dir data/regional_k21_cache/ \
--track regional \
--panel-id aadr_v66_ho \
--panel-version v66.0 \
--seeds 1,2,3,4,5
# 2. Project a target against an existing cache (fast).
admixture-cache project \
--target-bed target.bed \
--cache-dir data/regional_k21_cache/ \
--work-dir scratch/projection/
# 3. Check whether a cache matches the current panel/YAML/K config.
admixture-cache verify \
--panel-bed panel.bed \
--clusters-yaml clusters.yaml \
--k 21 \
--cache-dir data/regional_k21_cache/
# 4. Fetch a canonical published cache from GitHub Releases.
admixture-cache download --list # enumerate
admixture-cache download regional_k21_aadr_v66_ho # install
admixture-cache download regional_k21_aadr_v66_ho \
--cache-root ~/.admixture-cache/caches \
--cache-version v2 \
--force # pin + overwrite
Caches install at <cache-root>/<name>/ (default: ~/.admixture-cache/caches/, or $ADMIXTURE_CACHE_ROOT if set). The downloader streams the tarball, verifies its SHA-256, validates the extracted manifest, and atomically renames into place — partial downloads never leave a half-installed cache.
Publishing your own canonical caches: see docs/PUBLISH_CACHE.md for the tag convention + tarball format the discovery code expects.
The default SubprocessToolRunner runs the local admixture / plink2 binaries on PATH; override with --admixture-binary / --plink2-binary to point at a specific build.
build, project, and verify all surface a non-zero exit code on failure with a descriptive error: … line on stderr. project --json emits machine-readable JSON instead of human-readable text.
ToolRunner Protocol
When calling the library from Python (rather than via the CLI), pass any object satisfying the ToolRunner Protocol:
from collections.abc import Callable
from pathlib import Path
class MyToolRunner:
def run(
self,
*,
args: list[str],
cwd: Path,
log_dir: Path,
timeout_seconds: int = 600,
# The two kwargs below are OPTIONAL but REQUIRED for
# parallel `build_panel_cache` (max_parallel_restarts > 1):
log_name: str | None = None,
pid_callback: Callable[[int], None] | None = None,
) -> object:
...
log_name— admixture-cache passes the per-restart canonical log filename (e.g.restart_3.out). Honor it when set; fall back to your own naming scheme whenNone. Required for parallel mode (concurrent restarts sharelog_dirand need disambiguated filenames).pid_callback— call with the subprocess PID immediately after spawning. admixture-cache uses this to SIGTERM in-flight restarts on first-failure cancellation. Required for parallel mode.- Spawn subprocesses with
start_new_session=Trueso each child gets its own process group. The cancellation path signals the pgid (viaos.killpg) rather than the bare PID — avoids the classic UNIX PID-recycle race when a subprocess exits between PID capture and the cancellation pass.
Adapters that forward via **kwargs (e.g. def run(self, **kwargs): return self._inner.run(**kwargs)) are recognized as supporting both extensions — but the inner runner MUST actually honor them. A **kwargs forwarder that silently strips unknown kwargs will pass the parallel-mode guard but produce incoherent logs and broken cancellation.
For non-parallel use (max_parallel_restarts=1), both extensions are optional — only the four baseline kwargs are required.
Cache directory layout
After build_panel_cache succeeds, cache_dir contains:
cache_dir/
├── panel.K.P # Best-LL restart's allele freqs (M × K)
├── panel.K.Q # Best-LL restart's non-target Q (N × K)
├── panel.bim # Variant set + REF/ALT axes (alignment ref)
├── restart_sd.json # Per-cluster SD across restarts
├── cluster_order.json # K column → cluster name mapping
├── manifest.json # Panel SHA + YAML SHA + K + version pins
└── build_logs/ # ADMIXTURE stdout/stderr per restart
Cache validity is determined by manifest.json SHAs matching the current config (panel.bim, clusters_yaml, K, optional geo-filter YAMLs). Any mismatch → consumer code can fall back to a full ADMIXTURE training pass or rebuild the cache.
When to use this
- Multi-user services: cache once, project for every user (~5,000× per-target speedup at scale)
- Reproducibility: published canonical caches (forthcoming via GitHub Releases) give byte-identical P across consumers
- CI/CD: faster integration tests once you have a cache
When NOT to use this
- One-time analyses with a custom panel that won't be reused — full ADMIXTURE is simpler
- Novel methodologies requiring per-target P refinement — the projection assumes P is fully determined by the panel
Status
- v1.4 — Current. Drops the consumer-specific
trackenum constraint; thetrackandcontinentmanifest fields are now free-text provenance labels. - v1.3 — Adds the
admixture-cache downloadcommand +download_cache/list_available_caches/CacheReleasePython API. Canonical caches are published as GitHub Releases following thecache-<name>-<version>tag convention (seedocs/PUBLISH_CACHE.md). - v1.2 — End-to-end integration suite against real ADMIXTURE 1.4 + plink2.
- v1.1 — NUMA pinning, PGEN target format support, Hypothesis property tests.
- v1.0 — First PyPI release. Cache directory layout stable at schema v1; numerical parity validated against stock ADMIXTURE.
See CHANGELOG.md for the full per-release detail.
Contributing
See CONTRIBUTING.md for dev setup, the three local validation gates (pytest / ruff / mypy), commit conventions, and the tag → OIDC PyPI release procedure. See DEVELOPMENT.md for the architecture map, design rationale, and module-level walkthroughs.
Acknowledgments
This library was extracted from ancestry-pipeline's in-pipeline supervised-ADMIXTURE projection module (pop_automation/admixture_projection.py, ~744 LOC, validated against real-world workloads). The split lets sibling projects depend on the cache layer without pulling in the larger orchestrator.
License
MIT. See LICENSE.
Release files for admixture-cache 1.6.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| admixture_cache-1.6.0.tar.gz | 73.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| admixture_cache-1.6.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size:152.0 kB
Release files / admixture_cache-1.6.0.tar.gz
| Download URL | admixture_cache-1.6.0.tar.gz |
|---|---|
| Size | 73.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5374d5d779c413569c46d75ef637e6b83a84b92338ebdc30330e70e638db09cc
|
|
BLAKE2b-256 checksum How to use checksums |
2d0668cff2704b62877a4b5ad3bd626aebaf8bd24ba30ae95000d2bc282cba23
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 29, 2026.
Transparency logRelease files / admixture_cache-1.6.0-py3-none-any.whl
| Download URL | admixture_cache-1.6.0-py3-none-any.whl |
|---|---|
| Size | 78.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8afe7b895a45f350457596738018af3cfdcaf55f32ecd547c42fb3068ffd84ee
|
|
BLAKE2b-256 checksum How to use checksums |
f367408cf685ff6490189e3c28ca22745f3a4e5ccd57b9e2b7044cdd04083943
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 29, 2026.
Transparency log