PLINKFORMATTER
plinkformatter transforms genotype and phenotype inputs into PLINK-compatible
artifacts for downstream linear mixed-model workflows (primarily PyLMM).
This repository is based on the original R workflow implemented in:
plinkformatter/IGNORE_misc/pyLMM_utils.Rtests/IGNORE_MISC/hao_v2/pyLMM_analysis_NonDO.R
The Python implementation keeps the same core workflow while improving maintainability, testability, and performance for large PED/MAP datasets.
Prerequisites
- Python 3.8+
- Poetry
- PLINK 2.0 (required for PLINK integration tests and full pipeline runs)
Install dependencies:
poetry install
If plink2 is not on your PATH, set PLINK2_PATH explicitly.
PowerShell example:
$env:PLINK2_PATH = "C:\path\to\plink2.exe"
Running Tests
Run all commands from the repository root.
1) Fast test pass (no external PLINK dependency)
poetry run pytest -q tests/test_generate_pheno_plink_fast.py
2) PLINK utility tests
These include PLINK-facing behavior and may require a working plink2 binary.
poetry run pytest -q tests/test_plink_utils.py
3) Full test suite
poetry run pytest -q tests
Test output flags
-q: quiet output (compact summary)-s: show stdout/stderr (print, logs written to console)
Example:
poetry run pytest -s -q tests/test_generate_pheno_plink_fast.py
Workflow Parity with R
The Python pipeline mirrors the same logical stages as the R scripts:
- Extract/normalize phenotype rows for selected measure IDs.
- Generate per-measure, sex-specific
.ped/.map/.phenofiles. - Build
.bed/.bim/.famwith PLINK. - Align
.phenoordering to.famand recompute rank-Z on retained samples. - Compute kinship (PyLMM3 or PLINK-based path).
Performance Design (Large PED Files)
A major difference from older dataframe-heavy patterns is how PED is handled in
generate_pheno_plink_fast.py:
- It does not load the full PED into a pandas DataFrame.
- It builds a compact byte-offset index (
strain -> file position) once. - It seeks directly to needed PED rows and writes outputs in a streaming manner.
This avoids high memory usage and scales better for large genotype files than
pandas.read_csv() on full PED content.
Publishing to PyPI
- Update version:
poetry version patch
- Build:
poetry build
- Configure repository and token:
poetry config repositories.pypi https://upload.pypi.org/legacy/
poetry config pypi-token.pypi pypi-YourActualTokenHere
- Publish:
poetry publish
Metadata
Release files for plinkformatter 0.1.83
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| plinkformatter-0.1.83.tar.gz | 13.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| plinkformatter-0.1.83-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 29.4 kB
Release files / plinkformatter-0.1.83.tar.gz
| Download URL | plinkformatter-0.1.83.tar.gz |
|---|---|
| Size | 13.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
df613d69f4292760c4a3f376d9427b8338db38c7e11facf7df8abc055052af1d
|
|
BLAKE2b-256 checksum How to use checksums |
5cf422dd8ed8d9d94b6860d6fb2f6b1926f4125648832c82b0d8e6320db1a74b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
poetry/1.7.1 CPython/3.9.13 Windows/10
|
Release files / plinkformatter-0.1.83-py3-none-any.whl
| Download URL | plinkformatter-0.1.83-py3-none-any.whl |
|---|---|
| Size | 15.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a32ced203155d5c101e283af77fdc904f0cab2b3eff3800a266a2e357a80e8af
|
|
BLAKE2b-256 checksum How to use checksums |
232c0b0c212a16d5cf5364dd2933088b940727fe992a3c027af56791f64c48e5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
poetry/1.7.1 CPython/3.9.13 Windows/10
|