pymdkit
A single command-line tool that bundles a collection of atomistic / molecular-dynamics
structure scripts behind one executable: pmk. Instead of copying individual
scripts into each working folder and running python some_script.py, you install
pymdkit once and call any tool from anywhere as pmk <command> [options].
Every command exposes named -flags (no positional guessing), and each underlying
script is still runnable on its own.
Install
Create a clean conda environment, activate it, then install pymdkit with pip:
conda create -n pymdkit python=3.10
conda activate pymdkit
pip install pymdkit
This installs the pmk command into the active conda environment, together
with its direct dependencies, including ASE, pymatgen, py4vasp-core, dpdata, matplotlib, and UMAP. Some of these packages may install their own transitive dependencies.
Verify:
pmk -version
pmk -help # lists every command
pmk <command> -help # shows that command's flags
Commands
Commands that transform structures accept either a single file (-i/-o) or a
whole folder (-if/-of); commands that analyse VASP runs scan the current
directory for job sub-folders automatically.
| Command | What it does |
|---|---|
add-config-type |
Add Config_type to one XYZ structure or every trajectory frame |
gpumd-group |
Tag atoms with a GPUMD group index by element order |
gpumd-relax |
Write GPUMD energy-minimization jobs for a file, folder, or trajectory |
ewald |
Compute CIF electrostatic energy with pymatgen EwaldSummation |
ehull |
Auto-detect VASP job folders and compute E_hull vs Materials Project |
vasp-fe |
List VASP job final energies from lowest to highest |
gather-fs |
Recursively gather final structures from converged VASP or GPUMD jobs |
gpumd-thermo |
Export and plot GPUMD thermo data and detect sustained equilibrium changes |
gpumd-trjcat |
Merge normal or all GPUMD trajectories and extract the first abnormal region |
pca |
Reduce descriptor data to two dimensions; optional FPS sampling |
trj-extract |
Extract a GPUMD XYZ trajectory by time range |
trj-sparse |
Deterministically retain a requested fraction of XYZ trajectory frames |
umap |
Reduce descriptor data to two dimensions with UMAP; optional FPS sampling |
msd |
Diffusivity & conductivity from GPUMD MSD jobs (auto-scans <structure>/<temp>/) |
nep-rmse |
Compute NEP energy/force/stress RMSE with terminal plots and optional candidate selection |
perturb |
Generate perturbed structures with dpdata |
vasp2xyz |
Collect SCF-converged VASP job folders (any name) into one extxyz file |
submit-vasp |
Submit/resubmit VASP job folders while limiting active queue jobs |
rmsd |
Compare structures by RMSD and optionally remove higher-energy duplicates |
chemsys-entry |
Download stable Materials Project structures for a chemical system |
convert |
Convert structures, trajectories, or folders between ASE/pymatgen formats |
substitute |
Randomly substitute or remove selected atoms/sites from a structure |
supercell |
Build a supercell with cell lengths capped at a maximum (Angstrom); optional per-temperature GPUMD setup |
symmetrize |
Detect/refine symmetry for structure files, folders, or CSV rows -> CIF |
vasp-relax |
Write VASP relaxation inputs for a structure (or folder); INCAR tags overridable |
vasp-static |
Write VASP static / single-point inputs for a structure (or folder) |
Examples
pmk add-config-type -i single-structure.xyz -n P-3m1
pmk add-config-type -it trajectory.xyz -n P-3m1
pmk gpumd-group -i opted.cif -elements Li Y Cl -o model.xyz
pmk gpumd-group -if cifs/ -elements Li Y Cl -of cifs-grouped/ # whole folder
pmk gpumd-group -elements Li Y Cl # scan subfolders, tag each model.xyz in place
pmk convert -i opted.vasp -o opted.xyz # convert one structure
pmk convert -if vasp-opted -of cif-opted -oe cif # folder conversion target format
pmk convert -it train.xyz -ot train.extxyz # convert a complete trajectory
pmk supercell -i opted.vasp -o sc.vasp -max-abc 20 # cell lengths <= 20 A
pmk supercell -if vasp-opted -max-abc 20 -individual # per-structure ./<name>/<name>.<ext>
pmk supercell -if extxyz-opted -max-abc 24 -individual -temp 500 600 -md-if input-files -add-groups Li Y Cl
# GPUMD: ./<name>/model.xyz (grouped) + ./<name>/<T>/ jobs
pmk vasp-relax -i opted.vasp # relax inputs in current dir
pmk vasp-relax -if optimal_occupancy # one ./<name>/ job folder per structure
pmk gpumd-relax -if example -nep nep89_20250409.txt # model.xyz + run.in + NEP per job folder
pmk gpumd-relax -it train.xyz -nep nep.txt -in run.in # one ./frame_N/ GPUMD job per frame
pmk vasp-static -if cifs/ -custom-setting my_incar.txt # static inputs, custom INCAR
pmk vasp-static -it traj.xyz # one ./frame_N/ job per trajectory frame
pmk msd # scans <structure>/<temp>/ -> per-job msd/ + msd_summary.txt
pmk msd -diffuse_ion Li -ion_charge 1 # choose the mobile ion for conductivity
pmk chemsys-entry -s Li La Ta Cl # MP stable entries -> Li-La-Ta-Cl-stable-entries/
pmk ehull -mp-api-key $MP_API_KEY # scans ./ for VASP jobs -> ehull.txt
pmk ehull -local Li-La-Ta-Cl-stable-entries-opted
pmk vasp-fe # scans ./ for VASP jobs -> final-energy.txt with convergence status
pmk gather-fs -job vasp -fs-name CONTCAR -of vasp-opted
pmk gather-fs -job vasp -fs-name CONTCAR -of vasp-opted -ehull 0.028
pmk gather-fs -job gpumd -fs-name relaxed.xyz -of gpumd-opted
pmk vasp2xyz # scans ./ for VASP output folders -> scf-converged.xyz
pmk vasp2xyz -position-only # write positions only, without energy/forces/stress
pmk submit-vasp -subscript sub_vasp -queue slurm -max-job-num 30
nohup pmk submit-vasp -subscript sub_vasp -queue slurm -max-job-num 30 > submit-vasp.log 2>&1 &
pmk substitute -i Li3YCl6.cif -se Li -sn 3 -we Na -wn 3 -on 100
pmk substitute -i Li3YCl6.cif -se Li -sn 3 -we none -on 100
pmk substitute -i Li96Ta6La11Cl72.cif -se Li1 Li2 -sn 20 67 -we none -on 100
pmk substitute -i Li96Ta6La11Cl72.cif -se Li2 -we none -ref La Ta -d 1.01 1.02
pmk substitute -i Li96Ta6La11Cl72.cif -se Li2 -we none -ref La Ta
pmk ewald -i Li3YCl6.cif
pmk ewald -if Li3YCl6-all
pmk nep-rmse # writes energy/force/stress train txt files, rmse_value.txt, and terminal plots
pmk nep-rmse -select -xyz train.xyz # interactive candidate selection; writes candidate.xyz and accurate.xyz
pmk perturb -i example.xyz -atom 0.2 -lattice 0.03 -n 100 -o example-perturb-atom-0.2-lattice-0.03.xyz
pmk pca -i descriptor.out -o pca-descriptor.txt -it train.xyz # also split descriptors by Config_type
pmk pca -i descriptor.out -o pca-descriptor.txt -fps 0.01 -it train.xyz
pmk umap -i descriptor.out -o umap-descriptor.txt -fps 0.01 -it train.xyz
pmk gpumd-thermo # current job, or recursively scan job folders
pmk gpumd-thermo -t thermo.out -in run.in
pmk gpumd-trjcat # stable frames -> train.xyz; abnormal regions -> abnormal.xyz
pmk gpumd-trjcat -i traj.xyz -o train.xyz -ao abnormal.xyz
pmk gpumd-trjcat -all # merge every readable trajectory into train.xyz
pmk trj-extract -it traj.xyz -in run.in -b 100 -e 200 -o traj-100ps-200ps.xyz
pmk trj-sparse -it example.xyz -r 0.5 -o sparse-example.xyz
pmk rmsd -i structure-1.vasp structure-2.vasp # one pair -> rmsd.txt
pmk rmsd -if vasp-opted # every pair -> rmsd.txt
pmk rmsd -if example -rm-duplicate -of unique-example # RMSD <= 0.1: keep lowest energy
pmk symmetrize -i opted.cif -add_oxidation yes -o opted-symm.cif
pmk symmetrize -if my_cifs/ -symprec 0.1 -add_oxidation no -of my_cifs-symm
pmk symmetrize -csv output_Li2YCl3_struct.csv -of Li2YCl3
VASP input commands (vasp-relax, vasp-static) always produce individual
jobs (one structure per folder): -i writes into the current dir (or -o),
-if creates one ./<name>/ folder per structure, and -it creates one
./frame_N/ folder per trajectory frame - all directly in the current path.
They start from sensible default INCAR settings; override them by passing a
settings file with -custom-setting FILE. The file may be a Python-dict block
or KEY = VALUE lines (a None/blank value clears a tag):
custom_settings = {
"ENCUT": "600.0",
"ISIF": "3",
"MAGMOM": None
}
vasp-static -it traj.xyz (also available on vasp-relax) reads a
multi-structure trajectory and writes one job sub-folder per frame
(frame_1/, frame_2/, ..., prefix configurable via -frame-prefix). Each
folder also keeps a frame_N.xyz, so Config_type survives for a later
vasp2xyz.
gpumd-relax accepts the same -i, -if, and -it structure modes. -nep
is required and the selected potential is copied into every job folder together
with model.xyz and run.in. -in FILE supplies a custom GPUMD input; its
first potential filename is synchronized to the copied NEP. Without -in,
the generated job-local input is:
potential <NEP filename>
minimize fire 2.0e-2 1000000
ensemble nve
dump_xyz -1 0 1 relaxed.xyz force
run 1
gather-fs recursively scans all job folders below the current directory. For
VASP it follows the global vaspout.h5 > vasprun.xml > OUTCAR priority and
requires full convergence before copying the selected -fs-name as .vasp. The gathered VASP file's first line is replaced with energy=<final_energy> eV for later screening.
For GPUMD it requires gpumd.out to report a force tolerance and a final
f_max no greater than that tolerance before copying the final structure.
-fs-name defaults to CONTCAR for VASP and relaxed.xyz for GPUMD. The
-ehull and -ehull-file filters are available only with -job vasp.
rmsd -i STRUCTURE_1 STRUCTURE_2 compares one pair, while rmsd -if FOLDER
records every unique pair in rmsd.txt. With -rm-duplicate -of OUTPUT, files
are processed from lowest to highest energy; a structure is removed only when
its RMSD to an already retained lower-energy representative is no greater than 0.1. Duplicate removal supports .vasp,
.xyz, and .extxyz: VASP energy is read from the first line, and extended-XYZ
energy from the second-line energy= tag. The report records every RMSD,
duplicate decision, energy, group, and retained representative.
symmetrize uses pymatgen's spglib-backed SpacegroupAnalyzer for authoritative
space-group detection and conventional-cell refinement. Its default Cartesian
symmetry tolerance is 0.1 Angstrom (-symprec), with a 5 degree angle
tolerance (-angle-tolerance). CIF occupancies are retained by the pymatgen
structure and written directly rather than reconstructed after symmetry finding.
CSV mode reads serialized pymatgen Structure dictionaries from a required cif
column. -csv INPUT.csv -of NAME writes INDEX-NAME-SPACEGROUP.cif, for
example 1-Li2YCl3-66.cif. Every run also writes symmetry.txt with one row
per successfully generated CIF: filename, crystal system, space-group symbol,
and space-group number. Crystal-system and space-group summaries are sorted
from highest to lowest count, with alphabetical ordering for ties, followed by
the total number of structures. Single-file mode places the report beside the
output CIF; folder and CSV modes place it inside -of.
Every CIF exported by symmetrize, convert, substitute, chemsys-entry,
supercell, or gpumd-group is written through the same symmetry-aware helper.
Each CIF contains the IUCr-defined _space_group_crystal_system item as well as
both _atom_site_symmetry_multiplicity and _atom_site_Wyckoff_symbol in the
atom-site loop. Multiplicity-letter assignments use the conventional
International Tables setting from pymatgen/spglib and correspond to the
Bilbao Crystallographic Server WYCKPOS tables;
the provenance URL is also recorded inside each generated CIF.
CIF atom types and site labels use per-element inequivalent-site identifiers
such as Li1, Y1, Cl1, and Cl2; oxidation numbers remain in the separate
_atom_type_oxidation_number column. PMK does not add an
_audit_creation_method block or a generated-by header comment.
add-config-type updates its input atomically in place. With -i, the file must contain exactly one XYZ structure; use -it for a trajectory. Existing Config_type values are replaced, missing values are added, and all other extended-XYZ bytes remain unchanged.
nep-rmse -select, PCA FPS, and UMAP FPS copy selected extended-XYZ frame blocks directly from the input trajectory. No parser rewrites energy, stress, forces, positions, precision, or extra metadata.
When pca or umap receives -it train.xyz, it reads Config_type from every frame and writes one additional descriptor table per value, such as pca-descriptor-P-3m1.txt and pca-descriptor-Pnma.txt. With -fps, it also writes matching files such as fps-0.01-pca-descriptor-P-3m1.txt and fps-0.01-train-P-3m1.xyz. These files partition the single global FPS selection; FPS is not rerun independently for each Config_type. Frames without Config_type remain in the main outputs but do not receive a type-specific file.
For folder conversion, convert -if INPUT -of OUTPUT writes .xyz by default. Set another target extension with -oe, for example -oe cif or -oe vasp. Single-file and trajectory modes infer the target format from -o and -ot.
Because Wyckoff metadata describes one crystal structure, each .cif output file accepts exactly one frame.
gpumd-thermo writes temperature.txt, potential-energy.txt, pressure.txt, lattice-parameters.txt, volume.txt, lattice-angles.txt, and a headless thermo.png in every detected GPUMD job folder. For an abnormal run it also writes abnormal-time.txt, with one BEGIN ps to END ps interval per line; a normal rerun removes a stale file. Potential energy is the primary equilibrium signal. Detection requires a persistent change with a meaningful magnitude relative to the robust natural energy fluctuations, so ordinary correlated thermal noise and isolated spikes are not classified as abnormal. An abnormal classification always requires a mapped potential-energy change. When at least two lattice-parameter or volume signals show the same mapped transition, they may enable a more sensitive noise-relative energy test; this avoids a fixed eV or percentage-of-total-energy threshold that would depend on system size. Cell signals can corroborate and refine the energy interval, but can never classify a run by themselves. Temperature and all six pressure components remain diagnostics only and cannot classify an otherwise stationary run as abnormal by themselves. A lone cell-axis drift, an isolated spike, or an empty anomaly mask never triggers a whole-run fallback. Thermo plots use Arial when available and Matplotlib's built-in qualitative Dark2 color cycle. They plot all six pressure components, use frameless upper-right legends with added vertical headroom, preserve correct Å units, and show abnormal regions as shaded legend entries on potential energy, lattice parameters, and volume. Potential energy and volume are divided by 1000 and labeled ×10³ only when their absolute plotted values reach 1000. The analysis uses the last run segment, or the last 80% when run.in has no run record; inspect the plot before making a final scientific judgment.
The thermo parser follows the official GPUMD thermo.out format: 18 columns are required, pressure.txt exports Pxx Pyy Pzz Pyz Pxz Pxy, box vectors are interpreted as a full 3×3 matrix, time_step propagates between runs, and dump_thermo does not.
gpumd-trjcat recursively scans job subfolders for traj.xyz, thermo.out, and run.in. By default, normal trajectories are merged in deterministic folder order into train.xyz (or -o), while only frames from the earliest mapped abnormal interval of each abnormal job are written to abnormal.xyz (or -ao). With -all, every readable full trajectory is merged into train.xyz regardless of normal, abnormal, missing-thermo, or failed-thermo status; an analyzable abnormal job still contributes its earliest abnormal interval to abnormal.xyz. If no mapped abnormal frames are found, abnormal.xyz is not created and a stale file at the selected -ao path is removed. Frame times are obtained from time_step and dump_xyz in that job's run.in. In default mode, missing, unreadable, or unmappable jobs are skipped. When <job>/model.xyz contains Config_type, that value is applied to every exported frame while all other extended-XYZ bytes remain unchanged.
trj-sparse keeps max(1, floor(frame_count × ratio)) frames without rewriting their extended-XYZ content. Selection is deterministic and starts with frame zero; for 100 frames and -r 0.5, it writes source indices 0, 2, ..., 98.
Each command's full flag list is in pmk <command> -help.
For long VASP batch submission, run submit-vasp with nohup and & if you
want it to keep sleeping, checking the queue, and submitting new jobs after you
exit the terminal:
nohup pmk submit-vasp -subscript sub_vasp -queue slurm -max-job-num 30 > submit-vasp.log 2>&1 &
The command itself controls the loop: it submits until the active queue reaches
-max-job-num, sleeps when the queue is full, checks again, and continues until
all needed jobs are submitted. nohup ... & is what makes that loop continue in
the background after logout.
substitute -ref removes selected sites near reference sites and writes one
<input-stem>_substitute.cif in the current path. If -d is omitted, each
cutoff is 0.7 * (selected covalent radius + reference covalent radius) using
the covalent radii from Cordero et al., Dalton Trans., 2008, 2832-2838.
VASP-output readers (vasp2xyz, ehull, vasp-fe, and other future VASP-output
commands) use the global priority vaspout.h5 > vasprun.xml > OUTCAR.
ehull auto-detects every sub-folder of the current path that contains a
supported VASP output, groups them by chemical system (elements ordered by
electronegativity, e.g. Li-Y-Cl), and builds/reuses one mp_cache_<system>.json
per system - so a pure Li-Y-Cl batch yields a single mp_cache_Li-Y-Cl.json, while
a mixed Li-Y-Cl + La-O batch yields both mp_cache_Li-Y-Cl.json and
mp_cache_La-O.json. (Formation energy is reported alongside E_hull in
ehull.txt.)
Layout
pymdkit/
|-- pyproject.toml # package metadata + the `pmk` entry point
|-- README.md
`-- src/pymdkit/
|-- pymdkit_main.py # dispatcher: discovers and runs commands
`-- commands/ # one module per command
|-- _fileio.py # shared -i/-o/-if/-of helper (not a command)
|-- _gpumd.py # shared GPUMD run.in and thermo analysis
|-- _geometry.py # shared minimum-image geometry helper
|-- _vaspset.py # shared VASP input-set helper (not a command)
|-- gpumd_group.py
|-- gpumd_relax.py
|-- gpumd_thermo.py
|-- gpumd_trjcat.py
|-- compute_ehull.py
|-- compute_rmsd.py
|-- ewald.py
|-- vasp_fe.py
|-- perturb.py
|-- nep_rmse.py
|-- add_config_type.py
|-- _cifio.py
|-- chemsys_entry.py
|-- submit_vasp.py
|-- convert.py
|-- substitute.py
|-- supercell.py
|-- vasp2xyz.py
|-- vasp_relax.py
|-- vasp_static.py
|-- ...
`-- symmetrize.py
Modules whose name starts with _ are shared helpers and are skipped by the
dispatcher, so they never appear as commands.
Adding a new tool later
Drop a module in src/pymdkit/commands/ that defines four things:
COMMAND = "my-tool" # the subcommand name you'll type
HELP = "One-line description."
def add_arguments(parser): # register flags
parser.add_argument("-input", required=True)
def run(args): # do the work; return an exit code (0 = ok)
...
return 0
if __name__ == "__main__": # keeps the script runnable on its own
import argparse
_p = argparse.ArgumentParser(description=__doc__)
add_arguments(_p)
raise SystemExit(run(_p.parse_args()))
It will appear in pmk -help automatically - no central registration needed.
Put heavy imports (pymatgen, ase, ...) inside run() where practical; the dispatcher
reads each command's name and help without importing it, so pmk -help stays
fast and a missing optional dependency only affects the one command that needs it.
Running a script standalone
Every command module still works directly, which is handy for debugging:
python src/pymdkit/commands/supercell.py -i in.cif -max-abc 20 -o sc.vasp
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pymdkit-1.4.12.tar.gz.
File metadata
- Download URL: pymdkit-1.4.12.tar.gz
- Upload date:
- Size: 89.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.10.20
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
dd9a6431e9d9e28a4b3aad18a4c4e4044a08d824812c14751a195665f110c42a
|
|
| MD5 |
09b5b1d6033f5b039241d42a226fcd69
|
|
| BLAKE2b-256 |
2c61e31ff51a46f64f0580bdda44d00628a506adb1715d9d812ad30bbeac2fe1
|
File details
Details for the file pymdkit-1.4.12-py3-none-any.whl.
File metadata
- Download URL: pymdkit-1.4.12-py3-none-any.whl
- Upload date:
- Size: 100.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.10.20
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0355d7da1a9077aac65d9e16c4d1e0e6b0ff8d996d5f673e9692034c9c131170
|
|
| MD5 |
79c624a805d065a7db8a021ac8a7afe3
|
|
| BLAKE2b-256 |
e5c86a397bcfbd15a949bc8bb84144fabec3737dd1a2268d33d158ae08311b28
|