Skip to main content

🔬 PDB/MMCIF toolkit in Molecular Dynamic Simulation workflows

Pre- and post-processing toolkit for PDB/MMCIF structure files in molecular dynamics workflows.

Prerequisites

Always check for the forums of the MD simulation software you are using.

Tool Description Url
Amber Amber Mailing List Archive http://archive.ambermd.org/

https://cse.google.com/cse?cx=partner-pub-9700140137778662:8927431201&ie=UTF-8&sa=Search&ref=lists.ambermd.org/
Gromacs GROMACS community forums https://gromacs.bioexcel.eu/
计算化学公社- 高水平计算化学、理论化学交流论坛[CHINESE] http://bbs.keinsci.com/forum.php

Installation

pip install pdb-md

Usage

❯ pdb-md --help
                                                                                                                                                    
 Usage: pdb-md [OPTIONS] COMMAND [ARGS]...                                                                                                          
                                                                                                                                                    
 Pre- and post-processing toolkit for PDB/MMCIF structure files in molecular dynamics workflows.                                                    
                                                                                                                                                    
╭─ Options ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ --install-completion          Install completion for the current shell.                                                                          │
│ --show-completion             Show completion for the current shell, to copy it or customize the installation.                                   │
│ --help                        Show this message and exit.                                                                                        │
╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
╭─ Commands ───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ select        Select segments from a PDB or MMCIF structure file.                                                                                │
│ termini-rm5p  Remove terminal phosphate groups from nucleic acids, e.g. 5' phosphate group from DNA/RNA.                                         │
│ res-rename    Rename residues of specified chains and residue numbers, e.g. to the residue names a                                               │
│               force field expects for a given protonation state (HIS -> HIE/HID/HIP, CYS -> CYM).                                                │
╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

                                                               

1️⃣ select

Select segments from a PDB or MMCIF structure file

❯ pdb-md select --help
                                                                                                                                                                               
 Usage: pdb-md select [OPTIONS]                                                                                                                                                
                                                                                                                                                                               
 Select segments from a PDB or MMCIF structure file.                                                                                                                           
                                                                                                                                                                               
╭─ Options ───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ *  --input        -I      <path>  Input PDB or MMCIF file [required]                                                                                                        │
│    --output       -O      <path>  Output PDB file, defaults to <input stem>_selected.pdb in the current working directory                                                   │
│ *  --segment      -S      <str>   Segments to select, either 'chain' (whole chain, e.g. A) or 'chain:start-end' (e.g. A:1-10). Repeatable. The start-end range filters      │
│                                   polymer residues only; HETATM residues on the same chain are gated solely by --keep-hetero.                                               │
│                                   [required]                                                                                                                                │
│    --keep-hetero  -H      <str>   Keep HETATM residues. Repeatable. Each entry is 'spec' where spec is 'none','all', or a comma-separated list of residue names (e.g. 'ZN'  │
│                                   or 'ZN,HOH'). An entry without a colon/chain is the global default for every chain; an entry with a chain (e.g. 'A:all') overrides the    │
│                                   global set for that chain. Entries in the same scope are combined by union (-H ZN -H HOH == -H ZN,HOH); a chain entry replaces, not       │
│                                   merges with, the global set. A chain must also appear in --segment, otherwise it is dropped before -H is consulted.                       │
│    --help                         Show this message and exit.                                                                                                               │
╰─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

example:

pdb-md select -I znf263_all_hs7_30_0.54_0.37_model_0.cif -O znf263_all_hs7_30_0.54_0.37_model_0_selected.pdb -S A:369-683 -S B -S C -S D -S E -S F -S G -S H -S I -S J -S K:1-47 -S L:29-75 -H ZN

change from to


Selecting HETATM / heterogen residues

In short: polymer residue is governed by range + chain double filter, while HETATM is governed by chain only. The --keep-hetero option is the only way to keep HETATM residues.

HETATM residues (ions, ligands, waters, sugars) are selected at the chain level, not by residue range. The chain:start-end syntax filters polymer residues only.

Consequences:

  • A chain made entirely of HETATM (e.g. a single Zn ion, a glycan chain) is governed by the chain whitelist alone — -S B keeps the whole Zn.
  • On a mixed chain, chain:start-end does NOT narrow which HETATM are kept: -S A:1-3 --keep-hetero UNX still keeps UNX at resid 109, because the HETATM branch ignores regions.
  • --keep-hetero '' (default) drops every HETATM; all keeps them all; a comma list (e.g. ZN, HOH) is an exact resname whitelist (case-insensitive).

To pick specific HETATM on a mixed chain, split the chain in preprocessing (e.g. give the heterogen its own chain id) — the current CLI has no per-HETATM range selector.

--keep-hetero syntax

-H is repeatable. Each entry is [chain:]spec:

entry meaning
-H ZN global: every chain keeps HETATM named ZN
-H ZN,HOH global: every chain keeps ZN and HOH
-H all global: every chain keeps all HETATM
-H none / -H '' global: every chain drops all HETATM (default)
-H A:all chain A keeps all of its HETATM
-H A:none chain A drops all of its HETATM
-H E:ZN,HOH chain E keeps only ZN and HOH

Two rules govern how entries combine:

  1. Union within the same scope. Several entries for the same scope add up: -H ZN -H HOH ≡ -H ZN,HOH, and -H E:ZN -H E:HOH ≡ -H E:ZN,HOH. Adding an entry never removes something already whitelisted.
  2. A chain entry overrides the global entry. A chain that has its own [chain:] entry uses only that set and ignores the global one. This is what makes per-chain removal expressible: -H ZN -H J:none keeps ZN everywhere except chain J.

A -H chain id must also appear in --segment; otherwise the chain is dropped before -H is ever consulted, and a warning is printed.

2️⃣ preprocessing for MD simulation

Here we summarize several important easy to use application for fixing problems in Protein Data Bank files in preparation for simulating them.

Tool name Description Url Note
PDBFixer PDBFixer is an easy to use application for fixing problems in Protein Data Bank files in preparation for simulating them https://github.com/openmm/pdbfixer

https://htmlpreview.github.io/?https://github.com/openmm/pdbfixer/blob/master/Manual.html
General purpose tool for fixing PDB files
pdb2gmx gmx pdb2gmx reads a .pdb (or .gro) file, reads some database files, adds hydrogens to the molecules and generates coordinates in GROMACS (GROMOS), or optionally .pdb, format and a topology in GROMACS format. These files can subsequently be processed to generate a run input file https://manual.gromacs.org/current/onlinehelp/gmx-pdb2gmx.html Designed for preparing PDB files for GROMACS simulations, pdb2gmx is a subcommand of GROMACS tool gmx —— gmx pdb2gmx, it has several options for pdb file processing like -ignh to add the hydrogens, see in gmx pdb2gmx -h
pdb4amber Analyse PDB files and clean them for further usage, especially with the LEaP programs of Amber https://ambermd.org/AmberTools.php pdb4amber tool from the AmberTools MD package, designed for preparing PDB files for Amber simulations, also a command-line utility in the AmberTools suite —— pdb4amber, it also has several options for pdb file processing, Removing hydrogen or water atoms, see pdb4amber -h
BioForge BioForge is a pure-Rust toolkit for automated preparation of biological macromolecules. It reads experimental structures (PDB/mmCIF), reconciles them with high-quality residue templates, repairs missing atoms, assigns hydrogens and termini, builds topologies, and optionally solvates the system with water and ions—all without leaving the Rust type system https://github.com/TKanX/bio-forge
pdbtools A set of tools for manipulating and doing calculations on wwPDB macromolecule structure files https://github.com/harmslab/pdbtools see File/structure manipulation part for cleaning function in https://github.com/harmslab/pdbtools#filestructure-manipulation
pdb-tools A dependency-free cross-platform swiss army knife for PDB files https://github.com/haddocking/pdb-tools see https://www.bonvinlab.org/pdb-tools/

normal preprocessing

cited from pdbfixer manual

- If the structure was generated by X-ray crystallography, most or all of the hydrogen atoms will usually be missing.
- There may also be missing heavy atoms in flexible regions that could not be clearly resolved from the electron density. This may include anything from a few atoms at the end of a sidechain to entire loops.
- Many PDB files are also missing terminal atoms that should be present at the ends of chains.
- The file may include nonstandard residues that were added for crystallography purposes, but are not present in the naturally occurring molecule you want to simulate.
- The file may include more than what you want to simulate. For example, there may be salts, ligands, or other molecules that were added for experimental purposes. Or the crystallographic unit cell may contain multiple copies of a protein, but you only want to simulate a single copy.
- There may be multiple locations listed for some atoms.
- If you want to simulate the structure in explicit solvent, you will need to add a water box surrounding it.
- For membrane proteins, you may also need to add a lipid membrane.

you can

Add missing heavy atoms.
Add missing hydrogen atoms.
Build missing loops.
Convert non-standard residues to their standard equivalents.
Select a single position for atoms with multiple alternate positions listed.
Delete unwanted chains from the model.
Delete unwanted heterogens.
Build a water box for explicit solvent simulations.

Remove Chains
Identify Missing Residues
Replace Nonstandard Residues
Remove Heterogens
Add Missing Heavy Atoms
Add Missing Hydrogens
Add Water
Add Membrane

Removal of terminal phosphate group for Nucleic Acid

It is usually suggested that while preparing a nucleic acid system for simulation, 5' terminal phosphate group must be removed; or any terminal charged phosphate group should be removed...

Nucleic acids typically do not have 5’-phosphate groups. Force fields are parametrized to the most common use cases and do not necessarily cover all possible chemical space. Delete the phosphate atoms from the 5’-nucleotide and you can generate the topology such that it has a free 5’-hydroxyl group.

In most cases, e.g. for DNA simulations, 5’-phosphate groups are less cared about, so you can just delete them

Remove any phosphate group from terminal residues; these are often present synthetically but are not how force fields are typically parametrized (5’-OH terminus is typical).

In detail, we just need to remove the [O1P/OP1, O2P/OP2, O3P/OP3 if exists, P, Hydrogens attached to them] atoms from the 5' terminal residue of a nucleic acid chain.

You can use the termini-rm5p command to remove the 5' terminal phosphate group from nucleic acids as talked above.

❯ pdb-md termini-rm5p --help
                                                                                                                                                                                                          
 Usage: pdb-md termini-rm5p [OPTIONS]                                                                                                                                                                     
                                                                                                                                                                                                          
 Remove terminal phosphate groups from nucleic acids, e.g. 5' phosphate group from DNA/RNA.                                                                                                               
                                                                                                                                                                                                          
╭─ Options ──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ *  --input   -I      <path>  Input PDB or MMCIF file [required]                                                                                                                                        │
│    --output  -O      <path>  Output PDB file, defaults to <input stem>_rm5p.pdb in the current working directory                                                                                       │
│    --chain   -C      <str>   Chains to process. Repeatable. If not provided, all chains will be processed.                                                                                             │
│    --help                    Show this message and exit.                                                                                                                                               │
╰────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
❯ pdb-md termini-rm5p -I znf263_all_hs7_30_0.54_0.37_model_0_selected.pdb -C K -C L

Wrote znf263_all_hs7_30_0.54_0.37_model_0_selected_rm5p.pdb
  chain K: removed OP3, P, OP1, OP2
  chain L: removed P, OP1, OP2

protonation(add hydrogens) and residue renaming

protonation state of the protein is important for MD simulation. The protonation state of a protein can be determined by the pH of the environment, which can affect the charge and conformation of the protein.

Here we summarize several tools for predicting or assigning the protonation state of a protein:

Basically predicted/assigned based on comparison between Pka calculation and the pH of the environment.

Tool name Description Url Note
H++ H++ is an automated system that computes pK values of ionizable groups in macromolecules and adds missing hydrogen atoms according to the specified pH of the environment. Given a (PDB) structure file on input, H++ outputs the completed structure in several common formats (PDB, PQR, AMBER inpcrd/prmtop) and provides a set of tools for analysis of electrostatic-related molecular properties. http://newbiophysics.cs.vt.edu/H++/index.php
PDB2PQR APBS-PDB2PQR software suite, Use PROPKA to assign protonation states at provided pH https://server.poissonboltzmann.org/pdb2pqr
PROPKA PROPKA predicts the pKa values of ionizable groups in proteins and protein-ligand complexes based in the 3D structure. https://github.com/jensengroup/propka
PKA17 PKA17: the grid-based pKa calculator for proteins http://kaminski.wpi.edu/PKA17/pka_calc.html
pdb2gmx gmx pdb2gmx, see above https://manual.gromacs.org/current/onlinehelp/gmx-pdb2gmx.html pdb2gmx --ignh
pdbfixer see above https://github.com/openmm/pdbfixer

https://htmlpreview.github.io/?https://github.com/openmm/pdbfixer/blob/master/Manual.html
tleap/pdb4amber see ambertools above

Of course, in most cases, people still rely on prior knowledge to determine the protonation state of specific residues. For example, tetracoordinate zinc finger proteins have been thoroughly studied, and the protonation states of coordinating atoms within this motif are well known, so computational tools are generally not required for such determinations.

In many cases, successful assignment of protonation states is often coupled with residue renaming operations. This is because in most force fields, the protonation state of a residue is distinguished by its residue name. For instance, the three protonation states of HIS correspond to three residue names: HID, HIE, and HIP. Therefore, residue renaming is usually needed after the protonation state is determined.

Take the classic tetracoordinate zinc finger protein as an example. The coordinating CYS and HIS residues are commonly renamed to CYM, HID/HIE/HIP, or further processed (e.g., ZAFF), so that the force field can correctly recognize their protonation states.

Here we take ZAFF as an example, with reference to ZAFF.

alt text

For residue renaming, most scripts are based on raw text processing given that the PDB file is an 80-column, fixed-width text file. However, this approach is not robust and can easily introduce errors just as manual editing does. Therefore, we recommend using the 'biopython' library to read the PDB file, modify the residue names, and then write it back to a new PDB file. This method is more robust and less error-prone.

You can use the res-rename command to rename residues by chain and residue number. This applies a renaming you have already decided on; the protonation-state decision itself is made beforehand (by the tools above, or by hand).

❯ pdb-md res-rename --help
                                                                                                                                                                                                          
 Usage: pdb-md res-rename [OPTIONS]                                                                                                                                                                     
                                                                                                                                                                                                          
 Rename residues of specified chains and residue numbers, e.g. to the residue names a force field expects for a given protonation state (HIS -> HIE/HID/HIP, CYS -> CYM).                               
                                                                                                                                                                                                          
╭─ Options ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ *  --input   -I      <path>  Input PDB or MMCIF file [required]                                                                                                                                      │
│    --output  -O      <path>  Output PDB file, defaults to <input stem>_renamed.pdb in the current working directory                                                                                  │
│ *  --rename  -R      <str>   Residue to rename as 'chain:resnum:newresname' (e.g. A:20:HIE). Repeatable. The new name may be at most 4 characters. [required]                                        │
│    --help                    Show this message and exit.                                                                                                                                             │
╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

example:

pdb-md res-rename -I znf263_all_hs7_30_0.54_0.37_model_0_selected_rm5p_renum.pdb -R 

A note on 4-character residue names. A PDB residue name occupies columns 18-20, but some force fields use 4-character names (GROMOS HISA/HISB, CHARMM CYSH) that borrow the blank column 21 between the name and the chain id. Biopython's PDBIO formats the residue name with a minimum width of 3, which does not truncate: a 4-character name pushes every following column one to the right and emits an 81-column record with the chain id and residue number out of place -- a record fixed-column parsers read wrong. res-rename renders the file in memory and repairs those records before saving, so both 3- and 4-character names come out as well-formed 80-column lines. Names longer than 4 characters are rejected outright, since they have nowhere in the format to go.

This note covers writing only. Biopython's parser reads the residue name from three columns, so a 4-character name already in the input is silently truncated to three (no error) -- and that includes the file this command writes. Re-running any pdb-md command on res-rename's output will lose the 4th character, so the truncation is reported as a warning.

``

A typical workflow for preparing a PDB file for MD simulation

pdbfixer(fix missing atoms, remove heterogens and hydrogens) -> pdb-md select (select chains and residues) -> pdb-md termini-rm5p (remove 5' terminal phosphate group for nucleic acids) -> pdb4amber renum -> protonation (add hydrogens) -> pdb-md res-rename (rename residues to the force field's protonation-state names) -> pdb2gmx/pdb4amber (generate topology and coordinates for MD simulation)

Metadata

Release files for pdb-md 0.3.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for pdb-md 0.3.1
File Interpreter ABI Platform
pdb_md-0.3.1-py3-none-any.whl Python 3 none any Details

Release files / pdb_md-0.3.1-py3-none-any.whl

Download URL pdb_md-0.3.1-py3-none-any.whl
Size 21.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3e7f636214e5e3351fa7496d866f85847b250b09225c3b73427a3d0bf7b6ed5c
BLAKE2b-256 checksum
How to use checksums
17e5761391ca8d68ef475781a6bcb4c3b677a55bae850c845a37f9a4e66f62dc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

This release

0.3.1 This release

1 release file

0.3.0

1 release file

0.2.2

1 release file

0.2.1

1 release file

0.2.0

1 release file

0.1.2

1 release file

0.1.1

1 release file

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page