Ultra-Compact Macromolecular Format (UMF)
High-Performance, Zero-Copy Binary Container for Macromolecular Complexes (Proteins, DNA, RNA, Small-Molecule Ligands, & Ions) in Structural Bioinformatics & Deep Learning
🚀 Overview
The Ultra-Compact Macromolecular Format (.umf) is a next-generation structural container format designed to replace legacy ASCII formats (.pdb, .cif) and lossy monomer representations in modern machine learning pipelines, virtual screening, and structural biology archives.
Unlike traditional text-based formats that require expensive string deserialization and runtime distance tree construction, .umf leverages Google FlatBuffers to enable direct zero-copy memory mapping into CPU and GPU PyTorch/NumPy tensors.
Core Architecture Highlights
- Universal Multi-Entity Support: First-class representation for multi-chain protein complexes, nucleic acids (DNA/RNA), small-molecule drugs/cofactors, metal ions, and crystallographic waters.
- Validated Chemical Graph Topology: Stores explicit covalent bond orders (single, double, triple, aromatic), formal charges, element numbers, and stereochemical parity for small-molecule ligands.
- Drift-Free Coordinate Stability ($\le 0.0025\text{ \AA}$): Solves the catastrophic kinematic drift of pure dihedral compressors (such as Foldcomp) via periodic Cartesian anchor snapping ($\Delta = 15$ residues).
- All-Atom NeRF Reconstruction: Quantized 16-bit rotamer dihedrals ($\chi_{1-4}$) allow deterministic reconstruction of all sidechain heavy atoms within $\le 0.30\text{ \AA}$ RMSD.
- Zero-Copy PyTorch Geometric HeteroData Streaming: Precomputed sparse spatial graphs ($r < 8.0\text{ \AA}$) for intra-protein contacts and bipartite protein-ligand interfaces, eliminating CPU $O(N \log N)$ distance tree generation and cutting dataloading latency by 5.6×.
📊 Benchmark Comparisons
1. Direct Head-to-Head Comparison (benchmarks/benchmark_sota.py)
Evaluated against Foldcomp (.fcz), BinaryCIF (.bcif), and PDB (.pdb) across representative benchmarks:
| Format | Avg Storage (B/res) | Direct PyG Ingestion | Coordinate Drift RMSD | Multi-Chain & Ligands? | Pre-Indexed Graphs? |
|---|---|---|---|---|---|
Standard PDB (.pdb) |
~1,000 B/res | ~12 ms (text parse + KDTree) | Exact | Text only (no bonds) | ❌ None |
BinaryCIF (.bcif) |
~110 B/res | ~15 ms (decompression + parse) | Exact | Tabular | ❌ None |
Foldcomp (.fcz) |
~17.9 B/res | ~12 ms (decompress + KDTree) | 1.96 Å (up to 17 Å on large proteins) | ❌ Stripped / Crashes | ❌ None |
| UMF (Backbone Mode) | ~14.5 B/res | ~0.35 ms (zero-copy memory map) | ≤ 0.0025 Å (Drift-Free) | Monomer fold archive | ❌ None |
| UMF (All-Atom + Graphs) | ~35.0 B/res | ~0.35 ms (zero-copy PyG HeteroData) | ≤ 0.0025 Å (Backbone) / 0.30 Å (Sidechains) | ✅ Full RDKit fidelity | ✅ Precomputed $r < 8\text{ \AA}$ |
2. Downstream Virtual Screening GNN Training (benchmarks/benchmark_pdbbind_virtual_screening.py)
5-Epoch training benchmark on 436 co-crystallized complexes from PDBbind:
- Dataloading I/O Latency: 5.6× lower with UMF zero-copy streaming ($1.09\text{ s}$ vs. $6.06\text{ s}$ per epoch).
- End-to-End Epoch Wall-Clock Time: 3.3× faster ($2.14\text{ s}$ vs. $7.13\text{ s}$ per epoch).
- GPU Dataloader Compute Saturation: Increases active compute saturation from 15.0% up to 49.0% (reducing idle CPU wait from 85% to 51%).
📦 Installation & Build
Prerequisites
- CMake $\ge$ 3.20 & Ninja
- C++20 compiler (GCC/G++ $\ge$ 11, Clang $\ge$ 14, or MSVC 2019+)
- Python 3.10+
- Google FlatBuffers (
flatc)
Install from PyPI
pip install umf-format
Building from Source
# Clone the repository
git clone https://github.com/messiay/Universal_Macromolecular_formate.git
cd Universal_Macromolecular_formate
# Install locally with pip:
pip install -e .
🐍 Python Quickstart
import umf
# 1. Encode any PDB or mmCIF macromolecular complex into .umf binary format
umf.encode_complex("data/complexes/1STP.pdb", "data/complexes/1STP.umf", contact_cutoff=8.0, anchor_interval=15)
# 2. Ingest zero-copy tensors directly into PyTorch Geometric HeteroData
data = umf.to_hetero_tensors("data/complexes/1STP.umf")
# Directly ready for Equivariant GNNs & Virtual Screening
print(data['protein'].coords.shape) # (N_prot, 3) Backbone CA coordinates
print(data['protein', 'contact', 'protein'].edge_index) # (2, E_prot) Intra-protein contact graph
print(data['ligand'].coords.shape) # (N_lig, 3) Small-molecule coordinates
print(data['ligand', 'bond', 'ligand'].edge_index) # (2, E_bonds) Covalent chemical bonds
print(data['ligand', 'interacts', 'protein'].edge_index) # (2, E_bip) Bipartite pocket graph
🧪 Testing & Verification
Run the test suite:
# 1. Test universal macromolecular complex encoding (Biotin-Streptavidin, B-DNA, Hemoglobin)
uv run python -m pytest tests/test_complex_encoding.py -v
# 2. Test adversarial failure cases (disordered loops, negative residues, pure ligands)
uv run python tests/test_adversarial_failure_cases.py
# 3. Test all-atom NeRF reconstruction precision
uv run python tests/test_all_atom_and_features.py
# 4. Run SOTA head-to-head comparison
uv run python benchmarks/benchmark_sota.py 5
📄 License
Licensed under the Apache 2.0 License.
Metadata
Release files for umf-format 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| umf_format-0.2.0.tar.gz | 1.1 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| umf_format-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 2.2 MB
Release files / umf_format-0.2.0.tar.gz
| Download URL | umf_format-0.2.0.tar.gz |
|---|---|
| Size | 1.1 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
603758f838f96c99318c4c45ad2e2e246054fefba9e3ff6cae6c6887de2bcc57
|
|
BLAKE2b-256 checksum How to use checksums |
b852ef4f37908d88eab7a078c1a31b3141185b8f1e208156861e1cb376753fe7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|
Release files / umf_format-0.2.0-py3-none-any.whl
| Download URL | umf_format-0.2.0-py3-none-any.whl |
|---|---|
| Size | 1.1 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
329c036d19b2065a08fbc8d2bc08dae334ea0b42387db254495046ade30bde01
|
|
BLAKE2b-256 checksum How to use checksums |
8ba6e2a7e2c1e3d9c86708004e40e2a27042b2d4a6a7b766a7bb15f55682f51a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|