Skip to main content

Ultra-Compact Macromolecular Format (UMF)

High-Performance, Zero-Copy Binary Container for Macromolecular Complexes (Proteins, DNA, RNA, Small-Molecule Ligands, & Ions) in Structural Bioinformatics & Deep Learning

C++20 Python 3.10+ FlatBuffers NeRF RMSD Storage


🚀 Overview

The Ultra-Compact Macromolecular Format (.umf) is a next-generation structural container format designed to replace legacy ASCII formats (.pdb, .cif) and lossy monomer representations in modern machine learning pipelines, virtual screening, and structural biology archives.

Unlike traditional text-based formats that require expensive string deserialization and runtime distance tree construction, .umf leverages Google FlatBuffers to enable direct zero-copy memory mapping into CPU and GPU PyTorch/NumPy tensors.

Core Architecture Highlights

  1. Universal Multi-Entity Support: First-class representation for multi-chain protein complexes, nucleic acids (DNA/RNA), small-molecule drugs/cofactors, metal ions, and crystallographic waters.
  2. Validated Chemical Graph Topology: Stores explicit covalent bond orders (single, double, triple, aromatic), formal charges, element numbers, and stereochemical parity for small-molecule ligands.
  3. Drift-Free Coordinate Stability ($\le 0.0025\text{ \AA}$): Solves the catastrophic kinematic drift of pure dihedral compressors (such as Foldcomp) via periodic Cartesian anchor snapping ($\Delta = 15$ residues).
  4. All-Atom NeRF Reconstruction: Quantized 16-bit rotamer dihedrals ($\chi_{1-4}$) allow deterministic reconstruction of all sidechain heavy atoms within $\le 0.30\text{ \AA}$ RMSD.
  5. Zero-Copy PyTorch Geometric HeteroData Streaming: Precomputed sparse spatial graphs ($r < 8.0\text{ \AA}$) for intra-protein contacts and bipartite protein-ligand interfaces, eliminating CPU $O(N \log N)$ distance tree generation and cutting dataloading latency by 5.6×.

📊 Benchmark Comparisons

1. Direct Head-to-Head Comparison (benchmarks/benchmark_sota.py)

Evaluated against Foldcomp (.fcz), BinaryCIF (.bcif), and PDB (.pdb) across representative benchmarks:

Format Avg Storage (B/res) Direct PyG Ingestion Coordinate Drift RMSD Multi-Chain & Ligands? Pre-Indexed Graphs?
Standard PDB (.pdb) ~1,000 B/res ~12 ms (text parse + KDTree) Exact Text only (no bonds) ❌ None
BinaryCIF (.bcif) ~110 B/res ~15 ms (decompression + parse) Exact Tabular ❌ None
Foldcomp (.fcz) ~17.9 B/res ~12 ms (decompress + KDTree) 1.96 Å (up to 17 Å on large proteins) ❌ Stripped / Crashes ❌ None
UMF (Backbone Mode) ~14.5 B/res ~0.35 ms (zero-copy memory map) ≤ 0.0025 Å (Drift-Free) Monomer fold archive ❌ None
UMF (All-Atom + Graphs) ~35.0 B/res ~0.35 ms (zero-copy PyG HeteroData) ≤ 0.0025 Å (Backbone) / 0.30 Å (Sidechains) ✅ Full RDKit fidelity ✅ Precomputed $r < 8\text{ \AA}$

2. Downstream Virtual Screening GNN Training (benchmarks/benchmark_pdbbind_virtual_screening.py)

5-Epoch training benchmark on 436 co-crystallized complexes from PDBbind:

  • Dataloading I/O Latency: 5.6× lower with UMF zero-copy streaming ($1.09\text{ s}$ vs. $6.06\text{ s}$ per epoch).
  • End-to-End Epoch Wall-Clock Time: 3.3× faster ($2.14\text{ s}$ vs. $7.13\text{ s}$ per epoch).
  • GPU Dataloader Compute Saturation: Increases active compute saturation from 15.0% up to 49.0% (reducing idle CPU wait from 85% to 51%).

📦 Installation & Build

Prerequisites

  • CMake $\ge$ 3.20 & Ninja
  • C++20 compiler (GCC/G++ $\ge$ 11, Clang $\ge$ 14, or MSVC 2019+)
  • Python 3.10+
  • Google FlatBuffers (flatc)

Install from PyPI

pip install umf-format

Building from Source

# Clone the repository
git clone https://github.com/messiay/Universal_Macromolecular_formate.git
cd Universal_Macromolecular_formate

# Install locally with pip:
pip install -e .

🐍 Python Quickstart

import umf

# 1. Encode any PDB or mmCIF macromolecular complex into .umf binary format
umf.encode_complex("data/complexes/1STP.pdb", "data/complexes/1STP.umf", contact_cutoff=8.0, anchor_interval=15)

# 2. Ingest zero-copy tensors directly into PyTorch Geometric HeteroData
data = umf.to_hetero_tensors("data/complexes/1STP.umf")

# Directly ready for Equivariant GNNs & Virtual Screening
print(data['protein'].coords.shape)                   # (N_prot, 3) Backbone CA coordinates
print(data['protein', 'contact', 'protein'].edge_index) # (2, E_prot) Intra-protein contact graph
print(data['ligand'].coords.shape)                    # (N_lig, 3) Small-molecule coordinates
print(data['ligand', 'bond', 'ligand'].edge_index)    # (2, E_bonds) Covalent chemical bonds
print(data['ligand', 'interacts', 'protein'].edge_index) # (2, E_bip) Bipartite pocket graph

🧪 Testing & Verification

Run the test suite:

# 1. Test universal macromolecular complex encoding (Biotin-Streptavidin, B-DNA, Hemoglobin)
uv run python -m pytest tests/test_complex_encoding.py -v

# 2. Test adversarial failure cases (disordered loops, negative residues, pure ligands)
uv run python tests/test_adversarial_failure_cases.py

# 3. Test all-atom NeRF reconstruction precision
uv run python tests/test_all_atom_and_features.py

# 4. Run SOTA head-to-head comparison
uv run python benchmarks/benchmark_sota.py 5

📄 License

Licensed under the Apache 2.0 License.

Metadata

Release files for umf-format 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for umf-format 0.2.0
File Size Uploaded
umf_format-0.2.0.tar.gz 1.1 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for umf-format 0.2.0
File Interpreter ABI Platform
umf_format-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 2.2 MB

Release files / umf_format-0.2.0.tar.gz

Download URL umf_format-0.2.0.tar.gz
Size 1.1 MB
Tags Source
SHA-256 checksum
How to use checksums
603758f838f96c99318c4c45ad2e2e246054fefba9e3ff6cae6c6887de2bcc57
BLAKE2b-256 checksum
How to use checksums
b852ef4f37908d88eab7a078c1a31b3141185b8f1e208156861e1cb376753fe7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release files / umf_format-0.2.0-py3-none-any.whl

Download URL umf_format-0.2.0-py3-none-any.whl
Size 1.1 MB
Tags Python 3
SHA-256 checksum
How to use checksums
329c036d19b2065a08fbc8d2bc08dae334ea0b42387db254495046ade30bde01
BLAKE2b-256 checksum
How to use checksums
8ba6e2a7e2c1e3d9c86708004e40e2a27042b2d4a6a7b766a7bb15f55682f51a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page