Skip to main content

plinder

The Protein Ligand INteractions Dataset and Evaluation Resource


license publish website bioRxiv docs coverage

overview

📚 About

PLINDER, short for protein ligand interactions dataset and evaluation resource, is a comprehensive, annotated, high quality dataset and resource for training and evaluation of protein-ligand docking algorithms:

  • > 400k PLI systems across > 11k SCOP domains and > 50k unique small molecules
  • 750+ annotations for each system, including protein and ligand properties, quality, matched molecular series and more
  • Automated curation pipeline to keep up with the PDB
  • 14 PLI metrics and over 20 billion similarity scores
  • Unbound (apo) and predicted Alphafold2 structures linked to holo systems
  • train-val-test splits and ability to tune splitting based on the learning task
  • Robust evaluation harness to simplify and standard performance comparison between models.

The PLINDER project is a community effort, launched by the University of Basel, SIB Swiss Institute of Bioinformatics, Proxima (formerly VantAI), NVIDIA, MIT CSAIL, and will be regularly updated.

To accelerate community adoption, PLINDER will be used as the field’s new Protein-Ligand interaction dataset standard as part of an exciting competition at the upcoming 2024 Machine Learning in Structural Biology (MLSB) Workshop at NeurIPS, one of the field's premiere academic gatherings. More details about the competition and other helpful practical tips can be found at our recent workshop repo: Moving Beyond Memorization.

👋 Join the P(L)INDER user group Discord Server!

🔢 Plinder versions

We version the plinder dataset with two controls:

  • PLINDER_RELEASE: the month stamp of the last RCSB sync
  • PLINDER_ITERATION: value that enables iterative development within a release

We version the plinder application using an automated semantic versioning scheme based on the git commit history. The plinder.data package is responsible for generating a dataset release and the plinder.core package makes it easy to interact with the dataset.

🐛🐛🐛 Known bugs:

  • Source dataset contains incorrect entry_release_date dates, please, use query_index to get correct dates patched.
  • Complexes containing nucleic acid receptors may not be saved corectly.
  • ligand_binding_affinity queries have been disabled due to a bug found parsing BindingDB

Changelog:

Unreleased:

  • Public downloads for 2024-06/v2 now use https://plinderdata.org (Cloudflare R2). The consumer no longer supports GCS or downloading older releases.

  • Downloads stream to disk and verify size and MD5 before atomic replacement. Interrupted transfers retry from the start. Cached files with the expected size are reused; offline access retains the existing cache layout. Use a forced refresh to repair same-size local corruption.

  • Google storage dependencies are now optional: install plinder[data] for the GCS data-generation utilities.

  • 2024-06/v2 (Current):

    • New systems added based on the 2024-06 RCSB sync
    • Updated system definition to be more stable and depend only on ligand distance rather than PLIP
    • Added annotations for crystal contacts
    • Improved ligand handling and saving to fix some bond order issues
    • Improved covalency detection and annotation to reference each bond explicitly
    • Added linked apo/pred structures to v2/links and v2/linked_structures
    • Added binding affinity annotations from BindingDB (see known bugs!)
    • Added statistics requirement and other changes in the split to enrich test set diversity
  • 2024-04/v1: Version described in the preprint, with updated redundancy removal by protein pocket and ligand similarity.

  • 2024-04/v0: Version used to re-train DiffDock in the paper, with redundancy removal based on <pdbid>_<ligand ccd codes>

🏅 Gold standard benchmark sets

As part of PLINDER resource we provide train, validation and test splits that are curated to minimize the information leakage based on protein-ligand interaction similarity. In addition, we have prioritized the systems that has a linked experimental apo structure or matched molecular series to support realistic inference scenarios for hit discovery and optimization. Finally, a particular care is taken for test set that is further prioritized to contain high quality structures to provide unambiguous ground-truths for performance benchmarking.

test_stratification

Moreover, as we enticipate this resource to be used for benchmarking a wide range of methods, including those simultaneously predicting protein structure (aka. co-folding) or those generating novel ligand structures, we further stratified test (by novel ligand, pocket, protein or all) to cover a wide range of tasks.

👨💻 Getting Started

The PLINDER dataset is provided in two ways:

  • You can either use the files from the dataset directly using your preferred tooling by downloading the data from the public bucket,
  • or you can utilize the dedicated plinder Python package for interfacing the data.

Downloading the dataset

Install the Python package below, then run plinder_download to download the published 2024-06/v2 release from Cloudflare R2. Dataset APIs also download required files lazily. No Google Cloud credentials or SDK are needed.

The default endpoint is https://plinderdata.org; PLINDER_MIRROR_URL can select another HTTP mirror. This release no longer falls back to GCS. Other release versions are unavailable through the public downloader. Existing cache paths and PLINDER_OFFLINE behavior are preserved.

Installing the Python package

plinder is available on PyPI.

pip install plinder

License

Data curated by PLINDER are made available under the Apache License 2.0. All data curated by BindingDB staff are provided under the Creative Commons Attribution 4.0 License. Data imported from ChEMBL are provided under their Creative Commons Attribution-Share Alike 4.0 Unported License.

📝 Documentation

A more detailed description is available on the documentation website.

📃 Citation

Durairaj, Janani, Yusuf Adeshina, Zhonglin Cao, Xuejin Zhang, Vladas Oleinikovas, Thomas Duignan, Zachary McClure, Xavier Robin, Gabriel Studer, Daniel Kovtun, Emanuele Rossi, Guoqing Zhou, Srimukh Prasad Veccham, Clemens Isert, Yuxing Peng, Prabindh Sundareson, Mehmet Akdel, Gabriele Corso, Hannes Stärk, Gerardo Tauriello, Zachary Wayne Carpenter, Michael M. Bronstein, Emine Kucukbenli, Torsten Schwede, Luca Naef. 2024. “PLINDER: The Protein-Ligand Interactions Dataset and Evaluation Resource.” bioRxiv ICML'24 ML4LMS

Please see the citation file for details.

plinder_banner

Release files for plinder 0.2.27

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for plinder 0.2.27
File Size Uploaded
plinder-0.2.27.tar.gz 28.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for plinder 0.2.27
File Interpreter ABI Platform
plinder-0.2.27-py3-none-any.whl Python 3 none any Details

Total release size: 33.1 MB

Release files / plinder-0.2.27.tar.gz

Download URL plinder-0.2.27.tar.gz
Size 28.0 MB
Tags Source
SHA-256 checksum
How to use checksums
7ea3f989f08c0c5e02d127e28e7ed8326878b2ce3122da201b258411c56cceb6
BLAKE2b-256 checksum
How to use checksums
58736c2afa5de48b58ac03a9077ccb6a13812df0a88dcfd265a3491b516f4a5d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release files / plinder-0.2.27-py3-none-any.whl

Download URL plinder-0.2.27-py3-none-any.whl
Size 5.0 MB
Tags Python 3
SHA-256 checksum
How to use checksums
9490e0bd53be3bcc24b328d0c595a0916fb833fc2376627274aa9ce4f01b60cd
BLAKE2b-256 checksum
How to use checksums
2db0bcb1fec66bb35424a7088d8f84dd6d6adafdb7adc7b8b264993b68c3193b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.27 This release

2 release files

0.2.26

2 release files

0.2.24

2 release files

0.2.23

2 release files

0.2.20

2 release files

0.2.19

2 release files

0.2.18

2 release files

0.2.17

2 release files

0.2.16

2 release files

0.2.15

2 release files

0.2.14

2 release files

0.2.13

2 release files

0.2.12

2 release files

0.2.11

2 release files

0.2.10

2 release files

0.2.9

2 release files

0.2.8

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.19

2 release files

0.1.18

2 release files

0.1.17

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page