Skip to main content

MDBench

PyPI Python Documentation License

简体中文

MDBench evaluates whether an AI system can recover scientific laws and the mechanisms that produce them from equations or observations.

What mechanism discovery means

MDBench treats a phenomenological equation as the observable consequence of several simple, mutually consistent relationships. The phenomenological law describes what variables do; a mechanism explains why through physical relationships, assumptions, and intermediate variables.

For example, Kepler's third law for a circular orbit follows from gravitation, Newton's second law, and uniform circular motion. See problems/demo_problem.yaml.

Each mechanism relationship uses variable = formula, where the formula must be parseable by nd2py. Explicit relationships form a DAG:

a = f1(x)
b = f2(x, a)
y = f3(x, a, b)

Implicit systems are also supported. Relationships are collected until the unknown variables form a closed system, then solved symbolically or with a numerical root finder:

a = f1(x, a, b)
b = f2(x, a, b)
y = f3(x, a, b)

All variables are declared under variable_description as target, inputs, intermediates, or auxiliary_inputs. The latter are external variables used only by the mechanism and eliminated from the final law. The original relationships remain in Problem.mechanism; executable solution steps are stored in Problem.solution.

Tasks and evaluation

MDBench provides three tasks:

  1. Symbolic regression: (X, y) → phenomenological equation.
  2. Mechanism explanation: phenomenological equation → mechanism equations.
  3. Mechanism discovery: (X, y) → mechanism equations.

Mechanism evaluation reports independent metrics and deliberately has no overall score:

  • Prediction accuracy: for symbolic regression and mechanism discovery, Pearson correlation, R², MAE, RMSE, sMAPE, and tolerance accuracy on public training data (feedback) or train/ID/OOD data (final).
  • Derived-equation equivalence: final-only SymPy, numeric, and LLM cross-check against the private phenomenological equation.
  • Mechanism fundamentality: LLM assessment dominated by the least fundamental submitted relationship; no reference answer is required.
  • Ground-truth structure recovery: soft formula-AST and dependency-graph matching against the reference mechanism. Variable names and numeric literal values are ignored.
  • Mechanism description complexity: reference-free mean, maximum, and total nd2py AST nodes; lower values describe simpler submitted relationships.

Install MDBench with Python 3.12 or newer:

pip install mdbench
mdbench --help

The Sphinx documentation lives in docs/. Build it with:

cd docs
make html

Commands

Export bundled problems

MDBench includes its reference problem library. Export it to a local directory without downloading repository files:

mdbench export --output-dir problems/

The same YAML collection is also available as a prebuilt archive from GitHub Releases. Releases also provide a complete public mechanism-discovery dataset containing problem.json, answer.json, and the train, in-domain-test, and out-of-domain-test arrays for every problem.

Existing files require confirmation; use --force for non-interactive overwrites. Other files in the destination directory are never removed.

The lifecycle commands below accept one or more YAML files or directories through --problems; the default is ./problems.

Validate problems

Checks schemas, variable usage, units, sampling specifications, explicit and implicit equation solving, and derivation of the target law:

mdbench validate --problems problems
mdbench validate --problems problems/demo_problem.yaml

An optional LLM check evaluates whether every relationship is sufficiently fundamental. API or response failures are reported directly and do not fall back to heuristics.

mdbench validate --problems problems --check-fundamentality \
  --llm-provider deepseek --llm-model deepseek-v4-flash

Generate synthetic data

Creates reproducible train, ID-test, and OOD-test splits:

mdbench synthetic --problems problems/ --output-dir data/synthetic_data/

Each NPZ stores the three arrays, their row order in variables, and a JSON generation_config containing the seed and sample counts. Auxiliary inputs are generated here and may be hidden later during task preparation.

Prepare tasks

Synthetic data must already exist. Answers are private by default:

mdbench prepare \
  --problems problems/ \
  --synthetic-data-dir data/synthetic_data/ \
  --task mechanism_discovery \
  --format directory

Use --save-answer to include answers and test splits, --reveal-auxiliary to expose auxiliary inputs in mechanism tasks, and --force to approve planned overwrites. Existing directories are never cleared; redundant files are reported. --format directory writes flat files, while --format file packs the same logical artifacts into one NPZ.

Evaluate submissions

A submission may be an inline formula, semicolon-separated mechanism equations, or a plain-text file with one equation per non-empty line. JSON and YAML submissions are intentionally unsupported.

mdbench evaluate \
  --evaluation-mode feedback \
  --problem data/problem/PREPARED_TASK \
  --submission submission.txt \
  --verbose

Feedback mode uses only the public task and training data. Benchmark operators run final evaluation with --evaluation-mode final --answer answer.json, which also enables hidden ID/OOD tests and reference-mechanism recovery. For Agent runs, copy only the prepared public task into an isolated temporary working directory and require the Agent to remain there. Without source problem YAML or private answer artifacts, the other lifecycle commands and final evaluation cannot access the material they require. --verbose prints concise equation chains for explicit or implicit solution steps.

Fundamentality scoring automatically uses the configured external model and prints its provider and model:

mdbench evaluate \
  --evaluation-mode feedback \
  --problem data/problem/PREPARED_TASK \
  --submission submission.txt \
  --llm-provider deepseek \
  --llm-model deepseek-v4-flash

Standalone entry points with equivalent behavior are available in scripts/:

validate_problem_main.py   validate problem definitions
synthetic_data_main.py     generate synthetic datasets
prepare_problem_main.py    prepare public/private task artifacts
evaluate_result_main.py    evaluate a submission

scripts/visualize_mechanism_main.py renders a solved mechanism as DOT, SVG, PNG, or PDF. Non-DOT formats require Graphviz.

Directory conventions

problems/                  source problem YAML files
data/
  synthetic_data/         generated train/ID/OOD datasets
  problem/                prepared benchmark tasks
src/
  core/                   dependency-light data models
  features/               project-specific I/O, solving, validation, sampling
  metrics/                formula and mechanism metrics
  utils/                  reusable utilities and LLM clients
scripts/                   standalone command entry points
tests/                     unit tests and validation fixtures

A prepared directory contains:

problem.json               public task description
data_train.npy             public training data for data-input tasks
answer.json                optional private answer
data_id_test.npy           optional private ID test data for data-input tasks
data_ood_test.npy          optional private OOD test data for data-input tasks

Without --save-answer, data-input tasks contain problem.json and data_train.npy; mechanism-explanation tasks contain only problem.json. Public interchange types remain simple: units are Dict[str, int | float], formulas are nd2py-compatible strings, and arrays use NumPy formats.

Metadata

Release files for mdbench 0.4.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mdbench 0.4.1
File Size Uploaded
mdbench-0.4.1.tar.gz 86.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mdbench 0.4.1
File Interpreter ABI Platform
mdbench-0.4.1-py3-none-any.whl Python 3 none any Details

Total release size: 188.4 kB

Release files / mdbench-0.4.1.tar.gz

Download URL mdbench-0.4.1.tar.gz
Size 86.2 kB
Tags Source
SHA-256 checksum
How to use checksums
24090f7104257dcf69512efdb131d194a11491771c097b59a8f72191797fcf90
BLAKE2b-256 checksum
How to use checksums
e6e5474537f8d62f4ce7ab30f33a8e2ae2fef8f95e307fa8b124605fdcb9acc7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 13, 2026.

Transparency log

Release files / mdbench-0.4.1-py3-none-any.whl

Download URL mdbench-0.4.1-py3-none-any.whl
Size 102.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f805890caa2131a95b5a4f01f07f2ca83ebc0ad42372fdbd7876c755e3f7cac2
BLAKE2b-256 checksum
How to use checksums
24732e5abecdece867ac727c7b2669cf6f75ad507929eadfa21139188c13b700
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 13, 2026.

Transparency log

Release history Release notifications | RSS feed

0.7.0

2 release files

0.6.1

2 release files

This release

0.4.1 This release

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page