Skip to main content

MDBench

PyPI Python Documentation License

简体中文

MDBench evaluates whether an AI system can recover scientific laws and the mechanisms that produce them from equations or observations.

What mechanism discovery means

MDBench treats a phenomenological equation as the observable consequence of several simple, mutually consistent relationships. The phenomenological law describes what variables do; a mechanism explains why through physical relationships, assumptions, and intermediate variables.

For example, Kepler's third law for a circular orbit follows from gravitation, Newton's second law, and uniform circular motion. See demo_problem.yaml.

Each mechanism relationship uses left_expression = right_expression, where both sides must be parseable by nd2py. The loader stores the relationship as the residual left_expression - right_expression = 0. Explicit relationships form a DAG:

a = f1(x)
b = f2(x, a)
y = f3(x, a, b)

Implicit systems are also supported. At each step, the solver selects N equations containing exactly N unresolved variables; those equations need not be contiguous and forms such as 0 = F - m * a are valid. The selected system is then solved symbolically or with a numerical root finder:

a = f1(x, a, b)
b = f2(x, a, b)
y = f3(x, a, b)

All variables are declared under variable_description as target, inputs, intermediates, or auxiliary_inputs. The latter are external variables used only by the mechanism and eliminated from the final law. The original relationships remain in Problem.mechanism; executable solution steps are stored in Problem.solution.

Tasks and evaluation

MDBench provides three tasks:

  1. Symbolic regression: (X, y) → phenomenological equation.
  2. Mechanism explanation: phenomenological equation → mechanism equations.
  3. Mechanism discovery: (X, y) → mechanism equations.

Mechanism evaluation reports independent metrics and deliberately has no overall score:

  • Prediction accuracy: for symbolic regression and mechanism discovery, Pearson correlation, R², MAE, RMSE, sMAPE, and tolerance accuracy on public training data (feedback) or train/ID/OOD data (final).
  • Derived-equation equivalence: final-only SymPy, numeric, and LLM cross-check against the private phenomenological equation.
  • Mechanism fundamentality: LLM assessment dominated by the least fundamental submitted relationship; no reference answer is required.
  • Ground-truth structure recovery: soft formula-AST and dependency-graph matching against the reference mechanism. Variable names and numeric literal values are ignored.
  • Mechanism description complexity: reference-free mean, maximum, and total nd2py AST nodes; lower values describe simpler submitted relationships.

Install MDBench with Python 3.12 or newer:

pip install mdbench
mdbench --help

The Sphinx documentation lives in docs/. Build it with:

cd docs
make html

Commands

Export bundled problems

MDBench includes its reference problem library. Export it to a local directory without downloading repository files:

mdbench export --output-dir problems/

The same YAML collection is also available as a prebuilt archive from GitHub Releases. Releases also provide a complete public mechanism-discovery dataset containing problem.json, answer.json, and the train, in-domain-test, and out-of-domain-test arrays for every problem.

Existing files require confirmation; use --force for non-interactive overwrites. Other files in the destination directory are never removed.

The lifecycle commands below accept one or more YAML files or directories through --problems; the default is ./problems.

Validate problems

Checks schemas, variable usage, units, sampling specifications, explicit and implicit equation solving, and derivation of the target law:

mdbench validate --problems problems
mdbench validate --problems demo_problem.yaml

An optional LLM check evaluates whether every relationship is sufficiently fundamental. API or response failures are reported directly and do not fall back to heuristics.

mdbench validate --problems problems --check-fundamentality \
  --llm-provider deepseek --llm-model deepseek-v4-flash

Generate synthetic data

Creates reproducible train, ID-test, and OOD-test splits:

mdbench synthetic --problems problems/ --output-dir data/synthetic_data/

Each NPZ stores the three arrays, their row order in variables, and a JSON generation_config containing the seed and sample counts. Auxiliary inputs are generated here and may be hidden later during task preparation.

Prepare tasks

Synthetic data must already exist. Answers are private by default:

mdbench prepare \
  --problems problems/ \
  --synthetic-data-dir data/synthetic_data/ \
  --task mechanism_discovery \
  --format directory

Use --save-answer to include answers and test splits, --reveal-auxiliary to expose auxiliary inputs in mechanism tasks, and --force to approve planned overwrites. Existing directories are never cleared; redundant files are reported. --format directory writes flat files, while --format file packs the same logical artifacts into one NPZ.

Evaluate submissions

A submission may be an inline formula, semicolon-separated mechanism equations, or a plain-text file with one equation per non-empty line. JSON and YAML submissions are intentionally unsupported.

mdbench evaluate \
  --evaluation-mode feedback \
  --problem data/problem/PREPARED_TASK \
  --submission submission.txt \
  --verbose

Feedback mode uses only the public task and training data. Benchmark operators run final evaluation with --evaluation-mode final --answer answer.json, which also enables hidden ID/OOD tests and reference-mechanism recovery. For Agent runs, copy only the prepared public task into an isolated temporary working directory and require the Agent to remain there. Without source problem YAML or private answer artifacts, the other lifecycle commands and final evaluation cannot access the material they require. --verbose prints concise equation chains for explicit or implicit solution steps.

Fundamentality scoring automatically uses the configured external model and prints its provider and model:

mdbench evaluate \
  --evaluation-mode feedback \
  --problem data/problem/PREPARED_TASK \
  --submission submission.txt \
  --llm-provider deepseek \
  --llm-model deepseek-v4-flash

Standalone entry points with equivalent behavior are available in scripts/:

validate_problem_main.py   validate problem definitions
synthetic_data_main.py     generate synthetic datasets
prepare_problem_main.py    prepare public/private task artifacts
evaluate_result_main.py    evaluate a submission

scripts/visualize_mechanism_main.py renders a solved mechanism as DOT, SVG, PNG, or PDF. Non-DOT formats require Graphviz.

Directory conventions

problems/                  source problem YAML files
data/
  synthetic_data/         generated train/ID/OOD datasets
  problem/                prepared benchmark tasks
src/
  core/                   dependency-light data models
  features/               project-specific I/O, solving, validation, sampling
  metrics/                formula and mechanism metrics
  utils/                  reusable utilities and LLM clients
scripts/                   standalone command entry points
tests/                     unit tests and validation fixtures

A prepared directory contains:

problem.json               public task description
data_train.npy             public training data for data-input tasks
answer.json                optional private answer
data_id_test.npy           optional private ID test data for data-input tasks
data_ood_test.npy          optional private OOD test data for data-input tasks

Without --save-answer, data-input tasks contain problem.json and data_train.npy; mechanism-explanation tasks contain only problem.json. Public interchange types remain simple: units are Dict[str, int | float], formulas are nd2py-compatible strings, and arrays use NumPy formats.

Metadata

Release files for mdbench 0.7.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mdbench 0.7.0
File Size Uploaded
mdbench-0.7.0.tar.gz 107.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mdbench 0.7.0
File Interpreter ABI Platform
mdbench-0.7.0-py3-none-any.whl Python 3 none any Details

Total release size: 279.0 kB

Release files / mdbench-0.7.0.tar.gz

Download URL mdbench-0.7.0.tar.gz
Size 107.1 kB
Tags Source
SHA-256 checksum
How to use checksums
da1c9369b1413deafb1f166759f3e7a80e13b5de86e790a12aa109feddb24bcd
BLAKE2b-256 checksum
How to use checksums
e7fdf831f1a937c0e45af7ef75dd3d9a44eda18311e8df3357fe6b6173cb8183
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.

Transparency log

Release files / mdbench-0.7.0-py3-none-any.whl

Download URL mdbench-0.7.0-py3-none-any.whl
Size 172.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
409acacd253bbb1a5d4260f9da2092e0887a1795ce4a5d105cf7462fedcd22a3
BLAKE2b-256 checksum
How to use checksums
ecb64faad09ab1d5337f583baf7e7bed03e3e569c5311d0c9cb65b39c8ee21fa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.7.0 This release

2 release files

0.6.1

2 release files

0.4.1

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page