sreval
An evaluator for symbolic regression: triple-equivalence testing that reports its own failure rate.
Why this exists
Symbolic regression is supposed to find the equation, not merely a function that fits. The published benchmark results show those are very different achievements: a method can score above 0.999 on a coefficient of determination while recovering the correct expression structure zero percent of the time. An evaluation that reports only accuracy cannot tell a discovery from a good curve fit.
Testing structure directly is the fix, but no single structural test is reliable:
| Test | Strength | Failure mode |
|---|---|---|
| Symbolic simplification | Exact when it terminates | Times out, throws, or fails to reduce a genuinely zero difference. Scoring those as method failures charges the search for a defect in the scorer. |
| Numerical probing | Robust and cheap | Cannot distinguish "equivalent everywhere" from "agrees on the region we sampled", which is exactly what breaks under extrapolation. |
| Structural edit distance | The only one giving graded credit | Sensitive to algebraically equivalent rewrites that a human would call the same answer. |
So sreval runs all three, reports them separately, reports whether they agreed, and
publishes the symbolic-test failure rate as a first-class number. That last one is the
contribution: it is a property of the measurement apparatus that the field currently absorbs
silently into the scores of the methods being measured.
Install
pip install sreval # numpy only
pip install "sreval[symbolic]" # adds sympy, needed for the symbolic test
Use
import numpy as np
from sreval import check, summarise
verdict = check(
candidate_infix="x*(a + b)",
truth_infix="x*a + x*b",
variables=["x", "a", "b"],
candidate_fn=lambda X: X[:, 0] * (X[:, 1] + X[:, 2]),
truth_fn=lambda X: X[:, 0] * X[:, 1] + X[:, 0] * X[:, 2],
box=[(-3.0, 3.0), (-3.0, 3.0), (-3.0, 3.0)],
candidate_tokens=["mul", "v0", "add", "v1", "v2"],
truth_tokens=["add", "mul", "v0", "v1", "mul", "v0", "v2"],
)
verdict.recovered # True: an algebraic rewrite of the same expression
verdict.agreed # True: the symbolic and numerical tests concur
verdict.structural.distance # graded and non-zero: written differently
report = summarise([verdict])
report.symbolic_failure_rate # the number nobody else publishes
report.notes # plain-language statements of what the rates mean
Accuracy and recovery are reported separately, always
from sreval import accuracy_solution, description_length
accuracy_solution is named at length on purpose. An accuracy solution is not a recovery, and the
gap between those two rates is the measurement this package was built to expose. Nothing in the
API merges them for you.
description_length is the recommended rule for picking one point off a Pareto front. Selecting the
most accurate member reproduces the field's headline failure, because the most accurate member of a
front is routinely the most over-parameterised one.
Scope
- It does not perform symbolic regression. It evaluates results from any engine.
- It does not decide whether a discovered equation is true. Fitting is not discovering, and no equivalence test changes that.
- The structural distance is a sequence edit distance over the pre-order traversal, not a full tree edit distance. The former is O(n*m), the latter O(n^2 m^2), and at the expression sizes symbolic regression actually produces the two agree closely enough that the extra cost buys nothing. This is stated because the number is reported, and a reported metric with an implicit definition is not reproducible.
Provenance
Written from published specifications. The widely used reference implementation of this evaluation protocol is GPL-3.0 licensed and no part of it is reproduced here, which is what allows this package to be MIT.
Built for SymLab, a public research lab on symbolic regression, and extracted as a standalone package because the evaluation problem is not specific to that lab.
Developed by Felipe Santibanez-Leal.
Metadata
Release files for sreval 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| sreval-0.1.1.tar.gz | 13.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sreval-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 24.9 kB
Release files / sreval-0.1.1.tar.gz
| Download URL | sreval-0.1.1.tar.gz |
|---|---|
| Size | 13.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f91c3931ef87e032930527424d189570d5325e154258863e3aad123599ac62fd
|
|
BLAKE2b-256 checksum How to use checksums |
8f51e728a6810f32ca49830ef6226d67d12384e7cc94ac67aed64e522ae41c2e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.
Transparency logRelease files / sreval-0.1.1-py3-none-any.whl
| Download URL | sreval-0.1.1-py3-none-any.whl |
|---|---|
| Size | 11.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
57fa58fe58e85d8328b7b01fad27db4fee0a1c337b3db3396f6d75fea29b07c6
|
|
BLAKE2b-256 checksum How to use checksums |
bb8ba789be4ba5dc31e17e98bece0a5067838615e8c78fc712fbcd0f4687a8d6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.
Transparency log