evaldata
Evaluate AI-generated SQL with pytest.
evaldata runs text-to-SQL evals in your existing test suite.
It checks semantic equivalence of SQL queries, diffs result sets in your warehouse, and uses an LLM judge for ambiguous cases.
Why evaldata
- Semantic equivalence. Parse both queries, normalize their ASTs, and
compare canonical forms. It doesn't execute queries or call an LLM. When it
can't confirm equivalence, it returns
unknown. - Execution in your warehouse. Run the query on DuckDB, Postgres, Databricks, Snowflake, or BigQuery and compare the results, accounting for row order, NULLs, float tolerance, and types.
- It's just
pytest. Every eval is a test, run in your suite and your CI on every PR. No new runner, notebook, or dashboard. - An LLM judge when you need one. For ambiguous questions, missing reference answers, or explanations to grade, use a grader model with explicit criteria.
Quickstart
uv add evaldata # core, includes the DuckDB adapter
An eval is a pytest test: a case (a question and its expected answer), a solver
(the system under test that writes the SQL), and a scorer (how the answer is judged).
Below, the AI's SQL has reordered predicates and different casing, but means the same thing
as the reference query. observed_equivalence() confirms the match with AST normalization;
no query runs.
from evaldata import CallableSolver, EvalCase, assert_eval, eval_case, observed_equivalence
from evaldata.platforms import duckdb_platform
platform = duckdb_platform(name="shop", path="shop.duckdb")
@eval_case(
input="Name the US customers with an id above 1.",
expected={"kind": "gold_query", "sql": "SELECT name FROM customers WHERE country = 'US' AND id > 1"},
platform=platform,
)
def test_us_customers(case: EvalCase) -> None:
solver = CallableSolver(lambda c: "select NAME from customers where id > 1 and country = 'US'")
assert_eval(case, solver, scorers=[observed_equivalence()])
uv run pytest
case result detail
──────────────────────────────────
test_us_customers PASS
1 passed, 0 failed
The full runnable version is in
examples/01_deterministic/test_showcase.py.
To test a real model instead of fixed SQL, swap the solver for
PromptSolver(model="openai/gpt-4o-mini") (needs the evaldata[litellm] extra). To judge
equivalence without a warehouse, swap the scorer for judged_equivalence(model).
More use cases
- Add execution-based SQL scoring to Pydantic Evals.
- Evaluate dbt projects against gold SQL.
- Evaluate dbt Semantic Layer queries against gold MetricFlow queries.
- Evaluate Snowflake Cortex Analyst against gold SQL.
- Evaluate BigQuery queries against a live project.
- Reproduce dbt's Semantic Layer benchmark locally on DuckDB.
Install
uv add evaldata # core (includes the DuckDB adapter)
uv add "evaldata[postgres]" # + Postgres adapter
uv add "evaldata[databricks]" # + Databricks adapter
uv add "evaldata[snowflake]" # + Snowflake adapter
uv add "evaldata[bigquery]" # + BigQuery adapter
uv add "evaldata[cortex]" # + Snowflake Cortex Analyst solver
uv add "evaldata[litellm]" # + litellm, to call a model from PromptSolver
uv add "evaldata[pydantic-evals]" # + Pydantic Evals integration
DuckDB, Postgres, Databricks, Snowflake, and BigQuery are the adapters available today.
Documentation
Full documentation: monospaceai.github.io/evaldata
- Getting started: write and run your first eval.
- Concepts: cases, solvers, scorers, and platforms.
- Scoring guides: semantic equivalence, LLM judge, composing scorers.
- Model guides: local Ollama, hosted model.
- Integration guides: Pydantic Evals, dbt project, dbt Semantic Layer.
- Platform guides: Databricks, Snowflake, BigQuery, Cortex Analyst.
- Reproduce dbt's Semantic Layer benchmark.
- Run a text-to-SQL benchmark: load a Spider/BIRD dataset and measure execution accuracy.
- API reference: the public API, generated from docstrings.
Examples
Runnable examples in examples/:
| Example | Shows |
|---|---|
| Showcase | Compare SQL meaning and results on DuckDB; no setup |
| Deterministic | Score SQL with expected results, reference queries, and data expectations on DuckDB |
| Local AI | Text-to-SQL evaluation with a local Ollama model |
| Hosted AI | Text-to-SQL evaluation with a hosted model |
| Databricks | Score SQL with expected results, reference queries, and data expectations on Databricks SQL Warehouse |
| LLM judge | Use an LLM to judge SQL equivalence |
| Benchmark | Measure text-to-SQL execution accuracy with Spider or BIRD datasets |
| Snowflake | Score SQL with expected results, reference queries, and data expectations on Snowflake |
| Cortex Analyst | Score Snowflake Cortex Analyst queries against expected results |
| BigQuery | Score SQL with expected results, reference queries, and data expectations on BigQuery |
| dbt | Text-to-SQL evaluation for a dbt project with fixed model responses |
| dbt Semantic Layer | dbt Semantic Layer (MetricFlow) queries, scored locally on DuckDB |
| Pydantic Evals | Add execution-based SQL scoring to a Pydantic Evals dataset |
See examples/README.md for details.
Contributing
git clone https://github.com/monospaceai/evaldata.git
cd evaldata
uv sync # core + dev tooling
uv run pre-commit install
just check # lint + typecheck + local tests with coverage
Run just --list for other development commands.
Platform e2e tests
Adapter conformance for real platforms is marked e2e. CI provisions Postgres as a
service container and runs the suite on every push, so the Postgres adapter is exercised
against a real engine on every change.
Run it locally against Postgres with:
docker compose up -d # postgres:17 on localhost:5432
uv run --extra postgres pytest -m e2e # connection via POSTGRES_TEST_* env (defaults match compose)
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file evaldata-0.9.0.tar.gz.
File metadata
- Download URL: evaldata-0.9.0.tar.gz
- Upload date:
- Size: 116.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.11.31 {"installer":{"name":"uv","version":"0.11.31","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
acde93b1ae239ec8a5e62197332ec03bed618b53b345c50c7e5f43bb4acb9f15
|
|
| MD5 |
34c9cf57fd391e6bb74992297849cc9f
|
|
| BLAKE2b-256 |
8c2e1e73119c30d4b9ce8c414afae610f8f981ab3fe8def2293e0b0331898004
|
File details
Details for the file evaldata-0.9.0-py3-none-any.whl.
File metadata
- Download URL: evaldata-0.9.0-py3-none-any.whl
- Upload date:
- Size: 152.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.11.31 {"installer":{"name":"uv","version":"0.11.31","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3a38382763b1147b4bd18bd5e654014f6c8a72b42eee684b2b629afd446ea046
|
|
| MD5 |
1b8f202cf8754fa3ddd86135c3b02663
|
|
| BLAKE2b-256 |
d1ccd1e21bb760dd863add7e1bfcf371ee12c2ee6aceb9ce5666bfecd4f46982
|