Skip to main content

evaldata

CI Coverage License: Apache 2.0

Evaluate AI-generated SQL with pytest.

evaldata runs text-to-SQL evals in your existing test suite.

It checks semantic equivalence of SQL queries, diffs result sets in your warehouse, and uses an LLM judge for ambiguous cases.

Why evaldata

  • Semantic equivalence. Parse both queries, normalize their ASTs, and compare canonical forms. It doesn't execute queries or call an LLM. When it can't confirm equivalence, it returns unknown.
  • Execution in your warehouse. Run the query on DuckDB, Postgres, Databricks, Snowflake, or BigQuery and compare the results, accounting for row order, NULLs, float tolerance, and types.
  • It's just pytest. Every eval is a test, run in your suite and your CI on every PR. No new runner, notebook, or dashboard.
  • An LLM judge when you need one. For ambiguous questions, missing reference answers, or explanations to grade, use a grader model with explicit criteria.

Quickstart

uv add evaldata   # core, includes the DuckDB adapter

An eval is a pytest test: a case (a question and its expected answer), a solver (the system under test that writes the SQL), and a scorer (how the answer is judged).

Below, the AI's SQL has reordered predicates and different casing, but means the same thing as the reference query. observed_equivalence() confirms the match with AST normalization; no query runs.

from evaldata import CallableSolver, EvalCase, assert_eval, eval_case, observed_equivalence
from evaldata.platforms import duckdb_platform

platform = duckdb_platform(name="shop", path="shop.duckdb")


@eval_case(
    input="Name the US customers with an id above 1.",
    expected={"kind": "gold_query", "sql": "SELECT name FROM customers WHERE country = 'US' AND id > 1"},
    platform=platform,
)
def test_us_customers(case: EvalCase) -> None:
    solver = CallableSolver(lambda c: "select NAME from customers where id > 1 and country = 'US'")
    assert_eval(case, solver, scorers=[observed_equivalence()])
uv run pytest
 case               result   detail
 ──────────────────────────────────
 test_us_customers  PASS

 1 passed, 0 failed

The full runnable version is in examples/01_deterministic/test_showcase.py.

To test a real model instead of fixed SQL, swap the solver for PromptSolver(model="openai/gpt-4o-mini") (needs the evaldata[litellm] extra). To judge equivalence without a warehouse, swap the scorer for judged_equivalence(model).

More use cases

Install

uv add evaldata                # core (includes the DuckDB adapter)
uv add "evaldata[postgres]"    # + Postgres adapter
uv add "evaldata[databricks]"  # + Databricks adapter
uv add "evaldata[snowflake]"   # + Snowflake adapter
uv add "evaldata[bigquery]"    # + BigQuery adapter
uv add "evaldata[cortex]"      # + Snowflake Cortex Analyst solver
uv add "evaldata[litellm]"     # + litellm, to call a model from PromptSolver
uv add "evaldata[pydantic-evals]"  # + Pydantic Evals integration

DuckDB, Postgres, Databricks, Snowflake, and BigQuery are the adapters available today.

Documentation

Full documentation: monospaceai.github.io/evaldata

Examples

Runnable examples in examples/:

Example Shows
Showcase Compare SQL meaning and results on DuckDB; no setup
Deterministic Score SQL with expected results, reference queries, and data expectations on DuckDB
Local AI Text-to-SQL evaluation with a local Ollama model
Hosted AI Text-to-SQL evaluation with a hosted model
Databricks Score SQL with expected results, reference queries, and data expectations on Databricks SQL Warehouse
LLM judge Use an LLM to judge SQL equivalence
Benchmark Measure text-to-SQL execution accuracy with Spider or BIRD datasets
Snowflake Score SQL with expected results, reference queries, and data expectations on Snowflake
Cortex Analyst Score Snowflake Cortex Analyst queries against expected results
BigQuery Score SQL with expected results, reference queries, and data expectations on BigQuery
dbt Text-to-SQL evaluation for a dbt project with fixed model responses
dbt Semantic Layer dbt Semantic Layer (MetricFlow) queries, scored locally on DuckDB
Pydantic Evals Add execution-based SQL scoring to a Pydantic Evals dataset

See examples/README.md for details.

Contributing

git clone https://github.com/monospaceai/evaldata.git
cd evaldata
uv sync                       # core + dev tooling
uv run pre-commit install
just check                    # lint + typecheck + local tests with coverage

Run just --list for other development commands.

Platform e2e tests

Adapter conformance for real platforms is marked e2e. CI provisions Postgres as a service container and runs the suite on every push, so the Postgres adapter is exercised against a real engine on every change.

Run it locally against Postgres with:

docker compose up -d                  # postgres:17 on localhost:5432
uv run --extra postgres pytest -m e2e # connection via POSTGRES_TEST_* env (defaults match compose)

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

evaldata-0.9.0.tar.gz (116.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

evaldata-0.9.0-py3-none-any.whl (152.0 kB view details)

Uploaded Python 3

File details

Details for the file evaldata-0.9.0.tar.gz.

File metadata

  • Download URL: evaldata-0.9.0.tar.gz
  • Upload date:
  • Size: 116.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.31 {"installer":{"name":"uv","version":"0.11.31","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for evaldata-0.9.0.tar.gz
Algorithm Hash digest
SHA256 acde93b1ae239ec8a5e62197332ec03bed618b53b345c50c7e5f43bb4acb9f15
MD5 34c9cf57fd391e6bb74992297849cc9f
BLAKE2b-256 8c2e1e73119c30d4b9ce8c414afae610f8f981ab3fe8def2293e0b0331898004

See more details on using hashes here.

File details

Details for the file evaldata-0.9.0-py3-none-any.whl.

File metadata

  • Download URL: evaldata-0.9.0-py3-none-any.whl
  • Upload date:
  • Size: 152.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.31 {"installer":{"name":"uv","version":"0.11.31","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for evaldata-0.9.0-py3-none-any.whl
Algorithm Hash digest
SHA256 3a38382763b1147b4bd18bd5e654014f6c8a72b42eee684b2b629afd446ea046
MD5 1b8f202cf8754fa3ddd86135c3b02663
BLAKE2b-256 d1ccd1e21bb760dd863add7e1bfcf371ee12c2ee6aceb9ce5666bfecd4f46982

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.9.0 This release

2 files

0.8.0

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page