Eigrel
A compiler for tabular machine learning. You declare the pipeline; the compiler checks it before anything runs and generates the code for your laptop (pandas + scikit-learn) or your cluster (Spark).
You don't review ML code; you review a plan. The compiler proves what it can, the backend runs it.
Eigrel describes datasets, cleaning, features, models, training, evaluation and registration in one small language. Because the compiler sees the whole program, it rejects mistakes that a notebook runs without complaint, and the same program compiles to Python, Spark or SQL.
dataset customers from csv("data/customers.csv")
transform customers {
filter age >= 18
select age, income, purchases, churned
}
features customers {
age
income
purchases
}
model churn = random_forest {
trees = 100
}
train churn {
target = churned
}
evaluate churn {
metrics = [accuracy, precision, recall, f1]
}
The mistake a notebook misses
Put the target among the features by accident. In a notebook, nothing complains, and validation looks perfect:
X = customers[['age', 'income', 'purchases', 'churned']] # churned is also the target
X_train, X_test, y_train, y_test = train_test_split(X, customers['churned'], stratify=customers['churned'], random_state=42)
RandomForestClassifier(random_state=42).fit(X_train, y_train).score(X_test, y_test) # 1.0
The same pipeline in Eigrel (examples/mistakes/target_leakage.eig)
does not compile:
$ eigrel check examples/mistakes/target_leakage.eig
error: target 'churned' is also declared as a feature of 'customers'
--> examples/mistakes/target_leakage.eig:26:14
|
26 | target = churned
| ^
The honest model scores 0.80 on this data. Good engineers leak data silently; a compiler that sees the whole pipeline does not. The generated code is held to the same standard: encoders are fitted on the training split only, inside the model, so the validation split never leaks into them.
Guarantees across the model's life
- No leakage in the pipeline. Encoders are fitted on the training split only, inside the model; the target can never be a feature.
- Time-ordered validation.
assumptions customers { time = signup_date }validates on the latest rows;eigrel planwarns when dated data is split at random. - No train/serve skew.
predict churn { data = new_customers, output = csv("scored.csv") }must receive every training feature with the same types, and gets the training fills replayed; anything else is a compile error. Registered models predict the original classes. - Clear backend limits. A capability matrix says what each backend supports, and the plan checks programs against it.
See examples/serving.eig.
See the consequences before running: eigrel plan
eigrel plan reads the data, checks the program against the real column names and types, and
shows what every step does to it, without training anything:
$ eigrel plan examples/ml.eig
Plan for ml.eig (python backend)
dataset customers ← csv("data/customers.csv")
columns: customer_id bigint, age bigint, income double, purchases bigint, churned bigint
rows: 400
filter (age >= 18) 400 → 394 rows (-6)
select age, income, purchases, churned 394 rows
train churn: random_forest classification on customers
split: 315 train / 79 validation (20%, stratified)
features: age, income, purchases
target churned: 0 209 (53%), 1 185 (47%)
missing values: none
✓ no problems found
state: no ml.eigstate yet; `eigrel plan --save` records the data schema so later plans can detect drift
When something would go wrong, the plan says so and points at the line:
✗ error: line 9: features with missing values cannot be trained with logistic_regression: revenue (22); use fill or drop_missing first
! warning: line 9: class imbalance: 'enterprise' has 4 rows (1%) against 424 for 'basic'
! warning: line 9: less than one row(s) of class 'enterprise' expected in the validation split (20%); its metrics will be unstable
eigrel plan --save records each source's schema in PROGRAM.eigstate; commit it, and later
plans report drift (removed columns, changed types, row counts). --json gives the same plan to
tools and agents, --target spark applies Spark's limits, and --strict turns warnings into a
failing exit code for CI.
Status
Eigrel 0.8 reads CSV, Parquet, JSON, SQL databases and BigQuery, cleans missing values, trains
scikit-learn, XGBoost or Spark MLlib models, scores new data with the training pipeline and
registers models in MLflow, compiling the same program to Python, Spark or SQL. eigrel plan
shows the consequences before anything runs. Next: agents (an MCP server). See the
roadmap.
Getting started
Install from PyPI (Python 3.13+). The python extra adds pandas and scikit-learn, which
eigrel run needs; add sql to read databases or bigquery for BigQuery
(pip install "eigrel[python,sql]"):
pip install "eigrel[python]"
eigrel init churn # creates churn/main.eig and sample data
eigrel run churn/main.eig
churn: random_forest classification, trained on 315 rows, validated on 79
accuracy 0.7975
precision 0.7442
recall 0.8649
f1 0.8000
CLI
| Command | What it does |
|---|---|
eigrel run FILE [-t python|spark] |
Compiles the program and runs it (default: Python) |
eigrel check FILE... [--json] |
Reports syntax and semantic errors with line and column (--json: machine-readable) |
eigrel plan FILE [-t python|spark] [--json] [--save] [--strict] |
Checks the program against its data and shows row counts, classes, warnings and drift |
eigrel compile FILE [-t python|spark|sql] [-o PATH] |
Prints (or writes) the generated code |
eigrel ir FILE |
Prints the intermediate representation |
eigrel ast FILE |
Prints the syntax tree as JSON |
eigrel tokens FILE |
Prints the token stream |
eigrel init NAME |
Creates a project with a starter program and sample data |
The compiler catches mistakes before anything runs, and points at the exact spot:
error: column 'income' does not exist here; available columns: age, purchases, churned
--> churn.eig:9:5
|
9 | income
| ^
How it works
source → lexer → parser → AST → semantic analysis → IR → Python backend → pandas + scikit-learn
eigrel ir shows the graph the backends work from:
%0 = load csv("data/customers.csv") # customers
%1 = filter %0 (age >= 18) # customers
%2 = select %1 [age, income, purchases, churned] # customers
%3 = train %2 random_forest(trees=100, max_depth=8) classification features=[age, income, purchases] target=churned validation=0.2 seed=42 # churn
%4 = evaluate %3 [accuracy, precision, recall, f1] # churn
Data sources
dataset customers from csv("data/customers.csv")
dataset events from json("data/events.jsonl")
dataset orders from sql(env("DATABASE_URL"), "shop.orders")
dataset users from bigquery("my-project.analytics.users")
eigrel compile --target sql turns the data part of a program into a query for the engine each
source lives in:
-- dataset users (bigquery, bigquery dialect)
SELECT `age`, `income`, `country` FROM `project.dataset.users` WHERE ((`age` > 18) AND (`income` <> 0));
Cleaning data, XGBoost and MLflow
transform customers {
fill income = 0
drop_missing age, purchases
}
model churn = xgboost {
trees = 200
learning_rate = 0.05
}
register churn {
name = "customer-churn"
}
register logs the parameters, metrics and model to MLflow and registers a new version; the
registered model takes raw rows, because encoding is part of its pipeline. It needs the xgboost
and mlflow extras (pip install "eigrel[python,xgboost,mlflow]"; on macOS XGBoost also needs
OpenMP: brew install libomp). See
examples/mlflow.eig.
Spark
The same program runs on Spark with --target spark, which needs the spark extra and Java 17 or
newer:
pip install "eigrel[python,spark]"
eigrel run --target spark churn/main.eig
eigrel compile --target spark churn/main.eig -o churn_spark.py # e.g. for spark-submit
Files are read natively, sql() through JDBC and bigquery() through the spark-bigquery connector,
and models train with Spark MLlib. See the Spark backend for how
it differs from the Python backend.
Docker
The image on GitHub Container Registry (linux/amd64 and linux/arm64) includes the python and
sql extras. Mount your project at /work and run as your own user, so the container can read your
files and anything it writes stays yours:
alias eigrel='docker run --rm --user "$(id -u):$(id -g)" -v "$PWD:/work" ghcr.io/thentsation/eigrel'
eigrel init churn
eigrel run churn/main.eig
Paths must be inside the current directory, since only it is mounted. To build the image locally,
use make docker-build and make docker-run.
AI agents
Eigrel is designed to be written by AI agents and verified by the compiler. Point your agent at
llms.txt — the complete grammar in one file — and have it loop
eigrel check --json FILE.eig until "ok": true; every error comes back with an exact line and
column, so the fix is mechanical. What passes check is guaranteed to compile, and the same
program runs on a laptop or a Spark cluster.
eigrel check --json churn.eig
Language
The full syntax is in docs/LANGUAGE.md. Examples live in examples/.
Contributing
Bug reports, language proposals and pull requests are welcome. Read the contributing guide to get set up, and note that this project follows a Code of Conduct. Security issues go through SECURITY.md.
License
Licensed under the Apache License 2.0.
Metadata
Release files for eigrel 0.8.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| eigrel-0.8.0.tar.gz | 82.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| eigrel-0.8.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 144.2 kB
Release files / eigrel-0.8.0.tar.gz
| Download URL | eigrel-0.8.0.tar.gz |
|---|---|
| Size | 82.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8cb5bd165e8df934f510c2ea1323b7d234feb92a8fcfcc39e6ddda2095930feb
|
|
BLAKE2b-256 checksum How to use checksums |
cdbf10109566ba666c73d0c5f188187e088ef7c0aea61d04e9f68add0e4acc5b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency logRelease files / eigrel-0.8.0-py3-none-any.whl
| Download URL | eigrel-0.8.0-py3-none-any.whl |
|---|---|
| Size | 61.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7a4d9ffda5843ecd21740e362fbd82a6b14d4d148084de14f3f2ccf62f246992
|
|
BLAKE2b-256 checksum How to use checksums |
a2669eaa70eee8eddb6413f48ee0ea50ef6b147e94edb78045330d3ab9ca0022
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency log