Skip to main content

Eigrel

A programming language for Data, Machine Learning and AI.

CI PyPI License OpenSSF Scorecard

Write what you want. Let the compiler decide how to run it.

Eigrel is a declarative language for data engineering, machine learning and GenAI. You describe datasets, transformations, features and models in one language; the compiler decides whether each step becomes Python, SQL, Spark or something else.

dataset customers from csv("data/customers.csv")

transform customers {
    filter age >= 18
    select age, income, purchases, churned
}

features customers {
    age
    income
    purchases
}

model churn = random_forest {
    trees = 100
}

train churn {
    target = churned
}

evaluate churn {
    metrics = [accuracy, precision, recall, f1]
}

Status

Eigrel is at v0.5 — ML: programs read CSV, Parquet, JSON, SQL databases and BigQuery, clean missing values, train scikit-learn, XGBoost or Spark MLlib models and register them in MLflow. They are checked for meaning before anything runs and compiled to Python, Apache Spark or SQL, so the same program runs on your laptop or on a Spark cluster. See the roadmap.

Getting started

Install from PyPI (Python 3.13+). The python extra adds pandas and scikit-learn, which eigrel run needs; add sql to read databases or bigquery for BigQuery (pip install "eigrel[python,sql]"):

pip install "eigrel[python]"
eigrel init churn             # creates churn/main.eig and sample data
eigrel run churn/main.eig
churn: random_forest classification, trained on 315 rows, validated on 79
  accuracy   0.7975
  precision  0.7442
  recall     0.8649
  f1         0.8000

CLI

Command What it does
eigrel run FILE [-t python|spark] Compiles the program and runs it (default: Python)
eigrel check FILE... [--json] Reports syntax and semantic errors with line and column (--json: machine-readable)
eigrel compile FILE [-t python|spark|sql] [-o PATH] Prints (or writes) the generated code
eigrel ir FILE Prints the intermediate representation
eigrel ast FILE Prints the syntax tree as JSON
eigrel tokens FILE Prints the token stream
eigrel init NAME Creates a project with a starter program and sample data

The compiler catches mistakes before anything runs, and points at the exact spot:

error: column 'income' does not exist here; available columns: age, purchases, churned
 --> churn.eig:9:5
  |
9 |     income
  |     ^

How it works

source → lexer → parser → AST → semantic analysis → IR → Python backend → pandas + scikit-learn

eigrel ir shows the graph the backends work from:

%0 = load csv("data/customers.csv")  # customers
%1 = filter %0 (age >= 18)  # customers
%2 = select %1 [age, income, purchases, churned]  # customers
%3 = train %2 random_forest(trees=100, max_depth=8) classification features=[age, income, purchases] target=churned validation=0.2 seed=42  # churn
%4 = evaluate %3 [accuracy, precision, recall, f1]  # churn

Data sources

dataset customers from csv("data/customers.csv")
dataset events    from json("data/events.jsonl")
dataset orders    from sql(env("DATABASE_URL"), "shop.orders")
dataset users     from bigquery("my-project.analytics.users")

eigrel compile --target sql turns the data part of a program into a query for the engine each source lives in:

-- dataset users (bigquery, bigquery dialect)
SELECT `age`, `income`, `country` FROM `project.dataset.users` WHERE ((`age` > 18) AND (`income` <> 0));

Cleaning data, XGBoost and MLflow

transform customers {
    fill income = 0
    drop_missing age, purchases
}

model churn = xgboost {
    trees = 200
    learning_rate = 0.05
}

register churn {
    name = "customer-churn"
}

register logs the parameters, metrics and model to MLflow and registers a new version; the registered model takes raw rows, because encoding is part of its pipeline. It needs the xgboost and mlflow extras (pip install "eigrel[python,xgboost,mlflow]"; on macOS XGBoost also needs OpenMP: brew install libomp). See examples/mlflow.eig.

Spark

The same program runs on Spark with --target spark, which needs the spark extra and Java 17 or newer:

pip install "eigrel[python,spark]"
eigrel run --target spark churn/main.eig
eigrel compile --target spark churn/main.eig -o churn_spark.py   # e.g. for spark-submit

Files are read natively, sql() through JDBC and bigquery() through the spark-bigquery connector, and models train with Spark MLlib. See the Spark backend for how it differs from the Python backend.

Docker

The image on GitHub Container Registry (linux/amd64 and linux/arm64) includes the python and sql extras. Mount your project at /work and run as your own user, so the container can read your files and anything it writes stays yours:

alias eigrel='docker run --rm --user "$(id -u):$(id -g)" -v "$PWD:/work" ghcr.io/thentsation/eigrel'
eigrel init churn
eigrel run churn/main.eig

Paths must be inside the current directory, since only it is mounted. To build the image locally, use make docker-build and make docker-run.

AI agents

Eigrel is designed to be written by AI agents and verified by the compiler. Point your agent at llms.txt — the complete grammar in one file — and have it loop eigrel check --json FILE.eig until "ok": true; every error comes back with an exact line and column, so the fix is mechanical. What passes check is guaranteed to compile, and the same program runs on a laptop or a Spark cluster.

eigrel check --json churn.eig

Language

The full syntax is in docs/LANGUAGE.md. Examples live in examples/.

Contributing

Bug reports, language proposals and pull requests are welcome. Read the contributing guide to get set up, and note that this project follows a Code of Conduct. Security issues go through SECURITY.md.

License

Licensed under the Apache License 2.0.

Metadata

Release files for eigrel 0.6.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for eigrel 0.6.1
File Size Uploaded
eigrel-0.6.1.tar.gz 57.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for eigrel 0.6.1
File Interpreter ABI Platform
eigrel-0.6.1-py3-none-any.whl Python 3 none any Details

Total release size: 102.2 kB

Release files / eigrel-0.6.1.tar.gz

Download URL eigrel-0.6.1.tar.gz
Size 57.6 kB
Tags Source
SHA-256 checksum
How to use checksums
c1af0eae6917c073a8171680f030b2feb6f47dccc9d9d7de10d391eb926bce59
BLAKE2b-256 checksum
How to use checksums
c0ecd6424900e79c99754354cbef1853c3670d795ce55936b527069089538f90
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release files / eigrel-0.6.1-py3-none-any.whl

Download URL eigrel-0.6.1-py3-none-any.whl
Size 44.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4fcd6fb012ae0384961f904a71eba6e5d0e76323b2c6cf82dfcf5173004ac3d4
BLAKE2b-256 checksum
How to use checksums
58c0591afbc499d74e908d2dd065e74cb1e39a6fc4d7de91e5395cc4e64762a8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release history Release notifications | RSS feed

1.0.0

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.0

2 release files

This release

0.6.1 This release

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page