dagster-dataframely
Dataframely describes what a Polars frame should look like.
Dagster has first-class places to show that: the Columns tab and asset checks.
dagster-dataframely wires the two together, so you describe a table once and Dagster shows it everywhere.
import dagster as dg
import dataframely as dy
import polars as pl
import dagster_dataframely as dd
class Orders(dy.Schema):
order_id = dy.String(primary_key=True)
amount = dy.Float64(nullable=False, min=0.0)
@dd.asset(Orders)
def orders(raw_orders: pl.DataFrame) -> pl.DataFrame:
return raw_orders.select("order_id", "amount")
That is the whole integration. From that one declaration you get:
- The catalog's Columns tab, filled in before the asset has ever run: dtypes, descriptions, nullability, uniqueness, the primary key stated once at table level, and every remaining constraint listed beside it.
- One asset check per Dataframely rule, each with its own pass/fail history.
- A blocking column-schema check that compares the frame's columns and dtypes against the schema, before a single row is filtered.
- Somewhere for the rows that do not fit, if you want it.
Add
quarantine=Trueand the rows that fail validation are written beside the table rather than failing the run, as long as something survives.
The decorated function is an ordinary Dagster asset body.
Upstream assets bind as parameters, you declare context if you want it, and you can return any of five things: a frame, or a dg.MaterializeResult carrying one, eager or lazy, or None.
@dg.asset is the mechanism underneath, and the vocabulary.
Anything @dg.asset lets you say about one asset, you can say here under the same name, bar six parameters the decorator owns or rules out.
A test asserts that in both directions, and USER_GUIDE.md lists the six.
Package philosophy
Schema on write. Validation happens when a table is written, never when it is read. The checks run before the IO manager sees the frame, so the rows that fail never reach the table.
That puts this package after ingestion, at bronze to silver to gold, where the data is already on your side and the question is whether it is fit to publish.
Land raw records with dlt, which has good reasons to be permissive, and declare a schema at the first table someone else would query.
Ingestion-scale and larger-than-memory work belongs elsewhere.
The schema is the table's shape, and closing the gap is the asset's job.
A dtype that disagrees aborts the run rather than being coerced, because coercing quietly is how a wrong number reaches a table nobody re-reads.
There is no lenient mode to turn on.
Narrowing is free, though: Schema.filter drops the columns the schema never declared and returns the rest in the schema's order, so dtypes are the only thing ever yours to fix.
Consent to partial data is a declaration, not a setting. quarantine=True is the only dial, and no environment variable reaches it.
Leave it off and one failing row stops the write, so your last-known-good table stays in place.
The strictness belongs to the decorator, not to the package.
It is assembled from parts the package also exports under dd.wiring, and each one plugs a single feature into an asset the decorator does not fit: the Columns tab onto an asset that writes its own storage, or the checks onto a table something else already wrote.
Quick start
uv add dagster-dataframely
You will need Python 3.12 or newer.
dagster, dataframely and polars are the dependencies, plus universal-pathlib, which already arrives with dagster.
This package ships no IO manager, so bring one. dagster-polars writes Polars frames to a filesystem or object store, and dagster-duckdb-polars writes them to a warehouse.
Anything addressed by asset key works, because nothing here learns which manager you bound.
[!NOTE] Pre-1.0. The public surface is covered by a characterization test rather than held by convention, so it will not move quietly. It can still move: a
0.xminor release is where a breaking change lands. Pin to one minor if that matters to you:>=the version you installed,<the next minor. Coming from 0.6 or 0.7, readCHANGELOG.mdfirst.
Declare the schema and the asset as above, then tell the code location where to write:
from dagster_polars import PolarsParquetIOManager
defs = dg.Definitions(
assets=[raw_orders, orders],
resources={"io_manager": PolarsParquetIOManager(base_dir="data/warehouse")},
)
orders binds raw_orders, so whatever produces that goes in the list too.
Point dg dev at that module and materialize orders from the UI, or call dg.materialize([orders], resources=...) from a script.
Four things now exist that did not before:
- The catalog's Columns tab, filled from
Orders, before the first run. - One asset check per rule, each with its own history, evaluated on every run.
- A materialization carrying the row count, a row sample and per-dtype-group statistics.
- A table wherever your IO manager puts one.
To keep the rows that fail rather than failing the run, add quarantine=True:
@dd.asset(Orders, quarantine=True)
def orders(raw_orders: pl.DataFrame) -> pl.DataFrame:
return raw_orders.select("order_id", "amount")
The invalid rows go through the same IO manager, under the asset's own key with _quarantine on the end, carrying one column per rule saying why.
The checks then fail at WARN and the run succeeds, so downstream proceeds on the data that is fine.
Documentation
USER_GUIDE.mdis everything this package does: the failure policy, the quarantine, partitioning,LazyFrames, settings, testing, the errors, and hand-wiring.CONTEXT.mdis the glossary, and the rule that a word Dagster, Dataframely or Polars already owns wins.docs/adr/holds the decisions behind the behaviour.docs/research/holds the measurements behind them.CHANGELOG.mdis the upgrade log.
License
Apache 2.0.
See LICENSE.
Metadata
Release files for dagster-dataframely 0.8.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dagster_dataframely-0.8.0.tar.gz | 53.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dagster_dataframely-0.8.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 117.1 kB
Release files / dagster_dataframely-0.8.0.tar.gz
| Download URL | dagster_dataframely-0.8.0.tar.gz |
|---|---|
| Size | 53.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
da368b9a92be20a6e3ca8c6b00ffda7310c2b24f9c9713fc0d8da5253bdf58a6
|
|
BLAKE2b-256 checksum How to use checksums |
17ca9b114d1fdc7422271597ab772bdf5fec879b465a819f5136d4f5bce7f65f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.
Transparency logRelease files / dagster_dataframely-0.8.0-py3-none-any.whl
| Download URL | dagster_dataframely-0.8.0-py3-none-any.whl |
|---|---|
| Size | 64.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9dbe9872307839f104270720028cc882ff11fd302b87f682d6390ec4ef4bfb5c
|
|
BLAKE2b-256 checksum How to use checksums |
7434881da6ecd71e11bf4ff0c799a61e5d57bdbc326f44a61b9ce257d2c21f18
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.
Transparency log