Skip to main content

A statically type-safe DataFrame abstraction layer

Project description

Colnade

A statically type-safe DataFrame abstraction layer for Python.

Colnade replaces string-based column references (pl.col("age")) with typed descriptors (Users.age), so column misspellings, type mismatches, and schema violations are caught by your type checker — before your code runs.

Works with ty, mypy, and pyright. No plugins, no code generation.

Installation

pip install colnade colnade-polars

Colnade requires Python 3.10+. Install the backend adapter for your engine:

Backend Install
Polars pip install colnade-polars
Pandas pip install colnade-pandas
Dask pip install colnade-dask

Quick Start

1. Define a schema

from colnade import Column, Schema, UInt64, Float64, Utf8

class Users(Schema):
    id: Column[UInt64]
    name: Column[Utf8]
    age: Column[UInt64]
    score: Column[Float64]

2. Create or read typed data

from colnade_polars import from_rows, read_parquet

# From Python data — schema drives dtype coercion
df = from_rows(Users, [
    Users.Row(id=1, name="Alice", age=30, score=85.0),
    Users.Row(id=2, name="Bob", age=25, score=92.5),
])

# Or from files
df = read_parquet("users.parquet", Users)
# df is DataFrame[Users] — the type checker knows the schema

3. Transform with full type safety

# Column references are attributes, not strings
result = (
    df.filter(Users.age > 25)
      .sort(Users.score.desc())
      .select(Users.name, Users.score)
)

4. Bind to an output schema

class UserSummary(Schema):
    name: Column[Utf8]
    score: Column[Float64]

output = result.cast_schema(UserSummary)
# output is DataFrame[UserSummary]

Safety Model

Colnade catches errors at three levels:

  1. In your editor — misspelled columns, schema mismatches at function boundaries, and nullability violations are flagged by your type checker (ty, pyright, mypy) before code runs
  2. At data boundaries — runtime validation ensures files and external data match your schemas (columns, types, nullability) and that expressions reference columns from the correct schema
  3. On your data valuesField() constraints validate domain invariants like ranges, patterns, and uniqueness

Where static safety ends: Static checking covers column references and schema-preserving operations (filter, sort, with_columns). Schema-transforming operations (select, group_by) return DataFrame[Any]cast_schema() re-binds to a named schema and is a runtime trust boundary. No type checker plugin is needed, but this is a deliberate tradeoff: plugins could theoretically infer output schemas at the cost of type-checker coupling and maintenance burden. See Type Checker Integration for the full list of what is and isn't checked.

Key Features

Type-safe column references

Column references are class attributes verified by the type checker at lint time:

Users.name   # Column[Utf8] — valid
Users.naem   # ty error: Class `Users` has no attribute `naem`

Schema-preserving operations

Operations that don't change the schema (filter, sort, limit, with_columns) preserve the type parameter:

def process(df: DataFrame[Users]) -> DataFrame[Users]:
    return df.filter(Users.age > 25).sort(Users.score.desc())

Typed expressions

Column descriptors build an expression tree with typed operators:

Users.age > 18          # Expr[Bool] — comparison
Users.score * 2         # Expr[Float64] — arithmetic
(Users.age > 18) & (Users.score > 80)  # Expr[Bool] — logical
Users.name.str_starts_with("A")        # Expr[Bool] — string method

Aggregations

result = df.group_by(Users.name).agg(
    Users.score.mean().alias(UserStats.avg_score),
    Users.id.count().alias(UserStats.user_count),
)

Null handling

# Fill nulls, filter nulls, check nulls
df.with_columns(Users.score.fill_null(0.0).alias(Users.score))
df.filter(Users.score.is_not_null())
df.drop_nulls(Users.score)

Joins with typed output

joined = users.join(orders, on=Users.id == Orders.user_id)
# JoinedDataFrame[Users, Orders] — both schemas accessible

class UserOrders(Schema):
    user_name: Column[Utf8] = mapped_from(Users.name)
    amount: Column[Float64]

result = joined.cast_schema(UserOrders)

Schema-polymorphic utility functions

Write generic functions that work with any schema:

from colnade.schema import S

def first_n(df: DataFrame[S], n: int) -> DataFrame[S]:
    return df.head(n)

# Works with any schema — type preserved
users_subset: DataFrame[Users] = first_n(users_df, 10)

Struct and List support

class Address(Schema):
    city: Column[Utf8]
    zip_code: Column[Utf8]

class UserProfile(Schema):
    name: Column[Utf8]
    address: Column[Struct[Address]]
    tags: Column[List[Utf8]]

# Access nested data
df.filter(UserProfile.address.field(Address.city) == "New York")
df.with_columns(UserProfile.tags.list.len().alias(tag_count_col))

Value-level constraints

from colnade import Column, Schema, UInt64, Utf8, Float64, ValidationLevel
from colnade.constraints import Field, schema_check

class Users(Schema):
    id: Column[UInt64] = Field(unique=True)
    age: Column[UInt64] = Field(ge=0, le=150)
    email: Column[Utf8] = Field(pattern=r"^[^@]+@[^@]+\.[^@]+$")
    status: Column[Utf8] = Field(isin=["active", "inactive"])

    @schema_check
    def adult(cls):
        return Users.age >= 18

# Validate with df.validate() or auto-validate at the FULL level
colnade.set_validation(ValidationLevel.FULL)

Lazy execution

from colnade_polars import scan_parquet

lazy = scan_parquet("users.parquet", Users)
# LazyFrame[Users] — builds a query plan

result = lazy.filter(Users.age > 25).sort(Users.score.desc()).collect()
# Executes the optimized query plan

Untyped escape hatch

When you need to drop down to untyped operations:

untyped = df.untyped()  # UntypedDataFrame — string-based columns
retyped = untyped.to_typed(Users)  # Back to DataFrame[Users]

Performance

Colnade adds < 5% overhead for typical Polars operations and < 5% for single Pandas operations (10–25% for multi-step Pandas pipelines at large sizes). Dask overhead is a fixed ~200–300 us per operation on graph construction, negligible compared to compute time. Validation (STRUCTURAL, FULL) adds measurable cost at data boundaries — see the full benchmark results for details.

Type Checker Error Showcase

Colnade catches real errors at lint time. Here are actual error messages from ty:

Misspelled column name

x = Users.agee
error[unresolved-attribute]: Class `Users` has no attribute `agee`

Schema mismatch at function boundary

df: DataFrame[Users] = read_parquet("users.parquet", Users)
wrong: DataFrame[Orders] = df
error[invalid-assignment]: Object of type `DataFrame[Users]` is not assignable
to `DataFrame[Orders]`

Nullability mismatch in mapped_from

class Bad(Schema):
    age: Column[UInt8] = mapped_from(Users.age)  # Users.age is Column[UInt8 | None]
error[invalid-assignment]: Object of type `Column[UInt8 | None]` is not
assignable to `Column[UInt8]`

Comparison with Existing Solutions

Feature Colnade Pandera StaticFrame Patito Narwhals
Column refs checked statically Named attrs No Positional types No No
Schema preserved through ops Through ops¹ At boundaries² No No No
Works with existing engines Polars, Pandas, Dask Pandas, Polars, others Own engine Polars only Many engines
No plugins or code gen Yes Requires mypy plugin Yes Yes Yes
Generic utility functions Yes No No No No
Struct/List typed access Yes No No No No
Lazy execution support Yes No No No Yes
Value-level constraints Field() Check No Pydantic validators No

¹ Schema-preserving ops (filter, sort, with_columns) retain DataFrame[S]. Schema-transforming ops (select, group_by) return DataFrame[Any] — use cast_schema() to bind. ² Pandera's @check_types validates schemas at function boundaries via decorator, but column references within function bodies remain unchecked strings.

See Detailed Comparisons for a fuller discussion of tradeoffs.

Documentation

Full documentation is available at colnade.com, including:

Examples

Runnable examples are in the examples/ directory:

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

colnade-0.5.0.tar.gz (257.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

colnade-0.5.0-py3-none-any.whl (32.1 kB view details)

Uploaded Python 3

File details

Details for the file colnade-0.5.0.tar.gz.

File metadata

  • Download URL: colnade-0.5.0.tar.gz
  • Upload date:
  • Size: 257.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for colnade-0.5.0.tar.gz
Algorithm Hash digest
SHA256 812ce5f50db904f01b5f71d5c9f582b314ed961fc8aefd523afb8f4ad550d14b
MD5 b8fefa98ae5f387b5b8b8f00714a0adc
BLAKE2b-256 3d1bbf6da2e100d2c2a95b59438c13e7b1f050d0020b9f654d2ba25beccb7810

See more details on using hashes here.

Provenance

The following attestation bundles were made for colnade-0.5.0.tar.gz:

Publisher: publish.yml on jwde/colnade

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file colnade-0.5.0-py3-none-any.whl.

File metadata

  • Download URL: colnade-0.5.0-py3-none-any.whl
  • Upload date:
  • Size: 32.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for colnade-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 7232ff785595e027c17a180f74a6c5c5eceb3aa11530823a3304fcf58bcd459a
MD5 c49ee0f95a6b6612d4a83211ea2e96ab
BLAKE2b-256 63fd64eda160fa4ea70e0948b3faeb63bfc5707f8aacfbac77abb1e819a074ab

See more details on using hashes here.

Provenance

The following attestation bundles were made for colnade-0.5.0-py3-none-any.whl:

Publisher: publish.yml on jwde/colnade

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page