Skip to main content

Gamma Lake

High-performance feature store built on Delta Lake and Ray

Build Status codecov License PyPI

Overview

Gamma Lake is a feature store built on Delta Lake and Ray, designed for efficient storage, versioning, and retrieval of time-indexed features backed by Polars DataFrames.

It was built to solve the real-world pain points that teams encounter when using flat Parquet files for ML feature storage: expensive column additions, no versioning, and painful cross-team collaboration.

Why Gamma Lake?

The most common alternative — storing features in per-day Parquet files — breaks down quickly:

Problem Gamma Lake's answer
Adding a new feature group rewrites all existing files Gamma Lake writes one new Delta table per feature group — O(1) cost regardless of existing data
Multiple teams can't write features in parallel Each feature group is an independent Delta table — concurrent writes don't conflict
No versioning or audit trail Delta Lake's transaction log provides full time-travel and version history
Cross-group reads require expensive joins Gamma Lake maintains a master index; reads are aligned horizontal concatenations
Experimental features pollute production data Features carry owner and version metadata; read() accepts a filtered metadata frame

See performance for benchmark data comparing Gamma Lake against per-day Parquet at scale.

Key Features

  • Sortable index — a master index table ensures all feature groups stay aligned for efficient range-filtered reads
  • Parallel reads and writes — Ray remote functions parallelise across feature groups for large-scale workloads
  • Feature versioning — every add_features() call records owner and version; read any historical version by filtering the metadata frame
  • Multiple signal types — standard features, as-of features, sparse features, and runtime-computed features
  • Native Delta Lake storage — ACID transactions, schema enforcement, and time-travel out of the box
  • Local or cloudbase_path accepts a local directory, an on-prem object store path, or an s3:// URI
  • Configurable compressionzstd by default; pluggable via compression field

Quick Start

import tempfile
from datetime import datetime, timezone

import polars as pl
import pyarrow as pa
from ccflow import ArrowSchema

from gammalake import GammaFeatureLake

# 1. Create and initialise a feature lake
with tempfile.TemporaryDirectory() as tmp:
    lake = GammaFeatureLake(base_path=tmp, run_on_ray_cluster=False)
    lake.initialize(
        ArrowSchema.make(
            pa.schema([("timestamp", pa.timestamp("us", tz="UTC")), ("symbol", pa.large_string())])
        )
    )

    # 2. Build a toy feature DataFrame (timestamp × symbol × features)
    now = datetime(2024, 1, 1, tzinfo=timezone.utc)
    df = pl.DataFrame({
        "timestamp": pl.Series([now]).cast(pl.Datetime("us", "UTC")),
        "symbol":    ["AAPL"],
        "momentum":  [0.42],
        "vol_20d":   [0.18],
    })

    # 3. Write features
    lake.add_features(df, owner="my-team")

    # 4. Read features back
    result = lake.read(["momentum", "vol_20d"])
    print(result)

    # 5. Inspect metadata
    print(lake.feature_metadata_frame().collect())

Run the full annotated demo: examples/quickstart.py

Core API

from gammalake import GammaFeatureLake

lake = GammaFeatureLake(base_path="s3://my-bucket/features")

# One-time setup
lake.initialize(schema)

# Writing
lake.add_features(df, owner="team-a")          # standard features
lake.add_targets(df, owner="team-a")           # target/label columns
lake.add_as_of_features(df, params, owner=...) # point-in-time safe features
lake.add_sparse_features(df, owner=...)        # sparse / infrequently-updated features
lake.add_runtime_computed_features(exprs, ...) # features computed at read time

# Reading
lake.read(["feature_a", "feature_b"])                        # latest versions
lake.read(["feature_a"], start="2023-01-01", end="2024-01-01")  # date range
lake.read(lake.feature_metadata_frame()                      # specific owner/version
          .filter(pl.col("owner") == "team-a").collect())

# Metadata inspection
lake.feature_metadata_frame().collect()   # features, versions, owners
lake.table_metadata_frame().collect()     # Delta table details
lake.index_frame().collect()              # master index

Installation

gamma-lake can be installed via pip or conda, the two primary package managers for the Python ecosystem.

To install gamma-lake via pip, run this command in your terminal:

pip install gamma-lake

To install gamma-lake via conda, run this command in your terminal:

conda install gamma-lake -c conda-forge

Getting Started

See our wiki!

Development

Check out the contribution guide for more information.

License

Apache-2.0 — see LICENSE for details.

Release files for gamma-lake 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gamma-lake 0.1.0
File Size Uploaded
gamma_lake-0.1.0.tar.gz 52.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gamma-lake 0.1.0
File Interpreter ABI Platform
gamma_lake-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 109.0 kB

Release files / gamma_lake-0.1.0.tar.gz

Download URL gamma_lake-0.1.0.tar.gz
Size 52.7 kB
Tags Source
SHA-256 checksum
How to use checksums
78960e52d77ae59e7b0c9559158915fa6722327eddde457b3ebde8e12b2f1ffa
BLAKE2b-256 checksum
How to use checksums
4784f3d6e597da900e757ed028ebef813f656447778222a46e36547bb24674fe
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release files / gamma_lake-0.1.0-py3-none-any.whl

Download URL gamma_lake-0.1.0-py3-none-any.whl
Size 56.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2897bab74c2584d61cc7e90c07f6aab0c988c2d2bd4e11a019a14b0179f06e95
BLAKE2b-256 checksum
How to use checksums
68e9df4d6af1723f5764dac02df1237045220bf760324efb2803b4df6d1bcd19
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release history Release notifications | RSS feed

0.2.0

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page