Declarative DataFrame variable management with automatic DAG dependency resolution and ML model integration
Project description
VarFrame
Declarative DataFrame variable management with automatic DAG dependency resolution and ML model integration.
What is VarFrame?
VarFrame is a library that allows you to define DataFrame columns as Python classes rather than imperative scripts. It manages dependencies, types, and execution order automatically using a generic DAG (Directed Acyclic Graph) solver.
It is designed for complex, production-grade data pipelines where traceability, correctness, and structure are more important than raw implementation speed.
graph TD
Raw[Raw DataFrame] -->|extracts| Base[BaseVariable]
Base -->|inputs| Derived[DerivedVariable]
Base -->|features| Model[ML Model]
Derived -->|features| Model
Model -->|predicts| Pred[ModelVariable]
Pred -->|inputs| Ensemble[Ensemble Model]
Ensemble -->|predicts| Final[Final Prediction]
style Raw fill:#e1f5fe,stroke:#01579b,stroke-width:2px
style Base fill:#fff9c4,stroke:#fbc02d,stroke-width:2px
style Derived fill:#e0f2f1,stroke:#00695c,stroke-width:2px
style Model fill:#f3e5f5,stroke:#7b1fa2,stroke-width:2px
style Pred fill:#fce4ec,stroke:#c2185b,stroke-width:2px
style Ensemble fill:#f3e5f5,stroke:#7b1fa2,stroke-width:2px
Why VarFrame?
The Problem with Traditional Scripts
In traditional pandas scripts (df['b'] = df['a'] + 1), logic is often:
- Fragile: Reordering cells or lines breaks dependencies silently.
- Opaque: It's hard to tell exactly which columns are needed effectively.
- Hard to Test: You have to test "intermediate states" of a large dataframe.
The VarFrame Solution
VarFrame treats variables as definitions (Classes) rather than steps.
| Feature | VarFrame | Traditional Script |
|---|---|---|
| Dependency Resolution | Automatic (DAG). Order doesn't matter; the framework solves it. | Manual. You must order operations correctly yourself. |
| Logic Encapsulation | Logic, metadata, and types live in one Class. Self-documenting. | Distributed across scripts. logic often mixed with execution. |
| ML Integration | Models are just "Computed Variables". Predictions are treated like any other column. | Often separate "training" and "inference" pipelines. |
| Testing | Unit test single calculate(df) methods in isolation. |
Integration testing entire scripts is required. |
Best For
- Feature Stores: Reuse definitions across training and serving.
- Complex DAGs: When variable F depends on E, which depends on D, C, and B...
- Ensemble/Stacking: Where model predictions feed into other models (see
examples/ensemble_demo.py).
Installation
pip install varframe # Core only (pandas)
pip install varframe[ml] # + scikit-learn, joblib
pip install varframe[all] # Everything
Quick Start
1. Define Variables
from varframe import BaseVariable, DerivedVariable, VarFrame
# Map a raw column with type enforcement
class Lap(BaseVariable):
"""Current lap number."""
name = "lap"
raw_column = "lap_num"
dtype = "int"
class Gap(BaseVariable):
"""Gap to leader in seconds."""
name = "gap"
raw_column = "gap_to_leader"
dtype = "float"
# Create a computed column with dependencies
class GapDelta(DerivedVariable):
"""Change in gap from previous row."""
name = "gap_delta"
dependencies = [Gap]
@classmethod
def calculate(cls, df):
return df["gap"] - df["gap"].shift(1)
2. Create a VarFrame
import pandas as pd
# Raw data with original column names
df_raw = pd.DataFrame({
"lap_num": [1, 2, 3],
"gap_to_leader": [0.0, 1.2, 0.8]
})
# Create VarFrame - columns are computed automatically
# Dependencies are resolvd automatically!
vf = VarFrame(df_raw, [Lap, Gap, GapDelta])
print(vf)
# lap gap gap_delta
# 0 1 0.0 NaN
# 1 2 1.2 1.2
# 2 3 0.8 -0.4
3. Access Variables
# By name
vf["gap"]
# By class
vf[Gap]
# Multiple variables
vf[[Lap, Gap]]
# Filter by type
vf.filter_by_type(DerivedVariable) # Only computed columns
ML Model Integration
Define models declaratively and use predictions as variables:
from varframe import BaseModel, ModelVariable
from sklearn.ensemble import RandomForestRegressor
class GapPredictor(BaseModel):
"""Predicts future gap based on features."""
name = "gap_predictor"
inputs = [Lap, Gap]
target = GapDelta
model_class = RandomForestRegressor
hyperparameters = {"n_estimators": 100, "max_depth": 5}
# Train the model
GapPredictor.train(training_vf)
# Use predictions as a variable
class PredictedGapDelta(ModelVariable):
name = "predicted_gap_delta"
model_class = GapPredictor
vf.add_variables(PredictedGapDelta)
Optimization & Export
Lazy Loading
Optimize memory by marking variables as lazy = True. They are computed on-demand and not stored in the DataFrame.
class HugeFeature(DerivedVariable):
lazy = True
dependencies = [RawData]
@classmethod
def calculate(cls, df):
return df["raw"] * 1000
Flexible Views
Export specific subsets of data using vf.view():
# Export only base variables
df_base = vf.view(include=["base"])
# Export specific variables (computes lazy vars on demand)
df_custom = vf.view(variables=[HugeFeature])
API Reference
Variable Classes
| Class | Purpose |
|---|---|
BaseVariable |
Maps a raw column (with optional dtype conversion) |
DerivedVariable |
Computed from other variables. Set lazy=True for on-demand computation. |
ModelVariable |
Predictions from an ML model |
VarFrame Methods
| Method | Description |
|---|---|
add_variables(*vars, compute=True) |
Compute and add new variables (or register if compute=False) |
add_variable(*vars) |
Alias for add_variables(*vars) |
filter_by_type(type) |
Filter to BaseVariable or DerivedVariable only |
get_variable(name) |
Get variable class by name |
view(include=..., variables=...) |
Export DataFrame with specific variables (handles lazy computation) |
list_variables() |
List all variable names |
describe_variables() |
Summary DataFrame of all variables |
to_pandas() / to_ml() |
Convert to plain DataFrame for ML pipelines |
BaseModel Methods
| Method | Description |
|---|---|
train(vf) |
Train on a VarFrame |
predict(vf) |
Generate predictions |
evaluate(vf) |
Compute metrics |
save(path) / load(path) |
Persist and restore model |
License
MIT
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file varframe-1.2.1.tar.gz.
File metadata
- Download URL: varframe-1.2.1.tar.gz
- Upload date:
- Size: 24.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d83e56d56fbea320226120e6abf90e3d24342997da35e620f3214aad6f6dd159
|
|
| MD5 |
d98c1f7ae2d35f55db5fefcfba20ab0b
|
|
| BLAKE2b-256 |
47a60e38b4c5bb678d57ae9f03fb858b8949470d8e651725777a0e4ed1cf1d26
|
Provenance
The following attestation bundles were made for varframe-1.2.1.tar.gz:
Publisher:
publish.yml on Santi-49/varframe
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
varframe-1.2.1.tar.gz -
Subject digest:
d83e56d56fbea320226120e6abf90e3d24342997da35e620f3214aad6f6dd159 - Sigstore transparency entry: 813675923
- Sigstore integration time:
-
Permalink:
Santi-49/varframe@9c927acfc9a0342a22320649baa60fd57e0b8366 -
Branch / Tag:
refs/tags/v1.2.1 - Owner: https://github.com/Santi-49
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@9c927acfc9a0342a22320649baa60fd57e0b8366 -
Trigger Event:
release
-
Statement type:
File details
Details for the file varframe-1.2.1-py3-none-any.whl.
File metadata
- Download URL: varframe-1.2.1-py3-none-any.whl
- Upload date:
- Size: 24.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7554ef59845903ca2b1931291ef9b68042cf284bc3f12c20aeb3a9195e8168f3
|
|
| MD5 |
cd430e07299558673b9b6bba0e05d1b0
|
|
| BLAKE2b-256 |
dce31287532067a1d1656b95399a08ba12829b19479e584f908ce737ecffaf4a
|
Provenance
The following attestation bundles were made for varframe-1.2.1-py3-none-any.whl:
Publisher:
publish.yml on Santi-49/varframe
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
varframe-1.2.1-py3-none-any.whl -
Subject digest:
7554ef59845903ca2b1931291ef9b68042cf284bc3f12c20aeb3a9195e8168f3 - Sigstore transparency entry: 813675924
- Sigstore integration time:
-
Permalink:
Santi-49/varframe@9c927acfc9a0342a22320649baa60fd57e0b8366 -
Branch / Tag:
refs/tags/v1.2.1 - Owner: https://github.com/Santi-49
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@9c927acfc9a0342a22320649baa60fd57e0b8366 -
Trigger Event:
release
-
Statement type: