Declarative DataFrame variable management with ML model integration
Project description
VarFrame
A declarative, class-based framework for defining, computing, and managing variables (columns) in pandas DataFrames with optional ML model integration.
Features
- Declarative Variable Definitions – Define columns as Python classes with metadata (dtype, description)
- Automatic Dependency Resolution – DAG-based ordering ensures derived columns compute in the correct order
- ML Model Integration – Train models and add predictions as DataFrame columns seamlessly
- Pandas Compatible –
VarFramebehaves like a regular DataFrame - Configurable Warnings – Control implicit operations and warnings globally
Installation
pip install varframe # Core only (pandas)
pip install varframe[ml] # + scikit-learn, joblib
pip install varframe[all] # Everything
Quick Start
1. Define Variables
from varframe import BaseVariable, DerivedVariable, VarFrame
# Map a raw column with type enforcement
class Lap(BaseVariable):
"""Current lap number."""
name = "lap"
raw_column = "lap_num"
dtype = "int"
class Gap(BaseVariable):
"""Gap to leader in seconds."""
name = "gap"
raw_column = "gap_to_leader"
dtype = "float"
# Create a computed column with dependencies
class GapDelta(DerivedVariable):
"""Change in gap from previous row."""
name = "gap_delta"
dependencies = [Gap]
@classmethod
def calculate(cls, df):
return df["gap"] - df["gap"].shift(1)
2. Create a VarFrame
import pandas as pd
# Raw data with original column names
df_raw = pd.DataFrame({
"lap_num": [1, 2, 3],
"gap_to_leader": [0.0, 1.2, 0.8]
})
# Create VarFrame - columns are computed automatically
vf = VarFrame(df_raw, [Lap, Gap, GapDelta])
print(vf)
# lap gap gap_delta
# 0 1 0.0 NaN
# 1 2 1.2 1.2
# 2 3 0.8 -0.4
3. Access Variables
# By name
vf["gap"]
# By class
vf[Gap]
# Multiple variables
vf[[Lap, Gap]]
# Filter by type
vf.filter_by_type(DerivedVariable) # Only computed columns
4. Add Variables Later
class GapPct(DerivedVariable):
"""Gap as percentage of total race time."""
name = "gap_pct"
dependencies = [Gap]
@classmethod
def calculate(cls, df):
return df["gap"] / df["gap"].max() * 100
vf.add_variables(GapPct)
ML Model Integration
Define models declaratively and use predictions as variables:
from varframe import BaseModel, ModelVariable
from sklearn.ensemble import RandomForestRegressor
class GapPredictor(BaseModel):
"""Predicts future gap based on features."""
name = "gap_predictor"
inputs = [Lap, Gap]
target = GapDelta
model_class = RandomForestRegressor
hyperparameters = {"n_estimators": 100, "max_depth": 5}
# Train the model
GapPredictor.train(training_vf)
# Use predictions as a variable
class PredictedGapDelta(ModelVariable):
name = "predicted_gap_delta"
model_class = GapPredictor
vf.add_variables(PredictedGapDelta)
Model Registry
Manage multiple models:
from varframe import ModelRegistry
registry = ModelRegistry()
registry.register(GapPredictor)
registry.train_all(training_vf)
registry.save_all("./models")
# Later
registry.load_all("./models")
Configuration
Control warnings and implicit operations:
from varframe import VFConfig
# Disable all warnings
VFConfig.warnings_enabled = False
# Block implicit model training (raises error instead)
VFConfig.allow_implicit_train = False
# Temporary suppression
with VFConfig.suppress_warnings():
vf.add_variables(SomeVariable)
# Reset to defaults
VFConfig.reset()
API Reference
Variable Classes
| Class | Purpose |
|---|---|
BaseVariable |
Maps a raw column (with optional dtype conversion) |
DerivedVariable |
Computed from other variables via calculate() |
ModelVariable |
Predictions from an ML model |
VarFrame Methods
| Method | Description |
|---|---|
add_variables(*vars, compute=True) |
Compute and add new variables (or register if compute=False) |
add_variable(*vars) |
Alias for add_variables(*vars) |
filter_by_type(type) |
Filter to BaseVariable or DerivedVariable only |
get_variable(name) |
Get variable class by name |
list_variables() |
List all variable names |
describe_variables() |
Summary DataFrame of all variables |
to_pandas() / to_ml() |
Convert to plain DataFrame for ML pipelines |
BaseModel Methods
| Method | Description |
|---|---|
train(vf) |
Train on a VarFrame |
predict(vf) |
Generate predictions |
evaluate(vf) |
Compute metrics |
save(path) / load(path) |
Persist and restore model |
Version
1.1.0
Author
Santiago Romagosa
License
MIT
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file varframe-1.1.0.tar.gz.
File metadata
- Download URL: varframe-1.1.0.tar.gz
- Upload date:
- Size: 21.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
bfe613c42fedaff8c7e6afebdad67753c73aefdab9d036e3b8e51c00bb3b60c7
|
|
| MD5 |
e999a30fa663a18ed44acb3f8e21be8d
|
|
| BLAKE2b-256 |
323c035f3e06a2e8f368e3f4f1f3c04375b725b6113ff7994629f5b57e88e8ad
|
Provenance
The following attestation bundles were made for varframe-1.1.0.tar.gz:
Publisher:
publish.yml on Santi-49/varframe
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
varframe-1.1.0.tar.gz -
Subject digest:
bfe613c42fedaff8c7e6afebdad67753c73aefdab9d036e3b8e51c00bb3b60c7 - Sigstore transparency entry: 812111077
- Sigstore integration time:
-
Permalink:
Santi-49/varframe@1996093bdd6dcd52cd1007c541209b5bf3eda84c -
Branch / Tag:
refs/tags/v1.1.1 - Owner: https://github.com/Santi-49
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@1996093bdd6dcd52cd1007c541209b5bf3eda84c -
Trigger Event:
release
-
Statement type:
File details
Details for the file varframe-1.1.0-py3-none-any.whl.
File metadata
- Download URL: varframe-1.1.0-py3-none-any.whl
- Upload date:
- Size: 22.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cd93037e7b2e3827d899301b08a95c4b8dca0ed52cc69228a5077faff50b2e17
|
|
| MD5 |
c526bebd9b7b1876520bbb10a4fcb1ab
|
|
| BLAKE2b-256 |
bbade2ec26797ce01a662cc76d55a16887d97e29ad59c55ff791a42a0c51326a
|
Provenance
The following attestation bundles were made for varframe-1.1.0-py3-none-any.whl:
Publisher:
publish.yml on Santi-49/varframe
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
varframe-1.1.0-py3-none-any.whl -
Subject digest:
cd93037e7b2e3827d899301b08a95c4b8dca0ed52cc69228a5077faff50b2e17 - Sigstore transparency entry: 812111111
- Sigstore integration time:
-
Permalink:
Santi-49/varframe@1996093bdd6dcd52cd1007c541209b5bf3eda84c -
Branch / Tag:
refs/tags/v1.1.1 - Owner: https://github.com/Santi-49
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@1996093bdd6dcd52cd1007c541209b5bf3eda84c -
Trigger Event:
release
-
Statement type: