survey_kit_formula
A Python implementation of R's formula / model.matrix() syntax. Give it a
formula string and a dataset, and get back a design matrix, using the same
term-expansion, contrast-coding, and marginality rules as R's
terms()/model.matrix().
Quickstart
import polars as pl
from survey_kit_formula import model_matrix, model_frame
df = pl.DataFrame({
"y": [1.0, 2.0, 3.0, 4.0],
"x1": [1.0, 2.0, 3.0, 4.0],
"x2": ["a", "b", "a", "b"],
})
model_matrix("y ~ x1 + x2", df)
# -> numpy.ndarray, dense float64, shape (4, 3): (Intercept), x1, x2b
model_frame("y ~ x1 + x2", df)
# -> a dataframe with the same columns, each packed to a compact
# dtype (Boolean for dummy columns, smallest int for
# whole-number columns, Float64 otherwise)
Use model_matrix if you want a NumPy array (e.g. to hand to a solver
that expects one). Use model_frame if you don't — it has the exact same
columns and values, just as a dataframe instead of a dense float64 array.
data can be a Polars DataFrame or LazyFrame, a pandas DataFrame, a
PyArrow Table, or any other dataframe type supported by
narwhals — including engines
like DuckDB or Dask. model_frame returns a dataframe in the same format
you passed in — pandas in, pandas out; DuckDB in, DuckDB out; and so on.
Fit once, reapply to new data
model_matrix and model_frame figure out the formula's structure (factor
levels, contrasts, spline settings) from the data every time you call them.
If you want to fit that structure once and reuse it on other data — e.g.
train on one dataset, then transform test data the same way — build a
ModelSpec and reuse it instead:
from survey_kit_formula import ModelSpec
spec = ModelSpec.from_formula("y ~ x1 + poly(x2, degree=2)", train_df)
train_matrix = spec.get_model_matrix(train_df) # numpy.ndarray
test_matrix = spec.get_model_matrix(test_df) # same columns, using
# train_df's factor levels
# and spline settings
train_frame = spec.get_model_frame(train_df) # same format as train_df
test_frame = spec.get_model_frame(test_df) # same format as test_df
ModelSpec.from_formula accepts null_dummy and null_fill keywords:
by default, any null in a modeled column raises an error. Pass
null_dummy=True to instead fill nulls (default 0.0, override with
null_fill=) and add a companion 0/1 "was this null" indicator column for
any variable that had nulls when the spec was created.
Formula syntax
| Syntax | Meaning |
|---|---|
y ~ x1 + x2 |
response ~ predictors, separated by + |
y ~ x1 - x2 |
remove a term |
y ~ x - 1, y ~ x + 0, y ~ 0 + x |
drop the intercept |
y ~ a * b |
full factorial: a + b + a:b |
y ~ a:b |
interaction only (no main effects added) |
y ~ a / b |
nesting: a + b %in% a |
y ~ b %in% a |
nesting, same semantics as / |
y ~ (a + b + c)^2 |
all terms up to order 2 |
y ~ . |
all other columns in the data |
~ x1 + x2 |
one-sided formula (no response) |
Special functions recognized inside a formula:
| Function | Purpose |
|---|---|
factor(x) |
force categorical/dummy coding |
ordered(x, levels=[...], scores=[...]) |
ordered-factor coding (contr.poly by default) |
C(x, contr, base=...) |
override the contrast scheme for a variable |
poly(x, degree=n) |
orthogonal polynomial basis (also multivariate: poly(x1, x2, degree=n)) |
bs(x, df=, knots=, degree=, intercept=) |
B-spline basis |
ns(x, df=, knots=, intercept=) |
natural cubic spline basis |
offset(x) |
tracked separately, excluded from the design matrix |
I(expr) |
arithmetic escape hatch, e.g. I(log(x) + 1) |
Available contrast names for C(x, name): contr.treatment (default for
unordered factors), contr.sum, contr.helmert, contr.poly (default for
ordered factors), contr.SAS — or the bare shorthand (treatment, sum,
helmert, poly, SAS).
Development
uv sync
uv run pytest
Release files for survey-kit-formula 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| survey_kit_formula-0.1.0.tar.gz | 156.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| survey_kit_formula-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 213.1 kB
Release files / survey_kit_formula-0.1.0.tar.gz
| Download URL | survey_kit_formula-0.1.0.tar.gz |
|---|---|
| Size | 156.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
144969c570a16ef8d422fca9af9a70afe2885ad4110b0e72dd13926130c813e7
|
|
BLAKE2b-256 checksum How to use checksums |
51b474807a35bea688ecf0256479aba91146964ef2206f8b75487b6603db6f18
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 15, 2026.
Transparency logRelease files / survey_kit_formula-0.1.0-py3-none-any.whl
| Download URL | survey_kit_formula-0.1.0-py3-none-any.whl |
|---|---|
| Size | 56.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7b57508232bce993e9dd2a7767af2e9385be4d3fdfb2bf75167c8ed145938016
|
|
BLAKE2b-256 checksum How to use checksums |
b9ee9e4406b74dd3cc355c0d8609d14ecc6e19226f72658084012b84f7e8abdb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 15, 2026.
Transparency log