DataFolio
A small, self-documenting home for the data behind one analysis.
You already have the objects: a feature table, some labels, an array of embeddings, a fitted model, a parameter dict, a QC plot. DataFolio gives them one directory, one name each, and one sentence each explaining what they are — then hands them back by name months later, from any notebook.
from datafolio import DataFolio
folio = DataFolio("analysis/experiment-12")
folio.add("features", features, description="One row per neuron; normalized morphology")
folio.add("labels", labels, description="Manual labels after the March review")
folio.add("params", {"alpha": 0.1, "seed": 7})
folio.add_model("classifier", clf, inputs=["features", "labels"])
# A different notebook, six months later. One path, no filenames.
folio = DataFolio("analysis/experiment-12")
folio.describe()
features = folio.get("features")
clf = folio.get_model("classifier")
Install
pip install datafolio # core
pip install 'datafolio[polars]' # + lazy scans and Polars frames
Python 3.10+.
Who it's for
One researcher, or a small team sharing mostly read-only work, holding a few to a few dozen understandable objects per analysis. It assumes you are comfortable in pandas, Polars, numpy, and scikit-learn and do not want another framework between you and them.
Good fit when you want to stop writing
pd.read_parquet(BASE / "features_v3_reviewed.parquet") in every notebook,
remember which of five similar files the paper used, keep one analysis's
objects together, or hand a colleague a single path.
What you get
One write path, one read path. add() picks the storage format from the
object's type; get() reads the catalog and returns the matching Python object.
folio.add("table", dataframe) # pandas / Polars / LazyFrame -> Parquet
folio.add("counts", series) # pandas / Polars Series -> Parquet
folio.add("embeddings", array) # numpy -> .npy
folio.add("params", {"alpha": 0.1}) # dict / list / tuple / set / scalar -> JSON
folio.add_model("classifier", clf) # -> joblib (or skops)
folio.add_file("plots/qc.png") # -> the file, unchanged
folio.reference_table("raw", "gs://lab-data/raw.parquet") # link, don't copy
Descriptions live next to the data. The most valuable metadata is usually
one sentence saying which of several near-identical files this is. It is stored
in the catalog, shown by describe(), and readable by people who never install
the package.
Ordinary files, not a container format.
experiment-12/
├── items.json # the authoritative catalog (+ metadata + snapshots)
├── CONTENTS.md # derived, human-readable inventory
├── README.md # how to read this directory without datafolio
├── tables/features--r2.parquet
├── models/classifier--r8.joblib
└── artifacts/params--r5.json
Local or cloud, same code. DataFolio("gs://team-analysis/experiment-12").
Lineage, snapshots, and a CLI when you need them:
folio.add("predictions", preds, inputs=["classifier", "holdout"])
folio.create_snapshot("paper-v1", description="Figures 2–4 in the submission")
datafolio describe
datafolio snapshot status
What it is not
Not a database or query engine, not a workflow orchestrator, not data version control, not an experiment tracker, not a multi-writer system, not a backup. Many readers, one writer — treat a cloud folio as single-writer. Comfortable at a few dozen items per folio, not thousands.
DataFolio's rule is to be as lightweight as possible and hand off to better tools as soon as possible: big tables go to Polars as a lazy scan, or to you as a path. See What DataFolio is not for the full list, including the sharp edges worth knowing before you rely on it.
Documentation
Read in this order:
- Your first folio — ten minutes
- Everyday patterns
- Tables: big, external, and lazy
- Models
- Sharing a folio
- Snapshots
- Reading a folio without DataFolio
- What DataFolio is not
Reference: API cheat sheet · CLI · Migrating from 1.x
Development
uv sync
poe test # pytest with coverage
uv run ruff check src/ tests/
poe doc-preview # local docs server
License
MIT
Metadata
Release files for datafolio 2.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| datafolio-2.1.0.tar.gz | 108.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| datafolio-2.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 231.6 kB
Release files / datafolio-2.1.0.tar.gz
| Download URL | datafolio-2.1.0.tar.gz |
|---|---|
| Size | 108.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3458c5e7b8cd0b0a22442ad132dbdb9c2a38983868f48773e15a04f153ce8bc6
|
|
BLAKE2b-256 checksum How to use checksums |
bcd5cdf003656626e788eaa71d4c3cc4452e2888184252016f0cc126caba8c4c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.6 {"installer":{"name":"uv","version":"0.12.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / datafolio-2.1.0-py3-none-any.whl
| Download URL | datafolio-2.1.0-py3-none-any.whl |
|---|---|
| Size | 123.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9618b293129eef30c17b15f1f16d6d643a9e36794da572b296f08f7f7185cd4e
|
|
BLAKE2b-256 checksum How to use checksums |
d259cdbcb74945351679ad76213e06d1bcebd0b1007eb24a97e94870d2d7ae6c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.6 {"installer":{"name":"uv","version":"0.12.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|