DataFolio
A small, self-documenting home for the data behind one analysis.
You already have the objects: a feature table, some labels, an array of embeddings, a fitted model, a parameter dict, a QC plot. DataFolio gives them one directory, one name each, and one sentence each explaining what they are — then hands them back by name months later, from any notebook.
from datafolio import DataFolio
folio = DataFolio("analysis/experiment-12")
folio.add("features", features, description="One row per neuron; normalized morphology")
folio.add("labels", labels, description="Manual labels after the March review")
folio.add("params", {"alpha": 0.1, "seed": 7}, description="Settings used for the paper figures")
folio.add_model("classifier", clf, inputs=["features", "labels"],
description="Cell-type classifier on reviewed labels")
# A different notebook, six months later. One path, no filenames.
folio = DataFolio("analysis/experiment-12")
folio.describe()
DataFolio: analysis/experiment-12
=================================
Created: March 14, 2026 at 2:05 PM EDT
Updated: March 20, 2026 at 11:42 AM EDT
Tables (2):
• features: One row per neuron; normalized morphology
↳ size: 70.2 KB
• labels: Manual labels after the March review
↳ size: 8.3 KB
JSON Data (1):
• params: Settings used for the paper figures
↳ type: dict
↳ size: 31 B
Models (1):
• classifier: Cell-type classifier on reviewed labels
↳ size: 1.4 KB
↳ inputs: features, labels
Every item, what it is, and what it was made from — without opening a file. The same call works on a folio someone sends you as a cloud path.
features = folio.get("features")
clf = folio.get_model("classifier")
Install
pip install datafolio # core
pip install 'datafolio[polars]' # + lazy scans and Polars frames
Python 3.10+.
Who it's for
One researcher, or a small team sharing mostly read-only work, holding a few to a few dozen understandable objects per analysis. It assumes you are comfortable in pandas, Polars, numpy, and scikit-learn and do not want another framework between you and them.
Good fit when you want to stop writing
pd.read_parquet(BASE / "features_v3_reviewed.parquet") in every notebook,
remember which of five similar files the paper used, keep one analysis's
objects together, or hand a colleague a single path.
What you get
One write path, one read path. add() picks the storage format from the
object's type; get() reads the catalog and returns the matching Python object.
folio.add("table", dataframe) # pandas / Polars / LazyFrame -> Parquet
folio.add("counts", series) # pandas / Polars Series -> Parquet
folio.add("embeddings", array) # numpy -> .npy
folio.add("params", {"alpha": 0.1}) # dict / list / tuple / set / scalar -> JSON
folio.add_model("classifier", clf) # -> joblib (or skops)
folio.add_file("plots/qc.png") # -> the file, unchanged
folio.reference_table("raw", "gs://lab-data/raw.parquet") # link, don't copy
Descriptions live next to the data. The most valuable metadata is usually
one sentence saying which of several near-identical files this is. It is stored
in the catalog, shown by describe(), and readable by people who never install
the package.
Ordinary files, not a container format.
experiment-12/
├── items.json # the authoritative catalog (+ metadata + snapshots)
├── CONTENTS.md # derived, human-readable inventory
├── README.md # how to read this directory without datafolio
├── tables/features--r2.parquet
├── models/classifier--r8.joblib
└── artifacts/params--r5.json
Local or cloud, same code. DataFolio("gs://team-analysis/experiment-12").
Lineage, snapshots, and a CLI when you need them:
folio.add("predictions", preds, inputs=["classifier", "holdout"])
folio.create_snapshot("paper-v1", description="Figures 2–4 in the submission")
datafolio describe
datafolio snapshot status
What it is not
Not a database or query engine, not a workflow orchestrator, not data version control, not an experiment tracker, not a multi-writer system, not a backup. Many readers, one writer — treat a cloud folio as single-writer. Comfortable at a few dozen items per folio, not thousands.
DataFolio's rule is to be as lightweight as possible and hand off to better tools as soon as possible: big tables go to Polars as a lazy scan, or to you as a path. See What DataFolio is not for the full list, including the sharp edges worth knowing before you rely on it.
Documentation
Read in this order:
- Your first folio — ten minutes
- Everyday patterns
- Tables: big, external, and lazy
- Models
- Sharing a folio
- Snapshots
- Reading a folio without DataFolio
- What DataFolio is not
Reference: API cheat sheet · CLI · Migrating from 1.x
Development
uv sync
poe test # pytest with coverage
uv run ruff check src/ tests/
poe doc-preview # local docs server
License
MIT
Metadata
Release files for datafolio 2.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| datafolio-2.2.0.tar.gz | 124.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| datafolio-2.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 265.3 kB
Release files / datafolio-2.2.0.tar.gz
| Download URL | datafolio-2.2.0.tar.gz |
|---|---|
| Size | 124.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
aa2a743b526ff06d126ba9d0c2506ac538475e45baa7dbd15daada0e4f487182
|
|
BLAKE2b-256 checksum How to use checksums |
796269c24fc61708210dbb4e12f1a0251f615d49769dc6db043b0978fb96c85a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency logRelease files / datafolio-2.2.0-py3-none-any.whl
| Download URL | datafolio-2.2.0-py3-none-any.whl |
|---|---|
| Size | 140.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7934a6f387047c6356bc3dace29dff6c5dc98643e3b14eff166c95ce9e3e9788
|
|
BLAKE2b-256 checksum How to use checksums |
73caca554463163f6d3cabd7241617b6967409fe354dfbe64e87437346668c35
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency log