Skip to main content

DataFolio

Python 3.10+ Docs

A small, self-documenting home for the data behind one analysis.

You already have the objects: a feature table, some labels, an array of embeddings, a fitted model, a parameter dict, a QC plot. DataFolio gives them one directory, one name each, and one sentence each explaining what they are — then hands them back by name months later, from any notebook.

from datafolio import DataFolio

folio = DataFolio("analysis/experiment-12")

folio.add("features", features, description="One row per neuron; normalized morphology")
folio.add("labels", labels, description="Manual labels after the March review")
folio.add("params", {"alpha": 0.1, "seed": 7})
folio.add_model("classifier", clf, inputs=["features", "labels"])
# A different notebook, six months later. One path, no filenames.
folio = DataFolio("analysis/experiment-12")

folio.describe()
features = folio.get("features")
clf = folio.get_model("classifier")

Install

pip install datafolio            # core
pip install 'datafolio[polars]'  # + lazy scans and Polars frames

Python 3.10+.

Who it's for

One researcher, or a small team sharing mostly read-only work, holding a few to a few dozen understandable objects per analysis. It assumes you are comfortable in pandas, Polars, numpy, and scikit-learn and do not want another framework between you and them.

Good fit when you want to stop writing pd.read_parquet(BASE / "features_v3_reviewed.parquet") in every notebook, remember which of five similar files the paper used, keep one analysis's objects together, or hand a colleague a single path.

What you get

One write path, one read path. add() picks the storage format from the object's type; get() reads the catalog and returns the matching Python object.

folio.add("table", dataframe)            # pandas / Polars / LazyFrame -> Parquet
folio.add("counts", series)              # pandas / Polars Series -> Parquet
folio.add("embeddings", array)           # numpy -> .npy
folio.add("params", {"alpha": 0.1})      # dict / list / tuple / set / scalar -> JSON
folio.add_model("classifier", clf)       # -> joblib (or skops)
folio.add_file("plots/qc.png")           # -> the file, unchanged
folio.reference_table("raw", "gs://lab-data/raw.parquet")   # link, don't copy

Descriptions live next to the data. The most valuable metadata is usually one sentence saying which of several near-identical files this is. It is stored in the catalog, shown by describe(), and readable by people who never install the package.

Ordinary files, not a container format.

experiment-12/
├── items.json                    # the authoritative catalog (+ metadata + snapshots)
├── CONTENTS.md                   # derived, human-readable inventory
├── README.md                     # how to read this directory without datafolio
├── tables/features--r2.parquet
├── models/classifier--r8.joblib
└── artifacts/params--r5.json

Local or cloud, same code. DataFolio("gs://team-analysis/experiment-12").

Lineage, snapshots, and a CLI when you need them:

folio.add("predictions", preds, inputs=["classifier", "holdout"])
folio.create_snapshot("paper-v1", description="Figures 2–4 in the submission")
datafolio describe
datafolio snapshot status

What it is not

Not a database or query engine, not a workflow orchestrator, not data version control, not an experiment tracker, not a multi-writer system, not a backup. Many readers, one writer — treat a cloud folio as single-writer. Comfortable at a few dozen items per folio, not thousands.

DataFolio's rule is to be as lightweight as possible and hand off to better tools as soon as possible: big tables go to Polars as a lazy scan, or to you as a path. See What DataFolio is not for the full list, including the sharp edges worth knowing before you rely on it.

Documentation

Read in this order:

  1. Your first folio — ten minutes
  2. Everyday patterns
  3. Tables: big, external, and lazy
  4. Models
  5. Sharing a folio
  6. Snapshots
  7. Reading a folio without DataFolio
  8. What DataFolio is not

Reference: API cheat sheet · CLI · Migrating from 1.x

Development

uv sync
poe test                      # pytest with coverage
uv run ruff check src/ tests/
poe doc-preview               # local docs server

License

MIT

Metadata

Release files for datafolio 2.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for datafolio 2.1.0
File Size Uploaded
datafolio-2.1.0.tar.gz 108.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for datafolio 2.1.0
File Interpreter ABI Platform
datafolio-2.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 231.6 kB

Release files / datafolio-2.1.0.tar.gz

Download URL datafolio-2.1.0.tar.gz
Size 108.5 kB
Tags Source
SHA-256 checksum
How to use checksums
3458c5e7b8cd0b0a22442ad132dbdb9c2a38983868f48773e15a04f153ce8bc6
BLAKE2b-256 checksum
How to use checksums
bcd5cdf003656626e788eaa71d4c3cc4452e2888184252016f0cc126caba8c4c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.6 {"installer":{"name":"uv","version":"0.12.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / datafolio-2.1.0-py3-none-any.whl

Download URL datafolio-2.1.0-py3-none-any.whl
Size 123.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9618b293129eef30c17b15f1f16d6d643a9e36794da572b296f08f7f7185cd4e
BLAKE2b-256 checksum
How to use checksums
d259cdbcb74945351679ad76213e06d1bcebd0b1007eb24a97e94870d2d7ae6c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.6 {"installer":{"name":"uv","version":"0.12.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

2.2.0

2 release files

This release

2.1.0 This release

2 release files

2.0.0

2 release files

1.3.0

1 release file

1.2.0

1 release file

1.1.0

2 release files

1.0.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page