Skip to main content

DataFolio

Python 3.10+ Docs

A small, self-documenting home for the data behind one analysis.

You already have the objects: a feature table, some labels, an array of embeddings, a fitted model, a parameter dict, a QC plot. DataFolio gives them one directory, one name each, and one sentence each explaining what they are — then hands them back by name months later, from any notebook.

from datafolio import DataFolio

folio = DataFolio("analysis/experiment-12")

folio.add("features", features, description="One row per neuron; normalized morphology")
folio.add("labels", labels, description="Manual labels after the March review")
folio.add("params", {"alpha": 0.1, "seed": 7}, description="Settings used for the paper figures")
folio.add_model("classifier", clf, inputs=["features", "labels"],
                description="Cell-type classifier on reviewed labels")
# A different notebook, six months later. One path, no filenames.
folio = DataFolio("analysis/experiment-12")
folio.describe()
DataFolio: analysis/experiment-12
=================================

Created: March 14, 2026 at 2:05 PM EDT
Updated: March 20, 2026 at 11:42 AM EDT

Tables (2):
  • features: One row per neuron; normalized morphology
    ↳ size: 70.2 KB
  • labels: Manual labels after the March review
    ↳ size: 8.3 KB

JSON Data (1):
  • params: Settings used for the paper figures
    ↳ type: dict
    ↳ size: 31 B

Models (1):
  • classifier: Cell-type classifier on reviewed labels
    ↳ size: 1.4 KB
    ↳ inputs: features, labels

Every item, what it is, and what it was made from — without opening a file. The same call works on a folio someone sends you as a cloud path.

features = folio.get("features")
clf = folio.get_model("classifier")

Install

pip install datafolio            # core
pip install 'datafolio[polars]'  # + lazy scans and Polars frames

Python 3.10+.

Who it's for

One researcher, or a small team sharing mostly read-only work, holding a few to a few dozen understandable objects per analysis. It assumes you are comfortable in pandas, Polars, numpy, and scikit-learn and do not want another framework between you and them.

Good fit when you want to stop writing pd.read_parquet(BASE / "features_v3_reviewed.parquet") in every notebook, remember which of five similar files the paper used, keep one analysis's objects together, or hand a colleague a single path.

What you get

One write path, one read path. add() picks the storage format from the object's type; get() reads the catalog and returns the matching Python object.

folio.add("table", dataframe)            # pandas / Polars / LazyFrame -> Parquet
folio.add("counts", series)              # pandas / Polars Series -> Parquet
folio.add("embeddings", array)           # numpy -> .npy
folio.add("params", {"alpha": 0.1})      # dict / list / tuple / set / scalar -> JSON
folio.add_model("classifier", clf)       # -> joblib (or skops)
folio.add_file("plots/qc.png")           # -> the file, unchanged
folio.reference_table("raw", "gs://lab-data/raw.parquet")   # link, don't copy

Descriptions live next to the data. The most valuable metadata is usually one sentence saying which of several near-identical files this is. It is stored in the catalog, shown by describe(), and readable by people who never install the package.

Ordinary files, not a container format.

experiment-12/
├── items.json                    # the authoritative catalog (+ metadata + snapshots)
├── CONTENTS.md                   # derived, human-readable inventory
├── README.md                     # how to read this directory without datafolio
├── tables/features--r2.parquet
├── models/classifier--r8.joblib
└── artifacts/params--r5.json

Local or cloud, same code. DataFolio("gs://team-analysis/experiment-12").

Lineage, snapshots, and a CLI when you need them:

folio.add("predictions", preds, inputs=["classifier", "holdout"])
folio.create_snapshot("paper-v1", description="Figures 2–4 in the submission")
datafolio describe
datafolio snapshot status

What it is not

Not a database or query engine, not a workflow orchestrator, not data version control, not an experiment tracker, not a multi-writer system, not a backup. Many readers, one writer — treat a cloud folio as single-writer. Comfortable at a few dozen items per folio, not thousands.

DataFolio's rule is to be as lightweight as possible and hand off to better tools as soon as possible: big tables go to Polars as a lazy scan, or to you as a path. See What DataFolio is not for the full list, including the sharp edges worth knowing before you rely on it.

Documentation

Read in this order:

  1. Your first folio — ten minutes
  2. Everyday patterns
  3. Tables: big, external, and lazy
  4. Models
  5. Sharing a folio
  6. Snapshots
  7. Reading a folio without DataFolio
  8. What DataFolio is not

Reference: API cheat sheet · CLI · Migrating from 1.x

Development

uv sync
poe test                      # pytest with coverage
uv run ruff check src/ tests/
poe doc-preview               # local docs server

License

MIT

Metadata

Release files for datafolio 2.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for datafolio 2.2.0
File Size Uploaded
datafolio-2.2.0.tar.gz 124.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for datafolio 2.2.0
File Interpreter ABI Platform
datafolio-2.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 265.3 kB

Release files / datafolio-2.2.0.tar.gz

Download URL datafolio-2.2.0.tar.gz
Size 124.9 kB
Tags Source
SHA-256 checksum
How to use checksums
aa2a743b526ff06d126ba9d0c2506ac538475e45baa7dbd15daada0e4f487182
BLAKE2b-256 checksum
How to use checksums
796269c24fc61708210dbb4e12f1a0251f615d49769dc6db043b0978fb96c85a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release files / datafolio-2.2.0-py3-none-any.whl

Download URL datafolio-2.2.0-py3-none-any.whl
Size 140.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7934a6f387047c6356bc3dace29dff6c5dc98643e3b14eff166c95ce9e3e9788
BLAKE2b-256 checksum
How to use checksums
73caca554463163f6d3cabd7241617b6967409fe354dfbe64e87437346668c35
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

2.2.0 This release

2 release files

2.1.0

2 release files

2.0.0

2 release files

1.3.0

1 release file

1.2.0

1 release file

1.1.0

2 release files

1.0.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page