Skip to main content

A local dataset tracking tool.

Project description

Datashelf

datashelf logo

Datashelf is a lightweight local dataset tracking tool for data science projects.

It stores tabular datasets as immutable artifacts, tracks metadata, and lets you retrieve them later by name or hash. The goal is to make experiments easier to reproduce without introducing heavy infrastructure.

Example

$ datashelf init
Initialized DataShelf at .datashelf/

$ datashelf save data/people.csv people_raw --message "tiny dataset" --tag raw
Successfully saved 'people_raw' with hash c8a2f8e1

$ datashelf list

Hash      Name         Tag   Message
-----------------------------------------
c8a2f8e1  people_raw   raw   tiny dataset

Why Datashelf?

Many data science workflows struggle with dataset organization:

  • intermediate datasets get overwritten
  • multiple versions accumulate
  • experiments become difficult to reproduce

Instead of accumulating files like:

data.csv
data_clean.csv
data_final_v2.csv
data_final_really_final.csv

Datashelf stores datasets using content hashes and maintains a metadata registry so artifacts can always be located again.

Key ideas:

  • content-addressed storage (SHA256)
  • metadata registry for datasets
  • lookup by name or hash prefix
  • CLI + Python API
  • opinionated dataset tags based on Cookiecutter Data Science

It is intentionally local and lightweight, designed for individual projects rather than large data pipelines.

Features

  • Local dataset artifact storage
  • SHA256 content hashing
  • Metadata registry (name, tag, message, timestamp)
  • Lookup by dataset name or hash prefix
  • Optional dataset tags and messages
  • CLI + Python API
  • Automatic normalization to Parquet
  • Duplicate dataset detection
  • Basic unit test coverage

Installation

Clone the repository and install locally:

git clone <repo-url>
cd datashelf
pip install -e .

I am also working on getting it published on PyPi!

Quick Start

Initialize a Datashelf repository in your project directory:

datashelf init

This creates a hidden directory used to store artifacts and metadata:

.datashelf/
├── config.yaml
├── metadata.json
└── artifacts/

Example Workflow

Save a dataset:

datashelf save data/people.csv people_raw --message "tiny dataset" --tag raw

List stored datasets:

datashelf list

Inspect metadata:

datashelf show people_raw

Load the stored dataset path:

datashelf load people_raw

Load directly into pandas:

datashelf load people_raw --df

Export the artifact to another location:

datashelf checkout people_raw exports/people.parquet

Python API

Datashelf can also be used directly from Python:

import datashelf as ds

ds.init()

ds.save(
    data="data.csv",
    name="training_data",
    message="clean dataset",
    tag="processed"
)

df = ds.load("training_data", to_df=True)

Architecture

Datashelf separates user commands from internal system services.

User / CLI
    │
    ▼
Command Layer
(init, save, load, inspect, checkout)
    │
    ▼
Core Services
(directory, hashing, metadata, config)
    │
    ▼
.datashelf/
    artifacts + metadata registry

Command Layer

Handles user workflows such as saving, loading, inspecting, and exporting datasets.

Core Layer

Implements internal functionality including:

  • content hashing
  • metadata management
  • artifact storage
  • configuration management

This separation keeps command modules simple and makes core logic easier to test and maintain.

How Artifacts Are Stored

Datasets are stored using their SHA256 hash:

.datashelf/artifacts/<hash>.parquet

Metadata is stored in a registry:

{
  "file_hash": "c8a2f8e1...",
  "name": "people_raw",
  "tag": "raw",
  "message": "tiny dataset",
  "stored_path": "artifacts/c8a2f8e1.parquet",
  "datetime_added": "2026-03-10T12:30:00"
}

This ensures datasets can always be referenced reliably.

Comparison

Datashelf focuses on simple, local dataset tracking.

Tool Purpose
DataShelf Lightweight local dataset tracking for tabular data
DVC Full data version control with remote storage
Git LFS Large file versioning inside Git

Datashelf intentionally avoids:

  • Git integration
  • remote storage
  • pipeline orchestration

This keeps the tool simple and easy to use for smaller data science projects.

Running Tests

Run tests with:

pytest

Tests cover repository initialization, dataset saving, loading, metadata inspection, and artifact checkout.

Future Work

Possible extensions include:

  • dataset diffing
  • experiment tracking
  • dataset lineage tracking
  • remote artifact storage
  • richer filtering and search

License

MIT License.

About This Project

Datashelf was built as a personal project to make something that I thought would be useful in my day-to-day work at school.

The project demonstrates:

  • CLI tool development
  • artifact-based dataset management
  • modular Python package architecture
  • reproducible data pipelines
  • test-driven development

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

datashelf_py-0.1.0.tar.gz (17.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

datashelf_py-0.1.0-py3-none-any.whl (17.7 kB view details)

Uploaded Python 3

File details

Details for the file datashelf_py-0.1.0.tar.gz.

File metadata

  • Download URL: datashelf_py-0.1.0.tar.gz
  • Upload date:
  • Size: 17.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.10

File hashes

Hashes for datashelf_py-0.1.0.tar.gz
Algorithm Hash digest
SHA256 aa5b8b2560515115f72d032b4e745250105dd59b1667399e736e1ee961115a5d
MD5 f3337969f2bd3b2b16b0268efeefe957
BLAKE2b-256 664b4a2454b7051397dc6dbe4c8f9487694add938b84bdfc157cf17d1044412f

See more details on using hashes here.

File details

Details for the file datashelf_py-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: datashelf_py-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 17.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.10

File hashes

Hashes for datashelf_py-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b2e4df80a9502bfb7d2ae2863314555568619fa5488b5269bb79efc4b5534ff1
MD5 d03b2471e8f757685cce87aca617895b
BLAKE2b-256 7e75042e95803cf2e410c17f59b2b87803fbb6df920d328b75cb9514b2ca59de

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page