Skip to main content

A local dataset tracking tool.

Project description

Datashelf

Lightweight local dataset tracking for data science projects.

datashelf logo

Stop naming files data_final_really_final.csv. Datashelf stores tabular datasets as immutable artifacts and lets you retrieve them by name or hash — so your experiments stay reproducible without any heavy infrastructure.

$ datashelf save data/people.csv people_raw --message "initial load" --tag raw
Successfully saved 'people_raw' (hash: c8a2f8e1)

$ datashelf load people_raw --df
# Returns a pandas DataFrame, ready to use

Why Datashelf?

Data science projects often accumulates files like this:

data.csv
data_clean.csv
data_final_v2.csv
data_final_really_final.csv

Datashelf replaces that chaos with content-addressed storage: each dataset is hashed (SHA256), stored once as Parquet, and registered with metadata; and if you try to save a duplicate, Datashelf tells you. You can always get your data back by name or hash prefix.


Installation

pip install git+https://github.com/r0hankrishnan/datashelf.git

PyPI release coming soon.


Quick Start

# Initialize in your project directory
datashelf init

# Save a dataset
datashelf save data/people.csv people_raw --message "initial load" --tag raw

# List what's stored
datashelf list

# Load it back into pandas
datashelf load people_raw --df

Or use the Python API:

import datashelf as ds

ds.init()
ds.save("data/people.csv", name="people_raw", message="initial load", tag="raw")
df = ds.load("people_raw", to_df=True)

Commands

Command Description
datashelf init Initialize a .datashelf/ repo in the current directory (recommended to intialize in your project's root)
datashelf save <path> <name> Store a dataset artifact
datashelf list List all stored datasets
datashelf show <name> Inspect metadata for a dataset
datashelf load <name> Print the artifact path (use --df to load into pandas)
datashelf checkout <name> <dest> Export an artifact to another location

How It Works

When you save a dataset, Datashelf:

  1. Computes a SHA256 hash of the file contents
  2. Normalizes it to Parquet and stores it at .datashelf/artifacts/<hash>.parquet
  3. Registers metadata (name, tag, message, timestamp) in .datashelf/metadata.json

If you try to save the same data again under a different name, Datashelf detects the duplicate and asks if you want to update the metadata instead of storing a redundant copy.

.datashelf/
├── config.yaml
├── metadata.json
└── artifacts/
    └── c8a2f8e1...parquet

Design Philosophy

Datashelf deliberately tracks only tabular data. The core of the tool is duplicate detection and easy data organization: before storing anything, Datashelf checks whether you've already saved that data under a different name. For that check to work reliably, every dataset needs to be in a canonical format — you can't meaningfully compare a CSV and a Parquet of the same table without normalizing them first. I chose Parquet as the canonical format for its size benefits.

Accepting only tabular data is the direct consequence of that decision. It also makes future features like dataset diffing coherent — diffing only makes sense when you can compare rows and columns. Trying to extend Datashelf to handle images, audio, or arbitrary binary files would undermine both of those things without adding much value over a general-purpose tool like DVC.

The scope is intentionally narrow: Datashelf does one thing well for one kind of data.


Comparison

Tool Best for
Datashelf Lightweight local dataset tracking on a single project
DVC Full data version control with remote storage and pipeline orchestration
Git LFS Large file versioning inside a Git repository

Datashelf intentionally has no Git integration, no remote storage, and no pipeline orchestration. It's small and it stays out of your way.


Supported File Types

Datashelf accepts .csv, .parquet, .xlsx, and .json files and normalizes everything to Parquet internally.


Running Tests

pytest

Roadmap

  • PyPI release
  • Dataset diffing
  • Experiment tracking
  • Dataset lineage
  • Remote artifact storage

License

MIT — see LICENSE.


Contributing

See CONTRIBUTING.md.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

datashelf_py-0.1.2.tar.gz (17.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

datashelf_py-0.1.2-py3-none-any.whl (17.5 kB view details)

Uploaded Python 3

File details

Details for the file datashelf_py-0.1.2.tar.gz.

File metadata

  • Download URL: datashelf_py-0.1.2.tar.gz
  • Upload date:
  • Size: 17.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.10

File hashes

Hashes for datashelf_py-0.1.2.tar.gz
Algorithm Hash digest
SHA256 97d8063c97f71e2f77f708004561f8d0b80d423bde1575158a988b6fbf865058
MD5 9636556ce40a94f56f5534d127997243
BLAKE2b-256 7d7529ef58f6a67baa8e4de083d40013cd8b063fceff8614e920d42b11a7880f

See more details on using hashes here.

File details

Details for the file datashelf_py-0.1.2-py3-none-any.whl.

File metadata

  • Download URL: datashelf_py-0.1.2-py3-none-any.whl
  • Upload date:
  • Size: 17.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.10

File hashes

Hashes for datashelf_py-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 5725388259c13bb16da1d5ded84b37b4f506c82c226d3853a582feb046e05500
MD5 ad09056c058f49cc73dfb649c8a05345
BLAKE2b-256 acd10450db9077c49ac4a5227b4e8ca881e88b2b80792d2937011a20d73ca3f6

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page