Unified storage framework for machine learning datasets
Project description
Space: Unified Storage for Machine Learning
Unify data in your entire machine learning lifecycle with Space, a comprehensive storage solution that seamlessly handles data from ingestion to training.
Key Features:
- Ground Truth Database
- Store and manage data in open source file formats, locally or in the cloud.
- Ingest from various sources, including ML datasets, files, and labeling tools.
- Support data manipulation (append, insert, update, delete) and version control.
- OLAP Database and Lakehouse
- Iceberg style open table format.
- Optimized for unstructued data via reference operations.
- Quickly analyze data using SQL engines like DuckDB.
- Distributed Data Processing Pipelines
- Integrate with processing frameworks like Ray for efficient data transformation.
- Store processed results as Materialized Views (MVs); incrementally update MVs when the source is changed.
- Seamless Training Framework Integration
- Access Space datasets and MVs directly via random access interfaces.
- Convert to popular ML dataset formats (e.g., TFDS, HuggingFace, Ray).
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
space-datasets-0.0.4.tar.gz
(117.0 kB
view hashes)
Built Distribution
space_datasets-0.0.4-py3-none-any.whl
(172.5 kB
view hashes)
Close
Hashes for space_datasets-0.0.4-py3-none-any.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | b04c213975c849b990ff32342591d1c8745c2d78393f20d20dc5df08b3df13c7 |
|
MD5 | c338c84c059f68e32fe801af8d7939ab |
|
BLAKE2b-256 | f94636d5c85be0b211631b3b3c4677b44e8f6462797695b5df1dd2adf0d9b9d3 |