datasift-py
Zero-dependency data conversion, querying, manipulation, streaming, and schema validation for Python.
Why datasift-py?
datasift-py is a small, dependency-free Python library for working with structured data.
It combines:
- a fluent in-memory API;
- a small query language for nested data;
- CSV and JSONL streaming;
- schema validation using Python type hints;
- conversion between JSON, YAML, TOML, CSV, XML, and JSONL;
- a CLI built on the same library API.
Runtime dependencies are limited to Python's standard library.
Installation
pip install datasift-py
For development:
git clone https://github.com/sarahsalary/datasift-py.git
cd datasift-py
python -m pip install -e ".[dev]"
pytest
Quick start
Fluent API
from datasift import Data
result = (
Data("users.json")
.filter(age__gt=30)
.select("name", "email")
.sort("-age")
)
result.to("adults.csv")
For in-memory data:
from datasift import Data
users = [
{"name": "Alice", "age": 30},
{"name": "Bob", "age": 25},
{"name": "Carol", "age": 40},
]
adults = Data(users).filter(age__gte=30).to_list()
Streaming large files
Stream is single-use and lazy. It is intended for CSV and JSONL workloads that should not be fully materialized in memory.
from datasift import Stream
(
Stream.from_csv("huge.csv")
.filter(age__gt=30)
.select("name", "email")
.to_csv("filtered.csv")
)
CSV values remain strings when read. Numeric comparisons such as age__gt=30 normalize the compared value without changing the original record.
Query language
from datasift import query
data = {
"users": [
{"name": "Alice", "age": 30, "active": True},
{"name": "Bob", "age": 25, "active": False},
{"name": "Carol", "age": 40, "active": True},
]
}
query(data, "users[?age > 30].name")
# ["Carol"]
query(data, "users[?active == true].name")
# ["Alice", "Carol"]
query(data, "users[0].name")
# "Alice"
Supported operators include:
==,!=,>,>=,<,<=&&and||- dotted paths
- numeric indexes such as
[0] - list filters such as
[?age > 30] - list expansion with
[*]
Schema validation
from typing import List, TypedDict
from datasift import Data
class User(TypedDict):
name: str
age: int
email: str
Data("users.json").expect(List[User]).to("validated.yaml")
Validation supports common Python typing constructs including TypedDict, List, Dict, Tuple, Set, Optional, Union, and Literal.
Supported formats
| Format | Read | Write | Streaming |
|---|---|---|---|
| JSON | Yes | Yes | No |
| YAML | Yes | Yes | No |
| TOML | Yes | Yes | No |
| CSV | Yes | Yes | Yes |
| XML | Yes | Yes | No |
| JSONL / NDJSON | Yes | Yes | Yes |
YAML uses a small standard-library-compatible subset implemented by the project; it is not intended to be a full YAML 1.2 implementation.
XML model
XML uses a predictable mapping:
<users>
<user id="1">
<name>Alice</name>
</user>
</users>
becomes:
{
"users": {
"user": {
"@id": 1,
"name": "Alice",
}
}
}
Attributes use @name, mixed text uses #text, and repeated child tags become lists.
When serializing a top-level list, pass an explicit root tag:
Data(users).to_xml(root_tag="users")
CLI
Examples:
datasift convert users.json users.csv
datasift query users.json 'users[?age > 30].name'
datasift validate users.json --schema schema.py:User
Run:
datasift --help
for the complete command set.
Documentation
The documentation source is in docs/.
Build it locally with:
python -m pip install -e ".[docs]"
mkdocs serve
Development
Useful commands:
pytest
ruff check .
ruff format --check .
mypy
The CI workflow runs the supported Python versions and the project's quality checks.
Project structure
src/datasift/
├── core.py
├── data.py
├── stream.py
├── query.py
├── schema.py
├── schema_loader.py
├── cli.py
└── formats/
├── csv_io.py
├── json_io.py
├── jsonl_io.py
├── toml_io.py
├── xml_io.py
└── yaml_io.py
License
MIT. See LICENSE.
Release files for datasift-py 0.4.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| datasift_py-0.4.1.tar.gz | 33.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| datasift_py-0.4.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 63.9 kB
Release files / datasift_py-0.4.1.tar.gz
| Download URL | datasift_py-0.4.1.tar.gz |
|---|---|
| Size | 33.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
da7a6c19350264ff8faf42408e19c8a358ce67646a9b6c58e677963de80a1891
|
|
BLAKE2b-256 checksum How to use checksums |
ee6460216bf5f0db06e0becab8d4c69bce3e8522ce8ec19273c6759178db2d38
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.0
|
Release files / datasift_py-0.4.1-py3-none-any.whl
| Download URL | datasift_py-0.4.1-py3-none-any.whl |
|---|---|
| Size | 30.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c495bafad6111be30e09726f25e497cc5ef15bf66ffad5d7eb6c910e323a4edb
|
|
BLAKE2b-256 checksum How to use checksums |
728d866e4d564a4cc126d50e860584ec7d27640677bcbed9950e37dbaa481004
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.0
|