datasift-py
Zero-dependency data conversion, querying, manipulation, streaming, and schema validation for Python.
Why datasift-py?
datasift-py is a small, dependency-free Python library for working with structured data.
It combines:
- a fluent in-memory API;
- a small query language for nested data;
- CSV and JSONL streaming;
- schema validation using Python type hints;
- conversion between JSON, YAML, TOML, CSV, XML, and JSONL;
- a CLI built on the same library API.
Runtime dependencies are limited to Python's standard library.
Installation
pip install datasift-py
For development:
git clone https://github.com/sarahsalary/datasift-py.git
cd datasift-py
python -m pip install -e ".[dev]"
pytest
Quick start
Fluent API
from datasift import Data
result = (
Data("users.json")
.filter(age__gt=30)
.select("name", "email")
.sort("-age")
)
result.to("adults.csv")
For in-memory data:
from datasift import Data
users = [
{"name": "Alice", "age": 30},
{"name": "Bob", "age": 25},
{"name": "Carol", "age": 40},
]
adults = Data(users).filter(age__gte=30).to_list()
Streaming large files
Stream is single-use and lazy. It is intended for CSV and JSONL workloads that should not be fully materialized in memory.
from datasift import Stream
(
Stream.from_csv("huge.csv")
.filter(age__gt=30)
.select("name", "email")
.to_csv("filtered.csv")
)
CSV values remain strings when read. Numeric comparisons such as age__gt=30 normalize the compared value without changing the original record.
Query language
from datasift import query
data = {
"users": [
{"name": "Alice", "age": 30, "active": True},
{"name": "Bob", "age": 25, "active": False},
{"name": "Carol", "age": 40, "active": True},
]
}
query(data, "users[?age > 30].name")
# ["Carol"]
query(data, "users[?active == true].name")
# ["Alice", "Carol"]
query(data, "users[0].name")
# "Alice"
Supported operators include:
==,!=,>,>=,<,<=&&and||- dotted paths
- numeric indexes such as
[0] - list filters such as
[?age > 30] - list expansion with
[*]
Schema validation
from typing import List, TypedDict
from datasift import Data
class User(TypedDict):
name: str
age: int
email: str
Data("users.json").expect(List[User]).to("validated.yaml")
Validation supports common Python typing constructs including TypedDict, List, Dict, Tuple, Set, Optional, Union, and Literal.
Supported formats
| Format | Read | Write | Streaming |
|---|---|---|---|
| JSON | Yes | Yes | No |
| YAML | Yes | Yes | No |
| TOML | Yes | Yes | No |
| CSV | Yes | Yes | Yes |
| XML | Yes | Yes | No |
| JSONL / NDJSON | Yes | Yes | Yes |
YAML uses a small standard-library-compatible subset implemented by the project; it is not intended to be a full YAML 1.2 implementation.
XML model
XML uses a predictable mapping:
<users>
<user id="1">
<name>Alice</name>
</user>
</users>
becomes:
{
"users": {
"user": {
"@id": 1,
"name": "Alice",
}
}
}
Attributes use @name, mixed text uses #text, and repeated child tags become lists.
When serializing a top-level list, pass an explicit root tag:
Data(users).to_xml(root_tag="users")
CLI
Examples:
datasift convert users.json users.csv
datasift query users.json 'users[?age > 30].name'
datasift validate users.json --schema schema.py:User
Run:
datasift --help
for the complete command set.
Documentation
The documentation source is in docs/.
Build it locally with:
python -m pip install -e ".[docs]"
mkdocs serve
Development
Useful commands:
pytest
ruff check .
ruff format --check .
mypy
The CI workflow runs the supported Python versions and the project's quality checks.
Project structure
src/datasift/
├── core.py
├── data.py
├── stream.py
├── query.py
├── schema.py
├── schema_loader.py
├── cli.py
└── formats/
├── csv_io.py
├── json_io.py
├── jsonl_io.py
├── toml_io.py
├── xml_io.py
└── yaml_io.py
License
MIT. See LICENSE.
Release files for datasift-py 0.4.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| datasift_py-0.4.2.tar.gz | 31.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| datasift_py-0.4.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 62.5 kB
Release files / datasift_py-0.4.2.tar.gz
| Download URL | datasift_py-0.4.2.tar.gz |
|---|---|
| Size | 31.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
350afb515469fbb6f586afe40ffbd605ef89a870530a2e16610fb745b829e8fa
|
|
BLAKE2b-256 checksum How to use checksums |
7c4f2c3e983e9636ac995c5602cfb780300891e04fdecfe0935f8da87c62f5fa
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency logRelease files / datasift_py-0.4.2-py3-none-any.whl
| Download URL | datasift_py-0.4.2-py3-none-any.whl |
|---|---|
| Size | 30.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0b4d70a505a5a9f69c6029cf3e10a0318b0799767e3fbbc5c69e6ab8207e666a
|
|
BLAKE2b-256 checksum How to use checksums |
2adf4ba4bff5de38ab8b79c372e845d60943fefb152dbf3bbf2851ef5b96eccb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency log