Skip to main content

datasift-py

Zero-dependency data conversion, querying, manipulation, streaming, and schema validation for Python.

PyPI version Python versions License: MIT Tests

Why datasift-py?

datasift-py is a small, dependency-free Python library for working with structured data.

It combines:

  • a fluent in-memory API;
  • a small query language for nested data;
  • CSV and JSONL streaming;
  • schema validation using Python type hints;
  • conversion between JSON, YAML, TOML, CSV, XML, and JSONL;
  • a CLI built on the same library API.

Runtime dependencies are limited to Python's standard library.

Installation

pip install datasift-py

For development:

git clone https://github.com/sarahsalary/datasift-py.git
cd datasift-py
python -m pip install -e ".[dev]"
pytest

Quick start

Fluent API

from datasift import Data

result = (
    Data("users.json")
    .filter(age__gt=30)
    .select("name", "email")
    .sort("-age")
)

result.to("adults.csv")

For in-memory data:

from datasift import Data

users = [
    {"name": "Alice", "age": 30},
    {"name": "Bob", "age": 25},
    {"name": "Carol", "age": 40},
]

adults = Data(users).filter(age__gte=30).to_list()

Streaming large files

Stream is single-use and lazy. It is intended for CSV and JSONL workloads that should not be fully materialized in memory.

from datasift import Stream

(
    Stream.from_csv("huge.csv")
    .filter(age__gt=30)
    .select("name", "email")
    .to_csv("filtered.csv")
)

CSV values remain strings when read. Numeric comparisons such as age__gt=30 normalize the compared value without changing the original record.

Query language

from datasift import query

data = {
    "users": [
        {"name": "Alice", "age": 30, "active": True},
        {"name": "Bob", "age": 25, "active": False},
        {"name": "Carol", "age": 40, "active": True},
    ]
}

query(data, "users[?age > 30].name")
# ["Carol"]

query(data, "users[?active == true].name")
# ["Alice", "Carol"]

query(data, "users[0].name")
# "Alice"

Supported operators include:

  • ==, !=, >, >=, <, <=
  • && and ||
  • dotted paths
  • numeric indexes such as [0]
  • list filters such as [?age > 30]
  • list expansion with [*]

Schema validation

from typing import List, TypedDict
from datasift import Data

class User(TypedDict):
    name: str
    age: int
    email: str

Data("users.json").expect(List[User]).to("validated.yaml")

Validation supports common Python typing constructs including TypedDict, List, Dict, Tuple, Set, Optional, Union, and Literal.

Supported formats

Format Read Write Streaming
JSON Yes Yes No
YAML Yes Yes No
TOML Yes Yes No
CSV Yes Yes Yes
XML Yes Yes No
JSONL / NDJSON Yes Yes Yes

YAML uses a small standard-library-compatible subset implemented by the project; it is not intended to be a full YAML 1.2 implementation.

XML model

XML uses a predictable mapping:

<users>
  <user id="1">
    <name>Alice</name>
  </user>
</users>

becomes:

{
    "users": {
        "user": {
            "@id": 1,
            "name": "Alice",
        }
    }
}

Attributes use @name, mixed text uses #text, and repeated child tags become lists.

When serializing a top-level list, pass an explicit root tag:

Data(users).to_xml(root_tag="users")

CLI

Examples:

datasift convert users.json users.csv
datasift query users.json 'users[?age > 30].name'
datasift validate users.json --schema schema.py:User

Run:

datasift --help

for the complete command set.

Documentation

The documentation source is in docs/.

Build it locally with:

python -m pip install -e ".[docs]"
mkdocs serve

Development

Useful commands:

pytest
ruff check .
ruff format --check .
mypy

The CI workflow runs the supported Python versions and the project's quality checks.

Project structure

src/datasift/
├── core.py
├── data.py
├── stream.py
├── query.py
├── schema.py
├── schema_loader.py
├── cli.py
└── formats/
    ├── csv_io.py
    ├── json_io.py
    ├── jsonl_io.py
    ├── toml_io.py
    ├── xml_io.py
    └── yaml_io.py

License

MIT. See LICENSE.

Release files for datasift-py 0.4.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for datasift-py 0.4.2
File Size Uploaded
datasift_py-0.4.2.tar.gz 31.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for datasift-py 0.4.2
File Interpreter ABI Platform
datasift_py-0.4.2-py3-none-any.whl Python 3 none any Details

Total release size: 62.5 kB

Release files / datasift_py-0.4.2.tar.gz

Download URL datasift_py-0.4.2.tar.gz
Size 31.9 kB
Tags Source
SHA-256 checksum
How to use checksums
350afb515469fbb6f586afe40ffbd605ef89a870530a2e16610fb745b829e8fa
BLAKE2b-256 checksum
How to use checksums
7c4f2c3e983e9636ac995c5602cfb780300891e04fdecfe0935f8da87c62f5fa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release files / datasift_py-0.4.2-py3-none-any.whl

Download URL datasift_py-0.4.2-py3-none-any.whl
Size 30.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0b4d70a505a5a9f69c6029cf3e10a0318b0799767e3fbbc5c69e6ab8207e666a
BLAKE2b-256 checksum
How to use checksums
2adf4ba4bff5de38ab8b79c372e845d60943fefb152dbf3bbf2851ef5b96eccb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.4.2 This release

2 release files

0.4.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page