Big-file data prep that never runs out of memory - no SQL required. Powered by DuckDB.
Project description
kenze
Big-file data prep that never runs out of memory — no SQL required.
kenze is a tiny command-line tool for cleaning and reshaping data files
(CSV, Parquet, JSON) that are too big for pandas. It's a friendly front-end
over DuckDB: DuckDB does the heavy lifting (streaming,
disk-spill, all your CPU cores) and kenze makes it a one-liner — and
auto-configures memory so your job doesn't crash.
pip install kenze
One name for everything: pip install kenze → the kenze command → import kenze.
Feature highlights
- Process files bigger than your RAM without crashing — memory is auto-capped and DuckDB spills to disk.
- 26 commands for the everyday work:
keep,drop,filter,rename,cast,fillna,dedup,sample,join,diff,pivot,split,partition, and more — no SQL needed. - Readable recipes (
.dqfiles) that chain steps into one streaming pass, with${VAR}templating for scheduled jobs. - Read and write the cloud directly —
s3://,gs://,https://— nothing to download first. - PII masking (
mask --method hash), schema validation (validate), and a file-integrity pre-check (check) for production pipelines. - No lock-in —
ejectany recipe to raw DuckDB SQL or Python. - Use it from Python too —
import kenzeand callkenze.sift(...),kenze.sql(...). - One dependency at heart (DuckDB), atomic writes, and ASCII-clean output on any terminal.
Why
- It doesn't OOM. Memory is capped to a fraction of free RAM and DuckDB spills to disk instead of dying. Point it at a file bigger than your RAM; it's fine.
- No SQL, no pandas. Simple verbs, or a readable recipe file.
- One streaming pass. A whole recipe compiles to a single query — no intermediate files, so it's fast and light.
- Any format, local or cloud. CSV / Parquet / JSON, plain or
.gz, on disk or ons3:///gs:///https://— auto-detected.
One-liners
kenze profile sales.parquet # schema + row count, instantly
kenze peek sales.parquet # first rows + types + null counts
kenze stats sales.parquet # per-column min/max/nulls/unique
kenze check sales.csv # is the file valid? any bad rows?
kenze keep sales.parquet --cols id,city,amount -o small.csv
kenze drop users.csv --cols email,phone -o clean.parquet
kenze filter sales.parquet --where "amount > 100" -o big.csv
kenze rename sales.csv --map "amount:total" -o out.csv
kenze cast users.csv --types "zip:VARCHAR" -o out.parquet # keep leading zeros
kenze fillna users.csv --with "city:Unknown" -o out.csv
kenze mask users.csv --cols email,ssn --method hash -o safe.csv
kenze dedup users.csv --on id -o unique.parquet
kenze sample sales.parquet --n 50000 -o sample.csv
kenze clip points.parquet --bbox -10,35,5,45 -o region.parquet
kenze join orders.csv users.parquet --on user_id -o joined.parquet
kenze diff old.csv new.csv --on id # added / removed / changed
kenze pivot sales.csv --on city --values amount --agg sum --group region -o wide.csv
kenze split sales.parquet --by city -o by_city/ # one file per value
kenze partition sales.parquet --by year -o lake/ # hive year=2026/ folders
kenze convert sales.parquet -o sales.csv # just change format
kenze sql "SELECT *, lag(amount) OVER (ORDER BY ts) FROM 'sales.parquet'" -o out.csv
Read or write the cloud directly (nothing to download first):
kenze filter s3://bucket/huge.parquet --where "amount > 0" -o local.csv
Pipe like any Unix tool (use - for stdin/stdout):
cat data.csv | kenze filter - --where "x > 1" -o - | gzip > out.csv.gz
Recipes
Chain steps in a readable .dq file — they run as one streaming pass:
# clean.dq
input: data/sales_${DAY}.parquet # ${DAY} filled from --set or the environment
keep: [id, city, amount]
types: zip:VARCHAR
filter: amount > 0
fillna: city:Unknown
dedup: id
sample: 50000
output: out/clean.csv
kenze run clean.dq --set DAY=2026-07-14
kenze recipe # show every valid recipe step
kenze eject clean.dq --to sql # print the raw DuckDB SQL (no lock-in)
From Python
import kenze
kenze.sift("big.parquet", "clean.csv", keep=["id", "city"], filter="amount > 0", sample=50000)
rows = kenze.sql("SELECT city, count(*) FROM 'big.parquet' GROUP BY 1")
kenze.profile("big.parquet")
Handy flags
--memory-limit 8— pin the RAM budget (GB) for reproducible / SLA runs.--temp-dir D:/spill— put disk-spill where there's room.--skip-bad-lines— ignore malformed rows in a dirty CSV.--log run.json— write a run manifest (inputs, rows, timing).- Writes are atomic — a cancelled run never leaves a half-written file.
Commands
profile · peek · stats · check · validate · keep · drop · rename · cast ·
fillna · mask · filter · dedup · sample · head · clip · convert · join ·
diff · pivot · split · partition · sql · eject · run · recipe
Where it stops (on purpose)
kenze is one dependency and one machine — that's the whole point. It maxes out your cores and spills to disk so a single laptop or VM can chew through files far bigger than its RAM. It does not run a cluster. If you've genuinely outgrown one machine (multi-terabyte, distributed pipelines with SLAs and lineage tracking), reach for Spark/Dask — kenze is the tool you use before you need those.
Troubleshooting
'kenze' is not recognized / kenze: command not found?
pip installed kenze correctly — the command just landed in a folder that isn't on
your PATH (this affects every pip-installed CLI). Options:
- Use it now, no setup:
python -m kenze --help - Fix it for good: reinstall Python from python.org
with "Add python.exe to PATH" ticked, or use
python -m pipx install kenze.
MIT licensed.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file kenze-0.3.1.tar.gz.
File metadata
- Download URL: kenze-0.3.1.tar.gz
- Upload date:
- Size: 23.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1665dd2b1e442f8a947c33afec71f2499fcfbf1fafe37a41cc90c04de26cf7f6
|
|
| MD5 |
e0351b42fc68c5d6f9d0893860c2efef
|
|
| BLAKE2b-256 |
6db4d2a43d19e5ea444ad850395334dbb6fb9a843bac3e6cee2d4d70b3f4e3fa
|
File details
Details for the file kenze-0.3.1-py3-none-any.whl.
File metadata
- Download URL: kenze-0.3.1-py3-none-any.whl
- Upload date:
- Size: 22.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ad74f86122b4b1742ec8b539e867d20392b3d82b04e822ca9fbdb4384d7d8de2
|
|
| MD5 |
3e7a88b537e2717bdfd0a1281340fbed
|
|
| BLAKE2b-256 |
24c351691e44d5610c7d237a533211f3af745f12299a2b6c2ee5e1a8b1fe11da
|