nano-xet
A tiny, readable re-implementation of the ideas behind Xet — the content-defined-chunking storage layer Hugging Face uses for large files — in ~1000 lines of Python, on top of fsspec.
Write a file to nxet://, and it is cut into content-defined chunks, hashed, deduplicated
against what is already stored, packed into xorb objects (each xorb holds many chunks),
and recorded in small JSON index files at the root of any filesystem fsspec knows about.
It is a teaching/demo tool, not a production storage system: no compression, no encryption, no CAS server, no concurrent writers. Target size: files below ~300 MB.
nxet://my/path/to/data.csv::file:///Users/me/tmp/my-nxet-store
└──────┬─────────────────┘ └──────────────┬──────────────────────┘
the file you see where the xorbs live
Install
pip install nano-xet # or: pip install "nano-xet[fast]" for numpy chunking
Quick start
fsspec
import fsspec
store = "file:///Users/me/tmp/my-nxet-store"
with fsspec.open(f"nxet://data/train.csv::{store}", "wb") as f:
f.write(b"id,value\n0,1\n1,2\n")
with fsspec.open(f"nxet://data/train.csv::{store}", "rb") as f:
print(f.readline()) # b'id,value\n'
fs = fsspec.filesystem("nxet", fo=store)
fs.ls("data") # virtual dirs, nothing on disk
fs.cat_file("data/train.csv", start=9) # random access to any byte range
fs.pipe_file("data/train_v2.csv", new_bytes) # only new chunks are written
fs.stats() # deduplication statistics
Everything fsspec can do works: find, glob, walk, copy, mv, get, put,
tail, head, touch, append mode, text mode, and chained URIs against other
protocols (memory://, s3://, gs://, smb://, …).
Store API (no fsspec needed)
from nano_xet import NXetStore
with NXetStore.open(f"nxet://::{store}") as store_: # or just NXetStore.open(store)
store_.write_file("data/train.csv", data)
store_.read_range("data/train.csv", 1_000_000, 1_000_128)
print(store_.stats().summary())
1 file(s), 1681 chunk(s), 1 xorb(s)
logical : 106.7 MB
stored : 53.5 MB (842 unique chunk(s))
dedup : 53.2 MB saved (49.9%) [839 chunk(s) reused]
CLI
nxet put train.csv nxet://data/train.csv::file:///tmp/my-store
nxet put train.csv nxet://data/train_v2.csv::file:///tmp/my-store # dedup: only new chunks land
nxet ls -R nxet://::file:///tmp/my-store
nxet cat nxet://data/train.csv::file:///tmp/my-store --start 0 --end 100
nxet get nxet://data/train.csv::file:///tmp/my-store ./train.csv
nxet stats nxet://::file:///tmp/my-store
nxet xorbs nxet://::file:///tmp/my-store
nxet gc nxet://::file:///tmp/my-store
Demo
python examples/demo.py /tmp/nano-xet-demo
Writes two versions of an 11 MB CSV and shows that the second one costs 0.6 MB.
How it works
data.csv ──gear hash──▶ chunks ──blake2b──▶ chunk hashes
│ ┌──────────────┐
┌─────────────────────┴───────────┐ │ nxet.json │ header
▼ ▼ │ (format, │
┌──────────────────────────┐ dedup lookup │ chunk sizes)│
│ 000000-a1b2c3….xorb │◀── chunks packed in ──┬───yes──▶└──────┬───────┘
│ (many chunks, one file) │ │ reuse ┌──────┴───────┐
└──────────────────────────┘ │ │ nxet.files │
│ chunk a1b2 │ chunk 04de │ chunk f9… │ │ │ .jsonl │ path → hashes
no └──────┬───────┘
┌─────────────────────────────┘
▼ ┌──────────────────┐
new xorb │ nxet.xorbs.jsonl │ hash → xorb+offset
└──────────────────┘
- Chunking (
chunking.py) — content-defined chunking with the same gear hash, the same 256-entry lookup table, the same boundary mask and the same size limits as Xet: 64 KiB mean, 8 KiB minimum, 128 KiB maximum, boundary whenhash & mask == 0. Chunk boundaries therefore match whatxet-coreproduces for the same input (tests/test_chunking.pycompares against golden values generated by a Rust reference chunker). - Hashing (
hashing.py) — every chunk is identified by itsblake2b-256digest, used as the deduplication key (blake3from the standard library equivalent, no dependency). - Xorbs (
store.py) — chunks are appended to the current xorb until it reaches 64 MiB or 8192 chunks (Xet's own limits), then a new xorb starts. Chunks inside a xorb are sorted by hash and stored raw, so a xorb is a plain concatenation of chunk bytes and a chunk is read with a singlepread. - Index (
index.py) — three small files at the root of the underlying filesystem:nxet.json(header),nxet.files.jsonl(append-onlyput/rmrecords: path → chunk hashes),nxet.xorbs.jsonl(xorb →[hash, size, offset]per chunk). Directories are virtual: they exist because some file path has them as a prefix.
A file is a list of chunk hashes; reading is hash → (xorb, offset, size) → bytes, and
consecutive chunks in the same xorb are coalesced into one read.
Same as Xet / different from Xet
| nano-xet | Xet | |
|---|---|---|
| gear hash table, boundary mask, min/mean/max chunk size | identical | — |
xorb as a physical multi-chunk container, 64 MiB / 8192 chunks |
identical | — |
| chunk deduplication across files and versions | yes | yes |
| hashing | blake2b-256 |
keyed blake3 |
| xorb content | raw chunk bytes | byte-grouped, compressed, encrypted |
| index | JSON/JSONL at the root of the filesystem | sharded merkle tables in a CAS |
| metadata updates | last write wins, reload() to see others |
CAS + commit with rebase |
| storage | any fsspec filesystem | HF CAS (+ local cache) |
| language | Python | Rust |
Performance
On an Apple M-series laptop, Python 3.12, a 107 MB CSV (1681 chunks, mean 64 KiB):
| operation | cost |
|---|---|
| chunking | 2.2 s with numpy (49 MB/s), 11 s pure Python (10 MB/s) |
| write: chunk + hash + dedup + store | 2.4 s |
| write a second version with 1 byte inserted | 2.4 s, +1 chunk stored |
| read the whole file back | 0.05 s (~2 GB/s from OS cache) |
| 20 random 1 KB reads at different offsets | 1.1 ms |
| two identical 107 MB files | 107 MB stored, not 214 MB |
The numpy path is optional (use_numpy=False, or the fast extra) and gives ~4x faster
chunking; the pure Python path keeps nano-xet dependency-free apart from fsspec.
Limitations (by design)
- One writer at a time. Readers pick up other writers' changes on the next miss
(
reload_if_stale), but two processes writing at the same time can lose a file record. - No compression, no encryption, no partial-file corruption recovery: a xorb is raw bytes.
- Everything needed to rebuild a file is in the JSONL index, so huge datasets mean big index files (nano-xet is meant for ≤ 300 MB files, not for a whole repository).
gcmust be run when files are deleted; unreferenced xorbs are only garbage.- Chunking is Xet-compatible, but the file hash is not a Xet merkle hash, so nano-xet stores and Xet stores are not interchangeable.
memory://works as an underlying filesystem for tests and demos, but it is per-process: twonxetcommands do not share it.
Development
pip install -e ".[test,fast]"
pytest -q # 163 passed, 1 xfailed
python examples/demo.py
ruff check src tests examples
Sources of truth for the parts copied from Xet:
xet-core —
xet_data/src/deduplication/chunking.rs (chunker) and
xet_core_structures/src/xorb_object/constants.rs (xorb/chunk sizes), plus the
gearhash crate for the lookup table.
License
Apache-2.0. See LICENSE.
Metadata
Release files for nano-xet 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| nano_xet-0.1.0.tar.gz | 42.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| nano_xet-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 74.3 kB
Release files / nano_xet-0.1.0.tar.gz
| Download URL | nano_xet-0.1.0.tar.gz |
|---|---|
| Size | 42.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
9b52b295386b85ef52b18112ed2518abad2bbc5d78e8ab27fab5e029abf401d4
|
|
BLAKE2b-256 checksum How to use checksums |
b995874615165bae166c0ad3818c770c50d63853607765a7ae83d6045584eb11
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.2
|
Release files / nano_xet-0.1.0-py3-none-any.whl
| Download URL | nano_xet-0.1.0-py3-none-any.whl |
|---|---|
| Size | 32.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7c25e16a667293ea84349b26248781705876b903a1cfcdbb60fcd308d330bba9
|
|
BLAKE2b-256 checksum How to use checksums |
be090f13a9d80737315170d9b00943c1591d516e126d2fe4ae3bdd9923d1e017
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.2
|