konan
Like the paper angel of the Akatsuki, konan folds your documents into precise pieces — Rust chunkers behind a Python API.
Why konan?
- 🦀 Rust core — all chunking runs in native code, no Python-loop overhead
- ⚡ Multithreaded by default —
chunk_many()fans out across all cores via rayon and releases the GIL; large single documents split at paragraph breaks and fan out too - 🌀 Real async —
chunk_async/chunk_many_asyncare nativeasync def, not thread-pool wrappers - 🎯 Char-accurate offsets —
text[chunk.start:chunk.end] == chunk.text, always (Python slicing semantics, emoji-safe) - 🧠 Semantic chunking — splits on topic shifts using any OpenAI-compatible embeddings endpoint, or your own async embedder
- 🔌 Ports & adapters — the
Embedderport is injectable; bring your own backend
Strategies
| Chunker | Splits by | Best for |
|---|---|---|
NaiveChunker |
fixed word count | quick & dirty baselines |
FixedSizeChunker |
chars, sentence-aware, overlap | classic RAG pipelines |
RecursiveChunker |
separator hierarchy (\n\n → \n → … ) |
general text, LangChain-compatible |
SentenceChunker |
unicode sentence boundaries | prose, multilingual text |
MarkdownChunker |
document structure + heading breadcrumbs | docs, wikis, READMEs |
TokenChunker |
exact token counts (cl100k_base, o200k_base) |
embedding-model token limits |
SemanticChunker |
embedding similarity drops | topic-coherent chunks |
Installation
uv add konan # or: pip install konan
Quickstart
from konan import RecursiveChunker
chunker = RecursiveChunker(chunk_size=1000, chunk_overlap=200)
chunks = chunker.chunk(open("moby_dick.txt").read())
print(chunks[0].text, chunks[0].start, chunks[0].end, chunks[0].hash)
Parallel — all cores, one call
# rayon work-stealing across every core, GIL released:
all_chunks = chunker.chunk_many(documents)
Async — real async def, no thread-pool wrappers
chunks = await chunker.chunk_async(text)
batches = await chunker.chunk_many_async(documents)
Markdown with breadcrumbs
from konan import MarkdownChunker
chunks = MarkdownChunker(chunk_size=800).chunk(readme_text)
# chunk text is prefixed with its heading trail: "# Guide > ## Install\n\n..."
# code fences are never split
Token-exact chunks
from konan import TokenChunker
chunker = TokenChunker(chunk_size=512, chunk_overlap=64, encoding="o200k_base")
Semantic chunking
Point it at any OpenAI-compatible /embeddings endpoint (OpenAI, vLLM,
Ollama, LiteLLM, …):
from konan import OpenAIEmbedder, SemanticChunker
embedder = OpenAIEmbedder(
base_url="https://api.openai.com/v1",
model="text-embedding-3-small",
api_key="sk-...",
batch_size=128, # texts per request
timeout=30.0, # request timeout, seconds
max_retries=2, # exponential backoff on 429/5xx/connect errors
dimensions=512, # optional: shorten text-embedding-3-* vectors
)
chunker = SemanticChunker(embedder=embedder, threshold=0.75)
chunks = await chunker.chunk_async(article)
Or inject your own embedder — any async callable works:
async def my_embedder(texts: list[str]) -> list[list[float]]:
return await my_model.embed(texts)
chunker = SemanticChunker(embedder=my_embedder, percentile=95.0)
chunks = await chunker.chunk_async(article) # async-only for Python embedders
Python embedders must return
list[list[float]]— call.tolist()on numpy arrays. They are async-only:chunk()/chunk_many()raise aRuntimeErrorpointing you at the_asyncvariants.
The Chunk object
chunk.text # the chunk's text
chunk.start # char offset into the source (Python slicing semantics)
chunk.end # char offset, exclusive
chunk.index # 0-based position
chunk.hash # xxh3-64 content hash, as 16 hex digits
chunk.hash_int # the same digest as an int, if you'd only parse the hex back
Benchmarks
Benchmarked on Apple M3 Pro (arm64), Python 3.12.10. Decimal MB/s, median
of 5 runs, measured from Python (the numbers you actually get). Reproduce
both tables and plots with uv run --extra bench benchmarks/bench.py;
Rust-level numbers with cargo bench -p konan-core (criterion) or
cargo run --release -p konan-core --example bench_min (min-of-N, which is
what survives a throttling machine).
Expect ±10% between runs on a laptop that throttles — the token chunker swings up to 1.4×. These figures come from the most median-typical of 17 runs, not the best one; every row in a table is from that same run, so the comparisons hold even where the absolute numbers drift.
vs other libraries (same 1 MB document, identical configs)
| Strategy | Library | Throughput | Chunks |
|---|---|---|---|
| recursive | konan | 337 MB/s | 1298 |
| recursive | semantic-text-splitter | 185 MB/s | 1368 |
| recursive | chonkie | 89 MB/s | 1413 |
| recursive | langchain-text-splitters | 25 MB/s | 1468 |
| token | konan | 140 MB/s | 438 |
| token | semchunk | 24 MB/s | 616 |
| token | langchain-text-splitters | 22 MB/s | 438 |
| token | chonkie | 18 MB/s | 438 |
| token | semantic-text-splitter | 3 MB/s | 505 |
| sentence | konan | 1,119 MB/s | 1042 |
| sentence | chonkie | 3 MB/s | 2124 |
| recursive (unicode) | konan | 234 MB/s | 1092 |
| recursive (unicode) | semantic-text-splitter | 164 MB/s | 1150 |
| recursive (unicode) | chonkie | 57 MB/s | 1171 |
Throughput per strategy (1 MB document)
| Chunker | Config | Throughput | Chunks |
|---|---|---|---|
NaiveChunker |
200 words | 1,234 MB/s | 804 |
FixedSizeChunker |
1000 chars, 200 overlap | 2,417 MB/s | 1255 |
RecursiveChunker |
1000 chars, 200 overlap | 339 MB/s | 1298 |
SentenceChunker |
1000 chars, 1 overlap | 1,087 MB/s | 1130 |
MarkdownChunker |
1000 chars, 200 overlap | 669 MB/s | 1511 |
TokenChunker |
512 tokens, 64 overlap (cl100k) | 131 MB/s | 438 |
Parallel scaling — rayon goes brrr
64 docs × 256 KB through RecursiveChunker:
| Mode | Time | Throughput | Speedup |
|---|---|---|---|
sequential chunk() loop |
49 ms | 332 MB/s | 1.0× |
chunk_many() (rayon, GIL released) |
8 ms | 2,150 MB/s | 6.5× |
(Pure Rust puts the same workload at ~18 GB/s — 0.9 ms of the 8 ms above.
The rest is building Python Chunk objects, which now dominates: the
chunking itself is no longer the bottleneck.)
Caveats, honestly:
- The sentence and token rows are multi-core; the other libraries are
not. Above 128 KB (sentence) and 256 KB (token), konan splits the
document at paragraph breaks and segments/encodes the pieces on all cores
— a paragraph break is a mandatory boundary for both UAX #29 and the
cl100k/o200k pretokenizers, so the output is unchanged. Pinned to one
thread with
RAYON_NUM_THREADS=1the same 1 MB document gives 342 MB/s sentence and 72 MB/s token, with identical chunk counts. Read the table rows as wall-clock on a 12-core laptop, and those two numbers as per-core efficiency. - recursive/token use identical configs across libraries (1000 chars / 200 overlap; cl100k, 512 / 64 — note the identical token chunk counts for the libraries with the same windowing semantics). Sentence configs are not directly comparable (konan groups by chars, chonkie by tokens) — read those rows as per-library cost, not head-to-head.
- The
recursive (unicode)rows run mixed German/CJK/Cyrillic/emoji prose, off konan's ASCII fast path. - konan tokenizes with
bpe-openai(tiktoken-equivalent output, much faster encoder).
Development
uv sync # set up the venv (builds the extension)
uv run maturin develop --uv # rebuild after Rust changes
cargo test --workspace # rust unit tests
uv run pytest -q # python integration tests
The workspace is hexagonal: crates/konan-core is pure
Rust (no PyO3) with Chunker and Embedder ports;
crates/konan-py adapts it to Python.
License
MIT
Named after Konan of the Akatsuki — the only one who could fold paper into anything. 🗞️
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file konan-0.4.0.tar.gz.
File metadata
- Download URL: konan-0.4.0.tar.gz
- Upload date:
- Size: 60.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5448e7b6d05bc4aaae5b2e46a0ac7fde0de5a9c8a8b74c208511946937847971
|
|
| MD5 |
ab4bb4a03098ec5ac0c581a08ee4dd85
|
|
| BLAKE2b-256 |
952d354a9024aaeff31abceb7512e95a7658741548ca98b19bd098bd280c515b
|
Provenance
The following attestation bundles were made for konan-0.4.0.tar.gz:
Publisher:
release-please.yaml on 4thel00z/konan
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
konan-0.4.0.tar.gz -
Subject digest:
5448e7b6d05bc4aaae5b2e46a0ac7fde0de5a9c8a8b74c208511946937847971 - Sigstore transparency entry: 2281539025
- Sigstore integration time:
-
Permalink:
4thel00z/konan@e56fbd8b478eb910e9f312f4a2a58fb258850afc -
Branch / Tag:
refs/heads/master - Owner: https://github.com/4thel00z
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release-please.yaml@e56fbd8b478eb910e9f312f4a2a58fb258850afc -
Trigger Event:
push
-
Statement type:
File details
Details for the file konan-0.4.0-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.
File metadata
- Download URL: konan-0.4.0-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
- Upload date:
- Size: 28.5 MB
- Tags: CPython 3.12+, manylinux: glibc 2.17+ x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
dcc5a58ee4c5775c99567c96f7d95f25dff6de0d0a8c1b17c95e4a8a16947530
|
|
| MD5 |
56400e9a0cff567850341026b3a0575d
|
|
| BLAKE2b-256 |
ef531594863ec9c6a019ebd317bbd948716d4cbc9a9dbd5588502fa39a84287f
|
Provenance
The following attestation bundles were made for konan-0.4.0-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:
Publisher:
release-please.yaml on 4thel00z/konan
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
konan-0.4.0-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl -
Subject digest:
dcc5a58ee4c5775c99567c96f7d95f25dff6de0d0a8c1b17c95e4a8a16947530 - Sigstore transparency entry: 2281539055
- Sigstore integration time:
-
Permalink:
4thel00z/konan@e56fbd8b478eb910e9f312f4a2a58fb258850afc -
Branch / Tag:
refs/heads/master - Owner: https://github.com/4thel00z
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release-please.yaml@e56fbd8b478eb910e9f312f4a2a58fb258850afc -
Trigger Event:
push
-
Statement type:
File details
Details for the file konan-0.4.0-cp312-abi3-macosx_11_0_arm64.whl.
File metadata
- Download URL: konan-0.4.0-cp312-abi3-macosx_11_0_arm64.whl
- Upload date:
- Size: 28.6 MB
- Tags: CPython 3.12+, macOS 11.0+ ARM64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0d919c52c266bfd1ac0fde37a5856e63ba00c817df222bd938148a475e931ba1
|
|
| MD5 |
abf029273dc1b23ca4ccd475fc39cde2
|
|
| BLAKE2b-256 |
64998f89fc0fd336af466870ad5822630df4b98d848d38fcce758edbb693196a
|
Provenance
The following attestation bundles were made for konan-0.4.0-cp312-abi3-macosx_11_0_arm64.whl:
Publisher:
release-please.yaml on 4thel00z/konan
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
konan-0.4.0-cp312-abi3-macosx_11_0_arm64.whl -
Subject digest:
0d919c52c266bfd1ac0fde37a5856e63ba00c817df222bd938148a475e931ba1 - Sigstore transparency entry: 2281539036
- Sigstore integration time:
-
Permalink:
4thel00z/konan@e56fbd8b478eb910e9f312f4a2a58fb258850afc -
Branch / Tag:
refs/heads/master - Owner: https://github.com/4thel00z
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release-please.yaml@e56fbd8b478eb910e9f312f4a2a58fb258850afc -
Trigger Event:
push
-
Statement type: