Skip to main content

konan logo

konan

Like the paper angel of the Akatsuki, konan folds your documents into precise pieces — Rust chunkers behind a Python API.

PyPI CI python-ci Python 3.12+ Rust License: MIT


Why konan?

  • 🦀 Rust core — all chunking runs in native code, no Python-loop overhead
  • Multithreaded by defaultchunk_many() fans out across all cores via rayon and releases the GIL; large single documents split at paragraph breaks and fan out too
  • 🌀 Real asyncchunk_async / chunk_many_async are native async def, not thread-pool wrappers
  • 🎯 Char-accurate offsetstext[chunk.start:chunk.end] == chunk.text, always (Python slicing semantics, emoji-safe)
  • 🧠 Semantic chunking — splits on topic shifts using any OpenAI-compatible embeddings endpoint, or your own async embedder
  • 🔌 Ports & adapters — the Embedder port is injectable; bring your own backend

Strategies

Chunker Splits by Best for
NaiveChunker fixed word count quick & dirty baselines
FixedSizeChunker chars, sentence-aware, overlap classic RAG pipelines
RecursiveChunker separator hierarchy (\n\n\n → … ) general text, LangChain-compatible
SentenceChunker unicode sentence boundaries prose, multilingual text
MarkdownChunker document structure + heading breadcrumbs docs, wikis, READMEs
TokenChunker exact token counts (cl100k_base, o200k_base) embedding-model token limits
SemanticChunker embedding similarity drops topic-coherent chunks

Installation

uv add konan        # or: pip install konan

Quickstart

from konan import RecursiveChunker

chunker = RecursiveChunker(chunk_size=1000, chunk_overlap=200)

chunks = chunker.chunk(open("moby_dick.txt").read())
print(chunks[0].text, chunks[0].start, chunks[0].end, chunks[0].hash)

Parallel — all cores, one call

# rayon work-stealing across every core, GIL released:
all_chunks = chunker.chunk_many(documents)

Async — real async def, no thread-pool wrappers

chunks = await chunker.chunk_async(text)
batches = await chunker.chunk_many_async(documents)

Markdown with breadcrumbs

from konan import MarkdownChunker

chunks = MarkdownChunker(chunk_size=800).chunk(readme_text)
# chunk text is prefixed with its heading trail: "# Guide > ## Install\n\n..."
# code fences are never split

Token-exact chunks

from konan import TokenChunker

chunker = TokenChunker(chunk_size=512, chunk_overlap=64, encoding="o200k_base")

Semantic chunking

Point it at any OpenAI-compatible /embeddings endpoint (OpenAI, vLLM, Ollama, LiteLLM, …):

from konan import OpenAIEmbedder, SemanticChunker

embedder = OpenAIEmbedder(
    base_url="https://api.openai.com/v1",
    model="text-embedding-3-small",
    api_key="sk-...",
    batch_size=128,    # texts per request
    timeout=30.0,      # request timeout, seconds
    max_retries=2,     # exponential backoff on 429/5xx/connect errors
    dimensions=512,    # optional: shorten text-embedding-3-* vectors
)
chunker = SemanticChunker(embedder=embedder, threshold=0.75)
chunks = await chunker.chunk_async(article)

Or inject your own embedder — any async callable works:

async def my_embedder(texts: list[str]) -> list[list[float]]:
    return await my_model.embed(texts)

chunker = SemanticChunker(embedder=my_embedder, percentile=95.0)
chunks = await chunker.chunk_async(article)   # async-only for Python embedders

Python embedders must return list[list[float]] — call .tolist() on numpy arrays. They are async-only: chunk()/chunk_many() raise a RuntimeError pointing you at the _async variants.

The Chunk object

chunk.text      # the chunk's text
chunk.start     # char offset into the source (Python slicing semantics)
chunk.end       # char offset, exclusive
chunk.index     # 0-based position
chunk.hash      # xxh3-64 content hash, as 16 hex digits
chunk.hash_int  # the same digest as an int, if you'd only parse the hex back

Benchmarks

Benchmarked on Apple M3 Pro (arm64), Python 3.12.10. Decimal MB/s, median of 5 runs, measured from Python (the numbers you actually get). Reproduce both tables and plots with uv run --extra bench benchmarks/bench.py; Rust-level numbers with cargo bench -p konan-core (criterion) or cargo run --release -p konan-core --example bench_min (min-of-N, which is what survives a throttling machine).

Expect ±10% between runs on a laptop that throttles — the token chunker swings up to 1.4×. These figures come from the most median-typical of 17 runs, not the best one; every row in a table is from that same run, so the comparisons hold even where the absolute numbers drift.

vs other libraries (same 1 MB document, identical configs)

konan vs other libraries

Strategy Library Throughput Chunks
recursive konan 337 MB/s 1298
recursive semantic-text-splitter 185 MB/s 1368
recursive chonkie 89 MB/s 1413
recursive langchain-text-splitters 25 MB/s 1468
token konan 140 MB/s 438
token semchunk 24 MB/s 616
token langchain-text-splitters 22 MB/s 438
token chonkie 18 MB/s 438
token semantic-text-splitter 3 MB/s 505
sentence konan 1,119 MB/s 1042
sentence chonkie 3 MB/s 2124
recursive (unicode) konan 234 MB/s 1092
recursive (unicode) semantic-text-splitter 164 MB/s 1150
recursive (unicode) chonkie 57 MB/s 1171

Throughput per strategy (1 MB document)

konan throughput per strategy

Chunker Config Throughput Chunks
NaiveChunker 200 words 1,234 MB/s 804
FixedSizeChunker 1000 chars, 200 overlap 2,417 MB/s 1255
RecursiveChunker 1000 chars, 200 overlap 339 MB/s 1298
SentenceChunker 1000 chars, 1 overlap 1,087 MB/s 1130
MarkdownChunker 1000 chars, 200 overlap 669 MB/s 1511
TokenChunker 512 tokens, 64 overlap (cl100k) 131 MB/s 438

Parallel scaling — rayon goes brrr

chunk_many parallel scaling

64 docs × 256 KB through RecursiveChunker:

Mode Time Throughput Speedup
sequential chunk() loop 49 ms 332 MB/s 1.0×
chunk_many() (rayon, GIL released) 8 ms 2,150 MB/s 6.5×

(Pure Rust puts the same workload at ~18 GB/s — 0.9 ms of the 8 ms above. The rest is building Python Chunk objects, which now dominates: the chunking itself is no longer the bottleneck.)

Caveats, honestly:

  • The sentence and token rows are multi-core; the other libraries are not. Above 128 KB (sentence) and 256 KB (token), konan splits the document at paragraph breaks and segments/encodes the pieces on all cores — a paragraph break is a mandatory boundary for both UAX #29 and the cl100k/o200k pretokenizers, so the output is unchanged. Pinned to one thread with RAYON_NUM_THREADS=1 the same 1 MB document gives 342 MB/s sentence and 72 MB/s token, with identical chunk counts. Read the table rows as wall-clock on a 12-core laptop, and those two numbers as per-core efficiency.
  • recursive/token use identical configs across libraries (1000 chars / 200 overlap; cl100k, 512 / 64 — note the identical token chunk counts for the libraries with the same windowing semantics). Sentence configs are not directly comparable (konan groups by chars, chonkie by tokens) — read those rows as per-library cost, not head-to-head.
  • The recursive (unicode) rows run mixed German/CJK/Cyrillic/emoji prose, off konan's ASCII fast path.
  • konan tokenizes with bpe-openai (tiktoken-equivalent output, much faster encoder).

Development

uv sync                       # set up the venv (builds the extension)
uv run maturin develop --uv   # rebuild after Rust changes
cargo test --workspace        # rust unit tests
uv run pytest -q              # python integration tests

The workspace is hexagonal: crates/konan-core is pure Rust (no PyO3) with Chunker and Embedder ports; crates/konan-py adapts it to Python.

License

MIT


Named after Konan of the Akatsuki — the only one who could fold paper into anything. 🗞️

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

konan-0.4.0.tar.gz (60.9 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

konan-0.4.0-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (28.5 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.17+ x86-64

konan-0.4.0-cp312-abi3-macosx_11_0_arm64.whl (28.6 MB view details)

Uploaded CPython 3.12+macOS 11.0+ ARM64

File details

Details for the file konan-0.4.0.tar.gz.

File metadata

  • Download URL: konan-0.4.0.tar.gz
  • Upload date:
  • Size: 60.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for konan-0.4.0.tar.gz
Algorithm Hash digest
SHA256 5448e7b6d05bc4aaae5b2e46a0ac7fde0de5a9c8a8b74c208511946937847971
MD5 ab4bb4a03098ec5ac0c581a08ee4dd85
BLAKE2b-256 952d354a9024aaeff31abceb7512e95a7658741548ca98b19bd098bd280c515b

See more details on using hashes here.

Provenance

The following attestation bundles were made for konan-0.4.0.tar.gz:

Publisher: release-please.yaml on 4thel00z/konan

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file konan-0.4.0-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for konan-0.4.0-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 dcc5a58ee4c5775c99567c96f7d95f25dff6de0d0a8c1b17c95e4a8a16947530
MD5 56400e9a0cff567850341026b3a0575d
BLAKE2b-256 ef531594863ec9c6a019ebd317bbd948716d4cbc9a9dbd5588502fa39a84287f

See more details on using hashes here.

Provenance

The following attestation bundles were made for konan-0.4.0-cp312-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release-please.yaml on 4thel00z/konan

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file konan-0.4.0-cp312-abi3-macosx_11_0_arm64.whl.

File metadata

  • Download URL: konan-0.4.0-cp312-abi3-macosx_11_0_arm64.whl
  • Upload date:
  • Size: 28.6 MB
  • Tags: CPython 3.12+, macOS 11.0+ ARM64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for konan-0.4.0-cp312-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 0d919c52c266bfd1ac0fde37a5856e63ba00c817df222bd938148a475e931ba1
MD5 abf029273dc1b23ca4ccd475fc39cde2
BLAKE2b-256 64998f89fc0fd336af466870ad5822630df4b98d848d38fcce758edbb693196a

See more details on using hashes here.

Provenance

The following attestation bundles were made for konan-0.4.0-cp312-abi3-macosx_11_0_arm64.whl:

Publisher: release-please.yaml on 4thel00z/konan

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.4.0 This release

3 files

0.3.0

3 files

0.2.4

3 files

0.2.3

3 files

0.2.2

3 files

0.2.1

3 files

0.2.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page