chunkwise
Token-aware, boundary-respecting text and document chunking for RAG ingestion.
Part of the ragkit suite. Install with
pip install ragkit-chunkwise, thenimport chunkwise.
Naive character slicing shreds sentences and paragraphs and ignores model token
limits. chunkwise splits text along real structural boundaries (paragraphs →
lines → sentences → words → characters), packs the pieces up to a token budget,
and carries a configurable overlap between chunks — so your retrieval index gets
clean, self-contained passages instead of arbitrary fragments.
- Pure Python standard library at runtime. No required third-party dependencies.
- Python 3.8+.
- Pluggable token counting — plug in
tiktoken, a HuggingFace tokenizer, or anycallable(str) -> int. - Accurate character offsets back into the original text.
Install
pip install ragkit-chunkwise
Optional tokenizer extra (pulls in tiktoken):
pip install "ragkit-chunkwise[tokenizers]"
Local development (from chunkwise/):
pip install -e .
Quick Start
from chunkwise import chunk_text
paragraph = (
"Retrieval-augmented generation grounds a language model in your own data. "
"You first split documents into chunks, embed them, and store the vectors. "
"At query time you retrieve the most relevant chunks and feed them to the model."
)
for c in chunk_text(paragraph, chunk_size=15, chunk_overlap=4):
print(c.index, c.token_count, repr(c.text))
Each result is a Chunk with the text, its position, character offsets, and a
token count.
The effect of chunk_size and chunk_overlap
chunk_size is the maximum length of a chunk, measured by your
length_function (words by default). chunk_overlap is how much of the tail of
one chunk is repeated at the head of the next — overlap preserves context that
would otherwise be cut off at a boundary, which improves retrieval recall.
from chunkwise import RecursiveChunker
text = " ".join(f"word{i}" for i in range(60))
# Bigger chunks, no overlap → fewer, disjoint chunks.
print(len(RecursiveChunker(chunk_size=30, chunk_overlap=0).split_text(text))) # ~2
# Smaller chunks with overlap → more chunks that share context.
print(len(RecursiveChunker(chunk_size=15, chunk_overlap=5).split_text(text))) # more
chunk_overlap must be strictly less than chunk_size or a ValueError is
raised.
API Reference
RecursiveChunker
RecursiveChunker(
chunk_size=512,
chunk_overlap=64,
separators=None, # default: ["\n\n", "\n", ". ", " ", ""]
length_function=word_token_counter,
keep_separator=True,
)
Recursively splits text using an ordered list of separators, trying the largest
structural boundary first. If a piece still exceeds chunk_size, it recurses
with the next separator; the final "" separator splits by characters as a last
resort. Adjacent small pieces are greedily merged up to chunk_size, and
chunk_overlap tokens from the previous chunk's tail are carried forward.
Methods:
.split_text(text) -> List[str]— return chunk strings..chunk(text, metadata=None) -> List[Chunk]— returnChunkobjects with accurate character offsets, token counts, sequential indexes, andmetadatamerged into each chunk.
chunk_text(text, chunk_size=512, chunk_overlap=64, **kwargs) -> List[Chunk]
Convenience wrapper around RecursiveChunker. Extra keyword arguments
(separators, length_function, keep_separator) are forwarded to the
chunker. An optional metadata= keyword is attached to every chunk.
SentenceChunker
SentenceChunker(chunk_size=512, chunk_overlap=64, length_function=word_token_counter)
Splits text into sentences with a lightweight regex (handles ., !, ?
followed by whitespace or end-of-string), then packs whole sentences into chunks
up to chunk_size with sentence-level overlap. Sentences are never cut mid-way.
Exposes .split_text(text) and .chunk(text, metadata=None).
chunk_markdown(text, chunk_size=512, chunk_overlap=64, length_function=word_token_counter) -> List[Chunk]
Splits markdown into sections at headings (lines starting with #). The trail of
active headings (outermost → innermost) is attached to each chunk's metadata
under the "headings" key, and oversized sections are further split with
RecursiveChunker.
from chunkwise import chunk_markdown
md = "# Guide\nIntro.\n\n## Setup\nInstall the package and configure it."
for c in chunk_markdown(md, chunk_size=50):
print(c.metadata["headings"], "->", repr(c.text))
# ['Guide'] -> '# Guide\nIntro.\n'
# ['Guide', 'Setup'] -> '## Setup\nInstall the package and configure it.'
Custom length_function
Any callable(str) -> int works. Built-ins:
word_token_counter(text)— whitespace-split word count (default).char_token_counter(text)—len(text).
To chunk by real model tokens, plug in a tokenizer:
import tiktoken
enc = tiktoken.get_encoding("cl100k_base")
from chunkwise import RecursiveChunker
chunker = RecursiveChunker(
chunk_size=256,
chunk_overlap=32,
length_function=lambda t: len(enc.encode(t)),
)
chunks = chunker.chunk(my_document)
The Chunk dataclass
| Field | Type | Meaning |
|---|---|---|
text |
str |
The chunk text. |
index |
int |
Position in the output sequence (0-based). |
start |
int |
Character offset of the chunk's primary span in the original. |
end |
int |
End character offset (exclusive). |
token_count |
int |
Length of text per the length_function. |
metadata |
dict |
User metadata (defaults to {}). |
len(chunk) returns token_count.
Offsets and overlap: for RecursiveChunker (and the sentence/markdown
chunkers built on the same logic), original_text[chunk.start:chunk.end] == chunk.text holds even when overlap is used — successive chunks simply share an
overlapping character range. start/end always describe the contiguous
primary span a chunk covers.
Design notes / correctness
- No chunk exceeds
chunk_sizeby more than a single indivisible unit (e.g. one very long word or a single sentence longer than the budget). - Empty or whitespace-only input returns
[]. - Overlap never produces an infinite loop; an indivisible piece larger than the budget is isolated rather than re-seeding subsequent chunks.
Running the tests
python -m unittest discover -s tests -v
License
MIT
Metadata
Release files for ragkit-chunkwise 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ragkit_chunkwise-0.1.1.tar.gz | 13.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ragkit_chunkwise-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 23.7 kB
Release files / ragkit_chunkwise-0.1.1.tar.gz
| Download URL | ragkit_chunkwise-0.1.1.tar.gz |
|---|---|
| Size | 13.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0f681520398d6000dc080f6f04810923f9ca58e96b85bb14ec09c571fd5a7e2e
|
|
BLAKE2b-256 checksum How to use checksums |
a4d39b10cded6bbba025c7b4a2aedd293ab0df317945dcd7b3f57c930e238b4c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 25, 2026.
Transparency logRelease files / ragkit_chunkwise-0.1.1-py3-none-any.whl
| Download URL | ragkit_chunkwise-0.1.1-py3-none-any.whl |
|---|---|
| Size | 9.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d11a37673b94086cc415c4ac43e2670251dbcbdf5b8c37d776163b07d77bf6ff
|
|
BLAKE2b-256 checksum How to use checksums |
7b8062677972b84a22c90d5d807a42dad06de84ab25afdca1dddfc25fa246e4b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 25, 2026.
Transparency log