semantic-chunker
A strongly-typed semantic text chunking library that intelligently splits content while preserving structure and meaning. Built on top of semantic-text-splitter with an enhanced type-safe API.
Features
- 🎯 Multiple tokenization strategies:
- OpenAI's tiktoken models (e.g., "gpt-3.5-turbo")
- Hugging Face tokenizers (from objects, JSON strings, or files)
- Custom tokenization callbacks
- 📝 Three specialized chunking modes:
- Plain text
- Markdown (preserves structure)
- Code (preserves syntax via tree-sitter)
- 🔄 Configurable chunk overlapping
- ✂️ Optional whitespace trimming
- 💪 Full type safety with Protocol types
Installation
Basic installation (text and markdown support):
pip install semantic-chunker
With code chunking support:
pip install semantic-chunker[code]
With Hugging Face tokenizers support:
pip install semantic-chunker[tokenizers]
With all features:
pip install semantic-chunker[all]
Usage
Text Chunking
from semantic_chunker import get_chunker
plain_text = """Contrary to popular belief, Lorem Ipsum is not simply random text. ..."""
chunker = get_chunker(
"gpt-3.5-turbo",
chunking_type="text", # required
max_tokens=10, # required
trim=False, # default True
overlap=5, # default 0
)
chunks = chunker.chunks(plain_text) # list[str]
chunk_with_indices = chunker.chunk_with_indices(plain_text) # list[tuple[str, int]]
Markdown Chunking
from semantic_chunker import get_chunker
markdown_text = """# Lorem Ipsum Intro ..."""
chunker = get_chunker(
"gpt-3.5-turbo",
chunking_type="markdown",
max_tokens=10,
trim=False,
overlap=5,
)
chunks = chunker.chunks(markdown_text) # list[str]
chunk_with_indices = chunker.chunk_with_indices(markdown_text) # list[tuple[str, int]]
Code Chunking
from semantic_chunker import get_chunker
kotlin_snippet = """import kotlin.random.Random ..."""
chunker = get_chunker(
"gpt-3.5-turbo",
chunking_type="code",
max_tokens=10,
tree_sitter_language="kotlin", # required for code chunking
trim=False,
overlap=5,
)
chunks = chunker.chunks(kotlin_snippet) # list[str]
chunk_with_indices = chunker.chunk_with_indices(kotlin_snippet) # list[tuple[str, int]]
Error Handling
# Missing language for code chunking
try:
chunker = get_chunker("gpt-4", chunking_type="code", max_tokens=10)
except ValueError as e:
print(e) # "Language must be provided for code chunking."
# Missing required package for code chunking
try:
chunker = get_chunker("gpt-4", chunking_type="code", tree_sitter_language="python", max_tokens=10)
except ModuleNotFoundError as e:
print(e) # "tree-sitter-language-pack is required for 'code' style chunking..."
Chunking Type and Tokenization Options
-
get_chunkerrequires the first argument to be one of:- A tiktoken model name string (e.g.,
gpt-4o) - A function that takes a string and returns a token count (integer)
- A
tokenizers.Tokenizerinstance - A string path to a
tokenizerstokenizer JSON file
- A tiktoken model name string (e.g.,
-
Required kwargs:
chunking_type: Eithertext,markdown, orcode.max_tokens: Maximum tokens per chunk. Accepts an integer or a tuple (min, max).- If
chunking_typeiscode,tree_sitter_languageis required.
Contribution
This library is open to contribution. Feel free to open issues or submit PRs. Its better to discuss issues before submitting PRs to avoid disappointment.
Local Development
- Clone the repo
- Install the system dependencies
- Install the full dependencies with
uv sync - Install the pre-commit hooks with:
pre-commit install && pre-commit install --hook-type commit-msg
- Make your changes and submit a PR
License
This library uses the MIT license.
Metadata
Release files for semantic-chunker 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| semantic_chunker-0.2.0.tar.gz | 5.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| semantic_chunker-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 10.7 kB
Release files / semantic_chunker-0.2.0.tar.gz
| Download URL | semantic_chunker-0.2.0.tar.gz |
|---|---|
| Size | 5.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1c297289e338e253b046bd141938721bdcfca1537b1d79466b50a098048e0b43
|
|
BLAKE2b-256 checksum How to use checksums |
a1cfec861c40371852ee515bce9066ccd2e53b40f48bb7b063aec8b5892f00ca
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.8
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Feb 6, 2025.
Transparency logRelease files / semantic_chunker-0.2.0-py3-none-any.whl
| Download URL | semantic_chunker-0.2.0-py3-none-any.whl |
|---|---|
| Size | 5.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
cfe0f00cf5be8415d3250a9e7c84fbf1bd0d3b49ed398fa4fbbd71383bdd3727
|
|
BLAKE2b-256 checksum How to use checksums |
16a80c07b04b1d2136bfd454b4d37b5ea8a6b69bc04e0f0349042cc9de0071f5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.8
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Feb 6, 2025.
Transparency log