Skip to main content

llm-semantic-chunker

A Python library for LLM-based semantic chunking, designed for Retrieval-Augmented Generation (RAG) pipelines. Instead of splitting text at a fixed character count, it uses a local LLM to decide where a topic actually changes and groups sentences accordingly.

The library runs entirely against a local Ollama model — no API keys, no data leaving your machine.

Links: PyPI · Source on GitHub · Issues


Installation

pip install llm-semantic-chunker

That is the whole install — httpx, nltk and langdetect. No vector store, no torch. On first use the sentence splitter downloads NLTK's punkt data once.

Requires Ollama running locally with a compatible model, here:

ollama pull qwen3.5:4b

Optional extras, only if you want them:

pip install "llm-semantic-chunker[langchain]"    # LangChain adapter
pip install "llm-semantic-chunker[llamaindex]"   # LlamaIndex adapter (needs Python 3.10+)
pip install "llm-semantic-chunker[pdf]"          # read .pdf input

Quick Start

from llm_semantic_chunker import LLMChunker, OllamaClient

TEXT = (
    "The sun is a star at the center of the solar system. It is a nearly "
    "perfect sphere of hot plasma, heated to incandescence by nuclear fusion "
    "in its core. Its diameter is about 1.39 million kilometres, roughly 109 "
    "times that of Earth. "
    "Dolphins are highly intelligent marine mammals. They live in social "
    "groups called pods and use echolocation to navigate and hunt. Some "
    "species have been observed teaching their young to use tools."
)

chunker = LLMChunker(client=OllamaClient(), mode="incremental", max_chunk_chars=1200)

for i, chunk in enumerate(chunker.chunk(TEXT), 1):
    print(f"--- Chunk {i} ---")
    print(chunk)

The LLM splits between the two topics rather than at a character count:

--- Chunk 1 ---
The sun is a star at the center of the solar system. It is a nearly perfect
sphere of hot plasma, heated to incandescence by nuclear fusion in its core.
Its diameter is about 1.39 million kilometres, roughly 109 times that of Earth.
--- Chunk 2 ---
Dolphins are highly intelligent marine mammals. They live in social groups
called pods and use echolocation to navigate and hunt. Some species have been
observed teaching their young to use tools.

Use it inside LangChain

pip install "llm-semantic-chunker[langchain]"

The adapter implements LangChain's TextSplitter, so it goes wherever RecursiveCharacterTextSplitter goes — only the splitting step changes, the rest of the pipeline is untouched.

from langchain_core.documents import Document
from llm_semantic_chunker.integrations.langchain import LLMSemanticSplitter

TEXT = ("The sun is a star at the center of the solar system. It is a nearly "
        "perfect sphere of hot plasma. Dolphins are highly intelligent marine "
        "mammals. They live in social groups called pods.")


splitter = LLMSemanticSplitter(max_chunk_chars=1200)

docs: list[Document] = splitter.create_documents([TEXT])
for d in docs:
    print(d.page_content)

split_text(), split_documents() and transform_documents() work as well.


Use it inside LlamaIndex

pip install "llm-semantic-chunker[llamaindex]" — requires Python 3.10 or newer ; the rest of this library still runs on Python 3.9.

from llama_index.core import Document
from llm_semantic_chunker.integrations.llamaindex import LLMSemanticNodeParser

TEXT = ("The sun is a star at the center of the solar system. It is a nearly "
        "perfect sphere of hot plasma. Dolphins are highly intelligent marine "
        "mammals. They live in social groups called pods.")


parser = LLMSemanticNodeParser(max_chunk_chars=1200)

nodes = parser.get_nodes_from_documents([Document(text=TEXT)])
for n in nodes:
    print(n.text)

All three routes return the same chunks — only the object type differs.


How it works

The library ships two chunking strategies, selected via mode.

mode="incremental" (standard)

The LLM reads the document sentence by sentence and decides, for each new group of sentences, whether it still belongs to the chunk being built or starts a new one.

Four settings shape the result on top of that:

  • Heading awareness — heading_mode picks how section headings are found, and a detected heading forces a boundary. In "regex" mode the boundary sits exactly at the heading sentence; in "lines" and "hybrid" mode it sits in front of the sentence group (step_sentences) that contains the heading, so with step_sentences=2 the group's first sentence may precede the heading:
    • "regex" (default) — a sentence-level pattern, applied after sentence splitting
    • "lines" — a stronger line-based pattern applied to the raw text before sentence splitting, so numbered headings like "3.2. Error Handling" survive tokenisation
    • "hybrid" — the line-based pattern, plus an LLM check for short sentence groups (at most 90 characters and 12 words) the pattern did not flag. A heading that shares its group with a full sentence is not checked, so on well-formatted documents the LLM adds little over "lines"
  • Size cap — max_chunk_sentences / max_chunk_chars force a split once a chunk outgrows the limit, even if the topic continues. The cut never falls inside a sentence: the LLM picks the best sentence boundary, and the chunker walks it back until the piece fits. A single sentence longer than the cap therefore stays whole — the cap is a target, not a guarantee.
  • Low-info filter — a post-processing pass removes chunks that turned out to be near-empty boilerplate rather than actual content.
  • Topic enrichment — enrich=True prefixes every chunk with an LLM-generated [Topic: ...] line, so the embedding also carries where the chunk sits in the document. Off by default: it costs one extra LLM call per chunk.

mode="window" (legacy)

An earlier, two-pass approach: the text is pre-split into fixed-size mini-chunks, a sliding window over them proposes coarse boundaries. Only kept for comparison — without a size cap it degenerates into a few very large chunks.


Configuration

Settings live in ChunkerConfig. Pass one explicitly, or give the individual settings to LLMChunker and one is built for you — both are equivalent:

from llm_semantic_chunker import ChunkerConfig, LLMChunker, OllamaClient

# short form
LLMChunker(client=OllamaClient(), max_chunk_chars=1200)

# explicit — useful when you want to reuse, compare or log the settings
config = ChunkerConfig(max_chunk_chars=1200)
chunker = LLMChunker(client=OllamaClient(), config=config)
chunker.config.max_chunk_chars      # 1200

ChunkerConfig is frozen and validates itself, so a bad value fails before the run. Passing a config and individual settings at the same time is refused, because it would be ambiguous which one wins.

ChunkerConfig(
    mode="incremental",            # "incremental" (recommended) or "window" 

    # --- incremental mode ---
    step_sentences=3,              # sentences considered per boundary decision
    max_chunk_sentences=20,        # hard cap regardless of topic continuity
    max_chunk_chars=None,          # character cap, applied at sentence boundaries;
                                   # None = no cap
    respect_headings=True,         # force a boundary at detected section headings
    heading_mode="regex",          # "regex", "lines" or "hybrid"

    smart_split=True,              # at the size cap, let the LLM pick the split
                                   # point instead of cutting in the middle

    # --- window mode ---
    window_size=10,                # mini-chunks visible to the LLM per boundary decision
    step_size=5,                   # how far the window advances each iteration

    # --- shared ---
    filter_low_info=True,          # drop low-info chunks after assembly
    enrich=False,                  # prefix each chunk with an LLM-generated topic line
    language=None,                 # sentence-splitter language; auto-detected if None
    verbose=False,                 # attach a DEBUG console handler to the
                                   # package logger
)

OllamaClient

OllamaClient(
    model="qwen3.5:4b",
    base_url="http://localhost:11434",   # or set OLLAMA_BASE_URL
    temperature=0.0,                     # near-deterministic decoding
    seed=42,                             # temperature=0 alone is not bit-exact on
                                         # Ollama; the fixed seed made three full
                                         # runs identical (measured)
    timeout=600.0,                       # read timeout in seconds
    num_ctx=4096,                        # context window; lower saves RAM
)

The repository additionally holds the evaluation harness behind the bachelor thesis this library was written for, together with the documents and question sets it was measured on — see the README on GitHub.


License

MIT — see LICENSE.

Metadata

Release files for llm-semantic-chunker 1.0.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-semantic-chunker 1.0.2
File Size Uploaded
llm_semantic_chunker-1.0.2.tar.gz 31.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-semantic-chunker 1.0.2
File Interpreter ABI Platform
llm_semantic_chunker-1.0.2-py3-none-any.whl Python 3 none any Details

Total release size: 64.7 kB

Release files / llm_semantic_chunker-1.0.2.tar.gz

Download URL llm_semantic_chunker-1.0.2.tar.gz
Size 31.8 kB
Tags Source
SHA-256 checksum
How to use checksums
eb69b3b1efa2de2ce6aaf87bca19d034e63f05e20459558adf396f50e7142bae
BLAKE2b-256 checksum
How to use checksums
5cb326e020bc638dd128dbe01ddae59275b29b94ce21e13bf259753190839580
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.13

Release files / llm_semantic_chunker-1.0.2-py3-none-any.whl

Download URL llm_semantic_chunker-1.0.2-py3-none-any.whl
Size 32.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e649585c977736a202aa48ee6e3b567ce83e409ee50811103a922c3bef1fd537
BLAKE2b-256 checksum
How to use checksums
0484ff7b6a5802c3cde08e8f56358152e8ea13b872ff90e76b9cb75eca891dde
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.13

Release history Release notifications | RSS feed

This release

1.0.2 This release

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page