Skip to main content

nupunkt

A high-precision, high-throughput sentence boundary detection library optimized for legal text processing, with zero runtime dependencies.

0.7.0: deterministic output, a 25 KB bundled model that loads in milliseconds, a 4-5x faster tokenizer, and a standard segmentation interface for words, sentences and paragraphs. See the changelog for details and migration notes.

PyPI version Python Version License

Overview

nupunkt is a next-generation implementation of the Punkt algorithm specifically optimized for legal text processing. It accurately detects sentence boundaries in complex legal documents where periods are used for abbreviations, citations, and other non-sentence-ending contexts.

Key features:

  • Zero dependencies: Pure Python 3.11+ (tqdm optional for progress bars)
  • Adaptive mode: an optional confidence-based variant with a tunable precision/recall threshold
  • High precision: 91.1% precision on legal text benchmarks
  • High performance: tens of millions of characters per second on standard CPU hardware, with a 25 KB model
  • Pre-trained model: Ready to use with legal-optimized abbreviations
  • Trainable: Can be trained on domain-specific text
  • Paragraph detection: Split text into both sentences and paragraphs
  • CLI tools: Complete command-line interface for training and evaluation

Paper

For the research behind this implementation, see:

Precise Legal Sentence Boundary Detection for Retrieval at Scale: NUPunkt and CharBoundary
Michael J Bommarito, Daniel Martin Katz, Jillian Bommarito
arXiv:2504.04131 [cs.CL]
https://arxiv.org/abs/2504.04131

Interactive demo available at: https://sentences.aleainstitute.ai/

Installation

pip install nupunkt

Quick Start

from nupunkt import sent_tokenize

text = """
Employee also specifically and forever releases the Acme Inc. (Company) and the Company Parties (except where and 
to the extent that such a release is expressly prohibited or made void by law) from any claims based on unlawful 
employment discrimination or harassment, including, but not limited to, the Federal Age Discrimination in 
Employment Act (29 U.S.C. § 621 et. seq.). This release does not include Employee's right to indemnification, 
and related insurance coverage, under Sec. 7.1.4 or Ex. 1-1 of the Employment Agreement.
"""

# Tokenize into sentences
sentences = sent_tokenize(text)

for i, sentence in enumerate(sentences, 1):
    print(f"Sentence {i}: {sentence}\n")

Segmentation Interface

Words, sentences and paragraphs share one interface. For each level there are three list functions and a reusable segmenter object with generator forms:

Level Strings Spans Both
word words(text) word_spans(text) word_segments(text)
sentence sentences(text) sentence_spans(text) sentence_segments(text)
paragraph paragraphs(text) paragraph_spans(text) paragraph_segments(text)
from nupunkt import sentences, sentence_spans, sentence_segments, segmenter, contiguous

text = "Dr. Smith arrived.  He left at 5 p.m.\n\nThe end."

sentences(text)         # ['Dr. Smith arrived.', 'He left at 5 p.m.', 'The end.']
sentence_spans(text)    # [(0, 18), (20, 37), (39, 47)]
sentence_segments(text) # [Segment(text='Dr. Smith arrived.', start=0, end=18), ...]

# Reusable object with generators: iter_segments / iter_texts / iter_spans
seg = segmenter("paragraph")
for para in seg.iter_segments(text):
    print(para.start, para.end, para.text)

Segment is a named tuple (text, start, end) with .span for (start, end). Spans are tight: text[start:end] is exactly the segment with no surrounding whitespace, segments never overlap, and whitespace stays in the gaps. When you need gap-free coverage of the whole input, use contiguous(segments, text).

Sentence functions accept model= and adaptive=; segmenter("sentence", adaptive=True) returns the adaptive tokenizer. The older sent_tokenize, sent_spans, para_tokenize family remains available unchanged.

All levels in one pass

segment(text) runs sentence segmentation once and returns a Document tree (paragraphs, then sentences, then words). Paragraphs come from the same sentence boundaries, and words are only computed for a sentence when you read them:

from nupunkt import segment

doc = segment(text)
for para in doc.paragraphs:
    for sent in para.sentences:
        print(sent.start, sent.end, [w.text for w in sent.words])

doc.sentences   # flat list, == sentence_segments(text)
doc.words       # flat list, == word_segments(text)
doc.to_dict()   # JSON-ready nested dict; to_dict(words=False) omits words

Every node is a Segment (it unpacks as (text, start, end) and has .span), and each word lies inside its sentence, which lies inside its paragraph. For paragraphs plus sentences, segment() costs one sentence pass where paragraph_segments() + sentence_segments() cost two.

Adaptive Tokenization (New in v0.6.0)

Adaptive mode dynamically discovers abbreviation patterns and improves sentence boundary detection:

from nupunkt import sent_tokenize_adaptive

text = """Dr. Smith graduated from M.I.T. in 2020. She works at N.A.S.A. now.
Her colleague Mr. Johnson has a Ph.D. from U.C.L.A. and collaborates with researchers
at C.E.R.N. on quantum physics."""

# Use adaptive mode with abbreviation pattern detection
sentences = sent_tokenize_adaptive(text)

# Adjust confidence threshold (default: 0.7)
sentences = sent_tokenize_adaptive(text, threshold=0.8)

# Get confidence scores for each decision
sentences_with_scores = sent_tokenize_adaptive(text, return_confidence=True)
for sentence, confidence in sentences_with_scores:
    print(f"[{confidence:.2f}] {sentence}")

The adaptive tokenizer:

  • Automatically detects abbreviation patterns (M.I.T., Ph.D., etc.)
  • Uses context clues to make better decisions
  • Provides confidence scores for each boundary decision
  • Falls back to the robust base algorithm when uncertain

Sentence and Paragraph Spans

Get character-level spans for sentences and paragraphs:

from nupunkt import sent_spans, sent_spans_with_text, para_spans, para_spans_with_text

# Get sentence spans (start, end positions)
sentence_spans = sent_spans(text)

# Get sentences with their spans
sentences_with_spans = sent_spans_with_text(text)
for sentence, (start, end) in sentences_with_spans:
    print(f"[{start}:{end}] {sentence}")

# Same for paragraphs
paragraph_spans = para_spans(text)
paragraphs_with_spans = para_spans_with_text(text)

Adaptive Spans

Get spans using the adaptive algorithm for better abbreviation handling:

from nupunkt import sent_spans_adaptive, sent_spans_with_text_adaptive

# Get adaptive sentence spans
text = "Dr. Smith studied at M.I.T. in Cambridge."
spans = sent_spans_adaptive(text)
# Returns: [(0, 41)] - single sentence preserved

# Get sentences with spans
results = sent_spans_with_text_adaptive(text)
for sentence, (start, end) in results:
    print(f"[{start}:{end}] {sentence}")

# With confidence scores
results = sent_spans_with_text_adaptive(text, return_confidence=True)
for sentence, (start, end), confidence in results:
    print(f"[{confidence:.2f}] [{start}:{end}] {sentence}")

All span functions guarantee:

  • Contiguous spans with no gaps
  • Full coverage of the input text
  • Preservation of all whitespace

Paragraph Detection

from nupunkt import para_tokenize

# Get paragraph text
paragraphs = para_tokenize(text)

Command-line Interface

Basic usage

# Using Python directly
echo "Hello world. How are you?" | python -c "import sys; from nupunkt import sent_tokenize; print('\n'.join(sent_tokenize(sys.stdin.read())))"

# Or create a simple script
python -c "from nupunkt import sent_tokenize; import sys; [print(s) for s in sent_tokenize(sys.stdin.read())]"

Training models

# Train from text files
nupunkt train corpus.txt --output model.bin

# Train from HuggingFace datasets
nupunkt train hf:alea-institute/kl3m-data-usc -o legal_model.bin

# Memory-efficient training for large datasets
nupunkt train huge_corpus.txt --batch-size 1000000 --min-type-freq 5

Evaluating models

# Evaluate a model
nupunkt evaluate test_data.jsonl -m my_model.bin

# Compare multiple models
nupunkt evaluate test_data.jsonl --compare --models baseline.bin custom.bin

Model management

# Convert between formats
nupunkt convert model.json model.bin

# Get model information
nupunkt info model.bin

# Optimize hyperparameters
nupunkt optimize-params train.jsonl test.jsonl -o best_model.bin

Performance

nupunkt is designed for high-precision, high-throughput processing with a tiny footprint:

  • Bundled model: 25 KB, loads in about 2 ms, about 20 MB resident memory after load
  • Deterministic: the same input always produces the same output, in any order
  • Fast path processing for texts without sentence boundaries
  • Memoized token properties and a lean per-boundary decision path

Typical throughput on legal text is 15-25 million characters per second per core on a modern machine (measured on a 4 MB corpus with the default model; numbers vary with hardware and text). A reusable benchmark lives in scripts/profiling/.

Documentation

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

This project is licensed under the MIT License - see the LICENSE file for details.

Citation

If you use nupunkt in your research, please cite:

@article{bommarito2025precise,
  title={Precise Legal Sentence Boundary Detection for Retrieval at Scale: NUPunkt and CharBoundary},
  author={Bommarito, Michael J and Katz, Daniel Martin and Bommarito, Jillian},
  journal={arXiv preprint arXiv:2504.04131},
  year={2025}
}

Acknowledgments

nupunkt is based on the Punkt algorithm originally developed by Tibor Kiss and Jan Strunk.

Metadata

Release files for nupunkt 0.7.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for nupunkt 0.7.0
File Size Uploaded
nupunkt-0.7.0.tar.gz 137.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for nupunkt 0.7.0
File Interpreter ABI Platform
nupunkt-0.7.0-py3-none-any.whl Python 3 none any Details

Total release size: 279.4 kB

Release files / nupunkt-0.7.0.tar.gz

Download URL nupunkt-0.7.0.tar.gz
Size 137.9 kB
Tags Source
SHA-256 checksum
How to use checksums
9f9e303f946eb042ab7aa5fe56535dd53f74f0ce9bef0d87142ab8e4dbdc874b
BLAKE2b-256 checksum
How to use checksums
2e94c90a1d61794cb4a946ef558d41491a54646946ca300999780712c2bd404b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release files / nupunkt-0.7.0-py3-none-any.whl

Download URL nupunkt-0.7.0-py3-none-any.whl
Size 141.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8e739001bcc4d7daedfa4246f3c82b21a43469249b3402434340818884c2b743
BLAKE2b-256 checksum
How to use checksums
d662a25e4d8e77a3a3ef9390eaf8f7c09b6fc13a687094551cf34a6c75369f92
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

0.8.0

2 release files

This release

0.7.0 This release

2 release files

0.6.0

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page