Skip to main content

Real-Time Sentence Detection

Real-time processing and delivery of sentences from a continuous stream of characters or text chunks.

Hint: If you're interested in state-of-the-art voice solutions you might also want to have a look at Linguflex, the original project from which stream2sentence is spun off. It lets you control your environment by speaking and is one of the most capable and sophisticated open-source assistants currently available.

Demo

https://github.com/user-attachments/assets/0428d806-4b30-47fa-9cfe-8fb42c95b2e4

CLI demo code (reproduces the video above)

Table of Contents

Features

  • Generates sentences from a stream of text in real-time.
  • Customizable to finetune/balance speed vs reliability.
  • Option to clean the output by removing links and emojis from the detected sentences.
  • Easy to configure and integrate.

Installation

pip install stream2sentence

The base installation uses stream2sentence's built-in rule-based tokenizer and does not install NLTK or Stanza. Install either tokenizer only when you need it:

pip install "stream2sentence[nltk]"
pip install "stream2sentence[stanza]"
pip install "stream2sentence[nltk,stanza]"  # both, using standard pip extras
pip install "stream2sentence[all]"          # both

Usage

Pass a generator of characters or text chunks to generate_sentences() to get a generator of sentences in return.

Here's a basic example:

from stream2sentence import generate_sentences

# Dummy generator for demonstration
def dummy_generator():
    yield "This is a sentence. And here's another! Yet, "
    yield "there's more. This ends now."

for sentence in generate_sentences(dummy_generator()):
    print(sentence)

This will output:

This is a sentence.
And here's another!
Yet, there's more.
This ends now.

One main use case of this library is enable fast text to speech synthesis in the context of character feeds generated from large language models: this library enables fastest possible access to a complete sentence or sentence fragment (using the quick_yield_single_sentence_fragment flag) that then can be synthesized in realtime. The usage of this is demonstrated in the test_stream_from_llm.py file in the tests directory.

For English streams where the optional NLTK extra is installed, the recommended configuration is tokenizer="nltk+rule-based" with auto_context=True. This combines NLTK sentence splitting with stream2sentence's local boundary checks and allows safe sentence boundaries to be yielded earlier than the fixed context window when both checks support the split.

for sentence in generate_sentences(
    text_stream,
    tokenizer="nltk+rule-based",
    language="en",
    auto_context=True,
):
    print(sentence)

Configuration

The generate_sentences() function offers various parameters to fine-tune its behavior:

Core Parameters

  • generator: Iterator[str]

    • The primary input source, yielding chunks of text to be processed.
    • Can be any iterator that emits text chunks of any size.
  • context_size: int = 12

    • Number of characters considered for sentence boundary detection.
    • Larger values improve accuracy but may increase latency.
    • Default: 12 characters
  • context_size_look_overhead: int = 12

    • Additional characters to examine beyond context_size for sentence splitting.
    • Enhances sentence detection accuracy.
    • Default: 12 characters
  • auto_context: bool = False

    • Yields safe sentence boundaries before the full context_size delay when the boundary heuristic and tokenizer agree.
    • Falls back to the normal context_size and context_size_look_overhead behavior when the boundary is still ambiguous.
    • Recommended for English when used with tokenizer="nltk+rule-based".
    • Default: False
  • minimum_sentence_length: int = 10

    • Minimum character count for a text chunk to be considered a sentence.
    • Shorter fragments are buffered until this threshold is met.
    • Default: 10 characters
  • minimum_first_fragment_length: int = 10

    • Minimum character count required for the first sentence fragment.
    • Ensures the initial output meets a specified length threshold.
    • Default: 10 characters

Yield Control

These parameters control how quickly and frequently the generator yields sentence fragments:

  • quick_yield_single_sentence_fragment: bool = False

    • When True, yields the first fragment of the first sentence as quickly as possible.
    • Useful for getting immediate output in real-time applications like speech synthesis.
    • Default: False
  • quick_yield_for_all_sentences: bool = False

    • When True, yields the first fragment of every sentence as quickly as possible.
    • Extends the quick yield behavior to all sentences, not just the first one.
    • Automatically sets quick_yield_single_sentence_fragment to True.
    • Default: False
  • quick_yield_every_fragment: bool = False

    • When True, yields every fragment of every sentence as quickly as possible.
    • Provides the most granular output, yielding fragments as soon as they're detected.
    • Automatically sets both quick_yield_for_all_sentences and quick_yield_single_sentence_fragment to True.
    • Default: False

Text Cleanup

  • cleanup_text_links: bool = False

    • When True, removes hyperlinks from the output sentences.
    • Default: False
  • cleanup_text_emojis: bool = False

    • When True, removes emoji characters from the output sentences.
    • Default: False

Tokenization

  • tokenize_sentences: Callable = None

    • Custom function for sentence tokenization.
    • If None, uses the default tokenizer specified by tokenizer.
    • Default: None
  • tokenizer: str = "rule-based"

    • Specifies the tokenizer to use. Options: "nltk", "stanza", "rule-based", or "nltk+rule-based"
    • "nltk+rule-based" splits at boundaries accepted by both NLTK and the rule-based heuristic tokenizer, plus rule-based boundaries promoted by high-confidence local checks.
    • Recommended for English, especially with auto_context=True.
    • Default: "rule-based"
  • language: str = "en"

    • Language setting for the tokenizer.
    • Use "en" for English, "multilingual" for Stanza tokenizer, or a supported language code/name for rule-based heuristic data.
    • Default: "en"

Debugging and Fine-tuning

  • log_characters: bool = False

    • When True, logs each processed character to the console.
    • Useful for debugging or monitoring real-time processing.
    • Default: False
  • sentence_fragment_delimiters: str = ".?!;:,\n…)]}。-"

    • Characters considered as potential sentence fragment delimiters.
    • Used for quick yielding of sentence fragments.
    • Default: ".?!;:,\n…)]}。-"
  • full_sentence_delimiters: str = ".?!\n…。"

    • Characters considered as full sentence delimiters.
    • Used for more definitive sentence boundary detection.
    • Default: ".?!\n…。"
  • force_first_fragment_after_words: int = 15

    • Forces the yield of the first sentence fragment after this many words.
    • Ensures timely output even with long opening sentences.
    • Default: 15 words
  • fragment_lookahead_words: int = 0

    • Buffers quick-yield delimiter candidates until this many complete words have been observed.
    • At the threshold, a confirmed delimiter from full_sentence_delimiters is preferred over an earlier soft delimiter such as a comma; otherwise the latest soft delimiter is used.
    • If the input stream ends before the threshold, the remaining text is flushed normally.
    • This counts observed words, not words in the yielded fragment. For example, after eight observed words the splitter may yield an earlier two-word clause.
    • Default: 0, preserving immediate quick-yield behavior.

Time based strategy

Instead of a purely lexigraphical strategy, a time based strategy is available. A target tokens per second (tps) is input, and generate_sentences will yield the best available output (full sentence, longest fragment, or any available buffer, in that order) if it is approaching a "deadline" where what has been output would be slower than the input tps target. If LLM is more than two full sentences ahead of the target it will output a sentence even if it's ahead of the "deadline"

from stream2sentence.stream2sentence_time_based import generate_sentences_time_based

Parameters

  • generator (Iterator[str])
    • A generator that yields chunks of text as a stream of characters.`
  • target_tps: float = 4
    • the rate in tokens per second you want to use to calculate deadlines for output.
    • Default is 4. (approximately the speed of human speech)
  • lead_time: float = 1
    • amount of time in seconds to wait for the buffer to build for before returning values.
  • max_wait_for_fragments = [3, 2]
    • Max amount of time in seconds that the Nth sentence will wait beyond the "deadline" for a "fragment" (text preceeding a fragment delimiter), which is preferred over a piece of buffer.
    • The last value in the array is used for all subsequent checks.
  • min_output_lengths: int[] = [2, 3, 3, 4]
    • An array that corresponds to the minimum output size in words for the corresponding output sentence, the last value in the array is used for all remaining output.
    • For example [4,5,6] would mean the first piece of output must have 4 words, the second 5 words, and all subsequent 6.
  • preferred_sentence_fragment_delimiters: str[] = ['. ', '? ', '! ', '\n']
    • Array of strings that deliniate a sentence fragment. "Preferred" are checked first and always used if the fragment meets the length requirement over the other fragment delimiters.
    • Note the trailing spaces, added to differentiate between values like $3.5 and a proper sentence end
  • sentence_fragment_delimiters: str[] = ['; ', ': ', ', ', '* ', '**', '– ']
    • Array of strings that are checked after "preferred" delimiters
  • delimiter_ignore_prefixes: str[]
    • Array of strings that will not be considered "delimiters" if preceeded by a delimiter.
    • Used to ignore common abbreviations for things like Mr. Dr. and Mrs. where we don't want to split
    • Default is a long list documented in delimiter_ignore_prefixes
  • wait_for_if_non_fragment: str[]
    • Array of strings that the algorithm will not use as the last value if the whole buffer is being output (not a fragment or sentence).
    • Avoids awkward pauses on common words that are unnatural to pause at.
    • Default is a long list of common words documented in avoid_pause_words.py
  • deadline_offsets_static: float[] = [1]
    • Constant amount of time in seconds to subtract from the deadline for first n sentences.
    • Last value applied to all subsequent sentences
  • deadline_offsets_dynamic: float[] = [0]:
    • Added to account for the time it takes a TTS engine to generate output.
    • For example, if it takes your TTS engine around 1 second to generate 10 words, you can use a value of 0.1 so that the TTS generation time is included in the deadline.
    • Applied to first n sentences, last value applied to all subsequent

Contributing

Any Contributions you make are welcome and greatly appreciated.

  1. Fork the Project.
  2. Create your Feature Branch (git checkout -b feature/AmazingFeature).
  3. Commit your Changes (git commit -m 'Add some AmazingFeature').
  4. Push to the Branch (git push origin feature/AmazingFeature).
  5. Open a Pull Request.

License

This project is licensed under the MIT License. For more details, see the LICENSE file.


Project created and maintained by Kolja Beigel.

Metadata

Release files for stream2sentence 1.0.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for stream2sentence 1.0.4
File Size Uploaded
stream2sentence-1.0.4.tar.gz 44.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for stream2sentence 1.0.4
File Interpreter ABI Platform
stream2sentence-1.0.4-py3-none-any.whl Python 3 none any Details

Total release size: 85.0 kB

Release files / stream2sentence-1.0.4.tar.gz

Download URL stream2sentence-1.0.4.tar.gz
Size 44.1 kB
Tags Source
SHA-256 checksum
How to use checksums
ffd30fea05baf4e4a7502d1ab59b1dd83fdb97bce846842bf7b00388c3f7338a
BLAKE2b-256 checksum
How to use checksums
cc9350c780afe405145a0db86625b0d8e0e6b087abec21206c8c46cefda48123
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.5

Release files / stream2sentence-1.0.4-py3-none-any.whl

Download URL stream2sentence-1.0.4-py3-none-any.whl
Size 40.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
91350dfcf9661b1d179077016d5919dbcd68fecc1014305a0f88f5c3c9fea2b0
BLAKE2b-256 checksum
How to use checksums
2126a1a64f9a81e688418019d3203ec30fb0d3b37c74a74a23f23beb8b1cc554
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.5

Release history Release notifications | RSS feed

This release

1.0.4 This release

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.9

2 release files

0.2.8

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page