Skip to main content

LangSegment (Unofficial Backup)

PyPI version License: BSD-3-Clause

⚠️ This is an unofficial backup of LangSegment 0.3.5. The original repository has been removed or made private.

A multilingual text segmentation tool that automatically identifies and splits text by language. Particularly useful for TTS (Text-to-Speech) processing with mixed-language content.

Attribution & History

This package is published to PyPI to preserve access to this useful tool after the original repository was removed. All credit for the original work goes to the original author.

Installation

pip install langsegment-backup

Supported Languages

Primary support:

  • 🇨🇳 Chinese (zh)
  • 🇯🇵 Japanese (ja)
  • 🇬🇧 English (en)
  • 🇰🇷 Korean (ko)

Experimental support:

  • 🇫🇷 French (fr)
  • 🇻🇳 Vietnamese (vi)
  • 🇷🇺 Russian (ru)
  • 🇹🇭 Thai (th)

The tool can actually support up to 97 different languages through the underlying py3langid library.

Quick Start

from LangSegment import LangSegment

# Basic usage - segment mixed language text
text = "你好世界!Hello World! こんにちは!안녕하세요!"
results = LangSegment.getTexts(text)

for item in results:
    print(f"[{item['lang']}] {item['text']}")

Output:

[zh] 你好世界!
[en] Hello World! 
[ja] こんにちは!
[ko] 안녕하세요!

Features

Language Filtering

You can specify which languages to detect and in what priority order:

from LangSegment import LangSegment

# Set language filter (priority order: left = highest)
LangSegment.setfilters(["zh", "ja", "en", "ko"])

# Or for Chinese-English only
LangSegment.setfilters(["zh", "en"])

Language Statistics

Get statistics about the languages in your text:

from LangSegment import LangSegment

text = "你好世界!Hello World! こんにちは!"
LangSegment.getTexts(text)

# Get language counts (sorted by character count, descending)
counts = LangSegment.getCounts()
print(counts)  # [('zh', 10), ('en', 12), ('ja', 6)]

# Get the primary language
primary_lang, char_count = counts[0]
print(f"Primary language: {primary_lang}")

Manual Language Tags

You can manually specify language regions using tags:

text = "这是中文<ja>これは日本語です</ja>这又是中文"
results = LangSegment.getTexts(text)

SSML Support (Chinese)

The tool includes SSML-like tags for Chinese number/date processing:

# Number reading
text = "<number>12345</number>"  # → 一二三四五

# Phone number
text = "<telephone>13812345678</telephone>"  # → 幺三八幺二三四五六七八

# Currency
text = "<currency>12345</currency>"  # → 一万二千三百四十五

# Date
text = "<date>2024-08-24</date>"  # → 二零二四年八月二十四日

Configuration Options

from LangSegment import LangSegment

# Set Chinese/Japanese priority threshold (0-1, default: 0.89)
LangSegment.setPriorityThreshold(0.89)

# Enable/disable result merging
LangSegment.setLangMerge(True)

# Keep Chinese pinyin format
LangSegment.setKeepPinyin(False)

# Enable preview features (French, Vietnamese support)
LangSegment.setEnablePreview(True)

API Reference

Main Functions

Function Description
LangSegment.getTexts(text) Segment text and return list of {lang, text, score} dicts
LangSegment.classify(text) Alias for getTexts()
LangSegment.getCounts() Get language statistics as list of (lang, count) tuples
LangSegment.setfilters(list) Set language filter/priority list
LangSegment.getfilters() Get current language filters

Configuration Functions

Function Description
setPriorityThreshold(float) Set zh/ja disambiguation threshold (0-1)
getPriorityThreshold() Get current threshold
setLangMerge(bool) Enable/disable merging adjacent same-language segments
getLangMerge() Get merge setting
setKeepPinyin(bool) Keep Chinese pinyin format in parentheses
getKeepPinyin() Get pinyin setting
setEnablePreview(bool) Enable experimental language support
getEnablePreview() Get preview setting

Dependencies

  • numpy >= 1.19.5
  • py3langid >= 0.2.2

License

BSD 3-Clause License

Copyright (c) 2024 juntaosun. All rights reserved.

See LICENSE for full license text.

Contributing

Since this is a backup/preservation fork, major feature additions are not planned. However, bug fixes and compatibility updates are welcome. Please open an issue or pull request at https://github.com/MiniXC/LangSegment-0.3.5-backup.

Acknowledgments

Metadata

Release files for langsegment-backup 0.3.5.post1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for langsegment-backup 0.3.5.post1
File Size Uploaded
langsegment_backup-0.3.5.post1.tar.gz 25.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for langsegment-backup 0.3.5.post1
File Interpreter ABI Platform
langsegment_backup-0.3.5.post1-py3-none-any.whl Python 3 none any Details

Total release size: 51.4 kB

Release files / langsegment_backup-0.3.5.post1.tar.gz

Download URL langsegment_backup-0.3.5.post1.tar.gz
Size 25.5 kB
Tags Source
SHA-256 checksum
How to use checksums
60a4330f99acfcc98c055b406554bb1e7528002cc8ac5a2efd584ed6ae698cfd
BLAKE2b-256 checksum
How to use checksums
d7546570ba2597906c0372d984355d475c894d3a7c0eb12cb649f9ca31ccc9ab
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.2

Release files / langsegment_backup-0.3.5.post1-py3-none-any.whl

Download URL langsegment_backup-0.3.5.post1-py3-none-any.whl
Size 26.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2fe0993d3a02b2ef53e406a6f6abc993f6cec1a0fb7b42fe1b6e4b096c8a4e15
BLAKE2b-256 checksum
How to use checksums
f45c0a4e1c03ddc0f6f2bd0cad8611024cb747b46c9d3707f123170bc53589b9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.2

Release history Release notifications | RSS feed

This release

0.3.5.post1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page