A byte-level language detection model supporting 102 languages
Project description
Lark - Byte-Level Language Detection
Lark is a byte-level language detection model that supports 102 languages with high accuracy and efficiency.
🚀 Features
- 102 Languages: Supports a wide range of languages including English, Chinese, Japanese, Spanish, French, etc.
- Byte-Level Processing: No vocabulary limitations, handles any Unicode text
- High Accuracy: State-of-the-art performance on language detection tasks
- Fast Inference: Optimized for both CPU and GPU
- Easy Integration: Simple API for both batch and single text processing
📦 Installation
From PyPI (Recommended)
pip install lark-ld
From Source
git clone https://github.com/farshore-byte/LarkDetect.git
cd LarkDetect
pip install -e .
🎯 Quick Start
Basic Usage
from lark import LarkDetector
# Initialize detector
detector = LarkDetector()
# Detect language for single text
text = "Hello, how are you today?"
language, confidence = detector.detect(text)
print(f"Language: {language}, Confidence: {confidence:.4f}")
# Batch detection
texts = [
"Hello world!",
"今天天气真好",
"こんにちは、元気ですか?"
]
results = detector.detect_batch(texts)
for text, (lang, conf) in zip(texts, results):
print(f"'{text}' -> {lang} ({conf:.4f})")
Advanced Usage
from lark import LarkDetector
detector = LarkDetector()
# Get top-k predictions
text = "This is a sample text"
prediction, confidence, top_k = detector.detect_with_topk(text, k=5)
print(f"Prediction: {prediction} (Confidence: {confidence:.4f})")
print("Top 5 predictions:")
for i, item in enumerate(top_k):
print(f" {i+1}. {item['language']:8} - {item['probability']:.4f}")
# Confidence threshold
language, confidence, top_k = detector.detect_with_confidence(
text,
confidence_threshold=0.7
)
if language == "unknown":
print(f"Low confidence: {confidence:.4f}")
else:
print(f"Detected: {language} (Confidence: {confidence:.4f})")
📊 Supported Languages
Lark supports 102 languages including:
- European: English, Spanish, French, German, Italian, Russian, etc.
- Asian: Chinese, Japanese, Korean, Hindi, Arabic, Thai, etc.
- African: Swahili, Yoruba, Zulu, etc.
- Others: And many more...
See the full list in all_dataset_labels.json.
🏗️ Model Architecture
Lark uses a novel byte-level architecture:
- Byte Encoder: Converts raw bytes to contextual representations
- Boundary Predictor: Identifies segment boundaries using Gumbel-Sigmoid
- Segment Decoder: Processes segments for language classification
This architecture enables:
- No vocabulary limitations
- Robust handling of mixed-language text
- Efficient processing of long documents
📈 Performance
| Metric | Value |
|---|---|
| Overall Accuracy | 90.14% on validation set |
| Inference Speed | ~1ms per text (CPU) |
| Model Size | 9.28MB (float16) |
| Precision | float16 (CPU/GPU compatible) |
| Total Parameters | 4,866,919 |
| Supported Languages | 102 |
Training Dataset
The model was trained on a comprehensive dataset combining:
- opus-100 - Multilingual parallel corpus
- Mike0307/language-detection - Language detection dataset
- sirgecko___language_detection - Language detection dataset
- papluca/language-identification - Language identification dataset
- sirgecko/language_detection_train - Language detection training data
Dataset Statistics:
- Train samples: 109,636,748
- Validation samples: 385,306
- Total languages: 102
Evaluation Results
Detailed per-language evaluation results (accuracy, precision, recall, F1-score) will be available in the evaluation results file. The model achieves 90.14% overall accuracy on the validation dataset.
Note on Evaluation Data: While the model achieves strong overall performance, the current evaluation dataset has limited coverage for some languages (e.g., an, dz, hy, mn, yo). This is due to the validation split not containing samples for these languages. However, the training dataset provides comprehensive coverage.
Training Dataset Details
The model was trained on a comprehensive dataset combining:
- OPUS-100 - Multilingual parallel corpus containing 100 language pairs
- Mike0307/language-detection - Language detection dataset
- sirgecko___language_detection - Language detection dataset
- papluca/language-identification - Language identification dataset
- sirgecko/language_detection_train - Language detection training data
OPUS-100 Dataset Statistics:
- Contains approximately 55 million sentence pairs
- Covers 99 language pairs
- 44 language pairs have 1 million+ sentence pairs
- 73 language pairs have 100,000+ sentence pairs
- 95 language pairs have 10,000+ sentence pairs
- Each language has at least 10,000 training samples
Dataset Statistics:
- Train samples: 109,636,748
- Validation samples: 385,306
- Total languages: 102
🔧 API Reference
LarkDetector Class
class LarkDetector:
def __init__(self, model_path: str = None, labels_path: str = None):
"""Initialize the language detector"""
def detect(self, text: str) -> Tuple[str, float]:
"""Detect language for single text"""
def detect_batch(self, texts: List[str]) -> List[Tuple[str, float]]:
"""Batch language detection"""
def detect_with_topk(self, text: str, k: int = 5) -> Tuple[str, float, List[Dict]]:
"""Get top-k predictions with probabilities"""
def detect_with_confidence(self, text: str, confidence_threshold: float = 0.5) -> Tuple[str, float, List[Dict]]:
"""Detection with confidence threshold"""
🛠️ Development
Setup Development Environment
git clone https://github.com/farshore-byte/LarkDetect.git
cd LarkDetect
pip install -e ".[dev]"
Running Tests
python -m pytest tests/
Building from Source
python setup.py sdist bdist_wheel
📝 Citation
If you use Lark in your research, please cite:
@software{lark2025,
title={Lark: Byte-Level Language Detection},
author={Farshore AI},
year={2024},
url={https://github.com/farshore-byte/LarkDetect}
}
🤝 Contributing
We welcome contributions! Please see CONTRIBUTING.md for details.
📄 License
This project is licensed under the MIT License - see the LICENSE file for details.
🙏 Acknowledgments
- Thanks to the open-source community for datasets and tools
- Inspired by modern language detection approaches
- Built with PyTorch and Hugging Face ecosystem
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file lark_ld-1.0.1.tar.gz.
File metadata
- Download URL: lark_ld-1.0.1.tar.gz
- Upload date:
- Size: 15.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7fda38467002d782c80be4a2567ca2f0c16c1e9f219b921e9500134f5790ddff
|
|
| MD5 |
3bad62f40ba66133a3a04bfd101deffe
|
|
| BLAKE2b-256 |
55e19da9049a2da9d0738824f78f44a3a3c7fda06423882bdcbc377164adf5c9
|
File details
Details for the file lark_ld-1.0.1-py3-none-any.whl.
File metadata
- Download URL: lark_ld-1.0.1-py3-none-any.whl
- Upload date:
- Size: 12.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9848f1544349fdc569ebe9cefc9208546e831397440d4c3bc964a91902ace32f
|
|
| MD5 |
375b06febd73365cd1f1f16824c52d46
|
|
| BLAKE2b-256 |
ecac60bddb033450fac8a3b334e85418fe7b9c4e9009e264288cb62227151fbf
|