Advanced tokenizer with support for BPE, WordPiece, and Unigram algorithms
Project description
Ultra-Tokenizer
Ultra-Tokenizer is a high-performance, production-ready tokenizer that supports multiple subword tokenization algorithms including BPE, WordPiece, and Unigram. Designed for efficiency and flexibility, it's perfect for modern NLP pipelines.
Table of Contents
Why Choose Ultra-Tokenizer?
While there are several tokenization libraries available, Ultra-Tokenizer stands out in several key areas:
Performance Optimized
- Lightning Fast: Engineered for high throughput with minimal memory overhead
- Efficient Memory Usage: Intelligent caching mechanisms reduce memory footprint
- Parallel Processing: Built-in support for multi-core tokenization
Developer Friendly
- Simple API: Intuitive interface that's easy to integrate into any pipeline
- Full Type Hints: Better IDE support and code completion
- Comprehensive Documentation: Detailed guides and API references
- Minimal Dependencies: Lightweight with no unnecessary bloat
Flexible & Extensible
- Multiple Algorithms: Switch between BPE, WordPiece, and Unigram with a single parameter
- Custom Tokenization Rules: Easily add domain-specific tokenization rules
- Train on Your Data: Simple interface for training custom tokenizers on your corpus
Production Ready
- Robust Error Handling: Graceful handling of edge cases and malformed input
- 100% Test Coverage: Thoroughly tested across different scenarios
- Performance Benchmarks: Consistently outperforms alternatives in speed and memory usage
Comparison with Other Tokenizers
| Feature | Ultra-Tokenizer | HuggingFace Tokenizers | spaCy | NLTK |
|---|---|---|---|---|
| Multiple Algorithms | ✅ BPE, WordPiece, Unigram | ✅ BPE, WordPiece, Unigram | ❌ Mostly rule-based | ❌ Rule-based |
| Training Interface | ✅ Simple and intuitive | ✅ Comprehensive | ❌ Limited | ❌ No built-in training |
| Memory Efficiency | ✅ Excellent with smart caching | ⚠️ Good, but can be heavy | ✅ Good | ⚠️ Can be memory intensive |
| Performance | ⚡ Blazing fast | Fast | Fast | Slower |
| Dependencies | Minimal | Heavy (Rust) | Heavy (Cython) | Heavy |
| Type Hints | ✅ Full support | ⚠️ Partial | ⚠️ Partial | ❌ None |
| CLI Support | ✅ Built-in | ✅ Available | ❌ No | ❌ No |
| Learning Curve | Gentle | Steep | Moderate | Steep |
Features
- Multiple Tokenization Algorithms: Byte Pair Encoding (BPE), WordPiece, and Unigram support
- High Performance: Optimized for speed with efficient implementations
- Easy Integration: Simple API for training and using tokenizers
- Production Ready: Comprehensive test coverage and robust error handling
- Fully Typed: Complete type annotations for better development experience
- Multilingual Support: Excellent handling of various languages and scripts
- Special Token Support: Built-in handling of special tokens
- Memory Efficient: Low memory footprint with smart caching
- Customizable: Flexible configuration options for different use cases
- CLI Support: Command-line interface for easy usage
Installation
Install the latest stable version from PyPI:
pip install ultra-tokenizer
For the latest development version:
pip install git+https://github.com/pranav271103/Ultra-Tokenizer.git
For development:
git clone https://github.com/pranav271103/Ultra-Tokenizer.git
cd Ultra-Tokenizer
pip install -e ".[dev]" # Install in development mode with all dependencies
Quick Start
Basic Usage
from ultra_tokenizer import Tokenizer, TokenizerTrainer
# Initialize and train a tokenizer
trainer = TokenizerTrainer(
vocab_size=30000,
min_frequency=2,
lowercase=True,
strip_accents=True
)
tokenizer = trainer.train(
files=["path/to/your/text/file.txt"],
algorithm="bpe", # or "wordpiece" or "unigram"
num_workers=4
)
# Tokenize text
text = "This is an example sentence."
tokens = tokenizer.tokenize(text)
print(tokens)
Using Pre-trained Tokenizers
from advanced_tokenizer import Tokenizer
# Load a pre-trained tokenizer
tokenizer = Tokenizer.from_pretrained("your-pretrained-tokenizer")
# Encode and decode text
encoded = tokenizer.encode("Hello, world!")
decoded = tokenizer.decode(encoded.ids)
Documentation
Full documentation is available at https://pranav271103.github.io/Ultra-Tokenizer/
Key Components
- Tokenizer: Main class for tokenization
- TokenizerTrainer: For training new tokenizers
- Vocabulary: Manages token-to-ID mappings
- Pre/Post Processors: Handle text normalization and token processing
Development
Running Tests
pytest tests/
Code Style
We use black for code formatting and isort for import sorting:
black .
isort .
Building Documentation
cd docs
make html
Contributing
Contributions are what make the open-source community such an amazing place to learn, inspire, and create. Any contributions you make are greatly appreciated.
- Fork the Project
- Create your Feature Branch (
git checkout -b feature/AmazingFeature) - Commit your Changes (
git commit -m 'Add some AmazingFeature') - Push to the Branch (
git push origin feature/AmazingFeature) - Open a Pull Request
Please read CONTRIBUTING.md for details on our code of conduct and the process for submitting pull requests.
License
Distributed under the Apache 2.0 License. See LICENSE for more information.
Contact
Pranav Singh - pranav.singh01010101@gmail.com
Project Link: Ultra-Tokenizer
Acknowledgments
- Hugging Face Tokenizers - Inspiration for some design patterns
- YouTokenToMe - For BPE implementation reference
- All contributors who helped improve this project
Performance
The tokenizer is optimized for both training and inference:
- Training: Uses multiprocessing for faster vocabulary building
- Inference: Efficient lookup tables for fast tokenization
- Memory: Optimized to handle large vocabularies
Customization
You can extend the tokenizer by:
- Adding new pre-tokenization rules
- Implementing custom subword algorithms
- Adding support for additional languages
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ultra_tokenizer-0.1.1.tar.gz.
File metadata
- Download URL: ultra_tokenizer-0.1.1.tar.gz
- Upload date:
- Size: 40.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9b051404868eede1ce88522678e201c8e3cfb20efb3e80c5ae8752b416dec9ff
|
|
| MD5 |
561823623845b5ce4bc32b8657f3874e
|
|
| BLAKE2b-256 |
ad9f96cd4893bee561a882db4ac94f5816f9e48233205bf48696b8373eed3873
|
File details
Details for the file ultra_tokenizer-0.1.1-py3-none-any.whl.
File metadata
- Download URL: ultra_tokenizer-0.1.1-py3-none-any.whl
- Upload date:
- Size: 33.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4e26a6d6ac9f54635445d80c8d62547f79261d2622d9bfac9826559c06494c2d
|
|
| MD5 |
41c44209a208a0fc05efa52242faf4c1
|
|
| BLAKE2b-256 |
8e56c503f23cdfffa5396eb76811e6911af9ae3661cdfb1064b94f7b0210ff80
|