chat-tokenizer
Usage
from chat_tokenizer import ChatTokenizer
from transformers import AutoTokenizer
tokenizer = ChatTokenizer(AutoTokenizer.from_pretrained("qwen/Qwen2.5-0.5B-Instruct"))
audio_lens = [[4, 2], 3]
labels = [["今天天气不错", "哈哈"], "你好啊"]
input_ids, input_lens, label_ids, label_lens = tokenizer.batch_tokenize(audio_lens, labels)
input_ids = tokenizer.fill_labels(label_ids, input_ids)
Metadata
Release files for chat-tokenizer 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| chat_tokenizer-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Release files / chat_tokenizer-0.1.0-py3-none-any.whl
| Download URL | chat_tokenizer-0.1.0-py3-none-any.whl |
|---|---|
| Size | 8.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
67136dbab049eaad26c8044cbdf7e9cbcb7f633072404b51c6058ed9e7508b41
|
|
BLAKE2b-256 checksum How to use checksums |
b0be9f2e52006fefd5e61e360d8ed66493af838382ceb05497378c9677b0be11
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.12.3
|