Skip to main content

alice_tokenizer

the alice tokenizer is a project i developed as part of learning about the workings of transformers. this is a tokenizer i built for my own transformer and i decided to release as a standalone application to make building a byte-pair encoder on any dataset easy for anyone.

using alice_tokenizer is easy, it takes only a few steps to install and run.

installation

you can install alice_tokenizer using pip

pip install alice-tokenizer

usage

python

tokenizing

tokenizing a string in alice is easy. start by initialising the tokenizer.

import alice_tokenizer as alice

tokenizer = alice.Tokenizer()

to turn a string into tokens use the tokenize method.

tokens = tokenizer.tokenize("Tokenize me") # returns a list with integers

to turn tokens back into text use the detokenize method.

output = tokenizer.detokenize(tokens) # returns a string

for a visual look at how the tokens are split up, use the visualize_tokens method

tokenizer.visualize_tokens(tokens)
tokenizer.visualize_tokens("Tokenize me.") # works on str and token lists

making a tokenizer

to make a tokenizer, you must have a text file with text data inside. at this juncture alice only supports ascii, although this will be remedied in the future. for now ensure that your file uses utf-8 encoding as it will be read as utf-8.

start by initialising the tokenizer with the build=True condition.

import alice_tokenizer as alice

tokenizer = alice.Tokenizer(build=True)

then declare your text file (it is recommended to use a dataset of at least 20mb or larger, the larger the better).

tokenizer.set_data('data.txt')

then begin training

tokenizer.train(
    vocab_size = 32_000,
    chunks = 4_000,
    name = 'bob',
    save_format = 'json',
    save_to_file = True,
    pretokenize = True,
    do_tests = True,
    quiet = False
)

below is a description of each parameter

  • vocab_size: how large the resultant vocabulary should be. for a standard english tokenizer 32k is the recommended size but it can be as large or as small as you wish.
  • chunks: how many chunks the training data should be split into. ensure it is a factor of vocab_size to ensure that the exact vocab size is met.
  • name: the name of the tokenizer to save afterwards. when the vocabulary is saved to a file, this is the filename it will use.
  • save_format: the format of the final vocabulary file. only json is supported as of now, so don't change this
  • save_to_file: if False, the final vocab won't be saved to the disk after training.
  • pretokenize: if True, applys the GPT-2 pretokenization pattern to the text before tokenizing.
  • do_tests: if True, will perform some tests on selected excerpts from the training data after training is complete.
  • quiet: if True, will not print anything during training

these values are all the preset defaults, if not declared they will snap back to these defaults instead.

once the tokenizer is done building, you can use it immediately by using the tokenizer object methods.

tokens = tokenizer.tokenize("Tokenize me.")
output = tokenizer.detokenize(tokens)
tokenizer.visualize_tokens(tokens)
tokenizer.visualize_tokens("Tokenize me.")

to load your vocab into alice again, you can declare the vocab file when initialising the tokenizer

import alice_tokenizer as alice

tokenizer = alice.Tokenizer(vocab_file="bob.json")

Release files for alice-tokenizer 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for alice-tokenizer 0.2.0
File Size Uploaded
alice_tokenizer-0.2.0.tar.gz 278.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for alice-tokenizer 0.2.0
File Interpreter ABI Platform
alice_tokenizer-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 562.3 kB

Release files / alice_tokenizer-0.2.0.tar.gz

Download URL alice_tokenizer-0.2.0.tar.gz
Size 278.0 kB
Tags Source
SHA-256 checksum
How to use checksums
19d55e465614fee9c9f4aa0c10d67d47789d220472d949e3615eb7e4b96b6eb0
BLAKE2b-256 checksum
How to use checksums
bbc6baf3a3ce429c600c75c73df04a5ef98025925e622d0b445b205d62fd2e66
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.4

Release files / alice_tokenizer-0.2.0-py3-none-any.whl

Download URL alice_tokenizer-0.2.0-py3-none-any.whl
Size 284.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8da1538e81b78cfaaf3422c036bbb471ba5543024e125db34421d58d874b0d1b
BLAKE2b-256 checksum
How to use checksums
b42091c19120f4923324e1aae986c34f8a05fd2a3c9cd779a67e84c1daa7afd1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.4

Release history Release notifications | RSS feed

0.2.2

2 release files

0.2.1

2 release files

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page