Skip to main content

code_tokenizers

This library is built on top of the awesome transformers and tree-sitter libraries. It provides a simple interface to align the tokens produced by a BPE tokenizer with the tokens produced by a tree-sitter parser.

Install

pip install code_tokenizers

How to use

The main interface of code_tokenizers is the CodeTokenizer class. You can use a pretrained BPE tokenizer from the popular transformers library, and a tree-sitter parser from the tree-sitter library.

To specify a CodeTokenizer using the gpt2 BPE tokenizer and the python tree-sitter parser, you can do:

from code_tokenizers.core import CodeTokenizer

py_tokenizer = CodeTokenizer.from_pretrained("gpt2", "python")
None of PyTorch, TensorFlow >= 2.0, or Flax have been found. Models won't be available and only tokenizers, configuration and file/data utilities can be used.

You can specify any pretrained BPE tokenizer from the huggingface hub or a local directory and the language to parse the AST for.

Now, we can tokenize some code:

from pprint import pprint

code = """
def foo():
    print("Hello world!")
"""

encoding = py_tokenizer(code)
pprint(encoding, depth=1)
{'ast_ids': [...],
 'attention_mask': [...],
 'input_ids': [...],
 'is_builtins': [...],
 'is_internal_methods': [...],
 'merged_ast': [...],
 'offset_mapping': [...],
 'parent_ast_ids': [...]}

And we can print out the associated AST types:

Note

Note: Here the N/As are the tokens that are not part of the AST, such as the spaces and the newline characters. Their IDs are set to -1.

for ast_id, parent_ast_id in zip(encoding["ast_ids"], encoding["parent_ast_ids"]):
    if ast_id != -1:
        print(py_tokenizer.node_types[parent_ast_id], py_tokenizer.node_types[ast_id])
    else:
        print("N/A")
N/A
function_definition def
function_definition identifier
parameters (
N/A
N/A
N/A
N/A
call identifier
argument_list (
argument_list string
argument_list string
argument_list string
argument_list )
N/A

Metadata

Release files for code-tokenizers 0.0.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for code-tokenizers 0.0.5
File Size Uploaded
code_tokenizers-0.0.5.tar.gz 13.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for code-tokenizers 0.0.5
File Interpreter ABI Platform
code_tokenizers-0.0.5-py3-none-any.whl Python 3 none any Details

Total release size: 126.0 kB

Release files / code_tokenizers-0.0.5.tar.gz

Download URL code_tokenizers-0.0.5.tar.gz
Size 13.4 kB
Tags Source
SHA-256 checksum
How to use checksums
796b0dda0555bd5aea0a87643c9a062c444f57f61c92881a326a5242b5b2cdb4
BLAKE2b-256 checksum
How to use checksums
b8830d9323f6ea7fe953594392c12dd6c99231ff8e17c06f98853bcb66e7115b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.1 CPython/3.10.8

Release files / code_tokenizers-0.0.5-py3-none-any.whl

Download URL code_tokenizers-0.0.5-py3-none-any.whl
Size 112.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b4ce9108c840370cc8dd2582dc9a45451722806c9c8a3226f783870bbd4a9074
BLAKE2b-256 checksum
How to use checksums
1863ae91f45b305d413edd5f196658142778c5117385d4b633cb40ef8411cf2b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.1 CPython/3.10.8

Release history Release notifications | RSS feed

This release

0.0.5 This release

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page