Skip to main content

Consult the module API page at

https://engineering.purdue.edu/kak/distBabyGPT/babyGPT-1.1.4.html

for all information related to this module, including information related to the latest changes to the code. The page at the URL shown above lists all of the module functionality you can invoke in your own code.

Creating an instance of babyGPT:

        baby_gpt = babyGPT(
                            max_seq_length = max_seq_length,
                            batch_size = batch_size,
                            embedding_size = embedding_size,
                            num_basic_decoders = num_basic_decoders,
                            num_atten_heads = num_atten_heads,
                            optimizer_params = optimizer_params,
                            num_warmup_steps = num_warmup_steps,
                            masking = masking,
                            verify_text_corpus = False,
                            path_saved_model = {"decoder" : "./saved_decoder",
                                                "embedding_generator" : "./saved_embedding_generator",
                                               },
                          )

Since babyGPT calls on TransformerFG for language modeling, you must also construct an instance of that class:

        xformer = baby_gpt.TransformerFG(
                            max_seq_length = max_seq_length,
                            embedding_size = embedding_size,
                            tokenizer_json = tokenizer_json,
                            num_warmup_steps = num_warmup_steps,
                            optimizer_params = optimizer_params,
                  )

Within the TransformerFG module, it is the MasterDecoder class that is needed for the next token prediction for the purpose of self-supervised learning:

        master_decoder = baby_gpt.MasterDecoderWithMasking(
                            xformer,
                            num_basic_decoders = num_basic_decoders,
                            num_atten_heads = num_atten_heads,
                            masking = masking
                         )


Finally, here is an instance of the dataloader you're going to need:

        dataloader = baby_gpt.ArticleDatasetWithBufferedContext(
                            gpt = baby_gpt,
                            tokenizer_json = tokenizer_json,
                            context_window_size = context_window_size,
                            context_buffer_size = context_buffer_size,
                            articles_dir = articles_dir,
                     )

Metadata

Release files for babyGPT 1.1.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for babyGPT 1.1.4
File Size Uploaded
babygpt-1.1.4.tar.gz 1.0 MB Details

Release files / babygpt-1.1.4.tar.gz

Download URL babygpt-1.1.4.tar.gz
Size 1.0 MB
Tags Source
SHA-256 checksum
How to use checksums
4627b4f4322b2d2899e1c2e2e932485a24d4619c7fce8142c4789c68e7b6bc5a
BLAKE2b-256 checksum
How to use checksums
5e5c2f1ee8eff42f1d92aa6469d0c6f8b14c60870e718ff3f05dcb8b6bae56ba
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.10.12

Release history Release notifications | RSS feed

This release

1.1.4 This release

1 release file

1.1.3

1 release file

1.1.2

1 release file

1.1.1

1 release file

1.1.0

1 release file

1.0.9

1 release file

1.0.8

1 release file

1.0.7

1 release file

1.0.6

1 release file

1.0.5

1 release file

1.0.4

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page