one more transformers lib
Project description
Simple way to use transformer models
Installation
pip install plain-transformers
Usage
Multimodal transformer example with two tokenizers:
Step one: import model and some usefull staff;
import torch
from plain_transformers.models import MultimodalTransformer
from plain_transformers.layers import PostLNMultimodalTransformerDecoder
from plain_transformers.layers import PostLNTransformerEncoder
from plain_transformers import BPEWrapper
from plain_transformers.initializations import normal_initialization
from plain_transformers.samplers.nucleus_sampler import NucleusSampler
import youtokentome as yttm
Step two: train and load tokenizers;
# train your encoder tokenizer
yttm.BPE.train(..., model='encoder_tokenizer.model')
# train your decoder tokenizer
yttm.BPE.train(..., model='decoder_tokenizer.model')
# load tokenizers
encoder_tokenizer = BPEWrapper(model='encoder_tokenizer.model')
decoder_tokenizer = BPEWrapper(model='decoder_tokenizer.model')
Step three: init out model configuration;
cfg = {
'd_model': 768,
'first_encoder': {
'first_encoder_vocab_size': encoder_tokenizer.vocab_size(),
'first_encoder_max_length': 512,
'first_encoder_pad_token_id': encoder_tokenizer.pad_id,
'first_encoder_token_type_vocab_size': 2,
'first_encoder_n_heads': 8,
'first_encoder_dim_feedforward': 2048,
'first_encoder_num_layers': 3,
},
'second_encoder': {
'second_encoder_vocab_size': encoder_tokenizer.vocab_size(),
'second_encoder_max_length': 512,
'second_encoder_pad_token_id': encoder_tokenizer.pad_id,
'second_encoder_token_type_vocab_size': 2,
'second_encoder_n_heads': 8,
'second_encoder_dim_feedforward': 2048,
'second_encoder_num_layers': 3,
},
'decoder': {
'decoder_max_length': 512,
'decoder_vocab_size': decoder_tokenizer.vocab_size(),
'decoder_pad_token_id': decoder_tokenizer.pad_id,
'decoder_token_type_vocab_size': 2,
'decoder_n_heads': 8,
'decoder_dim_feedforward': 2048,
'decoder_num_layers': 3,
},
}
Step four: initialize model and apply weight initialisation (with default parameter std=0.02
);
model = MultimodalTransformer(
PostLNTransformerEncoder,
PostLNTransformerEncoder,
PostLNMultimodalTransformerDecoder,
cfg['d_model'],
**cfg['first_encoder'],
**cfg['second_encoder'],
**cfg['decoder'],
share_decoder_head_weights=True,
share_encoder_decoder_embeddings=False,
share_encoder_embeddings=True,
)
model.apply(normal_initialization)
Step five: train our model like ordinary seq2seq;
train(model, ...)
Step six: initialize Sampler and generate model answer;
sampler = NucleusSampler(model, encoder_tokenizer=(encoder_tokenizer, encoder_tokenizer), decoder_tokenizer=decoder_tokenizer)
sampler.generate('Hello Bob, what are you doing?', second_input_text='Fine, thanks!', top_k=5)
Example
You can find working example of NMT here.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Close
Hashes for plain-transformers-0.0.1.5rc2.tar.gz
Algorithm | Hash digest | |
---|---|---|
SHA256 | 375af967036752ba1c134742ae225800a029ee7ea4a8e9d1d9ac9a5a4a9345be |
|
MD5 | 20c1637af2676a4536d5a84e7ec4aff3 |
|
BLAKE2b-256 | 3cdcc80f04988e66ac71a8d968c4239c374ef718f7ae2ded7931feec595e7804 |
Close
Hashes for plain_transformers-0.0.1.5rc2-py3-none-any.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 0ced02d85a10b9e7f5f03c5fa0e7ca440e9b9c5c4b4f03fce1a21211908ce29d |
|
MD5 | f8f316476b6da0fa4e952aaf03f5abd7 |
|
BLAKE2b-256 | 79b7495c85109c5a33da34d7dcde5ca874e525bde573281d89e9dd6e30104596 |