rwkv · PyPI

The RWKV Language Model

These details have not been verified by PyPI

Project links

License
- OSI Approved :: Apache Software License
Operating System
- OS Independent
Programming Language
- Python :: 3

Project description

The RWKV Language Model

# !!! set these before import RWKV !!!

# os.environ["RWKV_V7_ON"] = '1' # ==> enable RWKV-7 mode
os.environ['RWKV_JIT_ON'] = '1' # '1' for better speed
os.environ["RWKV_CUDA_ON"] = '0' # '1' to compile CUDA kernel (10x faster), requires c++ compiler & cuda libraries

########################################################################################################
#
# Use '/' in model path, instead of '\'. Use ctx4096 models if you need long ctx.
#
# fp16 = good for GPU
# fp32 = good for CPU
# bf16 = supports CPU
# xxxi8 (example: fp16i8, fp32i8) = xxx with int8 quantization to save 50% VRAM/RAM, slower, slightly less accuracy
#
# We consider [ln_out+head] to be an extra layer, so L12-D768 (169M) has "13" layers, L24-D2048 (1.5B) has "25" layers, etc.
# Strategy Examples: (device = cpu/cuda/cuda:0/cuda:1/...)
# 'cpu fp32' = all layers cpu fp32
# 'cuda fp16' = all layers cuda fp16
# 'cuda fp16i8' = all layers cuda fp16 with int8 quantization
# 'cuda fp16i8 *10 -> cpu fp32' = first 10 layers cuda fp16i8, then cpu fp32 (increase 10 for better speed)
# 'cuda:0 fp16 *10 -> cuda:1 fp16 *8 -> cpu fp32' = first 10 layers cuda:0 fp16, then 8 layers cuda:1 fp16, then cpu fp32
#
# Basic Strategy Guide: (fp16i8 works for any GPU)
# 100% VRAM = 'cuda fp16'                   # all layers cuda fp16
#  98% VRAM = 'cuda fp16i8 *1 -> cuda fp16' # first 1 layer  cuda fp16i8, then cuda fp16
#  96% VRAM = 'cuda fp16i8 *2 -> cuda fp16' # first 2 layers cuda fp16i8, then cuda fp16
#  94% VRAM = 'cuda fp16i8 *3 -> cuda fp16' # first 3 layers cuda fp16i8, then cuda fp16
#  ...
#  50% VRAM = 'cuda fp16i8'                 # all layers cuda fp16i8
#  48% VRAM = 'cuda fp16i8 -> cpu fp32 *1'  # most layers cuda fp16i8, last 1 layer  cpu fp32
#  46% VRAM = 'cuda fp16i8 -> cpu fp32 *2'  # most layers cuda fp16i8, last 2 layers cpu fp32
#  44% VRAM = 'cuda fp16i8 -> cpu fp32 *3'  # most layers cuda fp16i8, last 3 layers cpu fp32
#  ...
#   0% VRAM = 'cpu fp32'                    # all layers cpu fp32
#
# Use '+' for STREAM mode, which can save VRAM too, and it is sometimes faster
# 'cuda fp16i8 *10+' = first 10 layers cuda fp16i8, then fp16i8 stream the rest to it (increase 10 for better speed)
#
# Extreme STREAM: 3G VRAM is enough to run RWKV 14B (slow. will be faster in future)
# 'cuda fp16i8 *0+ -> cpu fp32 *1' = stream all layers cuda fp16i8, last 1 layer [ln_out+head] cpu fp32
#
# ########################################################################################################

from rwkv.model import RWKV
from rwkv.utils import PIPELINE, PIPELINE_ARGS

# download models: https://huggingface.co/BlinkDL
model = RWKV(model='RWKV-x060-World-1B6-v2.1-20240328-ctx4096', strategy='cpu fp32')

pipeline = PIPELINE(model, "rwkv_vocab_v20230424") # for "world" models
# pipeline = PIPELINE(model, "20B_tokenizer.json") # for "pile" models, 20B_tokenizer.json is in https://github.com/BlinkDL/ChatRWKV

ctx = "\nIn a shocking finding, scientist discovered a herd of dragons living in a remote, previously unexplored valley, in Tibet. Even more surprising to the researchers was the fact that the dragons spoke perfect Chinese."
print(ctx, end='')

def my_print(s):
    print(s, end='', flush=True)

# For alpha_frequency and alpha_presence, see "Frequency and presence penalties":
# https://platform.openai.com/docs/api-reference/parameter-details

args = PIPELINE_ARGS(temperature = 1.0, top_p = 0.7, top_k = 100, # top_k = 0 then ignore
                     alpha_frequency = 0.25,
                     alpha_presence = 0.25,
                     alpha_decay = 0.996, # gradually decay the penalty
                     token_ban = [], # ban the generation of some tokens
                     token_stop = [], # stop generation whenever you see any token here
                     chunk_len = 256) # split input into chunks to save VRAM (shorter -> slower)

pipeline.generate(ctx, token_count=200, args=args, callback=my_print)
print('\n')

# !!! model.forward(tokens, state) will modify state in-place !!!

out, state = model.forward([187, 510, 1563, 310, 247], None)
print(out.detach().cpu().numpy())                   # get logits
out, state = model.forward([187, 510], None)
out, state = model.forward([1563], state)           # RNN has state (use deepcopy to clone states)
out, state = model.forward([310, 247], state)
print(out.detach().cpu().numpy())                   # same result as above
print('\n')

Project details

These details have not been verified by PyPI

Project links

License
- OSI Approved :: Apache Software License
Operating System
- OS Independent
Programming Language
- Python :: 3

Release history Release notifications | RSS feed

This version

0.8.29

May 2, 2025

0.8.28

Dec 7, 2024

0.8.27

Dec 6, 2024

0.8.26

Apr 26, 2024

0.8.25

Feb 10, 2024

0.8.24

Feb 1, 2024

0.8.23

Feb 1, 2024

0.8.22

Nov 17, 2023

0.8.21

Nov 15, 2023

0.8.20

Nov 3, 2023

0.8.19

Oct 31, 2023

0.8.18

Oct 31, 2023

0.8.17

Oct 30, 2023

0.8.16

Oct 8, 2023

0.8.15

Oct 7, 2023

0.8.14

Oct 5, 2023

0.8.13

Sep 27, 2023

0.8.12

Sep 4, 2023

0.8.11

Sep 2, 2023

0.8.10

Sep 2, 2023

0.8.9

Aug 5, 2023

0.8.8

Aug 4, 2023

0.8.7

Jul 29, 2023

0.8.6

Jul 29, 2023

0.8.5

Jul 29, 2023

0.8.0

Jun 26, 2023

0.7.5

Jun 10, 2023

0.7.4

May 19, 2023

0.7.3

Apr 4, 2023

0.7.2

Mar 30, 2023

0.7.1

Mar 22, 2023

0.7.0

Mar 19, 2023

0.6.2

Mar 18, 2023

0.6.1

Mar 18, 2023

0.6.0

Mar 17, 2023

0.5.0

Mar 15, 2023

0.4.2

Mar 13, 2023

0.4.1

Mar 13, 2023

0.4.0

Mar 13, 2023

0.3.1

Mar 12, 2023

0.3.0

Mar 12, 2023

0.2.1

Mar 11, 2023

0.2.0

Mar 8, 2023

0.1.0

Mar 7, 2023

0.0.9

Mar 7, 2023

0.0.8

Mar 6, 2023

0.0.7

Mar 5, 2023

0.0.6

Mar 1, 2023

0.0.5

Mar 1, 2023

0.0.4

Mar 1, 2023

0.0.3

Mar 1, 2023

0.0.2

Mar 1, 2023

0.0.1

Mar 1, 2023

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

rwkv-0.8.29.tar.gz (407.9 kB view details)

Uploaded May 2, 2025 Source

Built Distribution

rwkv-0.8.29-py3-none-any.whl (410.1 kB view details)

Uploaded May 2, 2025 Python 3

File details

Details for the file rwkv-0.8.29.tar.gz.

File metadata

Download URL: rwkv-0.8.29.tar.gz
Upload date: May 2, 2025
Size: 407.9 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.0.1 CPython/3.10.13

File hashes

Hashes for rwkv-0.8.29.tar.gz
Algorithm	Hash digest
SHA256	`8cf79910b8a852e7c4a616922f72664c8ea73c510055575f8e6c84265677d111`
MD5	`62d371af62a564821bdc7d1c73f87bfc`
BLAKE2b-256	`804845dded18c266f4ff9fc9b49c76ed29ad6f0ac6b49de8147d9daed5c28363`

See more details on using hashes here.

File details

Details for the file rwkv-0.8.29-py3-none-any.whl.

File metadata

Download URL: rwkv-0.8.29-py3-none-any.whl
Upload date: May 2, 2025
Size: 410.1 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.0.1 CPython/3.10.13

File hashes

Hashes for rwkv-0.8.29-py3-none-any.whl
Algorithm	Hash digest
SHA256	`4f79838f856196e662244ac4bc0360206e554029a2994c4c74fbe1ec0c153190`
MD5	`1d5045b7bb9b42011008d45f5c05982c`
BLAKE2b-256	`a4a442c237b378e6d94a0e5b2e89d098b00b7472d4a0c0d42da1f809ebb654dc`

See more details on using hashes here.

rwkv 0.8.29

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes