tilearn-llm

TILEARN for LLM

These details have not been verified by PyPI

Project links

Homepage

Project description

Tilearn.llm使用说明

1. CUDA Kernel（以LLAMA为例）

支持显卡：Ampere, Ada, or Hopper GPUs (e.g., A100, A800, H100, H800)

新版本

新版本Dependencies: pytorch >= 2.0.0

该版本完全兼容huggingface接口，不需要额外的转模型操作

LLAMA1/LLAMA2 A800 16GPU seq=1024相比deepspeed zero2训练加速约20%

cuda kernel使用方法-启动脚本修改如下

### TIACC CUDA Kernel
### Open: TIACC_TRAINING_CUDA_KERNEL=1
### Close: TIACC_TRAINING_CUDA_KERNEL=0
export TIACC_TRAINING_CUDA_KERNEL=1

cuda kernel使用方法-代码修改如下

### TIACC
TIACC_TRAINING_CUDA_KERNEL = int(os.getenv('TIACC_TRAINING_CUDA_KERNEL', '0'))
if TIACC_TRAINING_CUDA_KERNEL == 1:
    from tilearn.llm.transformers import LlamaForCausalLM

### 模型接口与标准huggingface一致
model = LlamaForCausalLM.from_pretrained(...)

### TIACC
TIACC_TRAINING_CUDA_KERNEL = int(os.getenv('TIACC_TRAINING_CUDA_KERNEL', '0'))
if TIACC_TRAINING_CUDA_KERNEL == 1:
    from tilearn.llm.transformers import AutoModelForCausalLM

### 模型接口与标准huggingface一致
model = AutoModelForCausalLM.from_pretrained(...)

旧版本

旧版本Dependencies: flash-attention 请安装https://github.com/Dao-AILab/flash-attention, 建议源码安装

### compile from source
git clone --recursive https://github.com/Dao-AILab/flash-attention
cd flash-attention && python setup.py install

### install layer_norm, fused_dense and rotary kernel
cd flash-attention/csrc/layer_norm && pip3 install .
cd flash-attention/csrc/fused_dense_lib && pip install .
cd flash-attention/csrc/rotary && pip install .

该版本不兼容huggingface接口，可直接读取huggingface模型和原始cuda kernel模型（训练保存的模型结构）

由于训练保存的模型为原始cuda kernel模型，非huggingface结构，若需要huggingface模型则手动执行脚本转换

LLAMA1/LLAMA2 A800 16GPU seq=1024相比deepspeed zero2训练加速约30%

cuda kernel使用方法-启动脚本修改如下

### TIACC CUDA Kernel
### Open: TIACC_TRAINING_CUDA_KERNEL_V0=1
### Close: TIACC_TRAINING_CUDA_KERNEL_V0=0
export TIACC_TRAINING_CUDA_KERNEL_V0=1
export TIACC_TRAINING_MODEL_FORMAT=llama-hf

# 若读取huggingface模型结构，则设置llama-hf
export TIACC_TRAINING_MODEL_FORMAT=llama-hf
# 若原始cuda kernel模型，则设置llama-origin
export TIACC_TRAINING_MODEL_FORMAT=llama-origin

cuda kernel使用方法-代码修改如下

### TIACC
TIACC_TRAINING_CUDA_KERNEL_V0 = int(os.getenv('TIACC_TRAINING_CUDA_KERNEL_V0', '0'))
if TIACC_TRAINING_CUDA_KERNEL_V0 == 1:
    from tilearn import llm

### LLAMA模型初始化
TIACC_TRAINING_MODEL_FORMAT = os.getenv('TIACC_TRAINING_MODEL_FORMAT', 'llama-origin')
model = llm.models.llama(model_args.model_name_or_path, model_format=TIACC_TRAINING_MODEL_FORMAT)

2. Static Zero

适用场景：在deepspeed zero1、zero2、zero3、offload、int8等不同优化状态间切换

启动脚本修改如下

### TIACC STATIC ZERO
### Open: TIACC_TRAINING_CUDA_KERNEL='O2' 
### support 'O2' / 'O2.5' / 'O3' / 'O3.5' / 'O3_Q8'(doing)
### Close: TIACC_TRAINING_CUDA_KERNEL='None'
export TIACC_TRAINING_STATIC_ZERO='None' #'O2'

代码修改如下

from transformers import HfArgumentParser

TIACC_TRAINING_STATIC_ZERO = os.getenv('TIACC_TRAINING_STATIC_ZERO', 'None')
if TIACC_TRAINING_STATIC_ZERO != 'None':
    from tilearn.llm.transformers import TrainingArguments
	
### 接口与标准huggingface一致
parser = HfArgumentParser((ModelArguments, DataTrainingArguments, TrainingArguments))

3. Dynamic Zero

适用场景：适用于zero3 + offload场景，大幅优化显存从而提升batchsize

启动脚本修改如下

### TIACC DYNAMIC ZERO
### Open: TIACC_TRAINING_DYNAMIC_ZERO=1 and set TIACC_ZERO_STAGE/TIACC_ZERO_STAGE/TIACC_PLACEMENT/TIACC_SHARD_INIT/TIACC_CPU_INIT
### Close: TIACC_TRAINING_DYNAMIC_ZERO=0
export TIACC_TRAINING_DYNAMIC_ZERO=0
export TIACC_ZERO_STAGE=3 #work when TIACC_TRAINING_DYNAMIC_ZERO=1
export TIACC_PLACEMENT='cpu' #'cuda' #work when TIACC_TRAINING_DYNAMIC_ZERO=1
export TIACC_SHARD_INIT=0 #work when TIACC_TRAINING_DYNAMIC_ZERO=1
export TIACC_CPU_INIT=1 #work when TIACC_TRAINING_DYNAMIC_ZERO=1

if [ ${TIACC_TRAINING_DYNAMIC_ZERO} = 0 ]; then
  #USE_DS="--deepspeed=./ds_config_zero3.json"
  USE_DS="--deepspeed=${deepspeed_config_file}"
else
  USE_DS=""
fi

torchrun --nnodes 1 --nproc_per_node 8 run_clm.py \
    ${USE_DS} \
	...

代码修改如下

TIACC_TRAINING_DYNAMIC_ZERO = int(os.getenv('TIACC_TRAINING_DYNAMIC_ZERO', '0'))
from contextlib import nullcontext
if TIACC_TRAINING_DYNAMIC_ZERO == 1:
    from tilearn.llm.trainer import TrainerTiacc as Trainer
    from tilearn.llm import init as llm_init
    from tilearn.llm import get_config as llm_get_config
	

	
### init in main func
def main():
    if TIACC_TRAINING_DYNAMIC_ZERO == 1:
        llm_config = llm_get_config()
        llm_init_context = llm_init(init_in_cpu=llm_config.cpu_init,
                                    shard_init=llm_config.shard_init,
                                    model_dtype=torch.half)
									
### add init_context when model init
    init_context = llm_init_context if TIACC_TRAINING_DYNAMIC_ZERO == 1 else nullcontext
    with init_context():
		### 接口与标准huggingface一致
        model = LlamaForCausalLM.from_pretrained(
            model_args.model_name_or_path,
            config=config,
            low_cpu_mem_usage=False #True,
			...
        )
		
		
### use trainer
    ### 接口与标准huggingface一致
    trainer = Trainer(
        model=model,
        ...
    )

Project details

These details have not been verified by PyPI

Project links

Homepage

Release history Release notifications | RSS feed

1.1.4

Mar 3, 2026

1.1.3

Jan 12, 2026

1.1.2

Jan 12, 2026

1.1.1

Oct 28, 2025

1.1.0

Sep 17, 2025

1.0.3.1

Jan 15, 2025

1.0.3

Jan 15, 2025

1.0.2

Jan 14, 2025

1.0.1

Jan 10, 2025

1.0.0.5

Jan 1, 2025

1.0.0.4

Dec 30, 2024

1.0.0.3

Dec 23, 2024

1.0.0.2.1

Jan 8, 2025

1.0.0.2

Dec 20, 2024

1.0.0.1

Dec 13, 2024

0.10.2.1

Dec 10, 2024

0.10.2

Dec 9, 2024

0.10.1.1

Nov 21, 2024

0.10.1

Nov 20, 2024

0.10.0.1

Nov 1, 2024

0.10.0

Oct 28, 2024

0.9.16.6

Sep 18, 2024

0.9.16.5

Sep 14, 2024

0.9.16.4

Sep 14, 2024

0.9.16.3

Sep 11, 2024

0.9.16.2

Sep 9, 2024

0.9.16.1

Sep 9, 2024

0.9.16

Sep 6, 2024

0.9.14

Sep 5, 2024

0.9.13

Sep 4, 2024

0.9.12

Sep 3, 2024

0.9.11

Aug 29, 2024

0.9.10

Aug 6, 2024

0.9.9

Jul 1, 2024

0.9.8

Jun 26, 2024

0.9.7

Jun 24, 2024

0.9.6

Jun 12, 2024

0.9.5

Jun 6, 2024

0.9.4

Jun 4, 2024

0.9.3.3

Jun 11, 2024

0.9.3.1

Jun 5, 2024

0.9.3

May 13, 2024

0.9.2

May 8, 2024

0.9.1

May 8, 2024

0.9.1.dev2 pre-release

May 8, 2024

0.9.0

May 8, 2024

0.8.7

Apr 19, 2024

0.8.6

Apr 11, 2024

0.8.5

Apr 10, 2024

0.8.4

Mar 29, 2024

0.8.3

Mar 27, 2024

0.8.2

Mar 23, 2024

0.8.1

Mar 23, 2024

0.8.0.3

May 25, 2024

0.8.0.2

May 25, 2024

0.8.0.1

May 25, 2024

0.8.0

Mar 23, 2024

0.7.15

Mar 4, 2024

0.7.14

Mar 4, 2024

0.7.13

Feb 29, 2024

0.7.12

Feb 2, 2024

0.7.11

Feb 2, 2024

0.7.10

Jan 17, 2024

0.7.9

Jan 11, 2024

0.7.7

Jan 3, 2024

0.7.6

Dec 27, 2023

0.7.5

Dec 27, 2023

0.7.3

Dec 18, 2023

0.7.2

Dec 5, 2023

0.7.1

Dec 1, 2023

0.7.0

Dec 1, 2023

0.6.4

Nov 3, 2023

0.6.3

Oct 23, 2023

0.6.2

Oct 12, 2023

0.6.1

Oct 11, 2023

0.6.0

Oct 11, 2023

0.5.12

Oct 10, 2023

0.5.11

Oct 8, 2023

0.5.10

Sep 22, 2023

0.5.9

Sep 15, 2023

0.5.8

Sep 12, 2023

0.5.7

Sep 12, 2023

0.5.6

Sep 5, 2023

0.5.5

Sep 5, 2023

0.5.4

Sep 5, 2023

0.5.2

Sep 1, 2023

This version

0.5.1

Aug 31, 2023

0.5.0

Aug 30, 2023

0.4.3

Aug 18, 2023

0.4.2

Aug 14, 2023

0.4.1

Aug 11, 2023

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

tilearn_llm-0.5.1-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.5 MB view details)

Uploaded Aug 31, 2023 CPython 3.10manylinux: glibc 2.17+ x86-64

tilearn_llm-0.5.1-cp39-cp39-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.5 MB view details)

Uploaded Aug 31, 2023 CPython 3.9manylinux: glibc 2.17+ x86-64

tilearn_llm-0.5.1-cp38-cp38-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.4 MB view details)

Uploaded Aug 31, 2023 CPython 3.8manylinux: glibc 2.17+ x86-64

File details

Details for the file tilearn_llm-0.5.1-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

Download URL: tilearn_llm-0.5.1-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Upload date: Aug 31, 2023
Size: 2.5 MB
Tags: CPython 3.10, manylinux: glibc 2.17+ x86-64
Uploaded using Trusted Publishing? No
Uploaded via: twine/4.0.2 CPython/3.9.13

File hashes

Hashes for tilearn_llm-0.5.1-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm	Hash digest
SHA256	`e39a8670ec8b574b11af374ca11d062f3a3971291ffe18a2088e3cac19bcc823`
MD5	`5ab51c12daab2a6e78d4e8b25bfd0811`
BLAKE2b-256	`0d62e22f7f893e5bfb7f6177ee8ef984567b41aad99ded71376931ebb832d605`

See more details on using hashes here.

File details

Details for the file tilearn_llm-0.5.1-cp39-cp39-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

Download URL: tilearn_llm-0.5.1-cp39-cp39-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Upload date: Aug 31, 2023
Size: 2.5 MB
Tags: CPython 3.9, manylinux: glibc 2.17+ x86-64
Uploaded using Trusted Publishing? No
Uploaded via: twine/4.0.2 CPython/3.9.13

File hashes

Hashes for tilearn_llm-0.5.1-cp39-cp39-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm	Hash digest
SHA256	`3d85cbe5675078730610dca1f1f3fe1a5246bc42cdd00272edcfe72a7912748e`
MD5	`5972cbc6076991105db4d2237ef27bce`
BLAKE2b-256	`3aaddbb4fdad853ab8afa3d50742b7d52f42a09541d9e5ae13bb930a11efde21`

See more details on using hashes here.

File details

Details for the file tilearn_llm-0.5.1-cp38-cp38-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

Download URL: tilearn_llm-0.5.1-cp38-cp38-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Upload date: Aug 31, 2023
Size: 2.4 MB
Tags: CPython 3.8, manylinux: glibc 2.17+ x86-64
Uploaded using Trusted Publishing? No
Uploaded via: twine/4.0.2 CPython/3.9.13

File hashes

Hashes for tilearn_llm-0.5.1-cp38-cp38-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm	Hash digest
SHA256	`cc456e867e7618bc5a462d764d792f606adfb218f6c3e67544fc069a9f58e324`
MD5	`88b392a4b6399544588801f1bc1eeb2b`
BLAKE2b-256	`db466266318f0e5908e02949d1cf255987dfac3717d59589cda9f4045e7f98d6`

See more details on using hashes here.

tilearn-llm 0.5.1

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Tilearn.llm使用说明

1. CUDA Kernel（以LLAMA为例）

新版本

旧版本

2. Static Zero

3. Dynamic Zero

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distributions

Built Distributions

File details

File metadata

File hashes

File details

File metadata

File hashes

File details

File metadata

File hashes