Skip to main content

LLM Falcon model

llm_falcon_model allows you to run a part of a Falcon model as a standalone PyTorch module. This enables you to run in distributed mode, using even old GPUs with less memory.

It only contains code needed for inference. The only dependencies are torch, tokenizers, and llm_sepweight.

The original implementation is available here.

Use it when you cannot fit the whole Falcon model into memory. If you have multiple old GPUs with less memory, you can run different parts of the Falcon model on each of them and when you make them communicate (using for example socket_rpc), you can run the full model on multiple heterogeneous hosts. For example, if you have 4 old gaming PCs with a 3090 card (~6000$), you can run Falcon 40B real-time (5-6 tokens/s)

You can also use it when you want to run Falcon on a large number of inputs and have insufficient memory for the model. You can serialize the intermediary results for all inputs and then continue with the next layers

Install with:

pip install llm_falcon_model

Downloads PyPi version PyPI license

Overview

The most important methods of this microlib are:

  1. llm_falcon_model.load_tokenizer() - which loads an instance of the Tokenizer for the models
  2. llm_falcon_model.init_part(model_name, spec, device) - which creates a part of a Falcon model, by a given name (7b, 40b or 180b), part specification (which layers you want to load, see sepweight part spec) and a PyTorch device.
  3. llm_falcon_model.generate - which allows you to generate text based on a prompt.
  4. llm_falcon_model.score_batch - which allows you to score a bunch of possible continuations based on a prompt.
  5. llm_falcon_model.run_part - which allows you to run a part of Falcon in a distributed mode using socket_rpc

Quick example

import torch
import llm_falcon_model

tokenizer = llm_falcon_model.load_tokenizer()

separated_weights_path = '<PATH TO SEPARATED WEIGHTS>'

model = llm_falcon_model.init_part(
    model_name='40b',
    spec='b 0-12', # Load begin and layers 0 to 12
    device='cuda:0'
)

input_text = "The world chess champion Magnus Carlsen"
input_ids = tokenizer.encode(input_text).ids
batch = torch.tensor(input_ids).unsqueeze(0)
x = model(batch)

# x is now the result after end layer 12, shaped:
# torch.Size([1, 7, 8192])

Release files for llm-falcon-model 0.7.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-falcon-model 0.7.0
File Size Uploaded
llm_falcon_model-0.7.0.tar.gz 795.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-falcon-model 0.7.0
File Interpreter ABI Platform
llm_falcon_model-0.7.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.6 MB

Release files / llm_falcon_model-0.7.0.tar.gz

Download URL llm_falcon_model-0.7.0.tar.gz
Size 795.6 kB
Tags Source
SHA-256 checksum
How to use checksums
9a7019a9feb82569d3290d33c2582a404a529483a0c26a5499b491c85e64c8f0
BLAKE2b-256 checksum
How to use checksums
c1246560aeef959ad8477337a3fbc10dff2fd20eac42454305c8eb740042b3fe
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.11.5

Release files / llm_falcon_model-0.7.0-py3-none-any.whl

Download URL llm_falcon_model-0.7.0-py3-none-any.whl
Size 809.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a74362aa748f6d531d6b5948623ac2e46b3ccaf8bed0f61fa2ef6df0465b59d1
BLAKE2b-256 checksum
How to use checksums
d008686b735e59d465e13c06967d903dd41907890c9db3f028e1f049f8f2a83b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.11.5

Release history Release notifications | RSS feed

This release

0.7.0 This release

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page