picoLLM Inference Engine Python Binding
Made in Vancouver, Canada by Picovoice
picoLLM Inference Engine
picoLLM Inference Engine is a highly accurate and cross-platform SDK optimized for running compressed large language models. picoLLM Inference Engine is:
- Accurate; picoLLM Compression improves GPTQ by significant margins
- Private; LLM inference runs 100% locally.
- Cross-Platform
- Runs on CPU and GPU
- Free for open-weight models
Compatibility
- Python 3.9+
- Runs on Linux (x86_64), macOS (arm64, x86_64), Windows (x86_64, arm64), and Raspberry Pi (3, 4, 5).
Installation
pip3 install picollm
Models
picoLLM Inference Engine supports the following open-weight models. The models are on Picovoice Console.
- DeepSeek-OCR-2
deepseek-ocr-2
- EmbeddingGemma
embeddinggemma-300m
- Gemma
gemma-2bgemma-2b-itgemma-7bgemma-7b-it
- Gemma3
gemma-3-270mgemma-3-270m-it
- Llama-2
llama-2-7bllama-2-7b-chatllama-2-13bllama-2-13b-chatllama-2-70bllama-2-70b-chat
- Llama-3
llama-3-8bllama-3-8b-instructllama-3-70bllama-3-70b-instruct
- Llama-3.2
llama3.2-1b-instructllama3.2-3b-instruct
- Mistral
mistral-7b-v0.1mistral-7b-instruct-v0.1mistral-7b-instruct-v0.2
- Mixtral
mixtral-8x7b-v0.1mixtral-8x7b-instruct-v0.1
- Phi-2
phi2
- Phi-3
phi3
- Phi-3.5
phi3.5
- Qwen3-VL
qwen3-vl-2b-it
AccessKey
AccessKey is your authentication and authorization token for deploying Picovoice SDKs, including picoLLM. Anyone who is using Picovoice needs to have a valid AccessKey. You must keep your AccessKey secret. You would need internet connectivity to validate your AccessKey with Picovoice license servers even though the LLM inference is running 100% offline and completely free for open-weight models. Everyone who signs up for Picovoice Console receives a unique AccessKey.
Usage
Text models
Create an instance of the engine and generate a prompt completion:
import picollm
pllm = picollm.create(
access_key='${ACCESS_KEY}',
model_path='${MODEL_PATH}')
res = pllm.generate(prompt='${PROMPT}')
print(res.completion)
Replace ${ACCESS_KEY} with yours obtained from Picovoice Console, ${MODEL_PATH} with the path to a model file
downloaded from Picovoice Console, and ${PROMPT} with a prompt string.
Instruction-tuned models (e.g., llama-3-8b-instruct, llama-2-7b-chat, and gemma-2b-it) have a specific chat
template. You can either directly format the prompt or use a dialog helper:
dialog = pllm.get_dialog()
dialog.add_human_request(prompt)
res = pllm.generate(prompt=dialog.prompt())
dialog.add_llm_response(res.completion)
print(res.completion)
To interrupt completion generation before it has finished:
pllm.interrupt()
Finally, when done, be sure to release the resources explicitly:
pllm.release()
Vision models
To run a VLM such as qwen3-vl-2b-it:
res = pllm.generate_with_image(
prompt='${PROMPT}',
image_width=${IMAGE_NUM_PIXELS_WIDTH},
image_height=${IMAGE_NUM_PIXELS_HEIGHT},
image=${IMAGE_DATA});
print(res.completion)
Replace ${PROMPT} with a text prompt. For the image, you will need to get image height and width in number of pixels and the raw pixel values of the image in 8-bit, RGB format.
OCR models
To run an OCR model such as deepseek-ocr-2:
res = pllm.generate_ocr(
image_width=${IMAGE_NUM_PIXELS_WIDTH},
image_height=${IMAGE_NUM_PIXELS_HEIGHT},
image=${IMAGE_DATA});
print(res.completion)
For the image, you will need to get image height and width in number of pixels and the raw pixel values of the image in 8-bit, RGB format.
Embedding models
To run an embedding model such as embeddinggemma-300m:
res = pllm.generate_embeddings(prompt='${PROMPT}');
for embedding in range(len(res)):
print(embedding)
Replace ${PROMPT} with a text prompt that you want to generate embeddings for.
Demos
picollmdemo provides command-line utilities for LLM completion and chat using picoLLM.
Metadata
Release files for picollm 2.1.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| picollm-2.1.4.tar.gz | 12.9 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| picollm-2.1.4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 25.8 MB
Release files / picollm-2.1.4.tar.gz
| Download URL | picollm-2.1.4.tar.gz |
|---|---|
| Size | 12.9 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6d903db9edc131b7f144431bc902c3929d1aba23bccc43469dd4c3840bb7f2e6
|
|
BLAKE2b-256 checksum How to use checksums |
37d6f0ddc85c133c22317858505f916abbda01ecf5e3550ba1c2f965cd6ab503
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|
Release files / picollm-2.1.4-py3-none-any.whl
| Download URL | picollm-2.1.4-py3-none-any.whl |
|---|---|
| Size | 12.9 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4e4928bc5ca394bfdb15bb357ae84f4d68eaf5008277e0ac0b623bfcd61ac0f1
|
|
BLAKE2b-256 checksum How to use checksums |
e850d631c7d00d4a5712e93a07231fe9c67fda89a59e17d30f49f61d0dbe828e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|