PolyLoRA
Minimal PyTorch runtime for batched LoRA inference where each row can use a different adapter.
PolyLoRA wraps an existing torch.nn.Module, replaces selected nn.Linear layers, and serves PEFT LoRA adapters from CPU, GPU, and optional disk caches.
Install
pip install .
With PEFT loading support:
pip install '.[peft]'
Usage
from polylora import PolyLoraConfig, PolyLoraModel
model = PolyLoraModel(
base_model,
PolyLoraConfig(
max_gpu_adapters=4,
max_rank=16,
target_modules=["query_proj", "key_proj", "value_proj", "dense"],
),
).eval()
model.load_adapter_from_disk("legal", "./adapters/legal")
model.load_adapter_from_disk("finance", "./adapters/finance")
outputs = model(**batch, adapter_ids=["legal", "finance"])
Omit adapter_ids to run the base model. Use __base__ for rows that should skip LoRA inside a mixed batch.
Caches
PolyLoRA uses three adapter tiers:
- GPU cache: fixed-size adapter slots for the active batch. Slot
0is reserved for__base__, so non-adapter rows share the same execution path. - CPU cache: LRU store for loaded adapter weights. GPU evictions can reload from CPU without touching disk.
- Disk cache: optional bounded PEFT adapter directory cache. CPU misses can reload adapters from this cold layer.
This makes small hot sets fast while still allowing a larger adapter catalog than GPU memory can hold.
Kernels
On CUDA, PolyLoRA uses Triton SGMV kernels for the LoRA A and B projections:
- Mixed batches can contain different adapter ids, including
__base__rows. - Different adapters may use different ranks, up to
max_rank. - Rank-0 rows skip adapter work, which is how base-only rows and missing layer weights are represented.
- The
Bprojection fuses scaling and add-back into the base linear output. - The implementation falls back to a PyTorch reference path on CPU or when Triton is disabled.
Adapter Layouts
Adapters do not need to cover every wrapped layer. If a model is wrapped with a larger target_modules set and an adapter only contains LoRA weights for some of those layers, missing layers are treated as rank-0 no-ops for that adapter.
PolyLoRA rejects adapters with weights outside the configured module set, which keeps mixed adapters predictable when different adapters target different subsets of the model.
Notes
- Supports standard PEFT LoRA adapters for inference.
- Does not support LoRA dropout, DoRA, RS-LoRA, or LoRA bias.
- Attention masks must be right padded when
enforce_right_padding=True.
Development
pip install -e '.[dev]'
pytest tests
Metadata
Release files for polylora 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| polylora-0.1.1.tar.gz | 18.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| polylora-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 34.4 kB
Release files / polylora-0.1.1.tar.gz
| Download URL | polylora-0.1.1.tar.gz |
|---|---|
| Size | 18.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bf06c7081999f06075a1faca1bf35d3404af69d7c7d6354a933b9c7e4846327a
|
|
BLAKE2b-256 checksum How to use checksums |
a6285ea8c6a67e57b90ce9cb300bab0f41645ad644c5850e0fcb31fe9db4c098
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|
Release files / polylora-0.1.1-py3-none-any.whl
| Download URL | polylora-0.1.1-py3-none-any.whl |
|---|---|
| Size | 16.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4a7fa4ee8b891a57d96ae595da562fd3a88f354aa44fe365a1c699c017a5f1ec
|
|
BLAKE2b-256 checksum How to use checksums |
7c7801e16c83aa8ccb016d3046899e5525f56f0c3aa75243ff93452989ad5b29
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|