This release has been yanked by its maintainers, and will be ignored by installers, except when explicitly specified.
GrillyInference
Native fp16 inference engine for Llama-family models — optional grilly extension.
Features
- Native fp16 inference — runs Llama 3.2 3B at ~6.4 GB VRAM, zero quality loss
- Paged KV-Cache — 256-token SRAM pages with LRU eviction, 4x context extension
- H2O Eviction — exponential decay on old KV heads, 32k context on 12GB
- VSA Multi-Scale Summaries — hypervector bind/bundle for 128k effective context
- SmoothQuant INT8 — per-group-64 weight quantization, <1% PPL loss
- 4-bit Block Quantization — run 100B models on 12GB VRAM with layer offloading
- Llama 3.2 Instruct — chat template, streaming generation, top-k/top-p sampling
Quick Start
pip install grillyinference
from grillyinference import LlamaForCausalLM, TextGenerator
from transformers import AutoTokenizer
model = LlamaForCausalLM.from_pretrained("meta-llama/Llama-3.2-3B-Instruct")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.2-3B-Instruct")
gen = TextGenerator(model, tokenizer)
# Simple generation
response = gen.generate("What is the meaning of life?", max_tokens=256)
print(response)
# Chat
response = gen.chat([
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain transformers in 3 sentences."},
])
print(response)
# Streaming
for token in gen.generate("Once upon a time", stream=True):
print(token, end="", flush=True)
Context Extension (12GB VRAM)
| Context | Decode Speed | PPL Hit | Technique |
|---|---|---|---|
| 2k | 9 t/s | 0% | Baseline |
| 8k | 8 t/s | 0% | PagedAttention |
| 32k | 7 t/s | 1.5% | + H2O eviction |
| 128k | 6 t/s | 3% | + VSA summaries |
from grillyinference import KVCache, LlamaConfig
config = LlamaConfig.from_pretrained("meta-llama/Llama-3.2-3B-Instruct")
kv = KVCache(config, raw_window=2048, h2o_lambda=0.0002, enable_vsa=True)
SmoothQuant INT8
from grillyinference.inference.quantize import SmoothQuantCalibrator, SmoothQuantizer
calibrator = SmoothQuantCalibrator(model, tokenizer)
stats = calibrator.calibrate()
quantizer = SmoothQuantizer(group_size=64)
quantized = quantizer.smooth_and_quantize(model._weights, stats)
Requirements
- Python 3.12+
- grilly >= 0.4.0
- numpy, safetensors
- Optional: huggingface_hub, transformers (for
from_pretrained)
License
MIT
Metadata
Release files for grillyinference 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| grillyinference-0.1.0.tar.gz | 31.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| grillyinference-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 54.8 kB
Release files / grillyinference-0.1.0.tar.gz
| Download URL | grillyinference-0.1.0.tar.gz |
|---|---|
| Size | 31.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c4558ddd0406f99ae012ff709dcd6da20a90dcbdd32db35caa55ccd282c57321
|
|
BLAKE2b-256 checksum How to use checksums |
b25c99b3791a328f79f0af05bc0c719de0f70ed1fa62f1268b3a9821ee4d299f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Feb 22, 2026.
Transparency logRelease files / grillyinference-0.1.0-py3-none-any.whl
| Download URL | grillyinference-0.1.0-py3-none-any.whl |
|---|---|
| Size | 23.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ad234a0f98003d34eaa35c7596242b245a73d47d3e2a27a2557943c8e6a5657b
|
|
BLAKE2b-256 checksum How to use checksums |
2b588c87ce6e29385bd9bf2963484319ce62977cde79b01c7452426f003fefe3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Feb 22, 2026.
Transparency log