Skip to main content
Yanked

This release has been yanked by its maintainers, and will be ignored by installers, except when explicitly specified.

GrillyInference

Native fp16 inference engine for Llama-family models — optional grilly extension.

Features

  • Native fp16 inference — runs Llama 3.2 3B at ~6.4 GB VRAM, zero quality loss
  • Paged KV-Cache — 256-token SRAM pages with LRU eviction, 4x context extension
  • H2O Eviction — exponential decay on old KV heads, 32k context on 12GB
  • VSA Multi-Scale Summaries — hypervector bind/bundle for 128k effective context
  • SmoothQuant INT8 — per-group-64 weight quantization, <1% PPL loss
  • 4-bit Block Quantization — run 100B models on 12GB VRAM with layer offloading
  • Llama 3.2 Instruct — chat template, streaming generation, top-k/top-p sampling

Quick Start

pip install grillyinference
from grillyinference import LlamaForCausalLM, TextGenerator
from transformers import AutoTokenizer

model = LlamaForCausalLM.from_pretrained("meta-llama/Llama-3.2-3B-Instruct")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.2-3B-Instruct")
gen = TextGenerator(model, tokenizer)

# Simple generation
response = gen.generate("What is the meaning of life?", max_tokens=256)
print(response)

# Chat
response = gen.chat([
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Explain transformers in 3 sentences."},
])
print(response)

# Streaming
for token in gen.generate("Once upon a time", stream=True):
    print(token, end="", flush=True)

Context Extension (12GB VRAM)

Context Decode Speed PPL Hit Technique
2k 9 t/s 0% Baseline
8k 8 t/s 0% PagedAttention
32k 7 t/s 1.5% + H2O eviction
128k 6 t/s 3% + VSA summaries
from grillyinference import KVCache, LlamaConfig

config = LlamaConfig.from_pretrained("meta-llama/Llama-3.2-3B-Instruct")
kv = KVCache(config, raw_window=2048, h2o_lambda=0.0002, enable_vsa=True)

SmoothQuant INT8

from grillyinference.inference.quantize import SmoothQuantCalibrator, SmoothQuantizer

calibrator = SmoothQuantCalibrator(model, tokenizer)
stats = calibrator.calibrate()
quantizer = SmoothQuantizer(group_size=64)
quantized = quantizer.smooth_and_quantize(model._weights, stats)

Requirements

  • Python 3.12+
  • grilly >= 0.4.0
  • numpy, safetensors
  • Optional: huggingface_hub, transformers (for from_pretrained)

License

MIT

Metadata

Release files for grillyinference 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for grillyinference 0.1.0
File Size Uploaded
grillyinference-0.1.0.tar.gz 31.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for grillyinference 0.1.0
File Interpreter ABI Platform
grillyinference-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 54.8 kB

Release files / grillyinference-0.1.0.tar.gz

Download URL grillyinference-0.1.0.tar.gz
Size 31.5 kB
Tags Source
SHA-256 checksum
How to use checksums
c4558ddd0406f99ae012ff709dcd6da20a90dcbdd32db35caa55ccd282c57321
BLAKE2b-256 checksum
How to use checksums
b25c99b3791a328f79f0af05bc0c719de0f70ed1fa62f1268b3a9821ee4d299f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Feb 22, 2026.

Transparency log

Release files / grillyinference-0.1.0-py3-none-any.whl

Download URL grillyinference-0.1.0-py3-none-any.whl
Size 23.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ad234a0f98003d34eaa35c7596242b245a73d47d3e2a27a2557943c8e6a5657b
BLAKE2b-256 checksum
How to use checksums
2b588c87ce6e29385bd9bf2963484319ce62977cde79b01c7452426f003fefe3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Feb 22, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page