Skip to main content

A lightweight, hackable LLM inference engine built from scratch.

Project description

liteinfer

PyPI

A lightweight, hackable LLM inference engine built from scratch — designed to make state-of-the-art inference techniques (paged KV cache, prefix caching, tensor parallelism, torch.compile, CUDA graphs, …) easy to read, test, and benchmark.

Goals

  1. Fast offline inference — throughput in the same league as vLLM on a single node.
  2. Readable codebase — clean, minimal, well-structured. The core engine should fit in your head.
  3. Optimization suite — a clear place for each technique (prefix caching, TP, torch.compile, CUDA graphs, …) with isolated, testable implementations.
  4. HuggingFace compatibility — load any compatible HF model from the Hub or a local safetensors directory.

Status

v0 — minimal end-to-end greedy/sampled inference on local safetensors. Static batching (B > 1), paged KV cache. No continuous batching yet. See docs/milestones.md for what is in, and docs/roadmap.md for what is queued.

Installation

pip install liteinfer

For development or benchmark comparisons:

git clone https://github.com/ValeGian/liteinfer.git
cd liteinfer
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
# Optional: install vLLM for benchmark comparisons
pip install -e ".[dev,bench]"

Quick start

from liteinfer import LLM, SamplingParams

llm = LLM(model="meta-llama/Llama-3.2-1B-Instruct")
params = SamplingParams(temperature=0.8, max_tokens=128)

outputs = llm.generate(["Explain paged attention in one paragraph."], params)
print(outputs[0].text)

Repository layout

liteinfer/
├── liteinfer/             # Library source
│   ├── llm.py             # User-facing LLM class
│   ├── config.py          # EngineConfig
│   ├── engine/            # Orchestration: scheduler, sequence, model runner
│   ├── models/            # Model loaders + per-architecture implementations
│   ├── layers/            # Reusable building blocks (attention, RMSNorm, …)
│   ├── cache/             # KV cache (paged, prefix-cached, …)
│   └── sampling/          # SamplingParams + Sampler
├── tests/                 # Unit / integration / e2e tests
├── benchmarks/            # vLLM comparison harness
└── pyproject.toml

Each module's __init__.py documents the contract it owns.

Architecture (brief)

User code calls LLM, a thin facade over LLMEngine, which owns:

  • Scheduler — picks which sequences run on the next forward pass (continuous batching; later, prefix-cache aware).
  • ModelRunner — runs the actual forward pass for the selected batch on the GPU. Tensor parallelism, torch.compile, and CUDA graph capture plug in here.
  • KVCache — paged blocks shared across sequences. Prefix caching is a KVCache variant.

Sampling is a separate stage so strategies (greedy, top-p, …) can be swapped without touching the engine.

Testing

liteinfer is test-first: every feature ships with the tests that pin its contract.

pytest                              # full suite — runs sequentially, GPU-safe
pytest -m "not gpu and not slow"    # fast suite (no model downloads, no GPU)
pytest -m gpu                       # GPU tests only — sequential, never use -n auto
pytest -n auto                      # parallel mode — CPU-only tests only
pytest tests/unit/                  # one directory

GPU tests: always run sequentially. The e2e tests (tests/e2e/) load real models onto GPU. Running them in parallel (-n auto) will cause OOM or cross-process interference. tests/e2e/test_vllm_runner.py additionally requires the bench extras (pip install -e ".[dev,bench]").

Test layout:

  • tests/unit/ — single-component tests. CPU-only, fast, no model loading.
  • tests/integration/ — multiple components wired together (still no HF download).
  • tests/e2e/ — load a small real model and verify generation against transformers.

See tests/README.md for conventions.

Performance

Llama-3.2-1B-Instruct · greedy · NVIDIA A40 — see docs/benchmarks.md for full setup and methodology, or open the interactive dashboard for a visual comparison.

Engines: liteinfer (no KV cache, RECOMPUTE) · liteinfer-kvcache (eager KV cache) · liteinfer-paged-kvcache (paged KV cache) · liteinfer-b4 (eager KV cache + static batching B=4) · vllm (B=1) · vllm-b4 (B=4).

Throughput — 32 requests submitted at once. E2E includes queue wait time.

Engine B req/s tok/s E2E p50 E2E p99
liteinfer 1 1.75 73 11607 ms 18249 ms
liteinfer-kvcache 1 1.81 72 10190 ms 17648 ms
liteinfer-paged-kvcache 1 1.57 63 11772 ms 20397 ms
liteinfer-b4 4 4.31 177 4220 ms 7421 ms
vllm 1 2.90 181 5796 ms 10999 ms
vllm-b4 4 10.77 656 1625 ms 2932 ms

Latency — sequential, no queue; each request sent only after previous finishes.

Engine B TTFT p50 TTFT p99 E2E p50 tok/s
liteinfer 1 13.7 ms 15.4 ms 1710 ms 72
liteinfer-kvcache 1 14.6 ms 16.8 ms 1541 ms 73
liteinfer-paged-kvcache 1 14.7 ms 16.7 ms 1829 ms 61
vllm 1 25.9 ms 28.8 ms 693 ms 182

Benchmarking against vLLM

Every engine implements the same EngineRunner interface, so comparing liteinfer against vLLM (or future variants of liteinfer itself) is a single command:

python -m benchmarks.compare \
    --model meta-llama/Llama-3.2-1B-Instruct \
    --engines liteinfer vllm \
    --workload throughput \
    --output benchmarks/results/throughput.json

Metrics: requests/sec, output tokens/sec, TTFT (p50/p99), inter-token latency, peak GPU memory. See benchmarks/README.md for adding workloads or new engines.

Roadmap

High-level direction:

  • Single-prompt greedy/sampled generation from local safetensors
  • Static batching (B > 1)
  • Paged KV cache
  • Continuous batching → prefix caching
  • torch.compile and CUDA graphs for decode
  • Tensor parallelism (single node)
  • Speculative decoding

Detailed, fine-grained backlog with scope and parity-test notes lives in docs/roadmap.md. Achieved milestones are tracked separately in docs/milestones.md.

License

Apache-2.0.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

liteinfer-0.1.4.tar.gz (86.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

liteinfer-0.1.4-py3-none-any.whl (46.2 kB view details)

Uploaded Python 3

File details

Details for the file liteinfer-0.1.4.tar.gz.

File metadata

  • Download URL: liteinfer-0.1.4.tar.gz
  • Upload date:
  • Size: 86.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for liteinfer-0.1.4.tar.gz
Algorithm Hash digest
SHA256 77d12c7341e98335d75a62c73d505190c9cc05c8adcd9ceb9864fd52c97f5428
MD5 597e75624dada9392a1d4f77ae06298d
BLAKE2b-256 186e2d38db9fde4af6df81bc47371af0e65d5e862a55000e51ad51bfbe2f463c

See more details on using hashes here.

Provenance

The following attestation bundles were made for liteinfer-0.1.4.tar.gz:

Publisher: publish.yml on ValeGian/liteinfer

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file liteinfer-0.1.4-py3-none-any.whl.

File metadata

  • Download URL: liteinfer-0.1.4-py3-none-any.whl
  • Upload date:
  • Size: 46.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for liteinfer-0.1.4-py3-none-any.whl
Algorithm Hash digest
SHA256 eaaec0e39d7ac3d4475048b74d00edaa7e7e99c0fc3a10e5cf55b4a01b241afd
MD5 dd28394b68f5233d5a64f197408a1db4
BLAKE2b-256 3367bb679d7bfc4f622e7d98cda54d980c5079f0434d092e513c2384654f7c7e

See more details on using hashes here.

Provenance

The following attestation bundles were made for liteinfer-0.1.4-py3-none-any.whl:

Publisher: publish.yml on ValeGian/liteinfer

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page