Skip to main content

A lightweight, hackable LLM inference engine built from scratch.

Project description

liteinfer

PyPI

A lightweight, hackable LLM inference engine built from scratch — designed to make state-of-the-art inference techniques (paged KV cache, prefix caching, tensor parallelism, torch.compile, CUDA graphs, …) easy to read, test, and benchmark.

Goals

  1. Fast offline inference — throughput in the same league as vLLM on a single node.
  2. Readable codebase — clean, minimal, well-structured. The core engine should fit in your head.
  3. Optimization suite — a clear place for each technique (prefix caching, TP, torch.compile, CUDA graphs, …) with isolated, testable implementations.
  4. HuggingFace compatibility — load any compatible HF model from the Hub or a local safetensors directory.

Status

v0 — minimal end-to-end greedy/sampled inference on local safetensors. Single-prompt at a time, no continuous batching, no paged cache. See docs/milestones.md for what is in, and docs/roadmap.md for what is queued.

Installation

pip install liteinfer

For development or benchmark comparisons:

git clone https://github.com/ValeGian/liteinfer.git
cd liteinfer
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
# Optional: install vLLM for benchmark comparisons
pip install -e ".[dev,bench]"

Quick start

from liteinfer import LLM, SamplingParams

llm = LLM(model="meta-llama/Llama-3.2-1B-Instruct")
params = SamplingParams(temperature=0.8, max_tokens=128)

outputs = llm.generate(["Explain paged attention in one paragraph."], params)
print(outputs[0].text)

Repository layout

liteinfer/
├── liteinfer/             # Library source
│   ├── llm.py             # User-facing LLM class
│   ├── config.py          # EngineConfig
│   ├── engine/            # Orchestration: scheduler, sequence, model runner
│   ├── models/            # Model loaders + per-architecture implementations
│   ├── layers/            # Reusable building blocks (attention, RMSNorm, …)
│   ├── cache/             # KV cache (paged, prefix-cached, …)
│   └── sampling/          # SamplingParams + Sampler
├── tests/                 # Unit / integration / e2e tests
├── benchmarks/            # vLLM comparison harness
└── pyproject.toml

Each module's __init__.py documents the contract it owns.

Architecture (brief)

User code calls LLM, a thin facade over LLMEngine, which owns:

  • Scheduler — picks which sequences run on the next forward pass (continuous batching; later, prefix-cache aware).
  • ModelRunner — runs the actual forward pass for the selected batch on the GPU. Tensor parallelism, torch.compile, and CUDA graph capture plug in here.
  • KVCache — paged blocks shared across sequences. Prefix caching is a KVCache variant.

Sampling is a separate stage so strategies (greedy, top-p, …) can be swapped without touching the engine.

Testing

liteinfer is test-first: every feature ships with the tests that pin its contract.

pytest                              # full suite — runs sequentially, GPU-safe
pytest -m "not gpu and not slow"    # fast suite (no model downloads, no GPU)
pytest -m gpu                       # GPU tests only — sequential, never use -n auto
pytest -n auto                      # parallel mode — CPU-only tests only
pytest tests/unit/                  # one directory

GPU tests: always run sequentially. The e2e tests (tests/e2e/) load real models onto GPU. Running them in parallel (-n auto) will cause OOM or cross-process interference. tests/e2e/test_vllm_runner.py additionally requires the bench extras (pip install -e ".[dev,bench]").

Test layout:

  • tests/unit/ — single-component tests. CPU-only, fast, no model loading.
  • tests/integration/ — multiple components wired together (still no HF download).
  • tests/e2e/ — load a small real model and verify generation against transformers.

See tests/README.md for conventions.

Performance

Llama-3.2-1B-Instruct · greedy · NVIDIA A40 · batch size 1 (all engines) liteinfer v0.0.6 vs vLLM 0.20.0 — see docs/benchmarks.md for full setup and methodology, or open the interactive dashboard for a visual comparison.

liteinfer = no KV cache (RECOMPUTE mode, v0 default) · liteinfer-kvcache = eager KV cache enabled

Throughput — 32 requests submitted at once, engine queues them (B=1 → sequential). E2E includes queue wait time.

Engine (B=1) req/s tok/s E2E p50 E2E p99
liteinfer 1.69 70 11956 ms 18914 ms
liteinfer-kvcache 1.77 71 10480 ms 18114 ms
vllm 2.89 180 5786 ms 11042 ms

Latency — sequential, no queue; each request sent only after previous finishes.

Engine (B=1) TTFT p50 TTFT p99 E2E p50 tok/s
liteinfer 14 ms 16 ms 1726 ms 72
liteinfer-kvcache 15 ms 17 ms 1541 ms 72
vllm 26 ms 31 ms 694 ms 182

Benchmarking against vLLM

Every engine implements the same EngineRunner interface, so comparing liteinfer against vLLM (or future variants of liteinfer itself) is a single command:

python -m benchmarks.compare \
    --model meta-llama/Llama-3.2-1B-Instruct \
    --engines liteinfer vllm \
    --workload throughput \
    --output benchmarks/results/throughput.json

Metrics: requests/sec, output tokens/sec, TTFT (p50/p99), inter-token latency, peak GPU memory. See benchmarks/README.md for adding workloads or new engines.

Roadmap

High-level direction:

  • Single-prompt greedy/sampled generation from local safetensors
  • Static batching (B > 1)
  • Paged KV cache → continuous batching → prefix caching
  • torch.compile and CUDA graphs for decode
  • Tensor parallelism (single node)
  • Speculative decoding

Detailed, fine-grained backlog with scope and parity-test notes lives in docs/roadmap.md. Achieved milestones are tracked separately in docs/milestones.md.

License

Apache-2.0.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

liteinfer-0.1.0.tar.gz (64.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

liteinfer-0.1.0-py3-none-any.whl (36.3 kB view details)

Uploaded Python 3

File details

Details for the file liteinfer-0.1.0.tar.gz.

File metadata

  • Download URL: liteinfer-0.1.0.tar.gz
  • Upload date:
  • Size: 64.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for liteinfer-0.1.0.tar.gz
Algorithm Hash digest
SHA256 c315fdaa01bd674149ce4aa945d06a9ecf696c61ebec22077d995c4b5126badd
MD5 7f784bbaf55813aeed38189b8c01429f
BLAKE2b-256 c90225e5fac570e084996247df7c30363c55ce0da1f41bb1cd44c0fe4c179c18

See more details on using hashes here.

Provenance

The following attestation bundles were made for liteinfer-0.1.0.tar.gz:

Publisher: publish.yml on ValeGian/liteinfer

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file liteinfer-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: liteinfer-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 36.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for liteinfer-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 0f8c0764d7b0396da6bba4c74957c755b50dcf702c577d5007960389439a26c7
MD5 ed3bca89a7989b10ae9097c08e06b7b1
BLAKE2b-256 cf81b009149575f68a1741ecb20670f39422207b61c99c0ff9b4cc9cff04600a

See more details on using hashes here.

Provenance

The following attestation bundles were made for liteinfer-0.1.0-py3-none-any.whl:

Publisher: publish.yml on ValeGian/liteinfer

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page