What is Narwhal?
Narwhal is the first open-source LLM inference framework that automatically hot-swaps prefill and decode roles on both NVIDIA and AMD GPUs. It moves engines between roles as demand changes, on a fixed GPU fleet with model weights already loaded. It scales from a single GPU to distributed multi-node deployments.
How engines change roles
| Role controller pass | Behavior |
|---|---|
| Inputs | Measured engine profiles, offered demand, and resident work |
| Candidates | The current role split and each adjacent split, one engine move away |
| Score | The worst projected service-level objective (SLO) ratio across time to first token (TTFT), time per output token (TPOT), and decode queueing |
| Demand | Measured window demand, using the larger of the short- and long-horizon decode estimates for decode-to-prefill candidates |
| Evidence window | Closes after controller.reactive.evidence_span_s and the minimum arrivals, or after controller.reactive.evidence_max_span_s under sparse traffic |
| Move | To the adjacent split that improves the score by at least the configured margin |
| Decode to prefill | Requires a closed evidence window and stable decode demand |
| Prefill to decode | Proceeds with the evidence window open |
| Guards | Pinned engines, role floors, cooldown, dwell time, the resident-stream ceiling on decode donors, and engine lifecycle holds |
| Floor repair | One engine per monitor pass while a phase sits below its configured floor |
| New requests | Follow the revised split |
| Resident requests | Finish on their assigned engines |
Evaluation
Read the full evaluation: Evaluating Narwhal
Get the commands
Install on Linux with Python 3.11 or newer:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install narwhal-inference
narwhal-serve --version
narwhal --help
The wheel installs these commands:
For each deployment:
- Record the
narwhal-serve --versionoutput with the fleet configuration, engine image, and profiles. - Pin that version on every router host.
Try it on one GPU
Narwhal dev runs a local NVIDIA CUDA fleet on Ubuntu or Ubuntu under WSL2.
narwhal dev init
narwhal dev up
narwhal dev verify
narwhal dev status
narwhal dev down
| GPU | Engines | Template |
|---|---|---|
| NVIDIA GPU with 8 GB of VRAM or less | 2 | |
| RTX 5090 | 4 |
Bring up a fleet
Run these gates from a management workstation:
Documentation
Contributing
Built on Arrow
Narwhal's scheduling algorithms derive from Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture by Wu et al. (2025).
Metadata
Release files for narwhal-inference 0.3.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| narwhal_inference-0.3.2.tar.gz | 3.0 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| narwhal_inference-0.3.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 3.3 MB
Release files / narwhal_inference-0.3.2.tar.gz
| Download URL | narwhal_inference-0.3.2.tar.gz |
|---|---|
| Size | 3.0 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
fbb30e1c5a13683cee05b86b76d7a23ab56df7e415f8746bc57746593d8e9df7
|
|
BLAKE2b-256 checksum How to use checksums |
faebbf209e2ebeb0423650d62ccc73ff911a5c351796280b8fcfa60ff9b2ffb1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency logRelease files / narwhal_inference-0.3.2-py3-none-any.whl
| Download URL | narwhal_inference-0.3.2-py3-none-any.whl |
|---|---|
| Size | 294.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c5d8718be1f78c6765ab36ad79859299aad87ac7b05e5936ab4a482974c689b6
|
|
BLAKE2b-256 checksum How to use checksums |
6967926e0ae7f54c941917fac495d3fb78a0a0a2439d463e1fb8c92d26add5fd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency log