What is Narwhal?
Narwhal is an adaptive, disaggregated inference framework that automatically hot-swaps prefill and decode roles as demand changes, and without having to reload model weights. It can scale from a single GPU to multi-node deployments.
How engines change roles
The role controller scores the current role split and each adjacent split, one engine move away. It projects each split from measured engine profiles, offered demand, and resident work. A split's score is its worst projected service-level objective (SLO) ratio across time to first token (TTFT), time per output token (TPOT), and decode queueing.
Projections use measured window demand. A decode-to-prefill candidate takes its decode demand from the larger of the short- and long-horizon estimates.
The controller moves to an adjacent split that improves the score by at least the configured margin. A decode-to-prefill move also needs stable decode demand and a closed arrival-evidence window. The window closes after controller.reactive.evidence_span_s with the minimum number of arrivals, or after controller.reactive.evidence_max_span_s under sparse traffic. A prefill-to-decode move with prefill load at or below controller.thresholds.shrink can proceed while the window is open.
When demand over the confirmation span shifts after a settled run, the controller moves one engine. The settled run is controller.reactive.evidence_span_s, or one confirmation span shorter when the shift reverses the controller's recent moves. The controller keeps moving engines in that direction on confirmation-span demand while the shift lasts: after its first move for a reversing shift, and after controller.reactive.evidence_span_s for any other shift. Under steady demand, the score chooses between adjacent splits once the arrival-evidence window has closed.
Every move passes guards for pinned engines, role floors, cooldown, dwell time, the resident-stream cap on decode donors, and engine lifecycle holds. While a role is below its configured floor, floor repair moves one engine per monitor pass.
New requests follow the revised split, and resident requests finish on their assigned engines.
Evaluation
Read the full evaluation: Evaluating Narwhal
Getting the commands
Install on Linux with Python 3.11 or newer:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install narwhal-inference
narwhal-serve --version
narwhal --help
The wheel installs these commands:
Trying it on one GPU
Narwhal dev runs a local NVIDIA CUDA fleet on Ubuntu, either directly or under WSL2.
narwhal dev init
narwhal dev up
narwhal dev verify
narwhal dev status
narwhal dev down
The installed template starts two engines on an NVIDIA GPU with 8 GB of VRAM or less. The RTX 5090 reference template starts four engines on an RTX 5090.
Bringing up a fleet
Run these gates in order from a management workstation:
- Freeze inputs and discover the deployment in Gate A.
- Package and install the approved revision in Gate B.
- Validate and start every engine in Gate C.
- Qualify the transfer fabric in Gate D.
- Attest the live engines in Gate E.
- Profile the engines and run preflight in Gate F.
- Start the router and validate capacity through an SSH tunnel in Gate G.
Documentation
Contributing
Built on Arrow
Narwhal's scheduling algorithms derive from Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture by Wu et al. (2025).
Metadata
Release files for narwhal-inference 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| narwhal_inference-0.5.0.tar.gz | 3.7 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| narwhal_inference-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 4.0 MB
Release files / narwhal_inference-0.5.0.tar.gz
| Download URL | narwhal_inference-0.5.0.tar.gz |
|---|---|
| Size | 3.7 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1dfd215cd2d8a7b35d48ecdec513f24d9f500fba40a541a226d61af93a939665
|
|
BLAKE2b-256 checksum How to use checksums |
b5d335b1186729c45ade462195969105d6b68ba16ee9322e2cae23c93ada4ebc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|
Release files / narwhal_inference-0.5.0-py3-none-any.whl
| Download URL | narwhal_inference-0.5.0-py3-none-any.whl |
|---|---|
| Size | 359.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7bcbcfe88420083e385f8548f55314e94af38b71c7d527f6a81f0c3ff01ced78
|
|
BLAKE2b-256 checksum How to use checksums |
315b5a3d2cf4f37093fccab8cb505233114e26f6075bc8b95882fc7d49aaf25c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|