Skip to main content

The Narwhal logo, a black narwhal with a teal spiral tusk above the wordmark

Apache-2.0 license Python 3.11 through 3.13 Lint and format by ruff Types checked with mypy Latest PyPI version

Documentation Deployment API reference Issues Contributing

Recognition

What is Narwhal?

Narwhal is the first open-source LLM inference framework that automatically hot-swaps prefill and decode roles on both NVIDIA and AMD GPUs. It moves engines between roles as demand changes, on a fixed GPU fleet with model weights already loaded. It scales from a single GPU to distributed multi-node deployments.

Capability Behavior Guide
Role hot-swap Reassigns prefill and decode roles across a fixed GPU fleet, with NIXL key-value (KV) transfer between them. Core concepts documentation
Serving Serves streaming and buffered completion and chat requests with latency-aware admission. HTTP API reference documentation
Fault tolerance Fails over to a warm-standby router and readmits engines against their live process generation. Operate Narwhal documentation
Measurement Profiles engines and runs ordered benchmark points with retained evidence. Measure a fleet documentation
Observability Exports router and engine metrics to Prometheus and a provisioned Grafana dashboard. Set up observability documentation
Operator tooling Validates fleet files offline and collects private diagnostic bundles. CLI reference documentation
Development mode Runs two to eight engines on one NVIDIA CUDA GPU under Ubuntu or WSL2. Narwhal dev documentation

How engines change roles

Role controller pass Behavior
Inputs Measured engine profiles, offered demand, and resident work
Candidates The current role split and each adjacent split, one engine move away
Score The worst projected service-level objective (SLO) ratio across time to first token (TTFT), time per output token (TPOT), and decode queueing
Demand Measured window demand, using the larger of the short- and long-horizon decode estimates for decode-to-prefill candidates
Evidence window Closes after controller.reactive.evidence_span_s and the minimum arrivals, or after controller.reactive.evidence_max_span_s under sparse traffic
Move To the adjacent split that improves the score by at least the configured margin
Decode to prefill Requires a closed evidence window and stable decode demand
Prefill to decode Proceeds with the evidence window open
Guards Pinned engines, role floors, cooldown, dwell time, the resident-stream ceiling on decode donors, and engine lifecycle holds
Floor repair One engine per monitor pass while a phase sits below its configured floor
New requests Follow the revised split
Resident requests Finish on their assigned engines

Role control and capacity floors documentation

Role controller changing engine roles with weights resident.

Evaluation

Read the full evaluation: Evaluating Narwhal

AIPerf v0.12.0 Kimi-K3 model Chat/document workload results Mixed-payload workload results Prefix caching enabled

Narwhal v0.1.0 Dynamo Planner Ray Serve LLM

Evaluation results for Narwhal, Dynamo Planner, and Ray Serve LLM.

Get the commands

Install on Linux with Python 3.11 or newer:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install narwhal-inference
narwhal-serve --version
narwhal --help

The wheel installs these commands:

narwhal narwhal-engine narwhal-serve narwhal-attest narwhal-profile narwhal-check

For each deployment:

  1. Record the narwhal-serve --version output with the fleet configuration, engine image, and profiles.
  2. Pin that version on every router host.

Install from PyPI documentation

Try it on one GPU

Narwhal dev runs a local NVIDIA CUDA fleet on Ubuntu or Ubuntu under WSL2.

narwhal dev init
narwhal dev up
narwhal dev verify
narwhal dev status
narwhal dev down
GPU Engines Template
NVIDIA GPU with 8 GB of VRAM or less 2 Installed template documentation
RTX 5090 4 RTX 5090 reference documentation

Bring up a fleet

Run these gates from a management workstation:

Step Gate
Freeze inputs and discover the deployment Gate A documentation
Package and install the approved revision Gate B documentation
Validate and start every engine Gate C documentation
Qualify the transfer fabric Gate D documentation
Attest the live engines Gate E documentation
Profile the engines and run preflight Gate F documentation
Start the router and validate capacity through an SSH tunnel Gate G documentation

Deploy a fleet documentation

Documentation

Architecture documentation Configuration documentation CLI documentation HTTP API documentation Measurement documentation Observability documentation Operations documentation Troubleshooting documentation

Contributing

Contributing Code of conduct Security policy

Built on Arrow

Narwhal's scheduling algorithms derive from Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture by Wu et al. (2025).

Arrow paper on arXiv Citation metadata for Arrow and Narwhal Apache-2.0 license

Metadata

Release files for narwhal-inference 0.3.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for narwhal-inference 0.3.2
File Size Uploaded
narwhal_inference-0.3.2.tar.gz 3.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for narwhal-inference 0.3.2
File Interpreter ABI Platform
narwhal_inference-0.3.2-py3-none-any.whl Python 3 none any Details

Total release size: 3.3 MB

Release files / narwhal_inference-0.3.2.tar.gz

Download URL narwhal_inference-0.3.2.tar.gz
Size 3.0 MB
Tags Source
SHA-256 checksum
How to use checksums
fbb30e1c5a13683cee05b86b76d7a23ab56df7e415f8746bc57746593d8e9df7
BLAKE2b-256 checksum
How to use checksums
faebbf209e2ebeb0423650d62ccc73ff911a5c351796280b8fcfa60ff9b2ffb1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release files / narwhal_inference-0.3.2-py3-none-any.whl

Download URL narwhal_inference-0.3.2-py3-none-any.whl
Size 294.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c5d8718be1f78c6765ab36ad79859299aad87ac7b05e5936ab4a482974c689b6
BLAKE2b-256 checksum
How to use checksums
6967926e0ae7f54c941917fac495d3fb78a0a0a2439d463e1fb8c92d26add5fd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.0

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

This release

0.3.2 This release

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page