Skip to main content

The Narwhal logo, a black narwhal with a teal spiral tusk above the wordmark

Apache-2.0 license Python 3.11 through 3.13 Lint and format by ruff Types checked with mypy Latest PyPI version

Documentation Deployment API reference Issues Contributing

Recognition

What is Narwhal?

Narwhal is an adaptive, disaggregated inference framework that automatically hot-swaps prefill and decode roles, without having to reload model weights. It scales from a single GPU to distributed multi-node deployments.

Capability Behavior Guide
Role hot-swap Reassigns prefill and decode roles across a fixed GPU fleet, with NIXL key-value (KV) transfer between them Core concepts documentation
Serving Serves streaming and buffered completion and chat requests with latency-aware admission HTTP API reference documentation
Fault tolerance Fails over to a warm-standby router and readmits engines against their live process generation Operating Narwhal documentation
Measurement Profiles engines and runs ordered benchmark points with retained evidence Measuring a fleet documentation
Observability Exports router and engine metrics to Prometheus and a provisioned Grafana dashboard Setting up observability documentation
Operator tooling Validates fleet files offline and collects private diagnostic bundles CLI reference documentation
Development mode Runs two to eight engines on one NVIDIA CUDA GPU under Ubuntu or WSL2 Narwhal dev documentation

How engines change roles

The role controller scores the current role split and each adjacent split, one engine move away. It works from measured engine profiles, offered demand, and resident work. The score is the worst projected service-level objective (SLO) ratio across time to first token (TTFT), time per output token (TPOT), and decode queueing.

Demand is the measured window demand, using the larger of the short- and long-horizon decode estimates for decode-to-prefill candidates. The evidence window closes after controller.reactive.evidence_span_s and the minimum arrivals, or after controller.reactive.evidence_max_span_s under sparse traffic.

The controller moves to the adjacent split that improves the score by at least the configured margin. A decode-to-prefill move requires a closed evidence window and stable decode demand, and a prefill-to-decode move proceeds with the window open.

Each move passes the guards for pinned engines, role floors, cooldown, dwell time, the resident-stream ceiling on decode donors, and engine lifecycle holds. Floor repair moves one engine per monitor pass while a phase sits below its configured floor.

New requests follow the revised split, and resident requests finish on their assigned engines.

Role control and capacity floors documentation

Role controller changing engine roles with weights resident.

Evaluation

Read the full evaluation: Evaluating Narwhal

AIPerf v0.12.0 Kimi-K3 model Chat/document workload results Mixed-payload workload results Prefix caching enabled

Narwhal v0.1.0 Dynamo Planner Ray Serve LLM

Evaluation results for Narwhal, Dynamo Planner, and Ray Serve LLM.

Getting the commands

Install on Linux with Python 3.11 or newer:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install narwhal-inference
narwhal-serve --version
narwhal --help

The wheel installs these commands:

narwhal narwhal-engine narwhal-serve narwhal-attest narwhal-profile narwhal-check

Installing from PyPI documentation

Trying it on one GPU

Narwhal dev runs a local NVIDIA CUDA fleet on Ubuntu or Ubuntu under WSL2.

narwhal dev init
narwhal dev up
narwhal dev verify
narwhal dev status
narwhal dev down

The installed template starts two engines on an NVIDIA GPU with 8 GB of VRAM or less. The RTX 5090 reference template starts four engines on an RTX 5090.

Installed template documentation RTX 5090 reference documentation

Bringing up a fleet

Run these gates in order from a management workstation:

  1. Freeze inputs and discover the deployment in Gate A.
  2. Package and install the approved revision in Gate B.
  3. Validate and start every engine in Gate C.
  4. Qualify the transfer fabric in Gate D.
  5. Attest the live engines in Gate E.
  6. Profile the engines and run preflight in Gate F.
  7. Start the router and validate capacity through an SSH tunnel in Gate G.

Deploying a fleet documentation

Documentation

Architecture documentation Configuration documentation CLI documentation HTTP API documentation Measurement documentation Observability documentation Operations documentation Troubleshooting documentation

Contributing

Contributing Code of conduct Security policy

Built on Arrow

Narwhal's scheduling algorithms derive from Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture by Wu et al. (2025).

Arrow paper on arXiv Citation metadata for Arrow and Narwhal Apache-2.0 license

Metadata

Release files for narwhal-inference 0.4.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for narwhal-inference 0.4.1
File Size Uploaded
narwhal_inference-0.4.1.tar.gz 3.4 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for narwhal-inference 0.4.1
File Interpreter ABI Platform
narwhal_inference-0.4.1-py3-none-any.whl Python 3 none any Details

Total release size: 3.7 MB

Release files / narwhal_inference-0.4.1.tar.gz

Download URL narwhal_inference-0.4.1.tar.gz
Size 3.4 MB
Tags Source
SHA-256 checksum
How to use checksums
cc18befab62000a2db1b06f620c72d4b6698336cd24f7a97126e3edc2e0ee96a
BLAKE2b-256 checksum
How to use checksums
dbe8f96c56cfa87d7251525df8d697404940dfb4c3e123ebbdf95c1dd37da7ba
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / narwhal_inference-0.4.1-py3-none-any.whl

Download URL narwhal_inference-0.4.1-py3-none-any.whl
Size 311.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5a40238aa6c65d93637d8cd12c9ef5f66a41389e54b22a998225808a926bbe11
BLAKE2b-256 checksum
How to use checksums
9fb47c42088ae3990a94cc7d902682828d0c43c5e8a0461ca0e1a6aefde771f2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.0

2 release files

0.5.0

2 release files

This release

0.4.1 This release

2 release files

0.4.0

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page