Kestrel
High-performance inference engine for multimodal models.
Kestrel is the inference engine behind Photon, Moondream's on-device deployment option. Most Moondream users should install via pip install moondream; this repository provides the engine directly and supports additional model families.
Kestrel provides async, micro-batched inference with streaming support, paged KV caching, and optimized CUDA and Metal kernels. It's designed for production deployments where throughput and latency matter.
Features
- Async micro-batching — Cooperative scheduler batches heterogeneous requests without compromising per-request latency
- Streaming — Real-time token and transcription progress
- Multi-task — Vision-language generation, spatial reasoning, and speech transcription
- Paged KV cache — Efficient memory management for high concurrency
- Prefix caching — Radix tree-based caching for repeated prompts and images
- LoRA adapters — Parameter-efficient fine-tuning support with automatic cloud loading
Requirements
- Python 3.10–3.14.
- One of:
- NVIDIA GPU on Linux x86_64 / aarch64 or Windows x86_64. Optimized kernels for SM80 (A100), SM86 (A10, RTX 30-series), SM87 (Jetson Orin), SM89 (L4, L40S, RTX 4090), SM90 (H100, H200, GH200), SM100 (B200), SM110 (Jetson Thor), SM120 (RTX PRO 6000). Other CUDA GPUs may work but have not been tested.
- Apple Silicon Mac (M-series) on macOS 13 (Ventura) or later, with native Metal kernels.
MOONDREAM_API_KEY(optional) — only needed for Moondream finetuned-model inference (get a key from moondream.ai)
Installation
pip install kestrel
For Jetson Orin (JetPack 6) or Jetson Thor (JetPack 7), see the Jetson setup guide.
Model Access
Kestrel supports these model families:
| Model | Repository | Notes |
|---|---|---|
| Moondream 2 | vikhyatk/moondream2 | Public, no approval needed |
| Moondream 3 | moondream/moondream3-preview | Public, no approval needed |
| Moondream 3.1 9B A2B | moondream/moondream3.1-9B-A2B | Public, no approval needed |
| Qwen 3.5 | Qwen 3.5 collection | 0.8B, 2B, 4B, 9B, 27B, and 35B-A3B; Base variants where published |
| Qwen 3.6 | Qwen 3.6 collection | 27B and 35B-A3B; BF16 and FP8 checkpoints |
| Gemma 4 | Gemma 4 collection | E2B, E4B, and 31B base/instruction variants |
| Whisper large-v3-turbo | openai/whisper-large-v3-turbo | Transcription, translation, long-form audio, and word timestamps |
Quick Start
import asyncio
from kestrel.config import RuntimeConfig
from kestrel.engine import InferenceEngine
async def main():
# Weights are automatically downloaded from HuggingFace on first run.
# Use a registered model name or Hugging Face repository ID.
cfg = RuntimeConfig(model="google/gemma-4-E2B-it")
# Create the engine (loads model and warms up). No API key needed for
# local inference; pass api_key="..." only for finetuned models.
engine = await InferenceEngine.create(cfg)
# Load an image (JPEG, PNG, or WebP bytes)
image = open("photo.jpg", "rb").read()
# Visual question answering
result = await engine.query(
image=image,
question="What's in this image?",
settings={"temperature": 0.2, "max_tokens": 512},
)
print(result.output["answer"])
# Clean up
await engine.shutdown()
asyncio.run(main())
Whisper transcription
Whisper uses its Hugging Face repository ID as the model name. The checkpoint is resolved at Kestrel's pinned revision, and execution uses the same packaged Kestrel kernels and generated-decode runtime as the other CUDA models.
import asyncio
from pathlib import Path
from kestrel.config import RuntimeConfig
from kestrel.engine import InferenceEngine
WHISPER_MODEL = "openai/whisper-large-v3-turbo"
async def main():
engine = await InferenceEngine.create(
RuntimeConfig(
model=WHISPER_MODEL,
max_batch_size=4,
)
)
whisper = engine.model(WHISPER_MODEL)
try:
result = await whisper.transcribe(
audio=Path("meeting.m4a"),
timestamps="word",
)
print(result.output["text"])
for segment in result.output["segments"]:
for word in segment.get("words", []):
print(
f"{word['start']:7.2f} {word['end']:7.2f} {word['word']}"
)
finally:
await engine.shutdown()
asyncio.run(main())
Kestrel accepts encoded paths, bytes, bounded binary streams, raw mono PCM, and asynchronous PCM iterators. Long paths are decoded incrementally. See Whisper transcription for supported formats, progressive and live input, translation, prompting, clipping, quality controls, and exact resource limits.
Tasks
Kestrel supports several vision-language tasks through dedicated methods on the engine.
Query (Visual Q&A)
Ask questions about an image:
result = await engine.query(
image=image,
question="How many people are in this photo?",
settings={
"temperature": 0.2, # Lower = more deterministic
"top_p": 0.9,
"max_tokens": 512,
},
)
print(result.output["answer"])
Caption
Generate image descriptions:
result = await engine.caption(
image,
length="normal", # "short", "normal", or "long"
settings={"temperature": 0.2, "max_tokens": 512},
)
print(result.output["caption"])
Point
Locate objects as normalized (x, y) coordinates:
result = await engine.point(image, "person")
print(result.output["points"])
# [{"x": 0.5, "y": 0.3}, {"x": 0.8, "y": 0.4}]
Coordinates are normalized to [0, 1] where (0, 0) is top-left. Point prompts can also include normalized spatial references:
result = await engine.point(
image,
"gaze",
spatial_refs=[[0.42, 0.18]], # e.g. the subject's head or eye location
)
Detect
Detect objects as bounding boxes:
result = await engine.detect(
image,
"car",
settings={"max_objects": 10},
)
print(result.output["objects"])
# [{"x_min": 0.1, "y_min": 0.2, "x_max": 0.5, "y_max": 0.6}, ...]
Bounding box coordinates are normalized to [0, 1].
Segment
Generate a segmentation mask (Moondream 3 only):
result = await engine.segment(image, "dog")
seg = result.output["segments"][0]
print(seg["svg_path"]) # SVG path data for the mask
print(seg["bbox"]) # {"x_min": ..., "y_min": ..., "x_max": ..., "y_max": ...}
Note: Segmentation requires Moondream 3 and separate model weights. Contact moondream.ai for access.
Streaming
For longer responses, you can stream tokens as they're generated:
image = open("photo.jpg", "rb").read()
stream = await engine.query(
image=image,
question="Describe this scene in detail.",
stream=True,
settings={"max_tokens": 1024},
)
# Print tokens as they arrive
async for chunk in stream:
print(chunk.text, end="", flush=True)
# Get the final result with metrics
result = await stream.result()
print(f"\n\nGenerated {result.metrics.output_tokens} tokens")
Streaming is supported for query and caption methods.
Response Format
All methods return an EngineResult with these fields:
result.output # Dict with task-specific output ("answer", "caption", "points", etc.)
result.finish_reason # "stop" (natural end) or "length" (hit max_tokens)
result.metrics # Timing and token counts
The metrics object contains:
result.metrics.input_tokens # Number of input tokens (including image)
result.metrics.output_tokens # Number of generated tokens
result.metrics.prefill_time_ms # Time to process input
result.metrics.decode_time_ms # Time to generate output
result.metrics.ttft_ms # Time to first token
Using Finetunes
If you've created a finetuned model through the Moondream API, you can use it by passing the adapter ID:
result = await engine.query(
image=image,
question="What's in this image?",
settings={"adapter": "01J5Z3NDEKTSV4RRFFQ69G5FAV@1000"},
)
The adapter ID format is {finetune_id}@{step} where:
finetune_idis the ID of your finetune jobstepis the training step/checkpoint to use
Adapters are automatically downloaded and cached on first use.
Configuration
RuntimeConfig
RuntimeConfig(
model="moondream3-preview", # or "moondream2" / "moondream3.1-9B-A2B"
max_batch_size=4, # Max concurrent requests
decode_path="auto", # "auto", "native", or fail-closed "generated"
)
decode_path="native" disables generated decode construction. For Qwen 3.5
and Gemma 4, decode_path="generated" requires compatible bundled programs
covering every active batch size up to max_batch_size; construction or decode
fails instead of falling back to native execution. Moondream currently supports
only the default "auto" policy.
To run from local files instead of the registered HuggingFace weights or
tokenizer, keep model set to the matching registered architecture and pass
local paths:
RuntimeConfig(
model="moondream3.1-9B-A2B",
model_path="/models/moondream/model.safetensors",
tokenizer_path="/models/moondream/tokenizer.json",
)
model_path points to the local checkpoint file and skips the automatic
HuggingFace weight download. tokenizer_path is optional and only applies to
models whose runtime uses a tokenizer; tokenizer-free models do not need one.
When provided, it can point directly to a tokenizer.json file or to a
directory containing tokenizer.json. When omitted, Kestrel uses the tokenizer
declared by the registered model. Local files must match the selected model
architecture and checkpoint format.
Environment Variables
| Variable | Description |
|---|---|
MOONDREAM_API_KEY |
Optional. Only needed for finetuned-model inference. Get this from moondream.ai. |
HF_HOME |
Override HuggingFace cache directory for downloaded weights (default: ~/.cache/huggingface). |
HF_TOKEN |
Hugging Face token for private or gated model repositories. Alternatively, run huggingface-cli login. |
Triton Inference Server
Kestrel can be deployed as a Triton Inference Server backend. See the Triton setup guide.
Benchmarks
Throughput and latency for the query skill are tracked in PERFORMANCE.md, with results broken out by GPU.
Telemetry
Kestrel reports basic usage telemetry to help us decide which hardware platforms to prioritize for support and optimization. Each report includes the model in use, your GPU type and memory, aggregate request/error and token counts, your machine's hostname, and timestamps. Prompts, images, and model outputs are never sent.
License
Local inference is free and requires no API key. Finetuned-model inference requires a Moondream API key — see moondream.ai/pricing.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file kestrel-0.6.0.tar.gz.
File metadata
- Download URL: kestrel-0.6.0.tar.gz
- Upload date:
- Size: 370.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
34f09f8fb019a2031393a42cf7a1fabee7968939e5be3451b51f0f8852c4517d
|
|
| MD5 |
8cae534db8190350070fbe615904d195
|
|
| BLAKE2b-256 |
81b31d7e7618ea3d39feba80751978458429215365cc8049445c5c179b801f60
|
File details
Details for the file kestrel-0.6.0-py3-none-any.whl.
File metadata
- Download URL: kestrel-0.6.0-py3-none-any.whl
- Upload date:
- Size: 427.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9335afe412c4757d8be6989d666235ab1d9197c61d2146632fe418bfb345d6cf
|
|
| MD5 |
9aa8ae8e81e173d7601a122e294b2707
|
|
| BLAKE2b-256 |
8f3d6bb2dbe4957d41651b84ffa049b763caa1c7bccc9b73fd55a9dad9fd3793
|