Skip to main content

Kestrel

Kestrel Overview

High-performance inference engine for multimodal models.

Kestrel is the inference engine behind Photon, Moondream's on-device deployment option. Most Moondream users should install via pip install moondream; this repository provides the engine directly and supports additional model families.

Kestrel provides async, micro-batched inference with streaming support, paged KV caching, and optimized CUDA and Metal kernels. It's designed for production deployments where throughput and latency matter.

Features

  • Async micro-batching — Cooperative scheduler batches heterogeneous requests without compromising per-request latency
  • Streaming — Real-time token and transcription progress
  • Multi-task — Vision-language generation, spatial reasoning, and speech transcription
  • Paged KV cache — Efficient memory management for high concurrency
  • Prefix caching — Radix tree-based caching for repeated prompts and images
  • LoRA adapters — Parameter-efficient fine-tuning support with automatic cloud loading

Requirements

  • Python 3.10–3.14.
  • One of:
    • NVIDIA GPU on Linux x86_64 / aarch64 or Windows x86_64. Optimized kernels for SM80 (A100), SM86 (A10, RTX 30-series), SM87 (Jetson Orin), SM89 (L4, L40S, RTX 4090), SM90 (H100, H200, GH200), SM100 (B200), SM110 (Jetson Thor), SM120 (RTX PRO 6000). Other CUDA GPUs may work but have not been tested.
    • Apple Silicon Mac (M-series) on macOS 13 (Ventura) or later, with native Metal kernels.
  • MOONDREAM_API_KEY (optional) — only needed for Moondream finetuned-model inference (get a key from moondream.ai)

Installation

pip install kestrel

For Jetson Orin (JetPack 6) or Jetson Thor (JetPack 7), see the Jetson setup guide.

Model Access

Kestrel supports these model families:

Model Repository Notes
Moondream 2 vikhyatk/moondream2 Public, no approval needed
Moondream 3 moondream/moondream3-preview Public, no approval needed
Moondream 3.1 9B A2B moondream/moondream3.1-9B-A2B Public, no approval needed
Qwen 3.5 Qwen 3.5 collection 0.8B, 2B, 4B, 9B, 27B, and 35B-A3B; Base variants where published
Qwen 3.6 Qwen 3.6 collection 27B and 35B-A3B; BF16 and FP8 checkpoints
Gemma 4 Gemma 4 collection E2B, E4B, and 31B base/instruction variants
Whisper large-v3-turbo openai/whisper-large-v3-turbo Transcription, translation, long-form audio, and word timestamps

Quick Start

import asyncio

from kestrel.config import RuntimeConfig
from kestrel.engine import InferenceEngine


async def main():
    # Weights are automatically downloaded from HuggingFace on first run.
    # Use a registered model name or Hugging Face repository ID.
    cfg = RuntimeConfig(model="google/gemma-4-E2B-it")

    # Create the engine (loads model and warms up). No API key needed for
    # local inference; pass api_key="..." only for finetuned models.
    engine = await InferenceEngine.create(cfg)

    # Load an image (JPEG, PNG, or WebP bytes)
    image = open("photo.jpg", "rb").read()

    # Visual question answering
    result = await engine.query(
        image=image,
        question="What's in this image?",
        settings={"temperature": 0.2, "max_tokens": 512},
    )
    print(result.output["answer"])

    # Clean up
    await engine.shutdown()


asyncio.run(main())

Whisper transcription

Whisper uses its Hugging Face repository ID as the model name. The checkpoint is resolved at Kestrel's pinned revision, and execution uses the same packaged Kestrel kernels and generated-decode runtime as the other CUDA models.

import asyncio
from pathlib import Path

from kestrel.config import RuntimeConfig
from kestrel.engine import InferenceEngine

WHISPER_MODEL = "openai/whisper-large-v3-turbo"


async def main():
    engine = await InferenceEngine.create(
        RuntimeConfig(
            model=WHISPER_MODEL,
            max_batch_size=4,
        )
    )
    whisper = engine.model(WHISPER_MODEL)
    try:
        result = await whisper.transcribe(
            audio=Path("meeting.m4a"),
            timestamps="word",
        )
        print(result.output["text"])
        for segment in result.output["segments"]:
            for word in segment.get("words", []):
                print(
                    f"{word['start']:7.2f}  {word['end']:7.2f}  {word['word']}"
                )
    finally:
        await engine.shutdown()


asyncio.run(main())

Kestrel accepts encoded paths, bytes, bounded binary streams, raw mono PCM, and asynchronous PCM iterators. Long paths are decoded incrementally. See Whisper transcription for supported formats, progressive and live input, translation, prompting, clipping, quality controls, and exact resource limits.

Tasks

Kestrel supports several vision-language tasks through dedicated methods on the engine.

Query (Visual Q&A)

Ask questions about an image:

result = await engine.query(
    image=image,
    question="How many people are in this photo?",
    settings={
        "temperature": 0.2,  # Lower = more deterministic
        "top_p": 0.9,
        "max_tokens": 512,
    },
)
print(result.output["answer"])

Caption

Generate image descriptions:

result = await engine.caption(
    image,
    length="normal",  # "short", "normal", or "long"
    settings={"temperature": 0.2, "max_tokens": 512},
)
print(result.output["caption"])

Point

Locate objects as normalized (x, y) coordinates:

result = await engine.point(image, "person")
print(result.output["points"])
# [{"x": 0.5, "y": 0.3}, {"x": 0.8, "y": 0.4}]

Coordinates are normalized to [0, 1] where (0, 0) is top-left. Point prompts can also include normalized spatial references:

result = await engine.point(
    image,
    "gaze",
    spatial_refs=[[0.42, 0.18]],  # e.g. the subject's head or eye location
)

Detect

Detect objects as bounding boxes:

result = await engine.detect(
    image,
    "car",
    settings={"max_objects": 10},
)
print(result.output["objects"])
# [{"x_min": 0.1, "y_min": 0.2, "x_max": 0.5, "y_max": 0.6}, ...]

Bounding box coordinates are normalized to [0, 1].

Segment

Generate a segmentation mask (Moondream 3 only):

result = await engine.segment(image, "dog")
seg = result.output["segments"][0]
print(seg["svg_path"])  # SVG path data for the mask
print(seg["bbox"])      # {"x_min": ..., "y_min": ..., "x_max": ..., "y_max": ...}

Note: Segmentation requires Moondream 3 and separate model weights. Contact moondream.ai for access.

Streaming

For longer responses, you can stream tokens as they're generated:

image = open("photo.jpg", "rb").read()

stream = await engine.query(
    image=image,
    question="Describe this scene in detail.",
    stream=True,
    settings={"max_tokens": 1024},
)

# Print tokens as they arrive
async for chunk in stream:
    print(chunk.text, end="", flush=True)

# Get the final result with metrics
result = await stream.result()
print(f"\n\nGenerated {result.metrics.output_tokens} tokens")

Streaming is supported for query and caption methods.

Response Format

All methods return an EngineResult with these fields:

result.output          # Dict with task-specific output ("answer", "caption", "points", etc.)
result.finish_reason   # "stop" (natural end) or "length" (hit max_tokens)
result.metrics         # Timing and token counts

The metrics object contains:

result.metrics.input_tokens     # Number of input tokens (including image)
result.metrics.output_tokens    # Number of generated tokens
result.metrics.prefill_time_ms  # Time to process input
result.metrics.decode_time_ms   # Time to generate output
result.metrics.ttft_ms          # Time to first token

Using Finetunes

If you've created a finetuned model through the Moondream API, you can use it by passing the adapter ID:

result = await engine.query(
    image=image,
    question="What's in this image?",
    settings={"adapter": "01J5Z3NDEKTSV4RRFFQ69G5FAV@1000"},
)

The adapter ID format is {finetune_id}@{step} where:

  • finetune_id is the ID of your finetune job
  • step is the training step/checkpoint to use

Adapters are automatically downloaded and cached on first use.

Configuration

RuntimeConfig

RuntimeConfig(
    model="moondream3-preview",  # or "moondream2" / "moondream3.1-9B-A2B"
    max_batch_size=4,            # Max concurrent requests
    decode_path="auto",          # "auto", "native", or fail-closed "generated"
)

decode_path="native" disables generated decode construction. For Qwen 3.5 and Gemma 4, decode_path="generated" requires compatible bundled programs covering every active batch size up to max_batch_size; construction or decode fails instead of falling back to native execution. Moondream currently supports only the default "auto" policy.

To run from local files instead of the registered HuggingFace weights or tokenizer, keep model set to the matching registered architecture and pass local paths:

RuntimeConfig(
    model="moondream3.1-9B-A2B",
    model_path="/models/moondream/model.safetensors",
    tokenizer_path="/models/moondream/tokenizer.json",
)

model_path points to the local checkpoint file and skips the automatic HuggingFace weight download. tokenizer_path is optional and only applies to models whose runtime uses a tokenizer; tokenizer-free models do not need one. When provided, it can point directly to a tokenizer.json file or to a directory containing tokenizer.json. When omitted, Kestrel uses the tokenizer declared by the registered model. Local files must match the selected model architecture and checkpoint format.

Environment Variables

Variable Description
MOONDREAM_API_KEY Optional. Only needed for finetuned-model inference. Get this from moondream.ai.
HF_HOME Override HuggingFace cache directory for downloaded weights (default: ~/.cache/huggingface).
HF_TOKEN Hugging Face token for private or gated model repositories. Alternatively, run huggingface-cli login.

Triton Inference Server

Kestrel can be deployed as a Triton Inference Server backend. See the Triton setup guide.

Benchmarks

Throughput and latency for the query skill are tracked in PERFORMANCE.md, with results broken out by GPU.

Telemetry

Kestrel reports basic usage telemetry to help us decide which hardware platforms to prioritize for support and optimization. Each report includes the model in use, your GPU type and memory, aggregate request/error and token counts, your machine's hostname, and timestamps. Prompts, images, and model outputs are never sent.

License

Local inference is free and requires no API key. Finetuned-model inference requires a Moondream API key — see moondream.ai/pricing.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

kestrel-0.6.1.tar.gz (371.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

kestrel-0.6.1-py3-none-any.whl (428.3 kB view details)

Uploaded Python 3

File details

Details for the file kestrel-0.6.1.tar.gz.

File metadata

  • Download URL: kestrel-0.6.1.tar.gz
  • Upload date:
  • Size: 371.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.0

File hashes

Hashes for kestrel-0.6.1.tar.gz
Algorithm Hash digest
SHA256 f7e271fb53f6d763d744d48aa01faa25f6f23f2d3cfb621445969a284cd41607
MD5 ff2f83cb28ce519638bf78e99c8c0ced
BLAKE2b-256 2bcdb95e5c17bd2dee2f4ca23e74f1c110e46ff026bfd7e813e176abf9359423

See more details on using hashes here.

File details

Details for the file kestrel-0.6.1-py3-none-any.whl.

File metadata

  • Download URL: kestrel-0.6.1-py3-none-any.whl
  • Upload date:
  • Size: 428.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.0

File hashes

Hashes for kestrel-0.6.1-py3-none-any.whl
Algorithm Hash digest
SHA256 ffd8abcb157ed5b864f9343f7fbce6fc0a806df36ddf8b219c9f5459ec32c29d
MD5 cff125678cc830f0ec09a3d05fd13579
BLAKE2b-256 585714a44e41c245c0df5b998a57db525b0645a29b69e7630ed146b3c1ff6443

See more details on using hashes here.

Release history Release notifications | RSS feed

0.7.0

2 files

This release

0.6.1 This release

2 files

0.6.0

2 files

0.5.0

2 files

0.4.2

2 files

0.4.1

2 files

0.4.0

2 files

0.3.1

2 files

0.3.0

2 files

0.2.1

2 files

0.2.0

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.0.2

2 files

0.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page