Skip to main content

Moondream Python Client Library

Official Python interface for Moondream Cloud and Photon local inference on NVIDIA GPUs, Apple Silicon Macs, or CPUs where the selected model supports them.

Capabilities

Photon exposes each selected model's capabilities through one model-bound client:

Method Description
caption Generate descriptive captions for images
query Ask questions about image content
chat Continue multi-turn conversations with text and images
detect Find bounding boxes around objects in images
point Identify the center location of specified objects
segment Generate an SVG path segmentation mask for objects
transcribe Transcribe or translate audio, including files and live PCM streams
synthesize Generate speech as complete or streamed PCM audio

Try it out on Moondream's playground.

Photon Models

Photon includes these local model families:

Family Models
Moondream Moondream 2, Moondream 3, Moondream 3.1 9B A2B
Qwen 3.5 0.8B, 2B, 4B, 9B, 27B, and 35B-A3B; Base variants where published
Qwen 3.6 27B and 35B-A3B; BF16 and FP8 checkpoints
Gemma 4 E2B, E4B, 26B-A4B, and 31B base/instruction variants
Whisper Whisper large-v3-turbo transcription and English translation
Qwen3-ASR 0.6B and 1.7B transcription and forced alignment
Parakeet TDT 0.6B v3 transcription; parakeet-redux, its ternary version for CPUs, Apple silicon, and CUDA; parakeet-ultra, the full-precision version trained further, for GPUs
Qwen3-TTS CustomVoice 0.6B and 1.7B text-to-speech on CUDA
Kokoro 82M phoneme-to-speech on CPU or CUDA

Use md.photon_models() to inspect the exact registered identifiers in the installed release. The returned client reports model_id, tasks, and supports(task) without importing the underlying runtime. Existing md.vl(local=True, model=..., ...) calls remain supported and delegate to md.photon(...).

Installation

pip install moondream

Quick Start

Choose how you want to run Moondream:

  1. Moondream Cloud — Get an API key from the cloud console
  2. Moondream Photon — High-performance local inference engine on NVIDIA GPUs (Linux / Windows), Apple Silicon Macs (macOS 13+), and CPUs for supported models. Base models run locally without an API key; an API key is only needed for finetuned models.
import moondream as md
from PIL import Image

# Initialize with Moondream Cloud
model = md.vl(api_key="<your-api-key>")

# Or initialize Photon local inference on a device supported by the model
model = md.photon("moondream3.1-9B-A2B")

# Load an image
image = Image.open("path/to/image.jpg")

# Generate a caption
caption = model.caption(image)["caption"]
print("Caption:", caption)

# Ask a question
answer = model.query(image, "What's in this image?")["answer"]
print("Answer:", answer)

# Stream the response
for chunk in model.caption(image, stream=True)["caption"]:
    print(chunk, end="", flush=True)

# Multi-turn chat accepts OpenAI-style messages
chat = model.chat([
    {"role": "user", "content": "My name is Alice."},
    {"role": "assistant", "content": "Nice to meet you, Alice!"},
    {"role": "user", "content": "What is my name?"},
])
print(chat["message"]["content"])

# Photon speech transcription uses the same model-bound interface
from pathlib import Path

with md.photon("openai/whisper-large-v3-turbo") as speech:
    transcript = speech.transcribe(
        audio=Path("meeting.m4a"),
        timestamps="word",
    )
    print(transcript["text"])

# Photon speech synthesis returns mono 24 kHz floating-point PCM
with md.photon("Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice") as voice:
    result = voice.synthesize(text="Good morning!", voice="Ryan")
    pcm = result["audio"]
    sample_rate = result["sample_rate"]

API Reference

Constructor

model = md.vl(api_key="<your-api-key>")                        # Cloud
model = md.photon("moondream3.1-9B-A2B")                       # Photon with Moondream 3.1
model = md.vl(api_key="<your-api-key>", model="moondream3-preview/ft_id@step")  # Finetune
qwen = md.photon("Qwen/Qwen3.5-4B")
gemma = md.photon("google/gemma-4-E2B-it")
speech = md.photon("openai/whisper-large-v3-turbo")
speech = md.photon("moondream/parakeet-redux")               # CPU, Apple silicon or CUDA
speech = md.photon("moondream/parakeet-ultra")               # the full-precision one, for GPUs
voice = md.photon("Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice")
voice = md.photon("hexgrad/Kokoro-82M", device="cpu")       # accepts phonemes, not text

Photon picks CUDA when it is available, then Apple silicon, then the CPU; pass device="cpu", "mps" or "cuda" to choose.

Photon clients share matching local engines. Call model.close() when an application is finished with a client, or use with md.photon("moondream3.1-9B-A2B") as model: for deterministic GPU and worker cleanup.

Methods

caption(image, length="normal", stream=False)

Generate a caption for an image.

Parameters:

  • image — Image.Image or EncodedImage
  • length — "normal", "short", or "long" (default: "normal")
  • stream — bool (default: False)

Returns: CaptionOutput — {"caption": str | Generator}

caption = model.caption(image, length="short")["caption"]

# With streaming
for chunk in model.caption(image, stream=True)["caption"]:
    print(chunk, end="", flush=True)

query(image, question, stream=False, spatial_refs=None)

Ask a question about an image.

Parameters:

  • image — Image.Image or EncodedImage
  • question — str
  • stream — bool (default: False)
  • spatial_refs — optional point or box hints, normalized to 0-1

Returns: QueryOutput — {"answer": str | Generator}

answer = model.query(image, "What's in this image?")["answer"]

# With streaming
for chunk in model.query(image, "What's in this image?", stream=True)["answer"]:
    print(chunk, end="", flush=True)

chat(messages, stream=False, reasoning=None)

Continue an OpenAI-style multi-turn conversation. Message content can be text or a list of text and base64 image_url parts. When reasoning is omitted, the selected model or Cloud service supplies its default.

result = model.chat([
    {"role": "user", "content": "Remember that my favorite color is green."},
    {"role": "assistant", "content": "Got it."},
    {"role": "user", "content": "What is my favorite color?"},
])
print(result["message"]["content"])

for chunk in model.chat(
    [{"role": "user", "content": "Write a short poem about the moon."}],
    stream=True,
)["message"]:
    print(chunk, end="", flush=True)

detect(image, object)

Detect specific objects in an image.

Parameters:

  • image — Image.Image or EncodedImage
  • object — str

Returns: DetectOutput — {"objects": List[Region]}

objects = model.detect(image, "car")["objects"]

point(image, object, spatial_refs=None)

Get coordinates of specific objects in an image.

Parameters:

  • image — Image.Image or EncodedImage
  • object — str
  • spatial_refs — optional point or box hints, normalized to 0-1

Returns: PointOutput — {"points": List[Point]}

points = model.point(image, "person")["points"]

segment(image, object, spatial_refs=None, stream=False)

Segment an object from an image and return an SVG path.

Parameters:

  • image — Image.Image or EncodedImage
  • object — str
  • spatial_refs — List[[x, y] | [x1, y1, x2, y2]] — optional spatial hints (normalized 0-1)
  • stream — bool (default: False)

Returns:

  • Non-streaming: SegmentOutput — {"path": str, "bbox": Region}
  • Streaming: Generator yielding update dicts
result = model.segment(image, "cat")
svg_path = result["path"]
bbox = result["bbox"]  # {"x_min": ..., "y_min": ..., "x_max": ..., "y_max": ...}

# With spatial hint (point)
result = model.segment(image, "cat", spatial_refs=[[0.5, 0.5]])

# With streaming
for update in model.segment(image, "cat", stream=True):
    if "bbox" in update and not update.get("completed"):
        print(f"Bbox: {update['bbox']}")  # Available in first message
    if "chunk" in update:
        print(update["chunk"], end="")  # Coarse path chunks
    if update.get("completed"):
        print(f"Final path: {update['path']}")  # Refined path
        print(f"Final bbox: {update['bbox']}")

encode_image(image)

Pre-encode an image for reuse across multiple calls.

Parameters:

  • image — Image.Image or EncodedImage

Returns: Base64EncodedImage

encoded = model.encode_image(image)

transcribe(audio=..., stream=False, **options)

Transcribe speech in its source language or translate it to English. Photon accepts encoded file paths or bytes, bounded binary streams, raw mono PCM, and asynchronous live PCM iterators. Live chunks must be nonempty one-dimensional NumPy arrays or CPU Torch tensors containing mono PCM. Options pass directly to the selected model.

from pathlib import Path

speech = md.photon("openai/whisper-large-v3-turbo")
result = speech.transcribe(
    audio=Path("interview.mp3"),
    timestamps="word",
)
print(result["text"])

# Progressive updates are replaceable transcript snapshots, not token deltas.
updates = speech.transcribe(
    audio=Path("meeting.m4a"),
    timestamps="segment",
    stream=True,
)
for update in updates:
    print(update["text"])
final = updates.result()
speech.close()

Live PCM producers remain asynchronous end to end through atranscribe:

import asyncio


async def microphone_chunks():
    while (chunk := await microphone.read()) is not None:
        yield chunk


async def main():
    with md.photon("openai/whisper-large-v3-turbo") as speech:
        updates = await speech.atranscribe(
            audio=microphone_chunks(),
            sample_rate=48_000,
            stream=True,
        )
        async for update in updates:
            print(update["text"])
        final = await updates.aresult()


asyncio.run(main())

Set task="translate" for English translation. Other options include language, sample_rate, initial_prompt, condition_on_previous_text, clip_start_seconds, clip_end_seconds, and model sampling settings.

moondream/parakeet-redux is the ternary Parakeet: 178 MB of weights, 25 languages, and local inference on CPUs, Apple silicon, and CUDA. It takes timestamps of "none", "segment", "word" or "character" and no language or prompt options.

with md.photon("moondream/parakeet-redux") as speech:
    print(speech.transcribe(audio="meeting.mp3")["text"])

synthesize(text=..., stream=False, **options)

Generate mono 24 kHz floating-point PCM with either Qwen3-TTS CustomVoice checkpoint. The 1.7B model also accepts instructions; both accept voice, language, and sampling settings. Streamed updates contain successive audio chunks, while result() returns the complete waveform.

with md.photon("Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice") as voice:
    stream = voice.synthesize(text="Good morning!", voice="Ryan", stream=True)
    for update in stream:
        play(update["audio"], update["sample_rate"])  # your audio output callback
    complete = stream.result()

For asynchronous callers, use await voice.asynthesize(...) and iterate the returned stream with async for. hexgrad/Kokoro-82M uses the same methods but accepts phonemes= rather than text=. Supply a Unicode phoneme string in Kokoro's vocabulary; text-to-phoneme conversion is the caller's responsibility. Kokoro does not require a G2P dependency in moondream. On Apple silicon, select device="cpu" for Kokoro because its Kestrel runtime does not use MPS.

Types

Type Description
Image.Image PIL Image object
EncodedImage Base class for encoded images
Base64EncodedImage Output of encode_image(), subtype of EncodedImage
Region Bounding box with x_min, y_min, x_max, y_max
Point Coordinates with x, y indicating object center
SpatialRef [x, y] point or [x1, y1, x2, y2] bbox, normalized to [0, 1]

Release files for moondream 2.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for moondream 2.5.0
File Size Uploaded
moondream-2.5.0.tar.gz 116.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for moondream 2.5.0
File Interpreter ABI Platform
moondream-2.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 228.0 kB

Release files / moondream-2.5.0.tar.gz

Download URL moondream-2.5.0.tar.gz
Size 116.8 kB
Tags Source
SHA-256 checksum
How to use checksums
905a2927d38a75470a92bddd7941b1b1cd965ca5d4dfe705699dbc80387c1848
BLAKE2b-256 checksum
How to use checksums
506529fdd11117838e52ac120b8be4f11ba81ef27924a52d7dd51196fea537f1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.9.0

Release files / moondream-2.5.0-py3-none-any.whl

Download URL moondream-2.5.0-py3-none-any.whl
Size 111.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6c05930fd38130d9c26c7b4233defac53798443ec468dd5f59b6fe8b1343bc15
BLAKE2b-256 checksum
How to use checksums
cf12725b5caf1fc72a9ea2b963d84235a7658694430121198ec812148f877290
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.9.0

Release history Release notifications | RSS feed

This release

2.5.0 This release

2 release files

2.4.1

2 release files

2.4.0

2 release files

2.3.0

2 release files

2.2.0

2 release files

2.1.1

2 release files

2.1.0

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.3.0

2 release files

1.2.2

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page