Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

mlx-h3

Pure MLX MiniMax-H3 text-to-video-and-audio inference for Apple Silicon.

Pre-release PyPI Python Apple Silicon MLX

appautomaton.renocrypt.com/mlx-h3

mlx-h3 is an independent, pure-MLX inference runtime for MiniMax-H3. It generates video and stereo audio jointly, keeps model residency phase-scoped, and targets large-memory Apple silicon systems without using PyTorch at runtime.

[!IMPORTANT] This project is pre-alpha. This package version is released as the v0.0.1a3 GitHub pre-release. Model files are not included in the repository or PyPI package.

Why mlx-h3

  • Joint audio and video — one DiT denoises both modalities in a shared sequence.
  • Pure MLX runtime — no PyTorch execution and no CUDA dependency.
  • Bounded model residency — text encoder, DiT, Video VAE, and Audio VAE load and release in separate phases.
  • Current sampling baseline — 20 simple schedule steps with the second-order res_multistep solver.
  • Dependency-light tokenizer — byte-level BPE implemented locally from tokenizer.json.
  • Fail-fast memory guard — configurable active-memory budget and swap detection.

Current scope

Capability Status
Text-to-video-and-audio (T2VA) Working
Synchronized H.264/AAC MP4 output Working
8-bit DiT and text encoder loading Working
First/last-frame conditioning (FL2VA) Working
Ordered image/video/audio references (Ref2VA) Working
Reference-video soundtrack conditioning Working
Context-IR and 2K regeneration Not available locally

Requirements

  • Apple silicon Mac
  • macOS with a recent MLX-compatible toolchain
  • Python 3.13 or newer
  • ffmpeg available on PATH
  • Local MiniMax-H3 tokenizer and checkpoints
  • Enough unified memory for the selected canvas and frame count

The default runtime memory budget is 70 GiB. It is a guardrail, not a promise that every system workload will remain swap-free.

Install

From a local checkout:

git clone https://github.com/appautomaton/mlx-h3.git
cd mlx-h3
uv sync

Install the current PyPI pre-release:

uv tool install --prerelease allow mlx-h3==0.0.1a3

Local model layout

Model files stay outside version control. The default paths are:

weights/
├── tokenizer/tokenizer.json
├── mlx-8bit/te_qwen3vl_a8g32.safetensors
├── mlx-8bit/dit_fl2va_a8g32.safetensors
├── mlx-8bit/dit_ref2va_a8g32.safetensors
└── bf16/vae/
    ├── minimax_h3_video_vae_fp16.safetensors
    └── minimax_h3_audio_vae_fp32.safetensors

Dense DiT and text-encoder weights may be retained locally for requantization, but inference never loads them. The dense Video VAE and Audio VAE checkpoints are runtime inputs.

Generate

Keep private input text in your shell environment rather than a tracked file:

uv run mlx-h3 "$MLX_H3_INPUT_TEXT" \
  --width 512 \
  --height 288 \
  --frames 124 \
  --steps 20 \
  --seed 42 \
  --output outputs/result.mp4

Long structured prompts can instead stay in an untracked UTF-8 file:

uv run mlx-h3 --prompt-file "$MLX_H3_PROMPT_FILE" \
  --width 768 \
  --height 448 \
  --frames 124 \
  --steps 10 \
  --output outputs/preview.mp4

Conditioning inputs are explicit. --first-frame and --last-frame select the FL2VA path. Repeat --ref-image, --ref-video, and --ref-audio in the order Ref2VA should read them; use --ref-video-silent to ignore embedded audio or --ref-video-with-audio VIDEO AUDIO to override a video's soundtrack.

Canvas dimensions must be multiples of 32 and may not exceed 768 * 1344 pixels. Frame requests are aligned to the Video VAE's 17n + 5 rule and capped at the released 15-second limit. Use --steps 10 for a faster preview; --steps 20 is the quality baseline.

Run uv run mlx-h3 --help for checkpoint path overrides and all generation options.

Memory model

The pipeline intentionally keeps only one large model phase resident at a time:

reference encoders -> release -> text/vision encode -> release
                   -> joint denoise -> release -> video decode -> release
                   -> audio decode -> release -> mux

Safety checks remain enabled in release runs. Scalar telemetry is emitted only when a callback is attached, so normal inference does not retain diagnostic tensors or model objects.

Development

uv run ruff check .
uv run pytest -q
python dev/check_public_tree.py
uv build --no-sources

The public-tree check rejects model files, media, private inputs, generated artifacts, large files, hidden local state, symlinks, and structured private prompt payloads. A local pre-commit hook runs the same check against staged files.

Reference notes live in docs/: architecture (what H3 is), weights (what is on disk), porting (validation and pitfalls).

Project identity

This project is not affiliated with or endorsed by MiniMax.

Release files for mlx-h3 0.0.1a3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mlx-h3 0.0.1a3
File Size Uploaded
mlx_h3-0.0.1a3.tar.gz 52.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mlx-h3 0.0.1a3
File Interpreter ABI Platform
mlx_h3-0.0.1a3-py3-none-any.whl Python 3 none any Details

Total release size: 113.3 kB

Release files / mlx_h3-0.0.1a3.tar.gz

Download URL mlx_h3-0.0.1a3.tar.gz
Size 52.5 kB
Tags Source
SHA-256 checksum
How to use checksums
0de0ca247a4899cdd1e9c632d4fbfbb384884a55abbb665dc0749a57ee0da759
BLAKE2b-256 checksum
How to use checksums
45f9c234b823979f9fe0fee647d8fe93be0fe91d4e17615b8a80256de8ff7f7e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 6, 2026.

Transparency log

Release files / mlx_h3-0.0.1a3-py3-none-any.whl

Download URL mlx_h3-0.0.1a3-py3-none-any.whl
Size 60.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
208366cfa66dbd2f88a13dddb84a9777bb4505d97ec6135768ffa55563f5bf5f
BLAKE2b-256 checksum
How to use checksums
dcebb16ae5e8142d67a462b47d38e383eac9eb713add8919af76206ec4113c96
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 6, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page