AFM — local AI infrastructure for Apple Silicon
Website · Documentation · GitHub releases
AFM turns an Apple Silicon Mac into a private, OpenAI-compatible AI server. Run Hugging Face MLX models or Apple’s on-device Foundation Model, then connect the clients and SDKs you already use.
- Native Swift executable—no Python runtime for serving
- Local inference—no cloud account or API key
- Chat, streaming, tools, structured output, reasoning, and logprobs
- Vision OCR, speech, embeddings, and a built-in WebUI
- Prefix caching, concurrent decode, speculative decoding, and metrics
- Importable Swift packages for apps that need in-process inference
AFM is for Apple Silicon Macs running current macOS/Xcode toolchains. MLX model weights download from Hugging Face the first time you use them.
Install
[!NOTE] Stable v0.9.16 is the recommended release. It adds Qwen 3.8 27B, Muse Glimmer 30B, and Gemma 4 support; improves Nemotron recurrent prefix reuse, Muse reasoning and tool calls, DwarfStar model resolution and reasoning separation; and requires a verified WebUI in every release package. Install
afm-nextonly to preview changes made after v0.9.16.The qualified nightly and v0.9.16 are essentially the same build. The nightly was promoted to this stable release after the full Qwen 3.8 qualification run; the remaining differences are release versioning and distribution packaging, not user-facing functionality. Use the stable release unless a newer nightly explicitly lists post-v0.9.16 changes you need.
| Stable (v0.9.16) | Nightly (afm-next) | |
|---|---|---|
| Homebrew | brew install scouzi1966/afm/afm |
brew install scouzi1966/afm/afm-next |
| pip | pip install macafm |
pip install --extra-index-url https://maclocal-ai.pages.dev/afm/wheels/simple/ macafm-next |
| Release notes | v0.9.16 | Latest nightly |
Install a previous version
Older stable releases are kept as pinned formulae in the Homebrew tap and as version-pinned wheels on PyPI. This is useful for reproducing an issue against a specific build or rolling back without waiting for a new release.
Homebrew (pinned stable formulae): afm@<version> — available for 0.9.0, 0.9.1, and 0.9.3–0.9.10.
brew install scouzi1966/afm/afm@0.9.10
brew uninstall afm
brew link afm@0.9.10
afm --version
Homebrew (pinned nightly formulae): afm-next@<full-version> — for example, afm-next@0.9.15-next.20260808.e70cc52. See the Homebrew tap for available pinned nightlies.
brew install scouzi1966/afm/afm-next@0.9.15-next.20260808.e70cc52
pip (version-pinned wheels): install any published release by version.
pip install macafm==0.9.10
pip install --extra-index-url https://maclocal-ai.pages.dev/afm/wheels/simple/ \
macafm-next==0.9.15.dev20260808
Start in two minutes
brew install scouzi1966/afm/afm
# Start a small MLX model and open the WebUI
afm mlx -m Qwen3-0.6B-4bit -w
AFM is now listening at http://127.0.0.1:9999/v1.
curl http://127.0.0.1:9999/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "Qwen3-0.6B-4bit",
"messages": [{"role": "user", "content": "Explain unified memory in one paragraph."}],
"stream": false
}'
Or use Apple’s on-device model:
afm -w
Choose your runtime
| Runtime | Best for | Start it |
|---|---|---|
| MLX | Open models, VLMs, agent controls, performance tuning | afm mlx -m <model> |
| Apple Foundation Models | Zero-download system model and .fmadapter LoRA adapters |
afm |
| DwarfStar | Compatible fixed-schedule Metal checkpoints | afm mlx -m <owner/repo> (auto-resolved) or afm mlx -m <checkpoint.gguf> --mlx-runtime dwarfstar |
| Gateway | One model list for Ollama, LM Studio, Jan, and other local servers | afm --gateway |
Model IDs without an organization default to mlx-community, so Qwen3-0.6B-4bit and mlx-community/Qwen3-0.6B-4bit both work.
Why AFM works well for agents
AFM is built for multi-turn, tool-using clients—not only chat demos.
| Capability | What it gives you |
|---|---|
| Native tool formats | Auto-detection for JSON, Qwen XML, Gemma, GLM, Kimi, MiniMax, LFM2, and related formats |
| Tool choice | auto, none, required, and named-function forcing |
| Streaming tool deltas | OpenAI-style tool-call chunks while ordinary content continues to stream |
| Structured output | json_object, json_schema, and token-level xgrammar enforcement when enabled |
| Reasoning extraction | <think> and harmony analysis channels mapped to reasoning_content |
| Determinism and inspection | seed, logprobs, top_logprobs, request IDs, tracing, and raw-parser mode |
| Long-running reliability | Cancellation, Retry-After, token counting, fair concurrent queues, and Prometheus metrics |
| Prefix reuse | Radix-tree KV caching for stable system prompts and multi-turn agent loops |
Pick a tool-calling mode
- Native (default): AFM detects the model’s own format and uses the narrowest parser. Use this for parity checks and benchmarks.
- Repair: add
--tool-call-parser afm_adaptive_xmlfor JSON-in-XML fallback, type coercion, nullable-schema handling, and fuzzy tool-name matching. Add--fix-tool-argswhen a model renames arguments. - Raw: add
--tool-call-parser noneto return the model’s tool markup as ordinary assistant content.
See MLX tool-calling modes for examples and benchmark guidance.
Connect an existing client
Most OpenAI-compatible clients need only a base URL and a placeholder API key:
Base URL: http://127.0.0.1:9999/v1
API key: x
Copy-ready guides:
OpenCode · OpenClaw · Cline · Continue · Aider · Cursor · Hermes
OpenClaw users can also generate a provider block directly:
afm mlx -m Qwen3-Coder-Next-4bit --openclaw-config
API surface
| Method | Endpoint | Purpose |
|---|---|---|
POST |
/v1/chat/completions |
Chat, SSE streaming, tools, reasoning, structured output, logprobs |
GET |
/v1/models |
Active model and gateway model discovery |
POST |
/v1/embeddings |
Apple NaturalLanguage embeddings for RAG and semantic search |
POST |
/v1/vision/ocr |
OCR, tables, barcodes, classification, saliency, and PDFs |
POST |
/v1/audio/transcriptions |
On-device speech-to-text |
POST |
/v1/audio/speech |
Text-to-speech using installed Apple voices |
POST |
/v1/tokenize |
vLLM-compatible tokens and counts for the loaded MLX model |
POST |
/v1/count_tokens |
Anthropic-style input token count |
POST |
/v1/batch/completions |
Multiplex up to 64 completions over SSE |
POST |
/v1/chat/completions/{id}/cancel |
Cancel an in-flight generation |
GET |
/metrics |
Prometheus queue, token, throughput, and timing metrics |
GET |
/openapi.json |
OpenAPI description |
GET |
/docs |
Interactive API reference served by AFM |
AFM also implements OpenAI-style file and batch-job endpoints under /v1/files and /v1/batches when the MLX batch service is active.
Apple-native tools
The CLI and HTTP server expose useful system frameworks without another service.
# OCR text or a table from an image/PDF
afm vision --file invoice.pdf --table
# Other Vision modes: text, table, barcode, classify, saliency, auto
afm vision --file photo.heic --mode classify --format json
# Speech recognition
afm speech transcribe --file meeting.wav --format srt
# Text to speech
afm speech synthesize "Hello from AFM" --voice nova --output hello.aac
# Dedicated OpenAI-compatible embeddings server (default port 9998)
afm embed
For vision-language models, add --vlm and pass one or more files with --media.
Performance controls
Defaults are a good starting point. Use these when the workload calls for them:
# Reuse prompt KV across requests
afm mlx -m <model> --enable-prefix-caching
# Save memory on long context
afm mlx -m <model> --kv-bits 8
# Fair-queue concurrent requests through one model
afm mlx -m <model> --concurrent 4
# Strict tool/JSON schemas with xgrammar
afm mlx -m <model> --enable-grammar-constraints
# Per-request device, memory, timing, and bandwidth estimates
afm mlx -m <model> --gpu-profile -s "Explain Metal kernels"
Supported checkpoints can also use speculative decoding:
--mtpfor Qwen3.6 checkpoints that include an MTP head--eagle3 <drafter-directory>for supported dense Gemma4 models--dspark-support <support.gguf>for compatible DwarfStar DSpark workflows
Read decode optimizations before choosing a checkpoint or interpreting benchmark results.
Sampling and response controls
The MLX backend supports temperature, top_p, top_k, min_p, repetition_penalty, presence_penalty, seed, stop, logprobs, and top_logprobs.
Useful server defaults:
# Apply one JSON schema when requests omit response_format
afm mlx -m <model> \
--guided-json '{"type":"object","properties":{"answer":{"type":"string"}},"required":["answer"]}' \
--enable-grammar-constraints
# Disable model reasoning/thinking
afm mlx -m <model> --no-thinking
# Pin chat-template keyword arguments
afm mlx -m <model> --chat-template-kwargs '{"enable_thinking":false}'
Use AFM as a Swift package
The repository publishes focused Swift Package Manager products:
AFMKitCore— provider contracts and core typesAFMOpenAICompat— OpenAI-compatible request/response typesAFMKitMLX— MLX model loading and inferenceAFMKitFoundationModels— Apple Foundation Models backendAFMKitFoundationModels27— macOS 27 provider protocol adaptersAFMKitFoundationModels27DwarfStar— opt-in DwarfStar macOS 27 adapterAFMKitDwarfStar— DwarfStar runtime integrationAFMKitServices— vision, speech, and embedding servicesAFMKit— high-level headless inference facadeAFMServer— Vapor HTTP layerafm— CLI executable
dependencies: [
.package(
url: "https://github.com/scouzi1966/maclocal-api.git",
branch: "main"
)
]
Start with the AFMKit public API guide and the consumer examples.
Build from source
git clone https://github.com/scouzi1966/maclocal-api.git
cd maclocal-api
./build.sh
The complete build initializes submodules, applies AFM-owned vendor patches, builds the WebUI, rebuilds Metal resources when the toolchain is available, and creates the release executable. Add --install to install it on your PATH.
Requirements
- Apple Silicon Mac
- macOS 26 or newer for the complete feature set
- Xcode 27 for development builds
- Disk and unified memory appropriate for the model you choose
Small 0.6B–4B quantized models are the easiest way to confirm a setup. Large 30B-class models need substantially more unified memory.
Documentation map
- Client setup guides
- MLX tool calling
- Vision OCR API
- Embeddings API
- Apple-native endpoints
- Model path resolution
- Decode optimizations
- AFMKit public API
- Parameter combinations and use cases
- Supported model architecture catalog
- Roadmap
Contributing
Issues, reproducible test cases, documentation improvements, and model-compatibility reports are welcome. Read AGENTS.md and CLAUDE.md before changing build, test, or vendored integration code.
If AFM is useful to you, star the repository. You may also like Vesta AI Explorer, a full-featured native macOS AI app.
License
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file macafm-0.9.16.tar.gz.
File metadata
- Download URL: macafm-0.9.16.tar.gz
- Upload date:
- Size: 21.3 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.27 {"installer":{"name":"uv","version":"0.9.27","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ab7e76ae766ec0f67435ecaa26497024f3c2b24a9457c4c62cec22bfb2446dd7
|
|
| MD5 |
82d7a82eef38b24579d635b8574b529b
|
|
| BLAKE2b-256 |
966a6ce1ab1f72db6828291d5c26011eb26f330737c50e46e6f8b38651abb9b2
|
File details
Details for the file macafm-0.9.16-py3-none-any.whl.
File metadata
- Download URL: macafm-0.9.16-py3-none-any.whl
- Upload date:
- Size: 21.4 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.27 {"installer":{"name":"uv","version":"0.9.27","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d5e0568ca38c39ae30c0fe0728160845a460fa07f31c29398453d3ab1d10cd4b
|
|
| MD5 |
ab20282697595a3189b219e88384348a
|
|
| BLAKE2b-256 |
bc024321fc259aec5bec47cb1a0997f02fdcfd33194bc9d2bf886aa24e76e041
|