Skip to main content

MacFleet

Distributed training for Apple Silicon fleets

Run PyTorch or MLX across multiple Macs with secure peer discovery, TLS/HMAC authentication, and framework-agnostic gradient synchronization.

pip install macfleet
macfleet join --bootstrap

Local hardware. Real data parallelism. No cloud bill.

MacFleet turns a room full of Apple Silicon machines into one training pool. Each Mac keeps a full model replica, processes its shard of the batch, and synchronizes gradients over a NumPy-only communication layer that never imports torch or MLX.

Why MacFleet

Apple Silicon is everywhere. Every researcher, student, and founder has a serious ML machine on their desk. What's missing is a way to team them up.

  • PyTorch on MPS has no distributed story. NCCL is CUDA-only. Gloo is broken on MPS. Single-GPU-on-MPS only.
  • MLX is native but most researchers' code is still PyTorch.
  • Cloud is expensive and the iteration loop is slow.

MacFleet fills that gap. Any two Macs on the same WiFi can pool their GPUs. Security is baked in (HMAC + TLS). Adaptive compression keeps WiFi viable for gradient sync. The framework-agnostic core lets you pick your engine (torch or mlx) per call.

Install

pip install macfleet                    # core
pip install "macfleet[torch]"           # + PyTorch
pip install "macfleet[mlx]"             # + Apple MLX
pip install "macfleet[all]"             # everything

The 5-minute path

1. Scaffold a starter script:

macfleet quickstart
# Wrote my_macfleet_demo.py

2. Run it:

python my_macfleet_demo.py
# Pool world size: 1
# Training done: {'loss': 0.31, 'epochs': 10, 'time_sec': 1.4}

3. Pair a second Mac:

On Mac #1:

macfleet join --bootstrap
# first run auto-generates a fleet token and prints a short-lived
# one-time pairing command. The permanent token is not printed.

On Mac #2:

macfleet pair --host <Mac-1-IP>:<enrollment-port> --code <one-time-code>
macfleet join

The enrollment code expires after 5 minutes and is single-use by default.

4. Set enable_pool_distributed=True and run the same script on both Macs — training now spans both: the pool forms a gradient mesh, rank 0 broadcasts initial weights, and every step's gradients are averaged across the fleet. The result dict's params_sha256 matches on both Macs when the fleet stayed in sync; degraded, unsynced_steps, and validation_fallback_steps tell you if any step fell back locally.

Features

  • Dual engine — PyTorch (MPS) and Apple MLX, same pool infrastructure
  • Zero config — mDNS discovery, no coordinator setup, no config files
  • Safe task dispatch@macfleet.task registry + msgpack args (no cloudpickle on the wire; local pickle fallback is explicit opt-in)
  • Adaptive compression — auto-selects TopK + FP16 based on link speed (locally; sparse-on-wire arrives in v2.3, see TODOS.md Issue 3)
  • Heterogeneous scheduling — faster Macs get bigger batches, adjusts for thermal throttling
  • Secure by default — auto-generated fleet tokens (scrypt-derived keys), client-first HMAC mutual auth (servers reveal nothing to unauthenticated peers), mandatory TLS with channel-bound handshakes (MITM-relay resistant), per-IP rate limiting
  • Framework-agnostic core — communication layer uses only numpy, never imports torch or mlx

Security

Security is on by default. The first macfleet join auto-generates a fleet token at ~/.macfleet/fleet-token (mode 0600). See the security reference for the full threat model.

Short version:

  • Fleet isolation — nodes with different tokens can't see each other on the network (mDNS service type is scoped by fleet hash)
  • Mutual authentication — HMAC-SHA256 challenge-response on every connection, plus signed hardware profile exchange (v2.2)
  • Encryption — TLS mandatory whenever auth is enabled
  • Rate limiting — 5 failed auth attempts per IP → 5-minute ban, exponential backoff in between (heartbeat read timeout tightened to 1s to stop slowloris)
  • No cloudpickle over the wire@macfleet.task routes registered callables by name, not by pickled closures
  • One-time pairingmacfleet join --bootstrap exposes only a short-lived enrollment code, not the permanent fleet token
  • Local audit trail — auth failures, enrollment, token rotation, legacy pickle use, and degraded training events are written to ~/.macfleet/audit.jsonl with credential fields redacted

CLI

macfleet join         Join the pool (auto-discovers peers)
macfleet pair         Pair with a one-time enrollment code
macfleet rotate-token Rotate the local fleet token
macfleet status       Show pool members and network info
macfleet info         Show local hardware profile
macfleet train        Run training (demo or custom script)
macfleet bench        Benchmark compute, network, or allreduce
macfleet doctor       System health check
macfleet quickstart   Write a starter training script

How it works

MacFleet uses data parallelism: every Mac holds a full copy of the model, trains on a weighted portion of the data, and averages gradients via Ring AllReduce after each step.

The compression layer (TopK + FP16) is applied locally before the allreduce; v2.2 transmits dense gradients on the wire (sparse allreduce is on the v2.3 roadmap as Issue 3). The bandwidth savings table below describes the target ratios once sparse-on-wire ships:

Network Compression 100 MB gradients (v2.3 target)
Thunderbolt 4 None 100 MB
Ethernet TopK 10% + FP16 ~5 MB
WiFi TopK 1% + FP16 ~500 KB

Requirements

  • macOS 14+ with Apple Silicon (M1/M2/M3/M4)
  • Python 3.11+
  • PyTorch 2.1+ or MLX 0.5+ (optional, pick your engine)

Documentation

Full docs: run mkdocs serve after pip install "macfleet[docs]", or read the Markdown source in docs/:

Development

git clone https://github.com/vikranthreddimasu/MacFleet.git
cd MacFleet
pip install -e ".[dev,all]"
make test       # 447 tests
make lint       # ruff + mypy

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

macfleet-2.2.1.tar.gz (147.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

macfleet-2.2.1-py3-none-any.whl (168.8 kB view details)

Uploaded Python 3

File details

Details for the file macfleet-2.2.1.tar.gz.

File metadata

  • Download URL: macfleet-2.2.1.tar.gz
  • Upload date:
  • Size: 147.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.12

File hashes

Hashes for macfleet-2.2.1.tar.gz
Algorithm Hash digest
SHA256 854d7cbf9739775cca6c8ee5a0558263463fb2d4bf3cf16b2ce4380d71c4da24
MD5 81af5682e58b913b890a0afa7441c4eb
BLAKE2b-256 45a4566f8904aaecef6b5dbf4200c1681bf9d8ad96395f391e110aa968ccc3a9

See more details on using hashes here.

File details

Details for the file macfleet-2.2.1-py3-none-any.whl.

File metadata

  • Download URL: macfleet-2.2.1-py3-none-any.whl
  • Upload date:
  • Size: 168.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.12

File hashes

Hashes for macfleet-2.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 51d90d15e000efb1230d1359c8864918b965dd904a852fa11a2d36ec1a5a25ce
MD5 bbc90c39eff1532c577a5a62a1e533aa
BLAKE2b-256 5ef27854fccf690aadce10395340491f9253b8f96a4b78f02e66222545a62f5f

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

2.2.1 This release

2 files

2.2.0

2 files

2.1.1

2 files

2.1.0

2 files

2.0.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page