Fleet manager for local LLM inference engines on Apple Silicon
Project description
asiai-inference-server
Fleet manager for local LLM inference engines on Apple Silicon.
Status: v0.3.0 — released on PyPI. Lifecycle, memory reclaim, presets, upgrades, fleet push and health probing are functional. See the roadmap below for what each version shipped.
asiai-inference-server is the control plane companion to
asiai (the observability/benchmark CLI). Where asiai
observes what's running on your Mac, this project manages it: install,
start, stop, unload, and orchestrate inference engines (llama.cpp,
Ollama, LM Studio, oMLX, TurboQuant, mlx-lm, vMLX, Rapid-MLX, MTPLX, …)
across one or several Apple Silicon machines.
It also fixes the long-standing macOS pain point: engine memory that never
gets freed because of the unified-memory compressor. Killing a process
doesn't release the VRAM. This tool combines per-engine unload APIs, full
LaunchDaemon restart, and sudo purge to reclaim memory deterministically —
and reports the actual delta measured, not a marketing promise.
Why
After a year of running multi-engine LLM inference on Apple Silicon (MacBook M1 Max, Mac Mini M4 Pro, MacBook M5 Max), the operational gap became obvious:
- Install/uninstall an engine should not require chasing brews, plists and firewall rules across READMEs.
- Switching profiles ("coding agent on Qwen-Coder 32B" → "70B chat on TurboQuant") should be a single command, not five.
- Memory unload should actually free the VRAM, not let the macOS compressor sit on it.
- A fleet of Macs should be a single dashboard, not three SSH sessions.
- AI agents (Claude Code, Cursor, Windsurf) should be able to manage the fleet via MCP, not just observe it.
asiai-inference-server ships these one at a time, building on the
asiai observability stack.
Roadmap
| Version | Scope | Status |
|---|---|---|
| v0.1 | Install/uninstall/start/stop/restart + unload + purge memory, pf firewall anchor, sudoers bootstrap | shipped |
| v0.2 | aisctl upgrade (brew, whitelisted), asiai versions provider, fleet Phase 2 (loopback serve + orchestrator fleet push) |
shipped |
| v0.3 | status --deep (generation probe, catches GPU-OOM zombies), install records + reinstall, single preset location, disable/enable (cold standby) |
shipped |
| v0.4 | asiai-priv privileged helper — a root-owned, content-validating helper replaces the wildcard sudoers: generate-don't-validate plists, a closed action allowlist, forced non-root run-as, and an append-only audit log; NOPASSWD is reduced to the helper alone |
shipped |
| v0.5 | SMAppService app bundle + asiai-launch (named entry with custom icon in Background Items) |
in review |
| later | Profile switching (apply/rollback), fleet inventory, web cockpit, MCP write tools | planned |
Architecture (high level)
- CLI:
aisctl <command>(standalone) orasiai engine <command>(auto-injected sub-CLI whenasiai-inference-serveris installed alongsideasiai). - Python stdlib only for the core (consistent with
asiai). The package depends onasiai>=1.15.0(auth.loopback+ the sharedfleet.command_spec). Optional extras:mcp(for future write tools). - macOS Apple Silicon only. We rely on
launchctl,vm_stat,sudo purge,pfctl,iogpu.wired_limit_mb. - HTTP agent for fleet operations (loopback
serve+fleet pushorchestrator, shipped v0.2). - TOML for human-edited files (engine manifests, profiles, fleet inventory). JSON for runtime state.
- Engine API keys stay out of plists and
ps: manifests declare[binary].api_key_file(path to an operator-created, mode-600 key file the engine reads at startup — e.g.llama-server --api-key-file) and[network].api_key_file(same file, used by the health/generation probes to authenticate). Only paths appear in manifests, generated plists, and process arguments — never key values.
License
Apache-2.0 © 2026 Jean-Marc Nahlovsky
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file asiai_inference_server-0.16.0.tar.gz.
File metadata
- Download URL: asiai_inference_server-0.16.0.tar.gz
- Upload date:
- Size: 427.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f32b0f52ccecc9e7c505c2726b3e8bc1143f09c3a96d65b112e42af2e2bff4cd
|
|
| MD5 |
b9260009e689701d36687ef0fff89a38
|
|
| BLAKE2b-256 |
9b470bd34a8cc663927b49d970e0e5ac47346bef3bacc1c39f215c07767f3b7f
|
Provenance
The following attestation bundles were made for asiai_inference_server-0.16.0.tar.gz:
Publisher:
release.yml on druide67/asiai-inference-server
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
asiai_inference_server-0.16.0.tar.gz -
Subject digest:
f32b0f52ccecc9e7c505c2726b3e8bc1143f09c3a96d65b112e42af2e2bff4cd - Sigstore transparency entry: 2187381985
- Sigstore integration time:
-
Permalink:
druide67/asiai-inference-server@dd6a0195d77065b36ab001543607b497be4454f5 -
Branch / Tag:
refs/tags/v0.16.0 - Owner: https://github.com/druide67
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@dd6a0195d77065b36ab001543607b497be4454f5 -
Trigger Event:
push
-
Statement type:
File details
Details for the file asiai_inference_server-0.16.0-py3-none-any.whl.
File metadata
- Download URL: asiai_inference_server-0.16.0-py3-none-any.whl
- Upload date:
- Size: 311.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b655f316f0dbb281f1e01c5b0e20e0ec33fd97ad2164b9e852734d9ca62deb14
|
|
| MD5 |
0d891e5ad9edf2802cf1cf7124e10726
|
|
| BLAKE2b-256 |
f65d4f4fc2eb7b52261186e861b9a39789d2cfbb0e93f51c2fda2e11b1f88a3c
|
Provenance
The following attestation bundles were made for asiai_inference_server-0.16.0-py3-none-any.whl:
Publisher:
release.yml on druide67/asiai-inference-server
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
asiai_inference_server-0.16.0-py3-none-any.whl -
Subject digest:
b655f316f0dbb281f1e01c5b0e20e0ec33fd97ad2164b9e852734d9ca62deb14 - Sigstore transparency entry: 2187382004
- Sigstore integration time:
-
Permalink:
druide67/asiai-inference-server@dd6a0195d77065b36ab001543607b497be4454f5 -
Branch / Tag:
refs/tags/v0.16.0 - Owner: https://github.com/druide67
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@dd6a0195d77065b36ab001543607b497be4454f5 -
Trigger Event:
push
-
Statement type: