Knurlogic
Knurlogic serves local language models on Apple Silicon. It works out the settings a model needs and whether it fits in memory before loading it, then runs it behind an OpenAI- and Anthropic-compatible server. People use it through a web page; agents use it through an MCP server. One model can also be split across two or more Macs.
Requirements
- A Mac with Apple Silicon and enough unified memory for the model.
- Python 3.10 or newer.
- mlx 0.31.2 and mlx-lm 0.31.3, pinned exactly (pip installs them).
Install
pip install knurlogic
From source:
git clone https://github.com/noahzelezny/Knurlogic
cd Knurlogic && pip install -e .
Quickstart
Get a model in MLX format: search and download it from the page's model
picker (Hugging Face), or with the hf CLI that comes with
huggingface_hub:
pip install -U huggingface_hub
hf download mlx-community/gemma-4-e4b-it-8bit \
--local-dir ~/Knurlogic/Models/gemma-4-e4b-it-8bit
knurlogic models lists the models already on this Mac (Knurlogic's own
folder, the Hugging Face cache, LM Studio, Ollama and exo folders).
Check it and serve it:
knurlogic doctor ~/Knurlogic/Models/gemma-4-e4b-it-8bit
knurlogic serve ~/Knurlogic/Models/gemma-4-e4b-it-8bit
Then ask it something:
curl http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model": "local", "messages": [{"role": "user", "content": "Hello"}]}'
The page is at http://127.0.0.1:8080/: what is loaded, where the memory
went, a chat and the settings. knurlogic ui opens the page at
http://127.0.0.1:8899/ without loading anything, with a Launch button for
every model on the disk, and a gear in the macOS menu bar (--no-menubar
skips it) that shows what is loaded and opens or quits the page.
Connect a harness
knurlogic connect --port 8080 prints these for a running server.
Claude Code (or any Anthropic Messages client):
ANTHROPIC_BASE_URL=http://127.0.0.1:8080 \
ANTHROPIC_API_KEY=x \
ANTHROPIC_DEFAULT_OPUS_MODEL=local \
ANTHROPIC_DEFAULT_SONNET_MODEL=local \
ANTHROPIC_DEFAULT_HAIKU_MODEL=local \
API_TIMEOUT_MS=3000000 \
claude
OpenAI-compatible clients (Zed, Cline, Continue, Open WebUI): base URL
http://127.0.0.1:8080/v1, any API key, model local.
Responses-API clients (the OpenAI SDK's client.responses, Codex-style
tools): the same base URL; POST /v1/responses with function tools,
streaming and reasoning summaries. previous_response_id and store: true
are refused: the server keeps no conversation, so resend the input.
Ollama clients (Open WebUI's Ollama mode, Continue, the ollama
libraries): OLLAMA_HOST=http://127.0.0.1:8080. /api/chat, /api/generate
(streaming NDJSON by default), /api/tags, /api/show and /api/version
are served, with images in chat messages on a vision model.
MCP, for an agent that should manage models rather than talk to one:
claude mcp add knurlogic -- knurlogic mcp
Its tools are models, fit, settings, drafting, ready, load,
state, unload and deps; knurlogic mcp --list describes each.
load refuses a model that does not fit.
Two or more Macs
To split one model across two or more Macs (up to 16) joined by Thunderbolt, install the same knurlogic version and the model on both, and run on each:
knurlogic ui --host cluster
The Macs find each other over Bonjour (--peer HOST names one directly, on each machine
(or link them with Thunderbolt);
knurlogic doctor --cluster says what is in the way). Then launch the
model from the page with the machines selected, or with the MCP load
tool's machines argument. The TCP ring works for any number of Macs;
RDMA (jaccl) needs every pair cabled with Thunderbolt 5, and beyond two
Macs it is experimental and untested.
Settings
Every setting is resolved from the model's config.json and the memory
available, and shown in the page's Settings panel and at /settings.json.
Two presets cover most needs: default (fastest) and lean (more
context in less memory: 512-token prompt chunks, MTP off, 8-bit KV).
serve --tune default|lean picks one, --kv-bits 8 stores the KV cache
in about half the memory, and --set KEY=VALUE overrides any single
setting. A context length past the model's native window turns on YaRN
for the Qwen families (up to 1,048,576 tokens).
Supported models
| Family | model_type |
Vision | Drafting (MTP) | 8-bit KV |
|---|---|---|---|---|
| Qwen 3.5 / 3.6 | qwen3_5, qwen3_5_moe |
yes | if the model has a head | yes |
| Qwen 3.8 | qwen4_exp |
yes | if the model has a head | yes |
| Gemma 4 | gemma4, gemma4_text |
yes | no | yes |
| GLM-5 | glm5_next |
yes | if the model has a head | yes |
| DeepSeek-V4 | deepseek_v4 |
no | no | no |
VQ-quantized models (published with their own model.py) are supported
for these families; each runs the model.py it ships. The page tags a
Hugging Face model "update" when the Hub has a newer revision (checked once
at start; knurlogic ui --offline or HF_HUB_OFFLINE=1 skips it).
Code
How the package is laid out and who may import whom: docs/architecture.md.
Status
Alpha. Expect rough edges and breaking changes before 1.0; see CHANGELOG.md.
License
Apache-2.0. Vendored third-party code is listed in NOTICE.
Metadata
Release files for knurlogic 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| knurlogic-0.1.0.tar.gz | 714.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| knurlogic-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.5 MB
Release files / knurlogic-0.1.0.tar.gz
| Download URL | knurlogic-0.1.0.tar.gz |
|---|---|
| Size | 714.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
52b45abdfa81aee224e1d01ebb493a0d3731f2bc9d430fc3f081d77032586eeb
|
|
BLAKE2b-256 checksum How to use checksums |
0116c46fd7e51fadb466bbb26a50969cc117d0454577d2e9f89df69492f37e1b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency logRelease files / knurlogic-0.1.0-py3-none-any.whl
| Download URL | knurlogic-0.1.0-py3-none-any.whl |
|---|---|
| Size | 806.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
096eede0a8072172d1cb36fac296ba4a24b01ec6610b6cfd9e59ce00b8128280
|
|
BLAKE2b-256 checksum How to use checksums |
9ab9faa81ffbe0cccd218a711b02686875376583077ee8b16fb3ee47541861c0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency log