Skip to main content

Knurlogic

Knurlogic serves local language models on Apple Silicon. It works out the settings a model needs and whether it fits in memory before loading it, then runs it behind an OpenAI- and Anthropic-compatible server. People use it through a web page; agents use it through an MCP server. One model can also be split across two or more Macs.

Requirements

  • A Mac with Apple Silicon and enough unified memory for the model.
  • Python 3.10 or newer.
  • mlx 0.31.2 and mlx-lm 0.31.3, pinned exactly (pip installs them).

Install

pip install knurlogic

From source:

git clone https://github.com/noahzelezny/Knurlogic
cd Knurlogic && pip install -e .

Quickstart

Get a model in MLX format: search and download it from the page's model picker (Hugging Face), or with the hf CLI that comes with huggingface_hub:

pip install -U huggingface_hub
hf download mlx-community/gemma-4-e4b-it-8bit \
  --local-dir ~/Knurlogic/Models/gemma-4-e4b-it-8bit

knurlogic models lists the models already on this Mac (Knurlogic's own folder, the Hugging Face cache, LM Studio, Ollama and exo folders).

Check it and serve it:

knurlogic doctor ~/Knurlogic/Models/gemma-4-e4b-it-8bit
knurlogic serve  ~/Knurlogic/Models/gemma-4-e4b-it-8bit

Then ask it something:

curl http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model": "local", "messages": [{"role": "user", "content": "Hello"}]}'

The page is at http://127.0.0.1:8080/: what is loaded, where the memory went, a chat and the settings. knurlogic ui opens the page at http://127.0.0.1:8899/ without loading anything, with a Launch button for every model on the disk, and a gear in the macOS menu bar (--no-menubar skips it) that shows what is loaded and opens or quits the page.

Connect a harness

knurlogic connect --port 8080 prints these for a running server.

Claude Code (or any Anthropic Messages client):

ANTHROPIC_BASE_URL=http://127.0.0.1:8080 \
ANTHROPIC_API_KEY=x \
ANTHROPIC_DEFAULT_OPUS_MODEL=local \
ANTHROPIC_DEFAULT_SONNET_MODEL=local \
ANTHROPIC_DEFAULT_HAIKU_MODEL=local \
API_TIMEOUT_MS=3000000 \
claude

OpenAI-compatible clients (Zed, Cline, Continue, Open WebUI): base URL http://127.0.0.1:8080/v1, any API key, model local.

Responses-API clients (the OpenAI SDK's client.responses, Codex-style tools): the same base URL; POST /v1/responses with function tools, streaming and reasoning summaries. previous_response_id and store: true are refused: the server keeps no conversation, so resend the input.

Ollama clients (Open WebUI's Ollama mode, Continue, the ollama libraries): OLLAMA_HOST=http://127.0.0.1:8080. /api/chat, /api/generate (streaming NDJSON by default), /api/tags, /api/show and /api/version are served, with images in chat messages on a vision model.

MCP, for an agent that should manage models rather than talk to one:

claude mcp add knurlogic -- knurlogic mcp

Its tools are models, fit, settings, drafting, ready, load, state, unload and deps; knurlogic mcp --list describes each. load refuses a model that does not fit.

Two or more Macs

To split one model across two or more Macs (up to 16) joined by Thunderbolt, install the same knurlogic version and the model on both, and run on each:

knurlogic ui --host cluster

The Macs find each other over Bonjour (--peer HOST names one directly, on each machine (or link them with Thunderbolt); knurlogic doctor --cluster says what is in the way). Then launch the model from the page with the machines selected, or with the MCP load tool's machines argument. The TCP ring works for any number of Macs; RDMA (jaccl) needs every pair cabled with Thunderbolt 5, and beyond two Macs it is experimental and untested.

Settings

Every setting is resolved from the model's config.json and the memory available, and shown in the page's Settings panel and at /settings.json. Two presets cover most needs: default (fastest) and lean (more context in less memory: 512-token prompt chunks, MTP off, 8-bit KV). serve --tune default|lean picks one, --kv-bits 8 stores the KV cache in about half the memory, and --set KEY=VALUE overrides any single setting. A context length past the model's native window turns on YaRN for the Qwen families (up to 1,048,576 tokens).

Supported models

Family model_type Vision Drafting (MTP) 8-bit KV
Qwen 3.5 / 3.6 qwen3_5, qwen3_5_moe yes if the model has a head yes
Qwen 3.8 qwen4_exp yes if the model has a head yes
Gemma 4 gemma4, gemma4_text yes no yes
GLM-5 glm5_next yes if the model has a head yes
DeepSeek-V4 deepseek_v4 no no no

VQ-quantized models (published with their own model.py) are supported for these families; each runs the model.py it ships. The page tags a Hugging Face model "update" when the Hub has a newer revision (checked once at start; knurlogic ui --offline or HF_HUB_OFFLINE=1 skips it).

Code

How the package is laid out and who may import whom: docs/architecture.md.

Status

Alpha. Expect rough edges and breaking changes before 1.0; see CHANGELOG.md.

License

Apache-2.0. Vendored third-party code is listed in NOTICE.

Metadata

Release files for knurlogic 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for knurlogic 0.1.0
File Size Uploaded
knurlogic-0.1.0.tar.gz 714.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for knurlogic 0.1.0
File Interpreter ABI Platform
knurlogic-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.5 MB

Release files / knurlogic-0.1.0.tar.gz

Download URL knurlogic-0.1.0.tar.gz
Size 714.0 kB
Tags Source
SHA-256 checksum
How to use checksums
52b45abdfa81aee224e1d01ebb493a0d3731f2bc9d430fc3f081d77032586eeb
BLAKE2b-256 checksum
How to use checksums
0116c46fd7e51fadb466bbb26a50969cc117d0454577d2e9f89df69492f37e1b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release files / knurlogic-0.1.0-py3-none-any.whl

Download URL knurlogic-0.1.0-py3-none-any.whl
Size 806.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
096eede0a8072172d1cb36fac296ba4a24b01ec6610b6cfd9e59ce00b8128280
BLAKE2b-256 checksum
How to use checksums
9ab9faa81ffbe0cccd218a711b02686875376583077ee8b16fb3ee47541861c0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page