Skip to main content

carmen: let the model cook. trust nothing it can't prove.

An AI writes GPU kernels for Apple Silicon. A judge it can't fool decides what's real.
The winners go into a real model.

tests

Why

AI can write GPU kernels now, and the papers are full of "3× faster". A lot of it doesn't survive a second look:

Writing kernels is cheap now. Trusting them is the hard part. I built carmen to learn how kernels work and to answer one question honestly: can an AI make a model on my Mac faster, with proof?

What happened

On a real model, yes, by a little. carmen's AI wrote two fused 4-bit kernels that replace 8 MLX kernels with 2 in every layer of Qwen 2.5 0.5B. Decode got 3% faster, with identical output.

Qwen 2.5 0.5B (4-bit), M4, decode vs stock MLX answers
mx.compile alone 1.00× [1.00–1.01] identical
carmen's kernels 1.03× [1.02–1.05] identical 256 tokens

30 paired turns, 95% bootstrap range. carmen e2e --modes stock,compile,mlp+compile --repeats 30 --gen 256

On single kernels, it wins where MLX runs several kernels and ties where Apple wrote one.

kernel vs MLX vs mx.compile
masked_softmax (MLX: 3 kernels) 1.82× 1.38×
add_rmsnorm (MLX: 2 kernels) 1.41× 1.43×
mlp_down, 4-bit, Qwen's shape (MLX: 2 kernels) 1.02× 1.06×
softmax · rmsnorm · layernorm (MLX: 1 tuned kernel) 0.95–0.99×
matmul · attention (MLX: heavily tuned) 0.93× · 0.54×

The judge caught all 75 bugs planted in seven kernels. A KernelBench-style check let 35 of them through. How the judge works →

What I learned

  • AI kernels don't beat hand-tuned code. They beat code nobody fused.
  • A kernel that wins alone can lose inside a model. One scored 1.15× in the judge and made Qwen 18% slower: it was timed at the wrong shapes, against a weak baseline, one launch at a time. Now carmen run <op> --for qwen0.5b judges at the model's real shapes and number format, against mx.compile, and checks the winner inside the model at the end (ideas from SoL-Pi).
  • Measuring honestly took more work than writing kernels. Running versions one after another, instead of taking turns, produced two "wins" (+40% prefill, 1.09× decode) that weren't real. A browser in the background swung speed by ±27%.
  • Small models on a Mac are launch-bound. Qwen 0.5B reads 0.28 GB of weights per word, so the M4 could do ~346 tok/s. Stock MLX reaches 51%. The rest is hundreds of tiny kernel launches per word. Bigger fused kernels are the way forward.

New to kernels? How it all works explains LLMs, GPUs and kernels from scratch.

How it works

  • Carmy (the model) writes the kernel. It's trusted to be smart and never trusted to grade. It can only return text.
  • The judge is plain Python. It checks every kernel against a float64 answer key on hundreds of inputs, including secret ones drawn fresh every round, races it against MLX, and says where a kernel is wrong, not just that it is.
  • The loop: three drafts in parallel, the best becomes the champion, the next round attacks one bottleneck. Lessons are kept only when the judge proved them.
flowchart LR
    C["Carmy × 3<br/>parallel drafts"] -->|kernel text| J{{"judge<br/>compile · float64 · invariants<br/>fresh fuzz · hidden draw · timing"}}
    J -->|"located error<br/>+ speed profile"| C
    J -->|verified & faster| K(["champion"])
    K -->|"one structural change"| C
    J -->|"proven lessons"| P[("playbook")]
    P --> C
    J -.->|"hidden results<br/>(never shown to Carmy)"| R["report"]

a cook in progress: three drafts, the judge's verdicts, the champion

Try it

On an Apple Silicon Mac, with uv:

uv tool install "carmen-kernels[metal,models]"   # or: pip install "carmen-kernels[metal,models]"
export ANTHROPIC_API_KEY=...        # or put it in a .env file where you run carmen
carmen                              # the app
carmen speedup                      # Qwen 2.5 0.5B, stock vs carmen's kernels, on your Mac (no API key)
command what it does
carmen the terminal app: what's proven, speed up a model, pick a kernel, watch it cook
carmen speedup [model] stock MLX vs mx.compile vs carmen's bundled kernels on a real model, paired runs; no API key (--quick for ~2 min)
carmen run <op> [--for qwen0.5b] let Carmy cook one kernel; --for judges it at a real model's shapes
carmen broken <op> plant bugs and check the judge catches every one, next to a KernelBench-style check
carmen judge <op> <file> judge a kernel you wrote
carmen e2e [model] run a real model stock vs with carmen's kernels: tok/s, noise range, same answers? (pip install 'carmen-kernels[models]')
carmen profile [model] where a model spends its time, step by step
carmen bench the feedback loop vs plain best-of-N, same budget

From source: git clone https://github.com/utsav1033/carmen, then pip install -e '.[metal,models,dev]', and put your key in .env (see .env.example).

Built on the shoulders of

Measuring the Checker · Gaming Without an Attacker · Metal-Sci · Hacker-Fixer Loops · Test-Input Generation for Tensor Programs · Reward Hacking in Self-Improving Code Agents · Meta-Harness · Auditing Harness Tampering · SAGE · Prime Agent · KernelBench-Verified · Contract-Grade Verifier · Building a C compiler with parallel Claudes

Named after Carmen "Carmy" Berzatto. Yes, chef. Banner set in Space Grotesk and JetBrains Mono (SIL Open Font License).

Metadata

Release files for carmen-kernels 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for carmen-kernels 0.1.0
File Size Uploaded
carmen_kernels-0.1.0.tar.gz 114.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for carmen-kernels 0.1.0
File Interpreter ABI Platform
carmen_kernels-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 239.2 kB

Release files / carmen_kernels-0.1.0.tar.gz

Download URL carmen_kernels-0.1.0.tar.gz
Size 114.5 kB
Tags Source
SHA-256 checksum
How to use checksums
8e3c5f65fd5d860dfdbf99e9ab97ac255ccccf7ca09ee11830dbe06045789b35
BLAKE2b-256 checksum
How to use checksums
131f272796dbf99086af418039fe33c9ad45e875b20f9078477eb0aff50012e5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / carmen_kernels-0.1.0-py3-none-any.whl

Download URL carmen_kernels-0.1.0-py3-none-any.whl
Size 124.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9d0e09b325d9f65e7b746e2f0bf5f610d8f51922d3e611fb2313b8d1e785553f
BLAKE2b-256 checksum
How to use checksums
ac5795ef1dd57ca2d3d91471899294b0f5a3247f81831d8fb3f9f4f4101ba4dd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page