coldpath
A linter and CI gate for whether your Arm AI binary actually uses the chip's matrix hardware.
The most popular local LLM runner ships its Arm build with the chip's matrix hardware switched off, and any portable Arm build that misses one compile flag does the same. On Azure Cobalt 100 (Neoverse N2, a cloud Arm CPU) the fix is 5.75x on prompt processing and ~2.2x on generation, which at a typical Arm-cloud rate is roughly $0.45 versus $0.08 per million prompt tokens. I built the tool that finds it in any binary, fixed the most popular offender in one line upstream, and gated it so a cold build can't reach your Arm cloud fleet.
$0.45 → $0.08 per 1M prompt tokens, from one build flag, measured live on Azure Cobalt 100 (Neoverse N2).
coldpath disassembles any AArch64 binary and proves whether it contains, and can dispatch, the chip's
matrix and dot-product instructions (SME/SME2, i8mm, bf16, dotprod). Absence is dispositive: zero smmla
means the binary cannot run an i8mm matmul on any core. It needs no Arm hardware and no profiler; it runs
on the x86 laptop you already have, and as a CI gate that stops a cold build reaching production.
The finding
$ coldpath ollama-windows-arm64/lib/ollama/ggml-cpu.dll
lib/ollama/ggml-cpu.dll
154,880 instructions 100.0% decoded COLD -- 0 matrix instructions
sha256:536ada0d3dd642f4
-- SME (ZA tile) 0
-- i8mm (smmla) 0
-- dotprod (sdot) 0
That is Ollama's official Windows-on-Arm build (v0.31.2): zero matrix, zero dot-product
instructions, every matmul in scalar/NEON. It is not a platform limit. llama.cpp's own Windows-on-Arm
build, same OS, same ggml source, ships the kernels (i8mm 244, dotprod 1,052). Ollama doesn't fork
ggml: it builds pinned upstream llama.cpp with one flag missing. This is a distribution hazard, not a
build default: a native cmake build detects the host and comes out warm, but any portable build
(cross-compiled, reproducible, or explicitly GGML_NATIVE=OFF for device compatibility, which is how
prebuilt binaries are made) must pick a target -march or fall back to baseline armv8-a. That is the
choice a distributor makes, and the one Ollama got wrong for Windows-on-Arm. (Distributors who use
GGML_CPU_ALL_VARIANTS=ON runtime dispatch avoid it, which is why Ollama's Linux arm64 build is warm.)
What it costs, measured on Azure Cobalt 100 (Neoverse N2), a cloud Arm CPU, on the free GitHub
runner. Every row is built GGML_NATIVE=OFF with a pinned -march to isolate the ISA effect; only
-march changes:
build (GGML_NATIVE=OFF, pinned -march) |
coldpath sees | pp512 tok/s | $ / 1M tokens | vs COLD |
|---|---|---|---|---|
COLD armv8-a, the baseline a portable build falls back to (Ollama's Windows build) |
i8mm 0, dotprod 0 | ~95 | ~$0.45 | 1.0x |
TEPID armv8.2-a+dotprod, the one-line fix I filed upstream |
dotprod 1,044 | ~545 | ~$0.08 | ~5.75x |
WARM armv8.6-a+i8mm |
i8mm 268, dotprod 1,044 | ~660 | ~$0.065 | ~6.9x |
The tok/s and the 5.75x ratio are measured on the 4-vCPU Neoverse N2 runner; the $/1M-tokens applies a
sample Arm-cloud on-demand rate (~$0.04/vCPU-hr, Graviton4-class), so only the dollar column assumes a
price. The workflow's compare job recomputes all of it live each run (figures vary a few percent). The
fix I filed (PR #17654) is the TEPID row, dot-product, which is safe on every shipped Arm device; the
6.9x WARM row needs i8mm, which server and newer mobile cores have but Windows-on-Arm's Cortex-A76-class
chips do not, so the PR intentionally ships dot-product only.
The fix, the root cause, and the reproducible measurement are in
examples/ollama-fix/. The benchmark is
.github/workflows/benchmark.yml, re-runnable by anyone (including a
judge) from the Actions tab on free Arm64 hardware. coldpath is the instrument that found this, and
the CI gate that stops it shipping again.
Why a static scan is sufficient (and why it works on Arm but not x86)
ggml, XNNPACK and most Arm kernel libraries bake their fast paths into each binary at compile time.
(ggml can ship several single-ISA binaries and pick one at load via GGML_CPU_ALL_VARIANTS, but each
binary is itself compile-time fixed, and coldpath scans each one, which is exactly what the missing
Windows-on-Arm flag turned off.) After llama.cpp
PR #10457 (which removed runtime ISA detection to fix
a ~15x regression, #10435), the i8mm/dotprod
intrinsics are #if-guarded on __ARM_FEATURE_MATMUL_INT8 / __ARM_FEATURE_DOTPROD.
So if the build used the wrong -march, the fast path is not merely skipped; it is compiled out,
physically absent from the binary. Disassembling .text and looking is therefore decisive.
Two properties make this sound:
- AArch64 is fixed-width: 4-byte instructions on a 4-byte grid. A data word embedded in
.text(a literal pool, a jump table) is local to its own 4 bytes and can never cascade into the surrounding code. coldpath decodes with a resyncing sweep, so one data word is skipped, never fatal. An x86 linear sweep would desynchronize and corrupt everything downstream; on Arm it cannot. The approach is sound here precisely because of a property x86 lacks. - Detection reads register operands, not capstone instruction groups. Capstone decodes SVE and SME
correctly but leaves
insn.groupsempty for both, so group-based detection silently reports zero, the exact false-negative class this tool exists to catch.
coldpath reports what is present and reachable in the binary. For ggml that equals what will execute, because dispatch is compile-time. For runtimes that dispatch at runtime (ONNX Runtime's MLAS, ACL), presence proves the kernel was shipped; coldpath does not claim to prove it is selected for a given shape. See Scope.
Install
pip install coldpath
Use
coldpath libggml-cpu.so # one binary: verdict + per-feature counts + coverage + sha256
coldpath ./ollama/lib/ollama/ # a whole release directory
coldpath onnxruntime-1.27.0-aarch64.whl # a pip wheel
coldpath app-release.apk # an Android APK
coldpath libfoo.so --json # machine-readable
Reads AArch64 ELF, PE and Mach-O (including universal binaries) and looks inside .whl, .apk,
.zip, .tar.gz and .tar.zst. Verdicts: HOT (SME), WARM (i8mm matrix), TEPID (dotprod
only), COLD (nothing), UNKNOWN (too little of .text decodable to judge, coldpath refuses to
assert absence it cannot back up).
As a CI gate
- uses: JonathanSolvesProblems/coldpath@v1
with:
path: ./build/
require: i8mm # fail the PR if the shipped binary has no matrix instructions
For a multi-variant release that dlopens the best of several single-ISA libraries at runtime, add
--any so the set is judged by its best member, not its armv8.0 fallback.
What it found
Scans of official, unmodified release binaries from each project's own release channel, with the
hardened tool. All decoded at 100% coverage, so the COLD/zero rows are sound, not truncation artifacts.
Re-verified on every push by .github/workflows/test.yml: if a project
fixes its build, that job fails on purpose.
| binary (official release) | verdict | SME | i8mm | dotprod | cov | sha256 |
|---|---|---|---|---|---|---|
| ONNX Runtime 1.27.0, aarch64 wheel | HOT | 469 | 800 | 1,642 | 100% | 9275ef52 |
| ExecuTorch 1.3.1, aarch64 wheel | HOT | 771 | 960 | 3,185 | 100% | 9208f5cf |
llama.cpp b10344, linux-arm64 (best variant, named armv9.2) |
WARM | 0 | 402 | 1,253 | 100% | 807564f5 |
| llama.cpp b10344, win-arm64 | WARM | 0 | 244 | 1,052 | 100% | b9a0dd0e |
| Ollama v0.31.2, linux-arm64 (best variant) | WARM | 0 | 384 | 1,110 | 100% | n/a |
| Ollama v0.31.2, win-arm64 | COLD | 0 | 0 | 0 | 100% | 536ada0d |
Three findings fall out:
1. Ollama's Windows-on-Arm build has no matrix or dot-product instructions. Cause and one-line fix
in examples/ollama-fix/; ~5.75x prefill from the dot-product fix I filed, up to
~6.9x with i8mm, measured on Cobalt 100. Its Linux build is fine (the row above), so this is
Windows-specific and build-flag-specific, not Ollama being incapable. Confirmed COLD on all 8 stable
releases (v0.32.2 was withdrawn) from v0.31.2 through the current v0.32.7, and
test.yml re-downloads the latest release and re-checks it on every push,
so this claim can't silently rot.
2. No ggml / llama.cpp CPU backend ships SME, including the backends named for it (ONNX Runtime and
ExecuTorch, in the table above, do ship SME by default; this is specific to the ggml stack). llama.cpp's
libggml-cpu-armv9.2_1.so / _armv9.2_2.so are named for the Armv9.2-A architecture level (where
SME is an optional extension), and contain zero ZA-tile instructions, zero smstart, zero
outer-products. Cause: SME reaches ggml only through KleidiAI, and GGML_CPU_KLEIDIAI defaults to
OFF, so stock builds contain no SME unless it is explicitly enabled at configure time. ggml's
runtime dispatcher still loads the highest variant a CPU supports, so on an SME-capable Android device
(Dimensity 9500, Snapdragon 8 Elite Gen 5) it selects the armv9.2 backend believing it is optimized,
and gets a library with no SME in it.
3. The cold path is a build-flag choice, not a hardware or ecosystem limit. ONNX Runtime and ggml
depend on the same KleidiAI. I disassembled the stock aarch64 onnxruntime wheel and counted 469 SME
instructions (208 of them MOPA outer-products); they originate in KleidiAI, which ORT compiles in with
onnxruntime_USE_KLEIDIAI=ON by default (opt-out). ggml ships zero by default (opt-in via
GGML_CPU_KLEIDIAI=ON). Same ISA, same dependency, opposite default.
Is it right?
The tool is validated against ground truth it did not author, and against a positive control.
Ground truth, llama.cpp's own ISA ladder. The official linux-arm64 release ships eight ggml-cpu
backends with the ISA level in the filename. A correct detector must reproduce that staircase exactly:
$ python scripts/verify_ladder.py llama-b10344/
variant dotprod sve i8mm sme verdict
armv8.0_1 0 0 0 0 ok
armv8.2_1 1,184 0 0 0 ok
armv8.2_2 1,184 0 0 0 ok
armv8.2_3 1,235 10,735 0 0 ok
armv8.6_1 1,253 12,102 402 0 ok
armv8.6_2 1,253 11,987 402 0 ok
armv9.2_1 1,253 12,087 402 0 ok
armv9.2_2 1,253 12,087 402 0 ok
coldpath reproduces llama.cpp's own 8-variant ISA ladder exactly.
Positive control, SME really is findable. The ORT wheel above proves a zero means a real absence,
not a broken detector. pytest (24 tests) covers each instruction family against hand-assembled
encodings, the resync-through-data property, the coverage gate, and the single-word corroboration floor,
so correctness is provable without any binary on disk.
This validation is the receipt, not the headline. The headline is the 5.75x on prefill (and ~2.2x on decode) that one build flag recovers on Arm cloud silicon.
Scope and honest limitations
- Static presence, not dynamic frequency. coldpath proves an instruction is present and reachable,
not how often it runs. Absence is proof (the kernel cannot execute); presence is necessary, not
sufficient. For the headline finding the benchmark closes that gap directly: the WARM build runs ~6.9x
faster than the byte-identical-source COLD build, and a binary that contained
smmlabut never dispatched to it would be no faster than COLD, so the speedup is live evidence the matrix path executes. For a runtime-dispatching library (ONNX Runtime's MLAS, ACL) coldpath proves the kernel shipped, not that it is selected for a given shape. - It sees the ISA path, not external matrix units. On macOS, ggml/llama.cpp can route matmul through Apple's Accelerate BLAS, which uses the AMX matrix unit, hardware acceleration that is not ISA-visible to a disassembler. So coldpath's headline is scoped to Linux/Neoverse, Windows-on-Arm, and Android, where the public ISA path is the only path. The specific findings here are safe from this: ggml has no runtime kernel generation, and Ollama ships no BLAS backend on Windows-on-Arm, so there is no external unit to miss.
- Three states, reported honestly. A kernel can be (i) absent, compiled out; (ii) present but runtime-gated behind a CPU-feature check; (iii) present and on the default path. coldpath distinguishes absent (its whole point) from present. It does not claim a present kernel is dispatched at runtime; for ggml that question is moot (dispatch is compile-time), for a runtime-dispatching library it is out of scope by design.
- No SME hardware was benchmarked. No shipping cloud Arm CPU has SME, not Graviton3/4/5, not Cobalt 100/200, not Axion, including the 2026 Neoverse-V3 parts. The SME findings are proven statically (the instructions are absent); their runtime cost is not measured here, and I do not quote Arm's "up to 6x" as if it were mine.
- Coverage is published next to every finding so the absence claims are auditable. Below 0.90 the verdict is UNKNOWN, never a false COLD.
- coldpath checks whether a binary is fast on Arm, not whether it builds on Arm, a different, already well-served question.
How it compares
| needs Arm hardware? | needs a running workload / profiler? | works on any x86 laptop? | one-command CI-gate verdict? | catches capstone's SME false-negative? | |
|---|---|---|---|---|---|
| coldpath | no | no | yes | yes | yes |
| Arm Streamline / Performix | yes | yes (live profiling) | no | no | n/a |
arm/mcp arm-readiness checks |
no | no | yes | partial | no (checks it builds, not that it's fast) |
objdump / llvm-objdump |
no | no | yes | no (raw bytes, DIY) | no |
runtime profilers (perf) |
yes | yes | no | no | n/a |
Streamline and Performix profile a running workload on real Arm silicon. arm-readiness tools check whether
code builds on Arm64. objdump shows raw bytes. None give a turnkey, hardware-free, CI-gateable answer
to "does my Arm AI binary actually use the matrix hardware."
Who this is for
- Platform / MLOps teams running LLM inference on Arm64 cloud (Graviton, Cobalt, Axion), the fast-growing default for cost-efficient inference. A cold build silently costs ~5.75x on prompt throughput and the matching share of the bill.
- Runtime and image distributors shipping portable AArch64 binaries (pip wheels, Docker images, prebuilt releases), where cross-compiling for device compatibility is exactly what strands the matrix path.
- Anyone evaluating Arm instances, to confirm the runtime they picked actually uses the silicon they are paying for.
Reusable artifacts come out of it: the upstream fix (PR #17654), the CI-gate Action others can drop into their own pipeline, the scan scoreboard of official Arm AI binaries, and a field guide that teaches why Arm builds ship cold and how to check and fix any of your own, in one command.
New here? Start with the field guide: the 30-second version, why it happens, how to check any build, and how to fix it.
Licence
MIT.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file coldpath-0.1.1.tar.gz.
File metadata
- Download URL: coldpath-0.1.1.tar.gz
- Upload date:
- Size: 25.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
091a3051b97a485af0d8f233c31fb3462caeb410daf59eaea086879b0379b87a
|
|
| MD5 |
db6bef760dfe5a9ca4a710f9edaecf53
|
|
| BLAKE2b-256 |
62638992e546a0c4b56b376cabbed586a8c1aaa1158fdb717028879c3c950176
|
Provenance
The following attestation bundles were made for coldpath-0.1.1.tar.gz:
Publisher:
publish.yml on JonathanSolvesProblems/coldpath
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
coldpath-0.1.1.tar.gz -
Subject digest:
091a3051b97a485af0d8f233c31fb3462caeb410daf59eaea086879b0379b87a - Sigstore transparency entry: 2419258947
- Sigstore integration time:
-
Permalink:
JonathanSolvesProblems/coldpath@aa78966308c8adb44b87c14b3cdd495c236eeb08 -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/JonathanSolvesProblems
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@aa78966308c8adb44b87c14b3cdd495c236eeb08 -
Trigger Event:
release
-
Statement type:
File details
Details for the file coldpath-0.1.1-py3-none-any.whl.
File metadata
- Download URL: coldpath-0.1.1-py3-none-any.whl
- Upload date:
- Size: 18.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
77ba57a94c8b76fb7a89bde5e40a07df1d9246564906952190e1d2eeabedd8b2
|
|
| MD5 |
1a7d2d30017dbbe87930159331915f7f
|
|
| BLAKE2b-256 |
45a187ce84bf30e272bf0d5de471248109dbda39cdb34c9e1cc8868789b829f7
|
Provenance
The following attestation bundles were made for coldpath-0.1.1-py3-none-any.whl:
Publisher:
publish.yml on JonathanSolvesProblems/coldpath
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
coldpath-0.1.1-py3-none-any.whl -
Subject digest:
77ba57a94c8b76fb7a89bde5e40a07df1d9246564906952190e1d2eeabedd8b2 - Sigstore transparency entry: 2419259628
- Sigstore integration time:
-
Permalink:
JonathanSolvesProblems/coldpath@aa78966308c8adb44b87c14b3cdd495c236eeb08 -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/JonathanSolvesProblems
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@aa78966308c8adb44b87c14b3cdd495c236eeb08 -
Trigger Event:
release
-
Statement type: