Skip to main content

Needle

A foundation model for mobiles, wearables, robots, smart home, automotive and microcontrollers. The whole model is a single 8-29 MB binary built on our Simple Attention Network, and we trade general chat capacity to beat models 10x its size on mobile tool calls and match 2-3x bigger models on extraction.

  • Tool calls: given the functions your app exposes, Needle picks the right ones and fills every argument from what the user said. Ask for two things and you get two calls in order; ask for something no tool covers and you get an empty list, not a guess.
  • Structured extraction: declare a shape, hand over messy text, get typed fields back: an invoice, a booking, a notification, a form. The decode grammar guarantees the output parses, and extraction generalises to classification.
  • Text embedding: the same model returns a vector for a sentence, so an app can search, match and route locally.

Needle 3 at a glance

Needle 3 is a Laddered Simple Attention Network: a Monarch Hadamard MLP in place of the FFN, GQA attention with causal conv taps, engram n-gram memory read by gather, and multi-lane hyper-connections, trained so that every depth from 2 to 20 layers is a deployable model. Most of its parameters sit in the engram, so the 121M model does the arithmetic of a 50M one. A byte-level grammar compiled from your schemas constrains every token, and every response carries a calibrated confidence score from a learned head. The architecture diagram is on the release page.

Benchmarks

Tool calling is exact-match accuracy on the full test splits, extraction is field micro-F1 on the full test splits.

Needle 3 against baselines on six benchmarks

The interactive frontier plot, the architecture and the fine-tuning results are at cactuscompute.com/needle.

Get started

pip install cactus-needle

Try it in the browser at cactuscompute.com/needle; the weights and every platform engine are on Hugging Face.

Decorate a function: the signature gives the argument types, the docstring is the tool description, and run() completes the loop, executing your function and returning its results.

import needle

@needle.tool
def get_weather(city: str):
    "Get the current weather for a city."
    return {"city": city, "temp_c": 27, "sky": "clear"}

agent = needle.Needle(tools=[get_weather])
print(agent.run("what's it like in Lagos right now?")["results"])
# [{'city': 'Lagos', 'temp_c': 27, 'sky': 'clear'}]

Every turn returns one JSON object with function_calls, the model's reasoning and a calibrated confidence; an off-topic request returns an empty list rather than a guess. needle.Needle(tools=[...], generation=2) keeps running Needle 2 for existing deployments.

Guides

llms.txt in this repo carries the same reference for AI coding assistants.

Customisation

Needle was designed to be customised. Its capacity is a ladder, and a subnetwork as small as 2 layers, fine-tuned on one product's tools, runs optimally on devices far smaller than the full model needs. Fine-tuning on DroidCall lifts every subnetwork by 18 to 36 points, and from 4 layers up the tuned subnetwork passes DeepSeek V4 Flash, starting at 29M parameters.

Every subnetwork before and after fine-tuning on DroidCall and on Mobile Actions

pip install "cactus-needle[train]"
needle finetune data.jsonl --epochs 10 --out adapter.safetensors
needle build --lora adapter.safetensors --layers 8 --out tuned.cact

Local fine-tuning trains and exports at 4 bits; the fine-tuning guide has the rest. The 2-bit post-training and quantisation behind the shipped model, enriched with Cactus proprietary datasets, run on the Cactus Platform.

Deploy

Every deployment target ships a prebuilt engine under 1 MB that loads the needle3.cact weights at start. needle build --platform <folder> [--layers N] fetches that engine and puts the weights beside it.

One engine per platform folder

needle build --platform macos-arm64
needle build --platform linux-arm64 --layers 8 --out ./pi
./macos-arm64/needle --model needle3.cact --tools tools.json --serve

The devices guide lists every folder and what ships in it.

Citation

Needle is built by the Cactus Compute team. If you use it in your work, please cite:

@misc{needle3_2026,
  title        = {Needle: Automation Foundation Model for Tiny Devices},
  author       = {Ndubuaku, Henry and Mosoyan, Karen and Mroz, Jakub and Cylich, Noah and
                  Kumar, Satyajit and Sandhu, Parkirat and Shemet, Roman and Lee, Justin H.},
  year         = {2026},
  organization = {Cactus Compute, Inc.},
  howpublished = {\url{https://github.com/cactus-compute/needle}}
}

Release files for cactus-needle 3.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cactus-needle 3.0.0
File Size Uploaded
cactus_needle-3.0.0.tar.gz 101.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cactus-needle 3.0.0
File Interpreter ABI Platform
cactus_needle-3.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 191.2 kB

Release files / cactus_needle-3.0.0.tar.gz

Download URL cactus_needle-3.0.0.tar.gz
Size 101.3 kB
Tags Source
SHA-256 checksum
How to use checksums
39301239e5a13ae650f91ec59d0a8262beb6e32850dbb58d666716268f83f5d8
BLAKE2b-256 checksum
How to use checksums
91c31190965e3db17aaf022526277ad8a2359d1e7b38a0c04719892b89d9b3d8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release files / cactus_needle-3.0.0-py3-none-any.whl

Download URL cactus_needle-3.0.0-py3-none-any.whl
Size 89.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
72757799e66c286ecea0f70d238dc3682d48cbfcf77d4df4508fb1f1dd36dfd1
BLAKE2b-256 checksum
How to use checksums
7f2f224fde52fc2c9f65a4d1e880ec785b4a947cb3994a3743fb8723a797f497
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release history Release notifications | RSS feed

3.0.5

2 release files

3.0.4

2 release files

3.0.3

2 release files

3.0.2

2 release files

3.0.1

2 release files

This release

3.0.0 This release

2 release files

2.0.15

2 release files

2.0.14

2 release files

2.0.11

2 release files

2.0.10

2 release files

2.0.9

2 release files

2.0.8

2 release files

2.0.7

2 release files

2.0.6

2 release files

2.0.5

2 release files

2.0.4

2 release files

2.0.3

2 release files

2.0.2

2 release files

2.0.1

2 release files

2.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page