Skip to main content

Needle

A foundation model for mobiles, wearables, robots, smart home, automotive and microcontrollers. The whole model is a single 8-29 MB binary built on our Simple Attention Network, and we trade general chat capacity to beat models 10x its size on mobile tool calls and match 2-3x bigger models on extraction.

  • Tool calls: given the functions your app exposes, Needle picks the right ones and fills every argument from what the user said. Ask for two things and you get two calls in order; ask for something no tool covers and you get an empty list, not a guess.
  • Structured extraction: declare a shape, hand over messy text, get typed fields back: an invoice, a booking, a notification, a form. The decode grammar guarantees the output parses, and extraction generalises to classification.
  • Text embedding: the same model returns a vector for a sentence, so an app can search, match and route locally.

Needle 3 at a glance

Needle 3 is a Laddered Simple Attention Network: a Monarch Hadamard MLP in place of the FFN, GQA attention with causal conv taps, engram n-gram memory read by gather, and multi-lane hyper-connections, trained so that every depth from 2 to 20 layers is a deployable model. Most of its parameters sit in the engram, so the 121M model does the arithmetic of a 50M one. A byte-level grammar compiled from your schemas constrains every token, and every response carries a calibrated confidence score from a learned head. The architecture diagram is on the release page.

Benchmarks

Tool calling is exact-match accuracy on the full test splits, extraction is field micro-F1 on the full test splits.

Needle 3 against baselines on six benchmarks

The interactive frontier plot, the architecture and the fine-tuning results are at cactuscompute.com/needle.

Get started

or bootstrap it with one line (creates an isolated venv, verifies the wheel
checksum, shims `neuralos` onto PATH):

```sh
curl -fsSL https://raw.githubusercontent.com/DrOlu/neuralOS/main/install.sh | sh
irm https://raw.githubusercontent.com/DrOlu/neuralOS/main/install.ps1 | iex

pip install neuralos


```bash
npm install neuralos       # Node.js — same version number, same API surface

One version, two channels. The PyPI wheels and the npm package set are built from the same release tag, so npm view neuralos version and the wheel filename always agree (currently 3.10.0). The npm package bundles the engine and weights for macOS arm64/x64, Linux x64/arm64 (glibc) and Windows x64; the wheels carry a broader engine matrix. The engine's own release line is independent of the product version.

neuralOS vs cactus-needle: the neuralos PyPI distribution is this project with the engine bundled - its platform wheels ship the engine shared library and the base weights (~36 MB), so it works fully offline after install. The needle command is installed as an alias alongside neuralos. pip install cactus-needle keeps the download-at-runtime behavior (engine fetched from HuggingFace on first use).


Try it in the browser at [cactuscompute.com/needle](https://cactuscompute.com/needle); the weights and every platform engine are on [Hugging Face](https://huggingface.co/Cactus-Compute/needle3).

Decorate a function: the signature gives the argument types, the docstring is the tool description, and `run()` completes the loop, executing your function and returning its results.

```python
import needle

@needle.tool
def get_weather(city: str):
    "Get the current weather for a city."
    return {"city": city, "temp_c": 27, "sky": "clear"}

agent = needle.Needle(tools=[get_weather])
print(agent.run("what's it like in Lagos right now?")["results"])
# [{'city': 'Lagos', 'temp_c': 27, 'sky': 'clear'}]

Every turn returns one JSON object with function_calls, the model's reasoning and a calibrated confidence; an off-topic request returns an empty list rather than a guess. needle.Needle(tools=[...], generation=2) keeps running Needle 2 for existing deployments.

Guides

  • How to design tools for Needle 3: one tool per action, names users would say, formats in descriptions, constraints in the grammar, triggers.
  • Leveraging Needle's confidence: what the score measures, what the engine withholds, and routing on act, confirm or refuse.
  • Structured JSON extraction with Needle: the record as the only tool, typed results, classification with enums.
  • Fine-tuning Needle: the data format, the commands, reading the loss, sizing the dataset.
  • Needle Python docs: the API, the response shape, the behaviour contract, system facts, tool retrieval, offline devices, environments, the CLI.
  • What devices are supported on Needle: every platform folder, the CLI runner, the C API, the browser, WASI, air-gapped setup.
  • The .cact format: the file the engine maps and reads in place, Cactus Quants at 2.125 bits per weight, and how to parse it yourself.
  • Porting Needle 3: notes for writing your own runtime, the oracle to test against, the tensor order the container promises, the prompt on the wire, the ladder rule, retrieval with needle_embed.

llms.txt in this repo carries the same reference for AI coding assistants.

Customisation

Needle was designed to be customised. Its capacity is a ladder, and a subnetwork as small as 2 layers, fine-tuned on one product's tools, runs optimally on devices far smaller than the full model needs. Fine-tuning on DroidCall lifts every subnetwork by 18 to 36 points, and from 4 layers up the tuned subnetwork passes DeepSeek V4 Flash, starting at 29M parameters.

Every subnetwork before and after fine-tuning on DroidCall and on Mobile Actions

Two ways to fine-tune, from the same package:

Local, needle finetune Platform, needle platform finetune
What trains LoRA adapters on the attention projections, base frozen, merged at export The full model, every depth from 2 layers up
What it keeps Your data only Your data reinforced with Needle's original dataset, so nothing already learned is unlearned
Confidence Head untouched; confidence is None Head fine-tuned with the model, calibrated on your tools
Precision 4-bit 2-bit, the same post-training as the shipped model
Data Your JSONL, query/answers or chat format Yours, or generated from your tool definitions, 100 to 10,000 examples per run
Scores Validation loss Validation and test accuracy for every depth
Compute Your machine, JAX on CPU, CUDA or Metal Cactus GPUs
Runs from The CLI The CLI, Python, the dashboard, or a coding agent holding your key

Local:

pip install "cactus-needle[train]"
needle finetune data.jsonl --epochs 10 --out adapter.safetensors
needle build --lora adapter.safetensors --layers 8 --out tuned.cact

Platform, with a key from the console in NEEDLE_API_KEY. One command uploads the files, trains and scores every size, and downloads the .cact files; once a job is submitted it can also be followed on the dashboard:

export NEEDLE_API_KEY=needle_ft_...
needle platform generate --tools tools.json --examples 1000 --out ./data
needle platform finetune data/train.jsonl data/validation.jsonl data/test.jsonl --suffix smart-home --out ./models
from needle.platform import Platform

client = Platform()
job = client.wait(client.finetune(["train.jsonl"], ["validation.jsonl"], ["test.jsonl"], suffix="smart-home"))
paths = client.download(job["fine_tuned_model"], "models", depth=8)

Or hand the key to Claude Code or Codex with cactuscompute.com/llms.txt and let the agent run the loop. needle platform jobs | models | files | billing list what the account holds, needle download model-<id> fetches a model by id, and the fine-tuning guide covers the data format and how to read the scores.

Deploy

Every deployment target ships a prebuilt engine under 1 MB that loads the needle3.cact weights at start. needle build --platform <folder> [--layers N] fetches that engine and puts the weights beside it.

One engine per platform folder

needle build --platform macos-arm64
needle build --platform linux-arm64 --layers 8 --out ./pi
./macos-arm64/needle --model needle3.cact --tools tools.json --serve

The devices guide lists every folder and what ships in it.

By default, telemetry is turned on in the binary. To turn it off, set environment variables NEEDLE_TELEMETRY=0 and DO_NOT_TRACK=1.

Citation

Needle is built by the Cactus Compute team. If you use it in your work, please cite:

@misc{needle3_2026,
  title        = {Needle: Automation Foundation Model for Tiny Devices},
  author       = {Ndubuaku, Henry and Mosoyan, Karen and Mroz, Jakub and Cylich, Noah and
                  Kumar, Satyajit and Sandhu, Parkirat and Shemet, Roman and Lee, Justin H.},
  year         = {2026},
  organization = {Cactus Compute, Inc.},
  howpublished = {\url{https://github.com/cactus-compute/needle}}
}

Metadata

Release files for neuralos 3.10.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for neuralos 3.10.2
File
neuralos-3.10.2-py3-none-win_amd64.whl Python 3 none Windows x86-64 Details
neuralos-3.10.2-py3-none-manylinux2014_x86_64.whl Python 3 none Linux glibc 2.17+ x86-64 Details
neuralos-3.10.2-py3-none-manylinux2014_aarch64.whl Python 3 none Linux glibc 2.17+ ARM64 Details
neuralos-3.10.2-py3-none-macosx_11_0_arm64.whl Python 3 none macOS 11.0+ ARM64 Details
neuralos-3.10.2-py3-none-any.whl Python 3 none any Details

Total release size: 136.7 MB

Release files / neuralos-3.10.2-py3-none-win_amd64.whl

Download URL neuralos-3.10.2-py3-none-win_amd64.whl
Size 34.2 MB
Tags Python 3 Windows x86-64
SHA-256 checksum
How to use checksums
d8a0cb94f16a098ae74728be45744611c1be4ebcc71e18138e59bff39c621c57
BLAKE2b-256 checksum
How to use checksums
7feb851e87f21a19b017a32ec3eb93784ac4c0d1c53ed8df0ce42a16a662fa97
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.12.1

Release files / neuralos-3.10.2-py3-none-manylinux2014_x86_64.whl

Download URL neuralos-3.10.2-py3-none-manylinux2014_x86_64.whl
Size 34.2 MB
Tags Linux glibc 2.17+ x86-64 Python 3
SHA-256 checksum
How to use checksums
e0de2931ee805781205e198ecc5c1c1039693c00f49a90fa5e76ca8f16335b7c
BLAKE2b-256 checksum
How to use checksums
ef63d1925a6daa8970b3224379640a0f5d9853a095db8c5cf7cb19aaf5c342db
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.12.1

Release files / neuralos-3.10.2-py3-none-manylinux2014_aarch64.whl

Download URL neuralos-3.10.2-py3-none-manylinux2014_aarch64.whl
Size 34.2 MB
Tags Linux glibc 2.17+ ARM64 Python 3
SHA-256 checksum
How to use checksums
ae968b3755ea4b9ed28e9b14039de36789c4f2ec0ba3646789f24f05c706074b
BLAKE2b-256 checksum
How to use checksums
9c7dcb0d4d6ca75419dd9b9a2663f8969e53c2fac007e601a9e4d562d0557d4d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.12.1

Release files / neuralos-3.10.2-py3-none-macosx_11_0_arm64.whl

Download URL neuralos-3.10.2-py3-none-macosx_11_0_arm64.whl
Size 34.0 MB
Tags Python 3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
8249ef39a4f29c54314f470c02ff494be8fbb059535e4ef5ebdc8cc6a68b48b3
BLAKE2b-256 checksum
How to use checksums
1163fa8fd9aa5767f75bcc423a11857c476caba7c5ff77417674ee8711952715
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.12.1

Release files / neuralos-3.10.2-py3-none-any.whl

Download URL neuralos-3.10.2-py3-none-any.whl
Size 100.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9ff8c128b817bede49a2f9b3b3130ab19a25b56417885d6d7af26d12bf2ad3c9
BLAKE2b-256 checksum
How to use checksums
b897b96e702b3d0468753c9941949d5fe81c64dd3a5a4ca242edb8ae0de27897
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.12.1

Release history Release notifications | RSS feed

This release

3.10.2 This release

5 release files

3.0.3

5 release files

3.0.2

5 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page