Needle 3: Tiny Tool-Calling Models for Phones, Raspberry Pi and Edge Devices
Needle 3 is an on-device model built for a narrow set of jobs: select tools, fill structured arguments, extract fields, and create embeddings. It deliberately gives up general chat ability to fit into a 9 to 29 MB quantized binary that can run on phones, Raspberry Pi systems, browsers, smart-home hardware and other constrained devices. The specifications and integration examples below follow the official Needle repository.
That scope is the key to understanding the release. Needle 3 is not a tiny substitute for a frontier assistant. It is an automation component for applications that need reliable JSON-shaped actions without sending every request to a cloud model.
Needle 3 specifications
| Field | Needle 3 |
|---|---|
| Architecture | Laddered Simple Attention Network |
| Parameter range | 29M to 121M, depending on depth |
| Shipped quantization | CQ2, 2-bit |
| Binary size | Approximately 9 to 29 MB by subnetwork |
| Inputs | Text and tool or extraction schemas |
| Outputs | Structured tool calls, extraction results or embeddings |
| Depth options | 2 to 20 layers |
| Python package | cactus-needle |
| Offline inference | Supported |
| License | Check the current repository and weight terms for each component |
Every supported depth is a usable subnetwork. A developer can choose fewer layers for a constrained device or more layers when accuracy matters more than compute. Most parameters live in the modelβs engram memory, so reducing active depth cuts compute more aggressively than file size.
What the model actually does
Needle exposes three main operations.
Tool calling: Given function definitions, it chooses an applicable tool and fills arguments from the userβs request. If no tool fits, strict mode can return no call rather than inventing one.
Structured extraction: Given a schema, it extracts typed fields from text. Examples include invoices, bookings, forms and notification routing.
Embeddings: The same model can return a vector for retrieval, matching or routing.
This differs from the open-ended agents described in our AI agent guide. Needle is closer to a local intent-and-structure engine. Your application still owns authorization, execution, retries and user confirmation.
Install and run it from Python
The official package is installed from PyPI:
python -m venv .venv
source .venv/bin/activate
pip install cactus-needle
Define tools with ordinary Python functions:
from typing import Literal
import needle
@needle.tool
def set_light(
room: str,
action: Literal["on", "off"],
brightness: int = 100,
):
"""Turn a room light on or off at a brightness from 0 to 100."""
return {
"room": room,
"action": action,
"brightness": brightness,
}
agent = needle.Needle(tools=[set_light])
result = agent.run("Turn the kitchen light on at 40 percent")
print(result["function_calls"])
The schema and description matter. Narrow tools with explicit formats are easier to route than one generic function that accepts arbitrary instructions. Our tool-calling guide explains why the modelβs JSON is only a proposal until application code validates and authorizes it.
Structured extraction
Needle can compile a Pydantic schema into constrained decoding:
from pydantic import BaseModel
import needle
class Invoice(BaseModel):
vendor: str
total: float
due_date: str
invoice = needle.extract(
"Invoice from Acme, total EUR 1200, due 2026-10-15",
Invoice,
)
print(invoice.vendor, invoice.total, invoice.due_date)
Grammar-constrained output helps ensure the response parses. It does not ensure that every extracted value is correct. Validate currencies, dates, totals and business rules separately, then route low-confidence or consequential cases to human review.
Phones, Raspberry Pi and native platforms
Cactus publishes prebuilt engines for multiple targets. The documented set includes macOS and Linux on ARM64, Linux x86-64, Windows, browser/WASM and a WASI component. Device-oriented targets extend to phones, embedded systems and other edge environments.
For a Raspberry Pi or similar ARM64 host, the build command fetches a platform engine and selected weights:
needle build --platform linux-arm64 --layers 8 --out ./needle-pi
./needle-pi/needle --model needle3.cact --tools tools.json --serve
The native engine is under 1 MB according to the project; model weights are supplied separately. Performance figures vary heavily by subnetwork, device, cooling and input size. Cactus reports 400 to 4,000 decode tokens per second on Raspberry Pi 5 across its tested configurations, but that is vendor testing rather than an independent guarantee.
Browser and WASI deployment
Needle ships a browser target and a WASI Preview 2 component. This enables local intent routing or extraction in applications that cannot assume Python or a full server runtime.
Browser deployment can improve privacy and latency because user text need not leave the device for inference. It also exposes the model and tools to the client environment. Never put reusable server credentials in browser tool definitions. Let the client propose an action, then authorize it on a trusted server.
WASI is useful for portable sandboxed runtimes. The component exposes load, initialize, complete, embed and reset operations. Test the exact host implementation because WASI support and resource limits vary.
Offline and air-gapped use
The official documentation includes an offline path. Download the Python package, engine and weights on a connected machine, copy them into the documented cache or package location, and set HF_HUB_OFFLINE=1 so missing artifacts fail instead of triggering a network request.
Inference can run without a network connection once the required files are present. That does not make the setup automatically compliant. Record model and engine hashes, review every copied dependency, and control how extracted data leaves the device after inference.
For broader local-model operations, see the Ollama guide and local AI security guidance.
Telemetry is on by default
The official repository states that telemetry is enabled by default in the binary. It documents these environment variables to disable it:
export NEEDLE_TELEMETRY=0
export DO_NOT_TRACK=1
Set and verify them before claiming a fully offline or air-gapped deployment. In regulated environments, monitor network attempts during acceptance testing rather than relying only on configuration. Also check whether Python wrappers, package managers or model-download libraries introduce their own telemetry or update checks.
The DeepSeek V4 Flash comparison
The headline claim needs careful wording. Cactus reports that a fine-tuned Needle 3 subnetwork can pass DeepSeek V4 Flash on specific downstream tool-calling tasks, starting at a particular tuned depth. The comparison does not establish parity in general coding, chat, research, reasoning or every tool set.
The projectβs tool benchmarks use exact-match or AST-style evaluation on suites such as Mobile Actions, DroidCall and BFCL. Extraction uses field-level F1 on structured datasets. Fine-tuning the Needle subnetworks on the downstream task materially changes the result. Therefore the fair conclusion is:
Needle 3 can be competitive with much larger hosted models on narrow, structured automation tasks after task-specific tuning.
It is not accurate to say that a 9 MB model generally matches DeepSeek V4 Flash.
Where Needle 3 fits
- Local smart-home command routing
- Form and invoice extraction on controlled devices
- Offline vehicle or robot commands
- Wearable interactions with a small tool catalog
- Browser-side intent classification
- Private embeddings for local search
- Low-latency routing before a larger model is called
It is a poor fit for open-ended writing, broad coding assistance, deep research or tasks requiring large factual knowledge. A hybrid architecture can route narrow actions locally and send genuinely complex work to a larger model with explicit consent.
Limitations
- It is specialized rather than a general conversational model.
- Vendor benchmark results need independent reproduction.
- Task-specific fine-tuning is central to the strongest comparison claims.
- Telemetry must be disabled explicitly for a no-network policy.
- Supported-device labels do not guarantee acceptable latency on every device.
- Schema-valid output can still contain wrong values.
- Application code remains responsible for permissions and side effects.
My take
Needle 3 is interesting because it narrows the problem instead of pretending a tiny model can do everything. For device commands, extraction and routing, 9 to 29 MB is operationally meaningful. The right pilot is one small tool catalog with a real refusal test, latency target and error budget. If that succeeds, expand carefully. Do not compare it with frontier models outside the specific structured tasks both systems actually ran.
FAQ
How large is Needle 3?
The project documents approximately 9 to 29 MB CQ2 binaries across subnetworks, with 29M to 121M parameters depending on depth.
Can Needle 3 run on Raspberry Pi?
Yes. Cactus publishes a Linux ARM64 target and documents Raspberry Pi 5 performance. Actual speed depends on depth, input and hardware conditions.
Does it work offline?
Yes, after the engine, weights and dependencies are downloaded. Set the documented offline and telemetry controls, then verify network behavior.
Does Needle 3 match DeepSeek V4 Flash?
Only in a narrow project-reported comparison involving fine-tuned subnetworks and specific downstream tool-calling tasks. It is not general model parity.
Is structured output always correct?
The grammar helps it parse correctly. Values can still be wrong, so applications need validation and human review for consequential actions.