πŸ€– AI Tools
Β· 7 min read

Needle 3: Tiny Tool-Calling Models for Phones, Raspberry Pi and Edge Devices


Needle 3 is an on-device model built for a narrow set of jobs: select tools, fill structured arguments, extract fields, and create embeddings. It deliberately gives up general chat ability to fit into a 9 to 29 MB quantized binary that can run on phones, Raspberry Pi systems, browsers, smart-home hardware and other constrained devices. The specifications and integration examples below follow the official Needle repository.

That scope is the key to understanding the release. Needle 3 is not a tiny substitute for a frontier assistant. It is an automation component for applications that need reliable JSON-shaped actions without sending every request to a cloud model.

Needle 3 specifications

FieldNeedle 3
ArchitectureLaddered Simple Attention Network
Parameter range29M to 121M, depending on depth
Shipped quantizationCQ2, 2-bit
Binary sizeApproximately 9 to 29 MB by subnetwork
InputsText and tool or extraction schemas
OutputsStructured tool calls, extraction results or embeddings
Depth options2 to 20 layers
Python packagecactus-needle
Offline inferenceSupported
LicenseCheck the current repository and weight terms for each component

Every supported depth is a usable subnetwork. A developer can choose fewer layers for a constrained device or more layers when accuracy matters more than compute. Most parameters live in the model’s engram memory, so reducing active depth cuts compute more aggressively than file size.

What the model actually does

Needle exposes three main operations.

Tool calling: Given function definitions, it chooses an applicable tool and fills arguments from the user’s request. If no tool fits, strict mode can return no call rather than inventing one.

Structured extraction: Given a schema, it extracts typed fields from text. Examples include invoices, bookings, forms and notification routing.

Embeddings: The same model can return a vector for retrieval, matching or routing.

This differs from the open-ended agents described in our AI agent guide. Needle is closer to a local intent-and-structure engine. Your application still owns authorization, execution, retries and user confirmation.

Install and run it from Python

The official package is installed from PyPI:

python -m venv .venv
source .venv/bin/activate
pip install cactus-needle

Define tools with ordinary Python functions:

from typing import Literal
import needle

@needle.tool
def set_light(
    room: str,
    action: Literal["on", "off"],
    brightness: int = 100,
):
    """Turn a room light on or off at a brightness from 0 to 100."""
    return {
        "room": room,
        "action": action,
        "brightness": brightness,
    }

agent = needle.Needle(tools=[set_light])
result = agent.run("Turn the kitchen light on at 40 percent")
print(result["function_calls"])

The schema and description matter. Narrow tools with explicit formats are easier to route than one generic function that accepts arbitrary instructions. Our tool-calling guide explains why the model’s JSON is only a proposal until application code validates and authorizes it.

Structured extraction

Needle can compile a Pydantic schema into constrained decoding:

from pydantic import BaseModel
import needle

class Invoice(BaseModel):
    vendor: str
    total: float
    due_date: str

invoice = needle.extract(
    "Invoice from Acme, total EUR 1200, due 2026-10-15",
    Invoice,
)

print(invoice.vendor, invoice.total, invoice.due_date)

Grammar-constrained output helps ensure the response parses. It does not ensure that every extracted value is correct. Validate currencies, dates, totals and business rules separately, then route low-confidence or consequential cases to human review.

Phones, Raspberry Pi and native platforms

Cactus publishes prebuilt engines for multiple targets. The documented set includes macOS and Linux on ARM64, Linux x86-64, Windows, browser/WASM and a WASI component. Device-oriented targets extend to phones, embedded systems and other edge environments.

For a Raspberry Pi or similar ARM64 host, the build command fetches a platform engine and selected weights:

needle build --platform linux-arm64 --layers 8 --out ./needle-pi
./needle-pi/needle --model needle3.cact --tools tools.json --serve

The native engine is under 1 MB according to the project; model weights are supplied separately. Performance figures vary heavily by subnetwork, device, cooling and input size. Cactus reports 400 to 4,000 decode tokens per second on Raspberry Pi 5 across its tested configurations, but that is vendor testing rather than an independent guarantee.

Browser and WASI deployment

Needle ships a browser target and a WASI Preview 2 component. This enables local intent routing or extraction in applications that cannot assume Python or a full server runtime.

Browser deployment can improve privacy and latency because user text need not leave the device for inference. It also exposes the model and tools to the client environment. Never put reusable server credentials in browser tool definitions. Let the client propose an action, then authorize it on a trusted server.

WASI is useful for portable sandboxed runtimes. The component exposes load, initialize, complete, embed and reset operations. Test the exact host implementation because WASI support and resource limits vary.

Offline and air-gapped use

The official documentation includes an offline path. Download the Python package, engine and weights on a connected machine, copy them into the documented cache or package location, and set HF_HUB_OFFLINE=1 so missing artifacts fail instead of triggering a network request.

Inference can run without a network connection once the required files are present. That does not make the setup automatically compliant. Record model and engine hashes, review every copied dependency, and control how extracted data leaves the device after inference.

For broader local-model operations, see the Ollama guide and local AI security guidance.

Telemetry is on by default

The official repository states that telemetry is enabled by default in the binary. It documents these environment variables to disable it:

export NEEDLE_TELEMETRY=0
export DO_NOT_TRACK=1

Set and verify them before claiming a fully offline or air-gapped deployment. In regulated environments, monitor network attempts during acceptance testing rather than relying only on configuration. Also check whether Python wrappers, package managers or model-download libraries introduce their own telemetry or update checks.

The DeepSeek V4 Flash comparison

The headline claim needs careful wording. Cactus reports that a fine-tuned Needle 3 subnetwork can pass DeepSeek V4 Flash on specific downstream tool-calling tasks, starting at a particular tuned depth. The comparison does not establish parity in general coding, chat, research, reasoning or every tool set.

The project’s tool benchmarks use exact-match or AST-style evaluation on suites such as Mobile Actions, DroidCall and BFCL. Extraction uses field-level F1 on structured datasets. Fine-tuning the Needle subnetworks on the downstream task materially changes the result. Therefore the fair conclusion is:

Needle 3 can be competitive with much larger hosted models on narrow, structured automation tasks after task-specific tuning.

It is not accurate to say that a 9 MB model generally matches DeepSeek V4 Flash.

Where Needle 3 fits

  • Local smart-home command routing
  • Form and invoice extraction on controlled devices
  • Offline vehicle or robot commands
  • Wearable interactions with a small tool catalog
  • Browser-side intent classification
  • Private embeddings for local search
  • Low-latency routing before a larger model is called

It is a poor fit for open-ended writing, broad coding assistance, deep research or tasks requiring large factual knowledge. A hybrid architecture can route narrow actions locally and send genuinely complex work to a larger model with explicit consent.

Limitations

  • It is specialized rather than a general conversational model.
  • Vendor benchmark results need independent reproduction.
  • Task-specific fine-tuning is central to the strongest comparison claims.
  • Telemetry must be disabled explicitly for a no-network policy.
  • Supported-device labels do not guarantee acceptable latency on every device.
  • Schema-valid output can still contain wrong values.
  • Application code remains responsible for permissions and side effects.

My take

Needle 3 is interesting because it narrows the problem instead of pretending a tiny model can do everything. For device commands, extraction and routing, 9 to 29 MB is operationally meaningful. The right pilot is one small tool catalog with a real refusal test, latency target and error budget. If that succeeds, expand carefully. Do not compare it with frontier models outside the specific structured tasks both systems actually ran.

FAQ

How large is Needle 3?

The project documents approximately 9 to 29 MB CQ2 binaries across subnetworks, with 29M to 121M parameters depending on depth.

Can Needle 3 run on Raspberry Pi?

Yes. Cactus publishes a Linux ARM64 target and documents Raspberry Pi 5 performance. Actual speed depends on depth, input and hardware conditions.

Does it work offline?

Yes, after the engine, weights and dependencies are downloaded. Set the documented offline and telemetry controls, then verify network behavior.

Does Needle 3 match DeepSeek V4 Flash?

Only in a narrow project-reported comparison involving fine-tuned subnetworks and specific downstream tool-calling tasks. It is not general model parity.

Is structured output always correct?

The grammar helps it parse correctly. Values can still be wrong, so applications need validation and human review for consequential actions.