An open 14MB agentic LLM for tool calling on tiny devices
Needle 2 is an open 45M-parameter agentic LLM designed for tool calling, device use, and structured extraction on tiny hardware. The entire model ships as a single 14 MB binary that runs a full session in about 28 MB of RAM, and it executes directly in the browser via WebAssembly. It is built on the Simple Attention Network findings, compressed to CQ2-bit with Cactus Quants, and baked into its own dependency-free C++ inference engine that runs from Cortex-M to x86 to WebAssembly. Needle 2 is licensed under Apache 2.0 with weights on Hugging Face.