Needle 2 Delivers 14MB LLM for Phones and Wearables
TL;DR. Needle 2, an open 45M-parameter LLM, runs as a 14MB binary on mobile, smart home, and robotics devices. - The model uses 28MB of RAM and competes with larger models like FunctionGemma 270M and Apple FM. - Needle 2 achieves up to 500 tokens/sec decode speed on a Raspberry Pi 5 and runs on low-cost hardware. - Its design focuses on tool calling and structured extraction, making it suitable for resource-constrained edge AI.
- Needle 2 is a 45M-parameter LLM packaged as a 14MB binary.
- It runs with 28MB session RAM on devices like phones, wearables, and Raspberry Pi.
- The model specializes in tool calling, device use, and structured data extraction.
- Performance rivals much larger models while operating on sub-$200 hardware.
- Needle 2 is open-source under Apache 2.0 with weights available on Hugging Face.
Sources
- Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots — cactuscompute.com