HenryNdubuaku
- Karma
- 0
- Created
- ()
- Submissions
- 0
- Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots (cactuscompute.com)
Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got…
-
Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-trained to know when…
-
Hey HN, Henry here from Cactus. We open-sourced Needle, a 26M parameter function-calling (tool use) model. It runs at 6000 tok/s prefill and 1200 tok/s decode on consumer devices. We were always frustrated by the little…
- DeepMind x YC x Cactus: Voice Agents Hack (win guaranteed YC interview) (events.ycombinator.com)
- Show HN: Maths, CS and AI Compendium (github.com)
Hey HN, I don’t know who else has the same issue, but: Textbooks often bury good ideas in dense notation, skip the intuition, assume you already know half the material, and get outdated in fast-moving fields like AI.…
-
Hey HN, Henry & Roman here, we are building Cactus (https://cactuscompute.com/), an AI inference engine specifically designed for phones. We're seeing a major push towards on-device AI, and for good reason: on-device AI…
- Show HN: Cactus – Ollama for Smartphones (github.com)
Hey HN, Henry and Roman here - we've been building a cross-platform framework for deploying LLMs, VLMs, Embedding Models and TTS models locally on smartphones. Ollama enables deploying LLMs models locally on laptops and…