⏱️ Reading time: 9 min
Cactus compressed a complete speech recognizer into 16.9 MB and got it running in 11 milliseconds on a regular CPU, with no GPU and without touching external servers.
📑 En este artículo
On October 2, 2026, the company released Cactus’s Whistle model, which transcribes seven languages and runs on the same C++ engine that already powers Needle, its text model for low-memory devices.
TL;DR
- Cactus released a 16.9 MB voice model on October 2, 2026, that runs on CPU with no dependencies.
- It transcribes seven languages (English, German, French, Spanish, Italian, Dutch, and Polish) in clips of up to 30 seconds.
- On an Apple M4 Pro, it reaches the first token in 11.1 ms, compared to 73.2 ms for Whisper base.
- It decodes at 1,319 tokens per second, roughly five times faster than Whisper base and Moonshine tiny v2.
- It shares the C++ engine and Simple Attention blocks with Needle, Cactus’s text model.
What Is Cactus’s Whistle Model
Whistle is the open source speech recognition model that Cactus released on October 2, 2026, for memory-constrained devices: phones, wearables, robots, cars, and microcontrollers. It weighs 16.9 MB, runs on CPU alone with no external dependencies, and transcribes 16 kHz audio in seven languages in a single pass.
The model was developed by Cactus, a startup that builds inference engines for hardware with limited memory and battery. Whistle isn’t an isolated project: it uses the same quantization, the same binary container, and the same attention blocks as Needle, the text model the company released earlier. A single binary can load both models and go from an audio clip to a tool call without ever touching a server.
What happened
The release came on October 2, 2026, signed by Jakub Mroz and Henry Ndubuaku, Cactus’s cofounders. The announcement describes three functions that run entirely on-device: single-pass transcription of 16 kHz audio up to 30 seconds long, per-word timestamps (start, end, and probability, calculated from the decoder’s attention), and a less common third output, the encoder embedding: one row for every 80 milliseconds of audio, without going through the decoder or generating a transcript.
Language detection is automatic unless one of the seven supported languages is specified: English, German, French, Spanish, Italian, Dutch, and Polish. The engine also measures the clip’s dynamic range before starting the decoder. If the audio falls below a silence threshold, it returns an empty transcript without even starting the beam search, which saves compute in the most common case for an always-on assistant: silence.
Cactus also published an in-browser sandbox where the 16.9 MB speech recognizer runs locally. The first use downloads the model, and from then on the audio never leaves the device. The demo accepts a hand-typed list of keywords, and the engine prioritizes them while decoding.
Context and history
Edge speech recognition comes from two distinct fronts. On one side, large models like Whisper, from OpenAI, which prioritize accuracy and multilingual support at the cost of size: the base variant weighs 145.3 MB. On the other, lightweight projects like Moonshine, from Useful Sensors, designed for microcontrollers but limited to English in its 41.9 MB tiny v2 version.
Whistle falls into a third category: multilingual and smaller than both, built on Cactus’s earlier bet with Needle. The decision to reuse Needle’s engine, quantization, and attention blocks, rather than building an audio engine from scratch, explains much of the size. Cactus’s transcription system doesn’t load its own runtime, just a different set of weights on top of the same C++ machinery.
That lineage also sets the ceiling on what Whistle can do today: the language list depends on the 8,192-piece text vocabulary the shared engine already uses, not on a tokenizer designed from scratch for audio.
Technical details
The front end converts audio into an 80-band log-mel spectrogram, with 25 ms windows and 10 ms hops, limited to the 250-3,500 Hz band and normalized per channel. Thirty seconds of audio end up as 3,000 frames. A convolutional stem with 128 channels and a kernel of 9 halves that number three times in a row, down to 375 frames, one every 80 ms: that’s the rate everything downstream operates at.
flowchart TD
A["Audio 16kHz mono"] --> B["Log-mel: 80 bands"]
B --> C["Conv stem: 375 frames"]
C --> D["Encoder: 8 Simple Attention blocks"]
D --> E["Gated cross-attention"]
E --> F["Decoder: 8 Laddered Attention blocks"]
F --> G["Beam search x5"]
G --> H["Final transcript"]
The encoder is eight Simple Attention blocks, the same ones Needle uses: four mHC residual lanes and a Monarch Hadamard MLP instead of the traditional feed-forward layer. Attention there isn’t causal, so a frame at 3 seconds can look directly at a frame at 12 seconds in the same clip.
The decoder is eight Laddered Simple Attention blocks with a width of 512, with 8 query heads against 2 key/value heads (GQA), dimensions of 48 for query and key and 64 for value, a 3-position causal convolution over Q, K, and V, and an engram memory in layers 3 and 7 with 18,432 positions. What’s different from Needle is a single per-layer addition: a gated cross-attention that reads the encoder.
x ← x + σ(g) · softmax(q̂ Kᵗ / √d) V
The gate g is learned at each layer, and K and V are taken from the audio clip. Those encoder projections are computed once per clip (375 frames across 8 layers) and stay fixed throughout decoding, so five search beams cost five short text caches, not five passes over the audio.
Decoding evaluates five beams with length-normalized log probability. An Aho-Corasick automaton biases the search toward a list of keywords the user can pass along with the audio. The vocabulary has 8,192 text pieces plus seven language tokens, one per supported language, and the transcript is capped at 320 pieces.
💭 Key point: Whistle’s decoder is laddered: every depth from 2 layers up was trained as an independent model, and the --audio-depth flag picks which one to load at startup. The encoder, by contrast, always runs its full 8 blocks, no exceptions.
Cactus didn’t publish a command or a call that independently confirms which depth of Cactus’s Whistle model was active in a given run. There’s no direct way to verify this from outside the process: the only documented signal is the one the engine itself prints when it loads the model.
Impact and analysis
Cactus compared the three models on their default official runtimes, on an Apple M4 Pro with ten seconds of audio: Whistle’s C++ engine with 5 beams, openai-whisper for Whisper base, and moonshine-voice in non-streaming mode over the full clip.
| Model | Size | First token | Decoding | Where it wins on accuracy (WER) |
|---|---|---|---|---|
| Whistle (Cactus) | 16.9 MB | 11.1 ms | 1,319 tokens/s | LibriSpeech, SPGISpeech, Earnings-22, FLEURS average |
| Whisper base (OpenAI) | 145.3 MB | 73.2 ms | 266 tokens/s | TED-LIUM, AMI, MLS average |
| Moonshine tiny v2 (Useful Sensors) | 41.9 MB | 22.8 ms | 262 tokens/s | English only; doesn’t report the above benchmarks |
The honest read on those numbers is that Cactus’s Whistle model doesn’t win across the board: Whisper base still leads on TED-LIUM, AMI, and the MLS average, three sets with meeting and conference recordings that are noisier and have more than one speaker than LibriSpeech’s audiobooks. Cactus built Whistle for the case where size and latency matter more than those last points of accuracy: a voice assistant on a wearable doesn’t have 145 MB of headroom or 73 ms to spend before the first token.
The speed comparison isn’t neutral either: Whisper base and Moonshine ran on their reference Python runtimes, while Whistle ran on Cactus’s C++ engine. Part of the 1,319-tokens-per-second edge over 266 comes from the model, but part of it comes from comparing a native binary against two interpreted runtimes, something the announcement itself doesn’t hide but also doesn’t isolate with a separate control.
What’s next
Cactus didn’t give a date for adding more languages or for a version of Whistle with more encoder layers, which today runs fixed at 8 blocks with no trimming option. It also didn’t confirm whether it will release weights for fine-tuning on custom vocabularies, something it does offer for Needle according to the same announcement. What is laid out is the pattern: a C++ engine and a quantization scheme, two models that share memory and boot from the same binary, ready for an agent to go from listening to reasoning without a trip to a server.
Try it yourself: open the sandbox on Cactus’s blog, hit the microphone, and see how long it takes for the first word to appear on screen.
Frequently Asked Questions
What languages does Whistle, Cactus’s voice model, recognize?
Seven: English, German, French, Spanish, Italian, Dutch, and Polish, in clips of up to 30 seconds, with automatic detection if none is specified.
How does Whistle’s size compare to Whisper base?
Whistle weighs 16.9 MB compared to Whisper base’s 145.3 MB, about an eighth of the size, according to the benchmarks Cactus published on an Apple M4 Pro.
What does the –audio-depth flag do in Cactus’s voice engine?
It picks how many decoder layers to load: every depth from 2 layers up was trained as an independent model. The encoder always runs its full 8 blocks.
Does Whistle work without an internet connection?
Yes. The model runs entirely on CPU and on-device; in the browser sandbox, after the first 16.9 MB download, the audio never leaves the machine.
How fast does the 16.9 MB speech recognizer Cactus released decode?
1,319 tokens per second on an Apple M4 Pro, compared to 266 tokens/s for Whisper base and 262 tokens/s for Moonshine tiny v2, according to Cactus’s tests.
References
- Cactus: official Whistle announcement: full architecture, benchmarks against Whisper base and Moonshine tiny v2, and an in-browser sandbox.
- Cactus Compute: official site of the company behind Whistle and Needle.
- OpenAI: official Whisper repository: code and weights of the model used as the comparison baseline.
- Useful Sensors: official Moonshine repository: code for the lightweight voice model used as the second baseline.
📱 Do you like this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.
Featured image: Foto de Daniel Sandvik en Unsplash
Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.
Leave a comment
0 Comments