Eigen RadarAI
Analysis

Whistle brings seven-language speech recognition onto the device

Cactus Compute released Whistle, a speech-recognition model that runs locally on a CPU and shares the Needle engine. It transcribes short audio clips in seven languages, with automatic language detection and word timestamps. The same engine can combine speech recognition with tool calls. Turkish is outside the supported language set, which includes English, German, French, Spanish, Italian, Dutch and Polish.

Artificial Intelligence··Night
An older bearded man in a beige cardigan speaks toward his laptop at a sunlit home desk.

Whistle runs speech recognition on the local processor

Cactus Compute released Whistle on October 2 as an open speech-recognition model designed to run locally on a central processor. It loads into Needle, the company’s existing engine for on-device language models. Speech input joins the same local software system used for text-based tasks.[1], [2]

Seven languages share automatic detection and word timings

Short, single-channel audio clips are accepted by Whistle, which supports English, German, French, Spanish, Italian, Dutch and Polish. Turkish is outside the supported set. It detects the language unless the application specifies one. The output includes the transcript and each word’s start, end and probability. Its encoder can also produce a numerical representation of speech without decoding the words. The browser demonstration downloads the model at first use and says audio stays on the device.[1]

Speech and tool calls can use the same engine

An example combines Whistle with Needle to turn a spoken clip into text and then structured tool calls. Both use the same C++ engine, container format and quantization, which represents model weights with fewer bits. The speech interface exposes functions for loading, transcription and speech embeddings. The engine and its platform folders are distributed through Cactus Compute’s code repository, while model weights are distributed through Hugging Face.[1]

Applications can supply keywords to raise the probability of selected phrases during decoding. A separate option changes decoder depth. Audio key and value projections are computed once and retained as the decoder produces the transcript, rather than recalculated for every step.[1]

References

  1. News sourceCactus ComputeWhistle transcribes seven languages with a 16.9 MB on-device model↩1↩2↩3↩4
  2. News sourceRuntimeWireCactus Compute releases Whistle for local speech recognition↩