GitHub radar

Turn speech into text on your device, no internet needed

Whistle is a tiny speech-to-text model under 17 MB that turns spoken audio into text right on your device, no server or GPU required.

01Cactus-Compute/whistle 2105.6k downloads/moautomatic-speech-recognition

It's an AI model that listens to audio and outputs text with the exact timing of each word.

When it helps

Useful for building a voice assistant, dictation app, or smart device in English and 6 other languages; skip it if you need Russian.

Pros

  • Just 16.9 MB — a single file
  • Runs offline, no server or GPU needed
  • Gives timestamps for every word
  • Free Apache 2.0 license for any use

Cons

  • Doesn't support Russian
  • Accuracy isn't backed by numbers, only charts

How to set it up — step by step

  1. 1Open your AI agent (Claude Code or Codex) inside your project folder.
  2. 2Send the agent the text from the block below.
  3. 3Let the agent install the cactus-needle package and fetch the model from huggingface.co/Cactus-Compute/whistle.
  4. 4Check that the agent outputs the transcribed text and word-level timestamps.
  5. 5Test your own English, German, French, Spanish, Italian, Dutch, or Polish audio.

Text for your agent

Copy this and send it to your agent — Claude Code, Codex, any of them:

Set up the Whistle speech-to-text model from Cactus-Compute: https://huggingface.co/Cactus-Compute/whistle — install the package, transcribe a test English voice recording, and show me the text with word-level timestamps.

Any regular laptop or even a phone — no graphics card needed, the whole model is under 17 MB.

Open on Hugging Face