GitHub radar

Free local model for instant decisions instead of paid Jev

Clef-Flash is a model from Cloudflare that instantly returns a probability for each answer option instead of writing text you'd have to parse yourself.

01Cloudflare/clef-flash 3836.4k downloads/mo9.4B paramsimage-text-to-text

It's a model that takes a situation — text, data, a photo, or video — plus a list of questions with answer options, and immediately returns how likely each option is, without extra text.

When it helps

Useful if you're paying a service like Jev for things like routing messages or tagging documents and want to do that yourself for free; for a normal chatbot, use a different model.

Pros

  • Responds in under a tenth of a second — far faster than Jev.
  • More accurate than Jev on everyday scenarios: 97.7% vs 52.3%.
  • Gives a ready probability number, no need to parse text.
  • Free and open under the Apache-2.0 license.

Cons

  • On real business cases like invoices, Jev is still more accurate (83.1% vs 73.3%).
  • Full functionality needs the author's Python code, not just a ready-made chat.
  • Without a GPU, the probability trick runs slowly or not fully.

How to set it up — step by step

  1. 1Open your agent (Claude Code, Codex, etc.) in your project folder.
  2. 2Send the agent the text from the block below.
  3. 3Wait for setup: you need at least 16 GB of RAM, and a GPU for full functionality.
  4. 4Give the agent your scenario and your list of questions with answer options.
  5. 5Check that the agent returned a probability number for each question, not plain text.
  6. 6Before trusting the model with important decisions, compare its answers to your own data or to Jev.

Text for your agent

Copy this and send it to your agent — Claude Code, Codex, any of them:

Set up Cloudflare's Clef-Flash model for me following its Hugging Face card: https://huggingface.co/Cloudflare/clef-flash — install it using the model's own Python code (not just Ollama), run my text and my questions through it, and show me the number it returns for each question.

The model has 9 billion parameters, and a ready quantized version for Ollama exists — a regular laptop with 16 GB of RAM or more can run it. But the real trick, the instant classifier that returns a number for every question, needs the author's own Python code, and it runs faster with a graphics card — Cloudflare's own tests used a data-center H200.

Open on Hugging Face