GitHub radar
A local model that turns a picture and a question into a probability
JEV-27B-VL is a model that takes text, an image, and a closed question, and instantly returns a probability for every answer — yes/no, a 0–5 score, or a pick from a list — with no reasoning, no extra words.
It's a decision model (a closed-question answering model): given text and images, it answers in one pass — yes/no, a 0–5 score, or a pick from a list — and returns a probability for each answer, not a paragraph.
When it helps
Useful as one small step inside your own project: have an agent decide whether to click a button from a screenshot, check an image for moderation, or rank short videos by their cover. Skip it if you want a conversation partner or long-form answers — it's not a chatbot.
Pros
- Ranks #1 of 20 vision decision models on a public leaderboard
- Beats a vision-language model about 15 times its size
- Returns a calibrated probability, not text you have to parse
- Runs on a 24GB GPU or a Mac with 32GB+ of memory
Cons
- No ready-made Ollama package — only vLLM or transformers
- Most numbers on its page are the author's own tests, not independent
How to set it up — step by step
- 1Open the model page from the link in the facts
- 2Check you have a Mac with 32GB+ memory or a 24GB GPU
- 3Open your AI agent (Claude Code or similar) in your project folder
- 4Send the agent the text from the block below
- 5Let the agent set up the run through vLLM or transformers
- 6Test it on your own image with a yes/no question and check the probability it returns
Text for your agent
Copy this and send it to your agent — Claude Code, Codex, any of them:
Figure out how to run AutoTrust's JEV-27B-VL model locally from https://huggingface.co/autotrust/JEV-27B-VL via vLLM or transformers, give it an image with a yes/no question, and show me the probability it returns.
You need a Mac with 32 GB of memory or more, or a GPU with 24 GB of VRAM — the model has about 28 billion parameters. There is no ready-made Ollama build; you would run it through vLLM or transformers (model-serving tools).
Open on Hugging Face▌ More finds