GitHub radar

Reasoning Classifier Beats Kev And Jev In Published Tests

PostHog released a reasoning version of a Jev-style classifier. The model thinks briefly before answering, and by its own tests it is more accurate than Kev-9B and Jev. A serious graphics card or a Mac with a large amount of memory is needed.

01PostHog/jeeves 391Python

This is a variant of the “small model doesn’t generate full response, but instead outputs probabilities for some predefined questions” approach we’ve seen before. The questions here are things like “which department does this request belong to?”, “should I escalate this issue?”, “how frustrated is the customer?”. But the twist with Jeeves is that it first “thinks” about its response, according to the authors, taking ~0.3 sec without reasoning vs. a median of 3.3 sec on a single H100 GPU with reasoning. The creators report that Jeeves outperforms both Kev-9B and Jev on unseen data (0.889 vs. 0.822 and 0.857) and on open tiers of JevBench (0.935 vs. 0.866 for Jev). Note that the Kev and Jev scores in the table above are self-reported, so take them with a grain of salt. This model was created by PostHog, a company in the product analytics space, who were inspired by the open source Kev project.

Why a vibe-coder should care

If you want high accuracy on edge cases and don’t mind using a GPU (or Mac with lots of RAM), then this might be worth checking out — according to the authors, it’s significantly better than previous models in terms of accuracy. But it won’t work on your laptop unless you have a ton of RAM (weights require 21 GB of memory) and access to a GPU that supports FP8 (fast mode requires a GPU). Also, keep in mind that it’s just a classifier, so it will only answer the predefined questions you ask about a given text — it won’t actually generate code or perform other tasks.

How to install

Copy this and send it to your agent — Claude Code, Codex, any of them:

Deploy the model from this repository: https://github.com/PostHog/jeeves — on my server with a graphics card, or on a Mac with 48+ GB of memory if I have one, connect it in place of a regular ticket classifier, and show me the answer on a test text.

A CUDA graphics card or a Mac with 48 gigabytes of memory or more is needed. The weights alone in bf16 weigh 21 gigabytes, it will not run on a regular laptop.

Open on GitHub