GitHub radar

$0.0007/task for evaluating AI reasoning skills with ARC Task Gen

Pathway’s ARC Task Gen creates new ARC-AGI-1-like reasoning tasks, alongside a paper demonstrating that a 150M parameter model outperforms its much larger competitors at 11x the efficiency.

01pathwaycom/arc-task-gen 3.8kPython

Pathway’s ARC Task Gen creates new ARC-AGI-1-like reasoning tasks. It is distribution matched to the publicly available evaluation data. This repo goes along with the paper on our new model architecture BDH-CQ, which incorporates iterative reasoning and in-context learning through evolving recurrent memory in a continuous latent space. We find that a 150M parameter BDH-CQ model achieves 29.5% pass@2 on the ARC-AGI-1 evaluation set at $0.0007/task, which is 11x more efficient than GPT-5.6 Luna Low. The generated tasks are in the same ARC JSON format as the originals, so they are compatible with existing evaluation frameworks.

Why a vibe-coder should care

This paper demonstrates that a model 10–100x smaller than state of the art can perform equally well on structured reasoning tasks while being much more efficient. If you have AI agents or need to evaluate models, this is a glimpse into the future of efficient reasoning and how to do so efficiently!

How to install

Copy this and send it to your agent — Claude Code, Codex, any of them:

Set up https://github.com/pathwaycom/arc-task-gen (pixi install), ask me for an API key, and generate a tasks.json set per instructions.md.

Runs on a regular laptop — only Python needed.

Open on GitHub