Berkeley: Agent Frameworks Drive Cost, Not Quality
UC Berkeley researchers tested 21 model–framework pairs: success rates varied by just ±2%, while costs differed by up to 5×. Claude Code ran twice as expensive as the minimalist Pi framework with only a 1.1% difference in results.
UC Berkeley researchers tested 21 model–framework pairs on SWE-bench Lite and Terminal-Bench 2.0: seven models, three frameworks — Claude Code, Codex CLI, and Pi. Success rates spread within ±2%, while costs for the same results differed by up to 5×. On Fable 5.1, Claude Code ran twice as expensive as the minimalist Pi — with a success rate gap of just 1.1%.
Pi operates with four tools — read, write, edit, bash — and still reaches the Pareto frontier on both benchmarks. The cost difference comes down to context: Claude Code pumps 10× more tokens into the start of a session than Pi does. For isolated tasks with a clear spec, the simpler framework is genuine savings; for long sessions with deep project context, the extra spend is justified.
Source: harnesstax.github.io
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →▌ Related guides

Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles