Nvidia reduces coding agent tokens by almost half with SoL-Pi
Nvidia’s SoL-Pi reduces coding agent token consumption by almost half by improving the harness surrounding the model instead of the model itself.
Nvidia has released SoL-Pi, a system which reduces coding agent token consumption by almost half by improving the harness surrounding the model instead of the model itself. SoL-Pi does not modify the underlying model, but rather modifies the harness, or set of instructions and policies around the model. They evaluated 152 configurations on nearly three thousand trials and found the optimal configuration reduced token consumption by 49% while maintaining 93.7% of the original quality. When applied to the EdgeBench benchmark, SoL-Pi achieved a 54.3% reduction in token consumption over Claude Code.
These reductions seem to be valid and didn’t come at too much of a cost to answer quality. There was also less backtracking and unnecessary steps taken by the agent. Keep in mind this is a research project from Nvidia and SoL-Pi can’t be used to improve Claude Code right now. These benefits aren’t as pronounced for other tasks such as math olympiad questions, but I appreciate the approach of modifying the harness surrounding the model rather than the model itself because in my experience there is a lot of bloat in the context the agent is carrying around that doesn’t do much.
Source: the-decoder.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →▌ Related guides

Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles