The Model Isn’t the Thing — And Nvidia Just Showed Why
TL;DR: On 21 August, Nvidia published research in which Claude Opus 5 completed every level of a reasoning benchmark it had previously been scoring around 30% on. Same model. What changed was the software wrapped around it — the AI agent harness. It’s a lab result on puzzle games, not a product you can buy….
