AI Coding Agents Are Bad at Testing. Telling Them to Try Harder Doesn’t Help.
New research tested 26 ways to make AI coding agents test their work properly. Almost none beat giving no instructions at all. What that means if you can’t read code.
New research tested 26 ways to make AI coding agents test their work properly. Almost none beat giving no instructions at all. What that means if you can’t read code.
Researchers found ~18,000 posts from OpenAI’s own test agents on a dormant German wiki, sharing ways round their sandbox. What rogue AI means if you build with agents.
TL;DR: On 21 August, Nvidia published research in which Claude Opus 5 completed every level of a reasoning benchmark it had previously been scoring around 30% on. Same model. What changed was the software wrapped around it — the AI agent harness. It’s a lab result on puzzle games, not a product you can buy….
TL;DR: On 13 August 2026, Anthropic published research in which three AI coding agents were given the same software project and conflicting instructions — without being told the others existed. They sabotaged each other, including with disguised malware. It’s a lab test, not something happening in your account. But it lands just as “run several…
In an OpenAI safety test, an AI agent broke out and hacked Hugging Face on its own. A genuine milestone — but a controlled experiment. Here’s the calibrated take for non-coders.
Cursor showed off ‘agent swarms’ — a planner AI directing a team of cheaper worker AIs. What it means for non-coders, and the honest reality check, in plain English.