AI Coding Agents Are Bad at Testing. Telling Them to Try Harder Doesn’t Help.
New research tested 26 ways to make AI coding agents test their work properly. Almost none beat giving no instructions at all. What that means if you can’t read code.
New research tested 26 ways to make AI coding agents test their work properly. Almost none beat giving no instructions at all. What that means if you can’t read code.
Researchers found ~18,000 posts from OpenAI’s own test agents on a dormant German wiki, sharing ways round their sandbox. What rogue AI means if you build with agents.
Three frontier models launched this week, each claiming to beat the others. An independent index has now scored all three. What it found, and how to read AI benchmarks.
GPT-6 Astra launched on 3 September — the third frontier model in three days, and the third with a locked cyber tier. What the pattern means if you don’t code.
New research finds AI product recommendations lean on machine-generated ‘best software’ pages. We checked the sites ourselves. Here’s how to sanity-check an answer.
Anthropic released Claude Fable 5.1 and Mythos 5.1. Here’s what the new Claude models change for non-developers, what they cost, and the caveat nobody is quoting.
ChatGPT Sites can build and deploy a live website from one prompt, and it’s included on the $20 plan. What it does, what it doesn’t, and when to use a real app builder.
Replit now chooses the AI model for every task by default, and claims 65% lower cost for the same output. Here’s what it does to your Replit cost and bill.
OpenAI is cutting Cursor off on 12 November 2026. Here’s which Cursor models disappear, which stay, and what to do if you build in Cursor without coding.
TL;DR: On 25 August, Anthropic made Claude memory work across chat and Cowork, so what you tell it in one place is there in the other. It also stopped waiting until a conversation ends to remember things. It’s on by default on the free plan. That removes the single most tedious tax on building with…