ai productivity lab numbers

Two AI Labs Published How Much AI Does Their Work. You Still Can’t Compare Them.

TL;DR: Anthropic has published internal numbers saying Claude now “leads” 26% of its AI research work, up from under 1% in February. OpenAI published its own a fortnight earlier, including that spending on coding agents reached roughly $600 per researcher per day. Both are careful, both are interesting, and — by Anthropic’s own admission — neither can be compared with the other, because there’s no shared methodology. If you want to know whether AI productivity claims hold up, this is the best evidence anyone has published — and it still doesn’t answer your question.

What Anthropic measured

On 17 September, Anthropic published a set of measurements putting a number on AI productivity inside its own walls — specifically, how much of its AI research is done by AI. The headline: Claude “leads” 26% of measured AI R&D work, against under 1% in February.

“Leads” has a precise meaning here, and it isn’t “does it alone”. Per Engadget, it means the AI “can complete most of [a] task end-to-end from a high-level prompt, while [a] human supervises.” Anthropic is explicit that Claude is not operating fully autonomously in any measured category — it tops out one rung below that on a scale from Epoch AI running from AL0 (no AI) to AL5 (no human in the loop).

More than 90% of its R&D reaches at least the “AI collaborates” level.

The method is worth knowing, because it shapes what the number means. Each week of July 2026, Anthropic randomly sampled 20% of staff in departments doing model R&D, producing roughly 15,000 individual tasks, which were then sorted into 542 categories. The sorting was done by Claude.

Two other figures from the same report: around 30,000 active research agents, with online monitors covering 100% of them and blocking about 0.002% of decisions — roughly one in 47,000. And 6% of AI R&D compute goes to safety work.

What OpenAI measured

OpenAI got there first, publishing “Research acceleration: The view inside OpenAI” on 6 September.

Its most quotable AI productivity figure is about money rather than percentages. Reading OpenAI’s own chart, Simon Willison records daily spend on coding agents per researcher rising from about $50 in April to $150 by June, then climbing steeply to roughly $600 by late August 2026.

The report also states that by mid-August, OpenAI was recording around 3.1 agent workdays for every human workday across its research organisation — and, to its credit, immediately warns against the obvious misreading: this does not mean a researcher became 3.1 times more productive, and agent runtime is not the same as useful output.

(A note on sourcing: `openai.com` blocks our fetcher, so OpenAI’s figures here come from Willison’s reading of the published chart and from secondary reporting rather than a direct read. The Anthropic figures are from Anthropic’s own document.)

The part that matters: these numbers don’t connect

Here’s what makes this more than two press releases.

Anthropic addresses cross-lab comparison directly, and rules it out: “Two obstacles stand in the way of cross-lab comparison on this type of reporting. First is the lack of a common methodology.” They propose third-party verification as the fix, and suggest other labs could replicate their approach.

So: one lab counts task categories and reports a percentage led by AI. The other counts agent workdays and dollars. There is no conversion between them. You cannot say which lab is more automated, whether 26% is high or low, or whether $600 a day is efficient — because there is nothing to measure it against.

This is the same problem we set out when an independent index finally scored the big models on one yardstick: a number without a shared method isn’t a measurement, it’s a claim with a decimal point.

Why these AI productivity numbers matter if you don’t code

The question underneath all of this is the one our readers actually have: does AI genuinely make you faster, or does it just feel like it?

These two reports are the best evidence anyone has published, because they come from the organisations with the most capable tools, the deepest expertise and every incentive to measure honestly for their own planning. And the honest reading is uncomfortable in three ways.

First, the conditions aren’t yours. The median OpenAI researcher was burning around $600 a day on inference. You are deciding between £20 and £200 a month. When a lab reports that agents accelerated its work, the setup producing that result costs more per day than your annual subscription. That doesn’t make the finding false — it makes it non-transferable.

Second, “leads” still means supervised. Anthropic’s 26% describes work where a human watches. Not one measured category reached full autonomy. The most AI-saturated research organisation in the world still has a person on every task — which is worth remembering when a tool suggests you needn’t be.

Third, even they can only estimate. Anthropic sampled a fifth of staff for one month and had Claude categorise the results. That’s a reasonable method, and it is still an estimate of something genuinely hard to measure. If measuring AI productivity is this difficult with full access to your own organisation, your own impression of how much time a tool saved you is not data.

What to do with AI productivity numbers

  1. Treat them as direction, not magnitude. “AI does more of our research than it did in February” is credible. “26%” is a specific answer to a question nobody else asks the same way.
  2. Check the definitions before the figure. Anthropic’s “leads” sounds like autonomy and explicitly isn’t. Definitions do more work than percentages in every report like this.
  3. Ask what it cost. Acceleration claims stated without spend are half a sentence. OpenAI published its spend, which is the most useful thing in either report.
  4. Measure your own, roughly. Time one task you do often, then time it again with the tool. It’s crude, it’s yours, and it beats any published figure for answering your question.
  5. Discount any claim with no method attached. If a vendor tells you its tool makes people 40% faster and won’t say how that was measured, you now know what a serious attempt at measuring looks like — and how carefully the serious ones hedge.

Who should care about these AI productivity figures (and who shouldn’t)

  • Deciding whether to pay more for a better tier: the cost angle is the useful part. Our piece on what Claude Max’s usage claims actually describe is the companion read.
  • Trying to justify AI spend to someone else: these are the most credible public numbers available — and you should present them with the same caveats the labs did, or you’ll be caught out.
  • Wondering if you’re using AI “wrong”: the labs need one human per task and spend hundreds a day. Whatever you’re doing, you are not the reason it isn’t magic.
  • Just building something: skip the numbers entirely. Step 4 answers your question better than either report.
  • Following AI safety: Anthropic’s oversight and compute-allocation figures are the more novel part, and the call for third-party verification is the substantive proposal.

Our take

We’d rather have these reports than not, and we want to be clear about why before criticising them.

Both labs published things that don’t flatter them. OpenAI warned readers off its own headline ratio. Anthropic stated that its number can’t be compared with anyone else’s and asked to be independently verified. That is close to the opposite of marketing, and after a month in which we’ve covered a benchmark from an interested party and a lawsuit over usage claims, it’s worth saying when disclosure is done well.

The limitation is structural rather than a failure of nerve. An AI productivity self-measurement that nobody can reproduce is interesting, not authoritative — and both companies know it, which is why both are asking for shared methods and outside checks. Until that exists, the correct posture toward any AI productivity figure, theirs included, is curiosity rather than belief.

What we take from it for our readers is smaller and more useful than any percentage: the people best equipped to answer “does AI make you faster” have spent real effort on it and produced numbers they themselves won’t over-claim. If they’re that careful, the confident 10x claim in a tool’s marketing email deserves considerably less of your trust than your own stopwatch.

Not sure which tool is worth paying for in the first place? Take the 60-second Vibe Coding Tool Finder quiz →

FAQ

Does AI actually make developers more productive?

The most credible public evidence comes from two frontier labs measuring themselves. Anthropic says Claude “leads” 26% of its AI R&D work with a human supervising; OpenAI reports roughly 3.1 agent workdays per human workday while warning that this is not a 3.1x productivity gain. Neither lab claims a clean productivity multiplier, and their methods aren’t comparable.

What does Anthropic mean by Claude “leading” 26% of its R&D?

That Claude can complete most of a task end-to-end from a high-level prompt while a human supervises. It is one level below full autonomy on a scale from Epoch AI, and Anthropic states Claude is not fully autonomous in any measured category of its AI R&D work.

How much do AI labs spend on coding agents?

OpenAI published a chart showing daily spend per researcher rising from around $50 in April 2026 to roughly $600 by late August, at API prices. That is per researcher, per day — useful context when comparing published AI productivity gains against a personal subscription costing a fraction of that per month.

Similar Posts