
Claude leads 26% of AI R&D at Anthropic. Who checks it?
Anthropic says Claude leads a quarter of its AI research work under human supervision. I look at what the figure measures and who still checks the results.
Every article tagged model-evaluation, newest first.

Anthropic says Claude leads a quarter of its AI research work under human supervision. I look at what the figure measures and who still checks the results.

DeepSWE puts Luna Max 2.2 points behind Sol High at roughly one-sixth the attempt cost. I explain why the models can still feel far apart in repository work.