Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.

Image: The New Stack
ad slot · in-content video 16:9
Coverage
More coverage
- AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems The post Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests. appeared first on The New Stack .