AnalysisAI ModelsSeptember 14, 2026

Claude Fable 5.1 tops Real-SWE coding benchmark at 38.8%

Real-SWE, a benchmark from Y Combinator-backed Specific Labs, tests coding agents on problems not pulled from public sources; the winning agent still failed more than six out of 10 tasks. Its 38.8% score was the best result.

1 source

More stories today

Open the live feed