AnalysisAI ModelsOctober 1, 2026

Arena study finds LLM judges favor their own answers

Read original source →arena.ai

Across 12 LLMs, judges picked their own answer 58% of the time versus 34% for humans — about 70% more self-preference. Astra backed itself in 88% of battles and Fable in 72%; OpenAI-family judges rated OpenAI models 37 points above human ratings.

1 source

More stories today

Open the live feed