AnalysisAI ModelsAugust 7, 2026

New papers probe LLM-as-a-judge reliability and benchmark validity

Read original source →arxiv.org

Multiple arXiv papers examine LLM-as-a-judge evaluation, including JudgeArena, a unified framework for reproducible LLM-judge evaluation, and studies on scoring bias, frozen interfaces, and temporal leakage in backtesting. One paper finds standard contamination checks are uninformative, as four flagship models fail them on questions they cannot have memorized.

How this story unfolded

2 weeks · 5 reports · from Aug 7

  1. Aug 7
  2. Aug 24

More stories today

Open the live feed