AnalysisAI ModelsSeptember 1, 2026

Ai2's BenchMIRT audits what LLM benchmarks actually measure

Read original source →huggingface.co

BenchMIRT analyzes benchmarks question by question, finding BBQ's questions distinguish models more by reasoning ability than safety. It extends Item Response Theory to separate signals within a benchmark.

2 sources

More stories today

Open the live feed