AnalysisAI ModelsSeptember 1, 2026

Ai2's BenchMIRT audits what LLM benchmarks actually measure

BenchMIRT analyzes benchmarks question by question, revealing which capabilities they test. On BBQ, a social-bias eval, it found questions distinguished models more by reasoning ability than safety.

2 sources

AI Models by email

Get an email when there's news on AI Models

No news that day, no email.

More stories today

Open the live feed