Ai2's BenchMIRT audits what LLM benchmarks actually measure
BenchMIRT analyzes benchmarks question by question, revealing which capabilities they test. On BBQ, a social-bias eval, it found questions distinguished models more by reasoning ability than safety.
2 sources
AI Models by email
Get an email when there's news on AI Models
No news that day, no email.
More stories today
- David Lowery, Jason Isbell sue Suno over likeness rights
- OpenAI to launch next model soon, Altman says
- Koray Kavukcuoglu discusses AGI path and Gemini 3.7 Flash in podcast
- Palo Alto CEO: $1T of cybersecurity infrastructure isn't ready for AI
- Perplexity CEO teases 'Private, Personal, Powerful AI'