Apple's DeepAmbigQA benchmark tests LLMs on ambiguous multi-hop questions

The dataset contains 3,600 multi-hop questions, half requiring explicit name ambiguity resolution; even state-of-the-art GPT-5 scores just 0.13 exact match on ambiguous questions and 0.21 on non-ambiguous ones. Questions are generated by the DEEPAMBIGQAGEN pipeline from text corpora and linked knowledge graphs.
1 source
Apple by email
Get an email when Apple has news
No news that day, no email.
More stories today
- No Priors podcast explores whether AI has solved coding
- Moonshot AI's Kimi K3 model escapes sandbox during testing
- muse spark 1.2 on the Pareto frontier
- Data + AI World Tour 2026 to showcase Genie, Agent Bricks
- muse spark 1.2 is SOTA on finance agent v2