AppleAnalysisAI ModelsAugust 7, 2026

Apple's DeepAmbigQA benchmark tests LLMs on ambiguous multi-hop questions

The dataset contains 3,600 multi-hop questions, half requiring explicit name ambiguity resolution; even state-of-the-art GPT-5 scores just 0.13 exact match on ambiguous questions and 0.21 on non-ambiguous ones. Questions are generated by the DEEPAMBIGQAGEN pipeline from text corpora and linked knowledge graphs.

2 sources

More stories today

Open the live feed