AppleAnalysisAI ModelsAugust 6, 2026

Apple's DeepAmbigQA benchmark tests LLMs on ambiguous multi-hop questions

The dataset contains 3,600 multi-hop questions, half requiring explicit name ambiguity resolution; even state-of-the-art GPT-5 scores just 0.13 exact match on ambiguous questions and 0.21 on non-ambiguous ones. Questions are generated by the DEEPAMBIGQAGEN pipeline from text corpora and linked knowledge graphs.

1 source

Apple by email

Get an email when Apple has news

No news that day, no email.

More stories today

Open the live feed