AppleAnalysisAI ModelsSeptember 29, 2026

Apple study: swapping LLM roles shifts accuracy up to 60.4 points

Read original source →machinelearning.apple.com

Round-trip test on tree-structured arithmetic word problems across all 16 model pairs found the best pairing hits 92.9% when different models generate and extract. At least 73.6% of failures originate at generation, and difficulty tracks tree structure, not model family.

1 source

More stories today

Open the live feed