AppleAnalysisAI ModelsOctober 1, 2026

Apple's RLTL;DR lifts Qwen 3.5 9B Pass@1 from 0% to 12-13%

Read original source →machinelearning.apple.com

Apple researchers introduce RLTL;DR: after each failed attempt the policy writes a one-line TL;DR insight from verifier output, and the next rollout is conditioned on all prior insights. On tool-calling and coding sets filtered to Pass@128 = 0, GRPO training of a Qwen 3.5 9B Thinking policy stays at 0-1% Pass@1, while RLTL;DR reaches 14-31% with insights in context and 12-13% without them at eval.

1 source

More stories today

Open the live feed