AnalysisAI ModelsSeptember 26, 2026

GEPA reflection beats GRPO RL with 3 examples vs 25,000 rollouts

Read original source →youtube.com

One round of reflection on three examples doubled the gains GRPO reached after 25,000 rollouts, per GEPA creator Lakshya A. Agrawal of UC Berkeley's Sky Computing Lab. GEPA optimizes prompts by reflecting on rollouts rather than collapsing them into a single reward score.

People · Lakshya A. Agrawal

1 source

More stories today

Open the live feed
GEPA reflection beats GRPO RL with 3 examples vs 25,000 rollouts