AnalysisAI ModelsSeptember 4, 2026

DRACO improves long-horizon agent training with dynamic rubrics

DRACO dynamically generates rubrics and redistributes trajectory-level scores into per-step advantages for RL without verifiers, improving long-horizon agent performance. It uses GRPO and is outcome-blind.

1 source

More stories today

Open the live feed
DRACO improves long-horizon agent training with dynamic rubrics