AnalysisAI ModelsJuly 17, 2026

On-policy value learning at 10000 frames per second

Talk covers REPO, an on-policy value learning method achieving 10,000 frames per second with resampling techniques. Shows when PPO beats value methods and when resampling matters.

Featured · Claas Voelcker

1 source