Scaling up Continual Learning — Ronak Malde, Trajectory

Ronak Malde of Trajectory describes how scaling policy self-distillation to trajectories with a hundred tool calls makes a model collapse into hedging, filling tokens with "wait, but, and maybe" — a failure he calls the "but wait problem".
Featured · Ronak Malde
How this story unfolded
5 days · 2 reports · 1 community post · from Aug 12
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Z.ai CEO Jie Tang: GLM 5.3 gains come from RL, not parameter count
- New tool adds 14 skills to Claude Code and Cursor for Markdown diagrams
- Tool turns Claude into a team of AI employees on your Mac
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills