AnalysisAI ModelsSeptember 11, 2026

T1: 122B MoE agent trained via RL for long-horizon terminal tasks

T1 is a 122B Mixture-of-Experts model trained with reinforcement learning to execute long-horizon terminal tasks in a cloud sandbox, reporting state-of-the-art results on Terminal-Bench. Training used stable actor-critic optimization, process rewards, and out-of-distribution training.

1 source

More stories today

Open the live feed