Apple's IVT framework cuts video reasoning latency by 5x

Apple researchers introduce Internalized Visual Thinking (IVT), a post-training framework that predicts latent future-frame representations during training, enabling direct inference without generating intermediate images. IVT matches or beats Visual CoT across six settings while reducing end-to-end latency by more than 5×.
1 source
Apple by email
Get an email when Apple has news
No news that day, no email.
More stories today
- Apple cuts jobs across Vision Pro, Siri teams to refocus on AI
- Tibo discusses Codex, ultrafast AI, and OpenAI vs Anthropic
- Instinct AI assistant raises privacy and security concerns
- MiniMax M3 and M2.7 models offer unlimited access on GMI Cloud
- Ethan Mollick: AI-written submissions are obvious to readers