AnalysisAI ModelsSeptember 23, 2026

Armen Aghajanyan on VLM/VLA limits and embodied agents

Read original source →youtube.com

Aghajanyan argues feeding a model an hour of video costs roughly a million visual tokens while loss lands on only about 2% of them, since ground truth is just a transcript or a few labeled frames. He calls that a humongous waste and critiques predicting every pixel.

People · Armen Aghajanyan

1 source

More stories today

Open the live feed