Beyond VLAs: How World Action Models Reshape Robot Manipulation

NVIDIA argues world action models (WAMs) — using a video world model backbone instead of a VLM's language model — improve physical generalization in robot policies. The post cites NVIDIA Cosmos 3 as a foundation and researcher Jim Fan's take: 'VLAs are dead, long live World Action Models.'
Featured · Jim Fan
1 source
NVIDIA by email
Get an email when NVIDIA has news
No news that day, no email.
More stories today
- Sequoia Capital invests in AI-native video platform Preview
- US Launches Effort to Speed Trade in AI Goods Between Allies
- DeepMind launches SL2T sign language-to-text model
- Liquid AI releases LFM2.5-VL-3B vision-language model for edge
- Grok and Meta's release discussed on ETN podcast episode