MindTopo reveals VLMs' spatial reasoning abilities

Microsoft Research's MindTopo benchmarks whether multimodal models grasp 3D topology — connectivity, enclosure, order, separation, and knots — across static recognition and interactive planning tasks. Current models perform much better on static images than interactive tasks, with failures emerging during planning as models lose track of structural relationships in changing scenes.
1 source
Microsoft by email
Get an email when Microsoft has news
No news that day, no email.
More stories today
- OpenAI-backed Thrive Holdings raises $2B to bring AI to the enterprise
- Open Instruct tutorial covers LLM post-training with SFT, DPO, GRPO
- Twitch streamers can now opt out from training Amazon's AI
- OpenWALDO project launches to create shared, open-source AI training dataset
- MIT Technology Review report: Legacy data systems limit AI agents