AnalysisRoboticsSeptember 27, 2026

Stanford's HomeBody humanoid explores, remembers, and acts on its own

Read original source →tml.stanford.edu

HomeBody uses a VLM to pick a skill and spatial target from the current ego view, map context, gripper state and recalled observations, then passes it as a structured tool call. Picking targets are image points normalized to 0-1000 with a chosen hand; depth comes from D435i stereo via Fast-FoundationStereo.

1 source

More stories today

Open the live feed