Grounded Action Model: 3D Grounding Forms Robotics Foundation
robotics
| Source: HF Papers | Original article
Researchers propose the Grounded Action Model, a 3D grounding approach that enables manipulation policies to identify relevant objects and locations, addressing gaps in current robot foundation models.
A new class of robot foundation models has been unveiled under the name Grounded Action Model (GAM). The research proposes “3‑D grounding” as the core building block for robot action learning, replacing the language‑generation and video‑generation backbones that dominate current vision‑language‑action (VLA) and world‑action (WAM) models. GAM can be prompted with natural language, point clicks or bounding‑box selections, each of which is translated into a shared, object‑centric representation of the selected items in metric 3‑D space.
The shift matters because manipulation policies need to know not only which objects are relevant but also precisely where they are. Existing pretrained backbones do not enforce this metric grounding, leading to brittle performance when scenes change or targets are moved. By anchoring policies to a pretrained, promptable 3‑D grounding model, GAM directly links visual observations and task prompts to concrete spatial coordinates. Early demonstrations show improved robustness to scene variations and target relocation, suggesting a path toward more reliable robot manipulation in unstructured environments.
The work also signals a broader move toward open‑source 3‑D vision foundations that natively understand motion, structure and appearance. As the model and its training pipeline are released publicly, the robotics community can test the approach on a range of platforms, from research labs to edge devices.
Going forward, observers should watch for benchmark results that compare GAM against established VLA and WAM systems, for integration efforts with existing robot stacks, and for any follow‑up releases that expand the promptable 3‑D grounding model’s capabilities. If the early gains hold, 3‑D grounding could become the standard substrate for future robot control, reshaping how manipulation policies are trained and deployed.
Sources
Back to AIPULSEN