GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robot Manipulation
training
| Source: HF Papers | Original article
Researchers unveil GE-Act 2.0, a world‑action model that leverages pretraining and scaling to improve robotic manipulation by predicting future states from both video and interaction data.
A new paper unveils Genie Envisioner Act 2.0 (GE‑Act 2.0), a world‑action model (WAM) built from the ground up to learn robotic manipulation at scale. Unlike most existing WAMs, which graft action heads onto video generators pretrained on web‑scale footage, GE‑Act 2.0 is pretrained from scratch on a mix of action‑free video and action‑labeled robot‑interaction data. The architecture combines a control‑oriented autoencoder, a single‑step visual planner and an inverse‑dynamics model, and employs knowledge‑aligned selective optimization to fuse heterogeneous inputs without relying on prior video‑generation weights.
The development matters because it tackles a long‑standing bottleneck in robot learning: the scarcity of scalable, zero‑shot manipulation capabilities. By learning dynamics directly from raw visual streams and interaction traces, GE‑Act 2.0 promises to generalise across a wide array of skills and environmental conditions without task‑specific fine‑tuning. This could accelerate the deployment of adaptable robots in manufacturing, logistics and service settings, reducing the data‑collection overhead that has limited broader adoption.
The research community will now watch for empirical results that benchmark GE‑Act 2.0 against prior WAMs such as Robust‑WAM and Atlas, especially in real‑world manipulation scenarios. Follow‑up work is likely to explore larger, more diverse datasets, tighter integration with hardware controllers, and extensions that combine the model’s predictive power with higher‑level planning systems. If the approach scales as claimed, it could reshape how roboticists think about pretraining, moving the field closer to truly versatile, vision‑driven robots that learn from the same visual streams that humans do.
Sources
Back to AIPULSEN