Original title: Gemini Robotics 2 brings whole body intelligence to robots
Article
Google presents Gemini Robotics 2 as a new embodied-intelligence stack built around three models: a vision-language-action model for whole-body motor control, an embodied reasoning model for long-horizon planning and human interaction, and an on-device VLA for local execution and quick adaptation to new robot bodies. The system is shown moving entire humanoids, using both hands and grippers for tasks like object transport, knot tying, bag sealing, and compact packing, and coordinating multiple robots on longer workflows. The updated reasoning model is positioned to handle multi-minute task sequences with self-correction, milestone tracking, uncertainty handling, and proactive requests for human help. The on-device version claims embodiment transfer in a few hours with fewer than 200 examples across varied hardware like Dexmate, SO101, and Trossen. The release also introduces safety and alignment work, including benchmarked uncertainty handling, safety-tool refusal behavior, and closer-human stopping behavior, while keeping deployment access limited to early partners and private previews plus a public AI Studio path for one model. Community discussion reflected both excitement and caution, praising the platform direction but pressing for evidence of real-world reliability and practical accessibility. They contrasted it with broader AI progress and noted potential impact in constrained industrial and care contexts, while warning that broader labor and trust implications remain open.
Comments are broadly mixed, combining enthusiasm about accelerated progress with doubts about immediate practicality. Several participants expect physical AI to follow recent language-model trends and eventually unlock major applications, while others argue humanoids are slow, cumbersome, or unnecessary for many home tasks that can be automated with simpler systems. Technical skepticism centers on whether full-language-model control can meet real-time requirements, with some arguing that vision can be ML-driven but actuation should remain traditional control. Safety concerns recur around force, proximity stopping behavior, sensor coverage, and trust in domestic deployments, with readers asking if robots can safely handle everyday interactions like doorknobs, falls, and obstacle-rich environments. Access and realism are questioned, including how to test models at home, what instrumentation is required, and whether released APIs deliver more than baseline benchmarks. Debates also reference wider industry dynamics, comparing Google to other AI labs, concerns about labor displacement, and uncertainty over whether humanoid-first designs limit innovation versus specialized robots.