Against the backdrop of continuously upgraded computing infrastructure, SIMA2 marks a substantial step forward for multimodal large models toward embodied intelligence. Unlike traditional approaches that rely on fixed training data, the system pioneers an autonomous cycle of “environment perception – task generation – trajectory optimization.” By calling on an independent Gemini model to batch-create training tasks and then using an internal reward model to filter high-quality interaction data, SIMA2 achieves continuous performance evolution without any human annotation. This mechanism has performed exceptionally well in open-world tests such as *No Man’s Sky*, where the agent can now parse environmental text, color codes, and even emoji-based instructions to complete abstract tasks like “destroy the blue marker.”\n\nNotably, DeepMind has for the first time deeply integrated SIMA2 with the generative world model Genie.

Within dynamically created, realistic 3D scenes, the AI can accurately recognize hundreds of objects—such as benches and butterflies—and execute interactions with them. This capability is underpinned by the underlying Transformer architecture’s unified representation of visual, linguistic, and action modalities, as well as chip-level computing power that supports real-time physics simulation. Project lead Jane Wang notes that such high-level decision-making modules represent the core challenge in migrating virtual intelligence to physical robots.\n\nClear divergences remain in the current technical path: SIMA2 focuses on cognitive-level reasoning, while DeepMind’s concurrently developed robot foundation models address low-level issues like motion control. No definitive plan yet exists for how these two lines will converge, reflecting the deep-seated contradiction in the AI industry between “brain” and “body” co-evolution. The lab has not announced a commercial timeline, but the open research preview reveals tech giants’ urgent need for ecosystem collaboration in the AGI race.