Physical AI is becoming a key frontier for global artificial intelligence in 2026. Over the past three years, large language models have pushed the boundaries of “speaking, thinking, and working” to the extreme. But as the industry begins to ask whether models can enter the real world and take action, a new door has opened. Recently, OpenAI’s GPT-6 Astra reignited discussion with its reasoning and multimodal capabilities, while Google DeepMind released Gemini Robotics 2 for embodied reasoning.
文章图片 2
One experiment illustrates where the gap lies. Robocurve deployed GPT-6 Astra on a real robotic arm and completed 19 of 20 “place the block in the bowl” tasks, while another model completed only 8. Yet on millimeter-level precision tasks such as “align and insert the blue puzzle piece into the slot,” GPT-6 Astra’s success rate fell from 95% to 10%. The general understanding accumulated by large models can bridge the gap from “seeing” to “coarse actions,” but it still stalls at “fine contact and precise insertion.” At this critical juncture, DeepWit, a company emerging from Zhongguancun in Haidian, Beijing, released its physical foundation model PhysBrain 1.5. Launched on September 9, 2026, it achieved a composite average score of 72.5 across 28 public benchmarks, ranking first among participating open-source models. Its gap with GPT-6 Astra (73.3) and Gemini 3.6 Flash (73.0) has narrowed to within 1 point. Across the full capability chain of spatial understanding, action generation, and future state prediction, PhysBrain 1.5 secured 14 first-place open-source rankings and 10 second-place open-source rankings.
文章图片 4
Notably, PhysBrain 1.5 has chosen to be fully open source: its technical report, 2B and 8B model weights, and evaluation toolkit are all open to the community. This aligns with its “build the foundation model solidly” approach—not competing on parameter scale, but deepening its focus on physical intelligence as a core direction. Open sourcing also means more teams can validate and iterate on the same foundation, lowering the barrier to embodied research.
文章图片 6
This connects with StarWar Technology’s focus on compute scheduling and agent deployment: physical AI training and evaluation are highly compute-intensive and require organizing perception, planning, execution, and other stages into a stable pipeline. Whether processing multimodal data or repeatedly running large-scale evaluations, underlying compute must be uniformly scheduled and elastically supplied. Otherwise, even a “smarter” model will struggle to become a “more reliable” robot. However, there is still a distance between benchmark scores and real-world productivity. Fine manipulation,