In recent months, OpenAI and Anthropic took turns pushing their flagship models forward, while Google remained relatively quiet. Flash was updated almost every three weeks—3.6, 3.7, and 3.8 arrived in succession—but the Pro line, which represents the upper limit of capability, showed little movement. Then, recently, a model labeled gemini-3.8-flash appeared in the blind-test arena Arena and clearly jumped ahead in coding, SVG generation, and Agent tasks. Developers quickly speculated that it was actually the unreleased Gemini 4 Pro. A benchmark comparison table circulating widely in the community described the model as fully leading in software engineering, terminal Agent, expert-level knowledge reasoning, and computer-use tests, even breaking 2,000 Elo. But neither the table nor the alleged internal screenshots have official backing. The only confirmable fact is the visible gap between this anonymous model and the official 3.8 Flash.
文章图片 2
More notable is that Google has clearly accelerated its effort to let AI participate in model R&D. Reuters reported in August this year that Google co-founder Sergey Brin has been pushing to speed up Gemini and has made recursive self-improvement a key focus. Google also published the Dream-RSI paper, examining how Agents can achieve recursive self-improvement by continuously improving exploration strategies. Anthropic has publicly said that, as of August this year, Claude could already lead 26% of its AI R&D tasks.
文章图片 4
Yet while the “deceleration initiative” has been loud, common rules have not been established. Slowing down alone means giving the lead window to rivals, so traces of next-generation model testing—such as Anthropic’s Opus 5.2 and OpenAI’s GPT-6 Sol—have appeared one after another. Several companies can jointly say they “should be cautious,” but once the question becomes “who slows first, and will the others slow too,” consensus rapidly disappears. This acceleration ultimately lands on computing power. The faster models iterate, the higher the demands on compute scheduling efficiency for training and inference across GPUs, AI chips, and heterogeneous clusters. StarWar GPU computing platform and heterogeneous cluster solutions focus on compute resource management and dynamic scheduling, helping enterprises improve GPU utilization in multi-card, multi-cluster environments so that continuously accelerating R&D and inference needs are not held back by inefficient infrastructure. When AI begins to participate in building stronger AI