On September 2, Google released Gemini 3.8 Flash, and the price list read as unchanged as ever: $0.75 per million input tokens and $3.75 per million output tokens. But Artificial Analysis' same-day benchmark told a different story — completing the same task now costs 40% more overall, with the model consuming 30% more output tokens per task. The price sheet didn't move, yet the bill did, and buried in that contradiction lies the true battleground for the next generation of model competition. First, the capability picture. Gemini 3.8 Flash on the DeepSWE long-horizon software engineering benchmark trails Claude Opus 5 by only 0.3 points while costing less than one-sixth as much. It posts similarly strong results on general reasoning and financial agent benchmarks, generating roughly 305 tokens per second — among the fastest commercial models ever independently measured. The model is already live across Gemini App, AI Studio, and coding platforms; Cursor and other developer tools integrated on launch day, following the volume-driven path through developer entry points.
文章图片 2
The core design philosophy behind this upgrade: the model tries harder. On complex tasks, the model executes more reasoning steps and iterates more frequently on tool calls. Each step's reasoning, tool invocation, and verification loop consumes output tokens — double the steps, and token consumption naturally climbs, with amplification effects more pronounced under high-compute settings. Artificial Analysis mapped the full chain: per-task output tokens rose about 30%, translating to a roughly 40% increase in per-task total cost. Even so, at approximately $0.58 per task, the cost remains on the leading edge of the intelligence-to-price frontier. The deeper shift lies in the pricing trajectory. Once promotional pricing expires, the model's rates will double to $1.50 per million input tokens and $7.50 per million output tokens — first usage scales up, then unit prices scale up; the two waves together form the complete pricing path. Google is meanwhile keeping 3.7 Flash as a low-cost option. This three-week iteration cadence for the Flash line targets coding agents and automated workflows — the fastest-growing source of token consumption today and the primary revenue battleground for API providers across the industry.
文章图片 4
The concurrently released cybersecurity edition takes a different, gated route: access is limited to roughly 650 institutions through a trusted access program, mirroring Anthropic's approach of locking frontier models inside restricted availability. This signals a broader trend: frontier capabilities are stratifying. Volume-oriented versions serve the developer ecosystem while controlled versions serve regulated, high-value tasks, and the degree of capability openness is becoming a new strategic lever. The boundaries of benchmark claims deserve scrutiny. The official headline benchmarks all show leadership, but third-party cross-checks of the full leaderboard reveal an uneven picture: on well-defined professional tasks with clear end states, it performs close to flagship models; on open-ended tasks requiring autonomous planning of next steps, the gap widens. Moreover, all comparison data reflects vendor-reported benchmark figures — a gap between launch-day benchmarks and production performance has historically always existed. The real-world cost impact also depends on task architecture: short Q&A is barely affected by token amplification, while long-horizon agentic tasks bear the brunt of rising costs.
文章图片 6
As per-token prices converge across providers, how many tokens a task burns has become the new pricing lever. Models that reason deeper and take more steps are naturally more expensive, even when price lists remain frozen; Anthropic's 75% cut to cache read prices two days ago runs the same arithmetic. For developers, anchoring model selection on per-task cost rather than unit price is fast becoming the universal rule in the large model API era. This likewise pushes aggregation gateways like large model API marketplaces toward comparing models by task-level outcomes rather than nominal per-token rates. The second phase of the price war is a contest over whose bill hides more. Model vendors have pushed their cost-reduction levers from headline prices down to the engineering layer of token consumption and cache reads. Enterprise procurement, in turn, must build finer-grained task-level metering. Whoever can turn model routing, caching, and task evaluation into a transparent, auditable system will stay clear-eyed in this veiled billing war — which is precisely the value StarWar Tech continues to build in its large model API marketplace.