NVIDIA yesterday unveiled benchmark results for its next-generation compute platform using SemiAnalysis's AgentX workload with the DeepSeek-v4-PRO 1.6T model. The evaluation showed Vera Rubin achieving 30x the per-megawatt throughput of Grace Blackwell, with the Vera Rubin NVL72's cost per million tokens coming in at roughly 1/35 of the GB300 NVL72. This marks the first time NVIDIA has publicly demonstrated a generational energy-efficiency leap using agentic programming inference workloads as the core measurement standard.
Tokens per Megawatt has become the defining efficiency metric for AI data centers, measuring how many tokens can be produced under a fixed power envelope. With power constraints tightening across the industry, more tokens per unit of electricity translates directly into larger agent deployments and lower operating costs. Energy efficiency is now displacing raw peak compute as the primary battleground in the AI infrastructure race.

The test data reveals accelerating generational improvements: under the DeepSeek-v4-PRO 1.6T workload, the GB300 NVL72 already delivers 15x the per-megawatt throughput of the H200 NVL8 solution, and Blackwell's cost per million tokens is roughly 1/10 of the previous generation. Vera Rubin then pushes per-megawatt throughput another 30x beyond GB300. The rate of decline in unit token costs is now approaching an order of magnitude per generation.
That said, these numbers warrant a measured read. The benchmark relies on SemiAnalysis's AgentX workload, which evaluates agentic programming inference scenarios through metrics such as end-to-end interactivity, standard interactivity, end-to-end latency, and time-to-first-token. This still differs from real-world production environments, and actual operational performance will need to be validated through large-scale deployment. But the direction is unmistakable: the energy-efficiency race has entered deep water.
The true value of this efficiency leap is moving agents from "able to run" to "affordable to run." When unit token costs fall by an order of magnitude, enterprises can afford to run agents continuously and at scale — not just in stop-and-go, per-invocation modes. StarWar Technology focuses on GPU compute platforms and compute scheduling capabilities, helping enterprises route training and inference workloads to the appropriate efficiency tier based on task characteristics, so that the cost advantages of next-generation platforms translate into real business outcomes.

For compute buyers, this efficiency shift means procurement logic must change. The industry used to compete on single-GPU peak performance; going forward, the metric that matters is effective tokens per unit of power and per unit of budget. Routing different workloads to the most suitable platforms will convert into cost competitiveness more directly than simply chasing the most powerful chip.
The narrative of the computing power industry is being rewritten. Compute is no longer a resource where more is always better — it is an investment that must be carefully optimized. Every refresh of the efficiency curve lowers the barrier to agent deployment and large-scale adoption, opening broader possibilities for the AI application ecosystem. For companies that have already positioned themselves in compute scheduling, this technology window is the moment to turn cost advantages into durable competitive strength.