Following the U.S. market close on August 26, NVIDIA published its Q2 FY2027 financial results. The company guided next-quarter revenue to $91 billion — robust by any measure — but the figure dominating market discussions was Jensen Huang's assertion during the earnings call: "In AI, compute is revenue." Underpinning this statement is a subtle yet telling structural change in the report: inference business growth surpassed training for the first time, with enterprise multi-agent and automated Agent operations explicitly identified as the core growth engine for inference computing demand.
Rewind two years, and AI computing demand was driven almost exclusively by model training. From the relentless expansion of GPT-scale large language models to the dense iteration of open-source models, the scale race among training clusters propped up the entire computing market. Cloud providers competed on who could secure more GPUs faster and build larger clusters. At that time, inference was little more than a byproduct of training — occupying a limited share of the demand structure, with computing supply largely oriented toward the training side.

Now, the industry is crossing an inflection point. According to the earnings analysis, inference now accounts for more than half of NVIDIA's data center revenue, with growth rates for the first time overtaking training. Enterprise multi-agent and automated Agent businesses are the primary catalyst behind this inference surge: agents are no longer confined to conversation — they are planning tasks, invoking tools, and executing across systems, with every execution translating into real inference token consumption. The AI industry's narrative is shifting from "training stronger models" to "deploying more capable agents."
Yet the explosion in inference demand poses new challenges for computing supply. Inference and training tasks have fundamentally different characteristics: training is large-scale, long-cycle batch processing that can be scheduled in bulk, while inference represents fragmented, latency-sensitive online workloads that fluctuate in real time with user requests. Traditional training-oriented scheduling systems are difficult to apply directly to inference scenarios. Elastic scaling, on-demand orchestration, and cost control all require a new methodology. For intelligent computing platforms, the ability to effectively schedule inference computing power is becoming a more critical differentiator than the sheer number of GPUs deployed.

This shift aligns closely with the capability-building direction of domestic computing service providers. StarWar Cloud, focused on GPU computing platforms and scheduling capabilities, has prioritized optimization of inference-side resource orchestration and scheduling strategies — dynamically allocating heterogeneous computing resources based on business loads to match the real-time inference demands of enterprise multi-agent and DiWorker operations. The goal is to transform computing power from "acquired" to "effectively utilized."
Looking at a longer horizon, the center of gravity in computing supply shifting from training to inference means the competitive logic across the entire industry chain is being rewritten. Chip makers are optimizing inference efficiency, cloud vendors are driving down inference costs, and computing platforms are building differentiated advantages in heterogeneous resource selection, inference scheduling, and cost structure. Competition in the inference era is no longer an arms race over peak compute — it is a systems-level capability defined by on-demand availability, schedulability, and cost controllability.
For enterprises and developers, the signal from NVIDIA's earnings is unmistakable: the main battlefield of AI deployment is migrating toward agents and inference. When selecting computing services, beyond chip models and cluster scale, attention must now be equally paid to scheduling flexibility, cost efficiency, and heterogeneous computing adaptability on the inference side. The flip side of "compute is revenue" is that computing power must genuinely and efficiently convert into business output.