On the night of July 19, Kimi chose to stop accepting new subscribers — a rare move for an AI company, and one that runs against conventional growth logic. Two days earlier, Moonshot AI had launched Kimi K3, and demand quickly outstripped expectations. With GPU capacity approaching its limit, the company redirected available compute to already paying users. For the new model, this is arguably a good sign: too many users want to pay. But for the company, it is a stark reminder that while model capability has successfully generated demand, the underlying compute infrastructure is not yet ready to serve all of it.
The numbers capture the tension. By mid-June, Kimi's annualized recurring revenue had surpassed $300 million, up from roughly $100 million in March. API revenue accounts for over 70% of the total, driven largely by overseas developers, coding tools, and agent-based invocations. Forty-five days later, multiple media outlets reported that Moonshot AI had confidentially filed its A1 documents with the Hong Kong Stock Exchange and was in the process of raising a new round at a pre-money valuation of approximately $50 billion — widely seen as the final private funding round before its IPO.

A company approaching listing typically moves away from private capital. Kimi appears to be doing the opposite: raising ever more money while accelerating toward the public markets. Its bet on K3 centers on long-horizon coding and agent tasks — precisely the most compute-intensive workloads in the industry. K3 carries 2.8 trillion total parameters and a 1-million-token context window, built on a sparse MoE architecture. The stronger the model, the higher the cost per task run; the flip side of users' willingness to pay is that every additional call burns more compute dollars.

The strain is visible across the entire supply chain. Bloomberg previously reported that Moonshot AI secured computing capacity equivalent to roughly 20,000 NVIDIA GPUs through Alibaba. Third-party cloud providers have begun hosting K3 and contributing to inference optimization. There are also reports that Moonshot AI has held discussions with Microsoft, Amazon, and Google about hosting K3, seeking service revenue shares of up to 30%. The strategic direction is clear: keep training in-house, and push as much inference as possible to the cloud.
Pausing new subscriptions is an unusual move for an internet-native product. When traffic floods a conventional server, the solution is to add more machines. But for frontier models, the bottleneck is not the number of machines — it is the supply of AI chips and the opportunity cost of pulling resources away from next-generation training runs. By late July, Moonshot AI's full-time team stood at roughly 300 people, with an average age under 30. That such a lean organization can support massive paid user traffic and API volume is itself a testament to offloading heavy assets onto the cloud side.

Zoom out, and Kimi is hardly alone. Anthropic has reportedly filed confidentially for a U.S. IPO and signed cloud computing agreements worth tens of billions of dollars. The logic running through the entire industry is consistent: capital buys GPUs and models; GPUs and models generate API revenue; and API revenue funds the next round of research — yet that final leg of the loop still relies on external capital to keep flowing. The question increasingly asked across the sector cuts deeper: is ARR truly recurring revenue, or is it simply an annualized run rate that hides how much capital is still being consumed?
That is why computing power is fast becoming the new competitive dimension among frontier model companies. Whoever can secure larger, more flexible compute at a lower cost will be the one able to absorb the demand that improved models create. Transforming GPU compute from heavy, self-built infrastructure into elastically scheduled, on-demand platform resources is the most direct way for fast followers to close the gap — exactly the challenge StarWar Cloud's GPU computing platform and scheduling capabilities are built to address.