On September 10, DeepSeek V4.1 Flash officially launched. It is a 552B-parameter MoE model built on a new Causal-Encoder-Decoder asymmetric architecture. Input activation is only 8B and output activation is 16B, making inference costs significantly lower than those of models of the same size. It also natively supports multimodal visual understanding without an external vision module.
Typically, smaller and faster models deliver lower performance, but V4.1 Flash comprehensively surpasses the previous generation on benchmarks. Through KV Cache compression, it reduces VRAM requirements to one quarter and SSD requirements to one eighth of the previous generation. Community tests show output speeds of 300 to 500 tokens per second. The simultaneous improvement in speed and cost is why it has drawn widespread discussion.

Prices are falling in parallel. From 12:00 on September 10, Flash off-peak cache-hit input is RMB 0.02, cache-miss input is RMB 1, and output is RMB 4 per million tokens; peak-hour prices double. DeepSeek has also adjusted its routing policy: V4 Pro requests will be routed to the higher-performance, faster V4.1 Flash and billed at Flash prices.
For developers, the variables in model selection are changing. As capability gaps among leading models narrow, token unit price, cache hit rate, inference speed, and context cost will directly determine whether an application makes economic sense. Competition among vendors is also shifting from parameter scale to unit cost and invocation experience.

This is exactly the direction StarWar Cloud focuses on in its LLM API marketplace: aggregating models of different specifications and price points into an API supply that can be compared, switched, and scheduled by invocation volume. This lets developers use models on demand based on a combination of cost and performance, rather than being locked into a single model. The more models there are, the more obvious the value of unified access and billing becomes.
Chinese vendors are not taking identical paths: some create stickiness and community buzz through irregular quota resets, while others expand share through open-source weights and low prices. Either way, the underlying judgment is the same—when capabilities are