Skip to main content
AI & Machine Learning

Chinese AI Models Gain Edge in Price War

Server room with screens displaying a price comparison chart and various AI models.

Z.ai charges US$1.92 per million output tokens for GLM-5.2; Anthropic charges $25 per million for Opus 4.8. That price gap, and others like it, are already reshaping who can afford to run the agentic AI systems companies now prize.

Z.ai’s GLM-5.2, Moonshot’s Kimi‑K3, and the new price axis

Two recent launches from Chinese start-ups — Z.ai’s GLM-5.2 and Moonshot’s Kimi‑K3 — have shown intelligence close to Western competitors while undercutting them on price. The ledger is stark: GLM‑5.2’s output tokens are priced at US$1.92 per million, compared with Anthropic’s Opus 4.8 at $25 per million. Kimi‑K3, released just last week, completed more of Artificial Analysis’s agentic tasks correctly than any other model tested, and at roughly one‑third the cost of Anthropic’s Fable 5 on the same benchmark.

Why tokens — and agentic workflows — magnify price differences

Token pricing matters because agentic AI consumes tokens at vastly greater rates than conventional LLM queries. An MIT–Stanford–Google DeepMind team estimated agentic tasks can use up to 3,500 times as many tokens as a simple reasoning task. Ordinary prompts may require a few thousand tokens; agentic coding or research tasks often run into the millions, and some users already burn tens or hundreds of millions of tokens per day. OpenRouter reported customers consuming more tokens in one week in June than in an entire previous three‑month period. The multiplication of token use is why a per‑million‑token price difference of a few dollars becomes a budget crisis at scale.

DeepSeek, DeepSeek‑V4 Flash, and the Artificial Analysis stress test

Price moves from DeepSeek illustrate the leverage in play. DeepSeek announced a permanent 75 percent cut to token prices for its API; Artificial Analysis noted this can make an agent driven by DeepSeek cost as little as 1/34 as much as one using OpenAI or Anthropic for large token volumes. In a controlled run where Artificial Analysis made leading models perform 657 agentic office and administrative tasks, Anthropic’s Opus 4.8 cost nearly $1,000 to complete the suite, Z.ai cost $270, and DeepSeek‑V4 Flash — which burned the most tokens of any model tested (1 billion) — cost only $14 for the same tasks. Performance, however, did not always track price: DeepSeek‑V4 Flash scored poorly on task correctness versus both Western and Chinese competitors.

Market reactions: Microsoft, Meta, Nvidia and Western model adjustments

High token bills have already influenced commercial decisions. The source reports that high token payments factored into Microsoft’s cancellation of its subscription to Claude Code and may relate to Microsoft’s sudden interest in using DeepSeek to power its Copilot Cowork agentic helper. Western suppliers are responding: Meta announced Muse Spark 1.1 at $4.25 per million output tokens on July 9, undercutting Anthropic and OpenAI’s higher sticker prices. OpenAI has rolled out GPT‑5.6 variants including Luna, described as a “cost‑efficient” model that is cheaper to run agentically than Kimi‑K3 and GLM‑5.2, albeit with lower performance — Kimi‑K3 scored nearly 10 percentage points higher than Luna on the Artificial Analysis agentic benchmark. Nvidia is also positioning a different approach: a new laptop designed to run agentic workflows locally, avoiding API token costs, reportedly priced around US$2,000.

President Xi Jinping, the Global South, and China’s deployment strategy

The pricing advantage has geopolitical implications. In a speech at the World AI Conference in Shanghai, President Xi Jinping said China would launch AI application cooperation centres within six regional intergovernmental organisations covering almost the whole Global South. An op‑ed in the People’s Daily last year argued that low cost and open weights of Chinese AI make the technology more accessible to developing countries compared with “hegemonic” Western AI. The consequence, the source suggests, is that cheaper Chinese models may become embedded in the AI ecosystems of developing countries, with attendant shifts in market reach and influence.

What this means for technologists, policymakers, and enterprises

  • Technologists and security teams: Watch token consumption as an operational variable. Models differ not only in per‑token price but in how many tokens they use for complex tasks (the Economist reported GLM‑5.2 uses more tokens on complex tasks than a Western equivalent), and some Chinese models can fail on demanding multi‑day tasks.
  • Policymakers and regulators: Expect procurement and development choices in the Global South to be sensitive to sticker price and hosting options. China’s planned cooperation centres and the People’s Daily narrative position lower‑cost offerings as more accessible to developing countries.
  • Enterprises and procurement leaders: High‑token agentic workloads make marginal price differences compound into large budget gaps. As Azeem Azhar noted about his own use, when agents operate at scale “even small cost differences compound into meaningful budget gaps.” Decisions about whether to host agents locally (Nvidia’s laptop option) or switch APIs will hinge on both price and end‑to‑end task performance.

The core tension is now explicit: a future dominated by agentic AI will prize the cheapest workable intelligence, not necessarily the smartest. Chinese firms have demonstrated they can undercut Western prices while approaching comparable performance; Western vendors are cutting prices and offering lower‑cost models, but often at some sacrifice in accuracy. If Western companies don’t want to be left behind by their cheaper Chinese counterparts, they must find a balance between intelligence and affordability.

Original story