American companies are quietly shifting more of their AI workloads to Chinese models, and the pace of that change is beginning to attract attention in both boardrooms and Washington. Rather than chasing the most powerful models available, many companies are discovering that “good enough” AI at a fraction of the cost delivers a better return on investment. The share of tokens used by U.S. companies on Chinese AI models via OpenRouter has sat above 30% each week since Feb. 8, with that figure rising as high as 46%. A year earlier, that figure barely registered. This is not a story about Chinese labs catching up on raw capability, though they have. It is a story about American companies deciding that raw capability isn’t always what they’re paying for.
The Numbers Behind the Shift
The average share of tokens consumed by U.S. companies through Chinese models on OpenRouter was just 11% over the previous 12 months, and only 4.5% during the first half of 2025. The jump since February has been sharp enough that, on OpenRouter, the four most widely used models are now all Chinese, with DeepSeek remaining the most popular. Vercel, a separate platform developers use to deploy applications, is seeing the same pattern from a different angle: DeepSeek’s share of token usage on the platform jumped from under 1% to 17% in May, while its share of revenue stayed near 1% a gap that itself tells a story about how these models are being used, in high volume, at very low margin per query.
The newest entrant has moved even faster. Z.ai’s GLM 5.2, released in June, saw the fastest adoption of any model Vercel tracked in 2026, with daily token volume growing roughly 27 times and its customer count growing roughly 80 times in its first full week. Numbers like that don’t reflect a niche experiment. They reflect engineering teams actively rerouting production traffic.
The Cost Equation Changes
Vercel’s Harpreet Arora put it plainly: when a task doesn’t need the best model, teams are increasingly routing it to the cheapest one that’s good enough, and Chinese models are winning that trade. The price gap explains why. A Citi note found leading Chinese models charging as little as 18 cents per million tokens, compared with roughly $4 per million tokens for top U.S. frontier models. DeepSeek specifically charges about 3% of the token price OpenAI charges for GPT-5.5.
The pattern shows up in individual developer decisions as much as platform-level statistics. San Diego developer Stu Clott found that an hour of coding that cost about $10 on Claude cost less than 50 cents on DeepSeek. Startup Lindy made a similar calculation at a larger scale: founder Flo Crivello said switching from Anthropic’s models to DeepSeek saved the company millions of dollars, arguing “you don’t need God to write your email.” Not every workload gets shifted wholesale one Dallas-based developer described splitting his usage, spending roughly $500 a month on Claude and ChatGPT for sophisticated planning while spending another $200 on Chinese models that handle about 90% of routine coding and voice-recognition work. That split is probably a more accurate picture of enterprise behavior than any single anecdote: premium models for the hard problems, cheap models for everything routine.
Capability Gap Narrows, But Doesn’t Close
The economics only work because the performance gap has shrunk enough to make the trade defensible. Brookings fellow Kyle Chan estimates Chinese models are currently six to nine months behind the top American frontier systems, while typically running at a fraction of the cost. GLM 5.2 is the clearest example: it landed within a percentage point of Anthropic’s Opus 4.8 on one closely watched agentic benchmark, at roughly a fifth of the cost. On more general capability rankings, it holds its own too Artificial Analysis ranks GLM 5.2 fifth globally, behind three Anthropic models and one OpenAI model, while outperforming Google’s offerings.
None of this means Chinese models are winning on merit alone. OpenRouter’s Justin Summerville frames it more narrowly: open-source Chinese models can run 60% to 90% cheaper than the leading Anthropic and OpenAI models. He added that the newer open-source models perform well and handle everything short of the most complex language-model tasks. That “short of the most complex” qualifier is doing real work it’s the reason enterprises are routing tasks selectively rather than replacing their premium subscriptions outright.
The Politics of Adoption
Cheap and capable does not mean uncomplicated. Some of the highest-profile adoption stories have come with political consequences attached. Lawmakers launched investigations into Airbnb and Anysphere, the owner of coding platform Cursor, after both companies disclosed using Chinese open models like Qwen and Kimi to build their AI infrastructure. That scrutiny hasn’t stopped the underlying economics from pulling major firms in the same direction Microsoft is reportedly exploring using DeepSeek, or another open-source model, as a lower-cost alternative for Copilot Cowork, which currently runs on Anthropic and OpenAI systems.
Security researchers have raised separate concerns that have nothing to do with cost. Independent testing has found that DeepSeek, Qwen, GLM, and Kimi all decline to answer, or provide government-aligned responses, on topics including Taiwan’s political status, the Tiananmen Square protests, and Xinjiang. Data jurisdiction adds another layer: API calls to Chinese providers route through Chinese-jurisdiction servers unless a company uses an intermediary like OpenRouter or Azure, or self-hosts the model. Enterprises trying to avoid that exposure increasingly go through Western-managed infrastructure instead of hitting Chinese-hosted APIs directly, which explains why cloud intermediaries have become as important to this story as the model developers themselves.
The pressure pushing companies toward cheaper options isn’t abstract, either. Uber reportedly exhausted its entire 2026 AI budget in just four months after employees rapidly adopted AI coding tools, forcing management to introduce usage limits. That kind of budget shock is exactly the scenario Chinese providers are positioned to solve, whatever the geopolitical baggage attached to the solution.
The Bottom Line
What’s happening here isn’t a single company or model winning a benchmark war. It’s enterprises recalibrating how they think about AI spend altogether treating capability as something to be rationed carefully rather than maximized by default, and routing the bulk of routine work to whichever model clears the bar most cheaply. Chinese labs didn’t create that mindset, but they built the product line that rewards it, and rising token costs at the top U.S. labs are pushing more companies to test that trade every month.
The open question is whether this settles into a durable two-tier market premium Western models for the hardest problems, cheap Chinese models for everything else or whether the security and political friction eventually caps how far that adoption can go. Right now, the cost pressure is winning more arguments than the geopolitical caution is, and every enterprise still weighing that trade-off is effectively deciding how much of its AI budget is worth defending on principle alone.







Leave a Reply