qwen-tts) — we dogfood the stack we sell.Welcome to the NovAI Podcast, the show where we help developers build with the world's best AI models without the eye-watering bill. In this first episode we are answering the question we get asked more than any other: why are Chinese AI models so much cheaper than the big Western ones, and can you actually trust them in production?
Let us start with the numbers, because the numbers are genuinely shocking. On the Western side, a frontier model like GPT-5.5 can cost around five dollars per million input tokens, and Claude Opus sits in a similar premium bracket. On the Chinese side, DeepSeek V4 comes in at roughly eight to twenty seven cents per million tokens depending on the tier, and Alibaba's Qwen and Moonshot's Kimi are in the same neighbourhood. That is not a ten or twenty percent discount. That is often a ten to twenty times lower price for a model that, on a large share of tasks, performs in the same league.
So why is the gap so wide? There are three big reasons. The first is a brutal domestic price war. China has dozens of well funded AI labs, DeepSeek, Alibaba, Moonshot, MiniMax, Zhipu, ByteDance, all competing for the same developers. When that many strong players fight for the same market, price collapses toward cost, and margins get thin. The second reason is subsidised and efficient compute. These labs optimise aggressively for inference cost, using techniques like mixture of experts architectures that only activate a fraction of the parameters per token, which makes each request dramatically cheaper to serve. The third reason is the open weight ecosystem. Many of these models are released with open weights, which means the whole ecosystem can build on them, distil them and serve them efficiently, and that competition keeps prices honest.
Now the honest question: is cheaper the same as worse? Sometimes, yes. If you need the absolute bleeding edge on the hardest reasoning problems, the top Western models can still pull ahead, and there are cases where you will happily pay the premium. But for the vast majority of production workloads, chatbots, summarisation, code generation, classification, translation, agents, retrieval, the quality difference has narrowed to the point where it is no longer the deciding factor. On many public coding and long context benchmarks the top Chinese models are trading blows with models that cost ten times as much.
Here is the practical part. The historic barrier was never quality, it was access. To use these models officially you often needed a Chinese phone number, a Chinese payment method, and sometimes business verification. That is exactly the problem an aggregator gateway solves. With a single OpenAI compatible endpoint, you point your base URL at the gateway, keep the exact same client code you already use, and suddenly you can call DeepSeek, Qwen, Kimi, GLM and MiniMax with one API key, paying by card, PayPal or crypto. No rewrite, no Chinese phone number, one bill.
So what is the takeaway? If you are spending real money on AI APIs in 2026 and you have only ever used the premium Western models, you are very likely overpaying for workloads where a Chinese model would perform just as well at a tenth of the cost. The smart play is not to pick a side. It is to route. Use the frontier model where it genuinely earns its price, and route the high volume, cost sensitive work to the models that give you ninety five percent of the quality for five percent of the bill.
That is it for episode one. If you want to see the full price comparison table with real per million token numbers, there is a link in the show notes to the written version on the NovAI blog. Thanks for listening, and go build something for less.
Prefer to read? The full written version with live price tables and code examples: