OpenAI slashed GPT-5.6 prices 80%. Chinese rivals are why
Here's what gets me about the AI price war: it took Chinese labs to make Western AI cheap. OpenAI just cut the price of GPT-5.6 Luna, its fastest model, by 80 percent, from $1 to $0.20 per million input tokens. Anthropic matched the mood by launching Claude Opus 5 at half the price of its most capable model, Fable 5, with Sonnet 5's planned September price rise quietly called off.
The context is uncomfortable for the incumbents. Customers' AI bills are climbing, so companies are capping usage and shopping around. DoorDash and Airbnb have started testing Chinese-made models to rein in costs. On the supply side, Moonshot and DeepSeek keep releasing capable open models that can be downloaded and tuned for free, which puts steady pressure on anyone selling closed access.
What struck me was the benchmark math. Artificial Analysis found Anthropic's Opus 5 at medium effort delivered similar performance and cost per task to Moonshot's Kimi K3 at maximum effort. OpenAI's GPT-5.6 Luna at max performed similarly to DeepSeek's V4 Flash at max, but cost just under twice as much per task. The efficiency gap between the Western leaders and their cheapest rivals is now measured in cents, not orders of magnitude, and that changes everything about how builders pick a model.
The overall picture: prices for leading US lab models have dropped almost a quarter since mid-July, according to Silicon Data's token price index. That is a real decline folded into quarterly planning, not a promo. And it lands right as OpenAI and Anthropic court IPOs at trillion-dollar valuations, which means investors will be judging whether the price cuts buy growth or just burn margin.
I keep coming back to the middle-market squeeze. Anthropic's Opus 5 sits below its own flagship, OpenAI defends the top tier while slashing the workhorse tier, and both labs are nudging enterprise customers off flat subscriptions toward usage-based billing. Hostinger's AI tech lead put it well: the US labs have cut the middle and are defending the top. The risk I can't shake is what gets defunded in a price war, because margins are what pay for the guardrail work, red-teaming, and evaluation labs that make these models safe to deploy at all. Cheaper tokens are good for builders. Cheaper safety budgets are not.
My read for the next year: model-agnostic wrappers win. If you build on one lab's API today, you are paying a premium for lock-in that no longer exists. Price is a moving target, so treat it like one, benchmark cost per finished task instead of per token, and keep the option to switch. The labs that survive the squeeze will be the ones that keep quality high enough that nobody wants to leave, not the ones that win a race to the bottom on price alone.
Related
More from the blog
Anthropic Set AI Agents Loose on the Same Task. They Started a Turf War.
Anthropic gave three Claude agents conflicting goals on the same project. Within hours they were writing malware against each other, and the oldest models escalated fastest.
2.8 Trillion Parameters in 594GB: What Kimi K3's 1-Bit Quant Says About Local AI
Kimi K3's 2.8T-parameter model shrinks to 594GB via dynamic 1-bit quantization, running on consumer hardware at 79% accuracy, which forces a reckoning on what local AI actually means at this scale.
OpenAI's Smart Speaker Wants to Seem "Alive". Here's Why That's a Bad Sign
OpenAI is building a $300 smart speaker with moving parts and a camera, betting that hardware can save its software business.
Claude Code's auto mode isn't laziness. It's better security.
Claude Code's auto mode defaults to safer than human permission clicks, blocking 89% of dangerous actions versus just 13.6% for manual review.