Does cheaper AI pricing actually reduce total costs for users?
Core argument: Headline token prices are an unreliable cost proxy, as higher-capability models completing tasks in fewer tokens or attempts can deliver lower total cost despite carrying a higher per-token list price.
OpenAI cut the price of GPT-5.6 Luna from $1 to $0.20 per mn input tokens and from $6 to $1.20 per mn output tokens. Anthropic launched Opus 5 at $5 per mn input tokens and $25 per mn output tokens — half the price of its Fable 5 model. This week, the company called off a planned rise in prices for its Sonnet 5 model, which had been due to take effect from September. Headline token prices do not provide a straightforward comparison between AI models, however. More capable models can sometimes complete a task using fewer tokens or with fewer attempts, meaning a model that appears more expensive based on the headline price of tokens can ultimately cost less.

