Does scaling inference revenue eliminate the need for high profit margins?
Core argument: Chinese model competition poses limited near-term economic threat to frontier labs: inference margins currently fund training runs, and an inference market growing faster than training costs preserves lab viability even as prices compress.
Right now demand exceeds supply for frontier models, and supply is limited by a lack of compute. This computer shortage doesn't just mean that a compute supplier like Nvidia makes very large margins, but also that Nvidia's customers, like SpaceXAI, can turn around and resell compute at high margins as well to a company like Anthropic. Anthropic, meanwhile, can pay the markup because they can sell tokens with a higher markup still. I think the reaction to Kimi and Chinese models generally is pretty over-blown, at least from an economic perspective. As long as training consumed more GPUs than inference, it was critical to maximize inference revenue to help fund the next training run, which meant charging very high prices for inference. Going forward, however, I expect the inference market to grow much faster than training costs, which means they really can make it up in volume... frontier labs should have more confidence that they can not just survive but thrive with lower prices (once they have sufficient compute).


