Are smaller AI models making cloud computing less essential for artificial intelligence?
Core argument: Inference migration from data centers to devices results in a fundamentally different supply chain, threatening capex-heavy frontier model providers lacking.
A quiet story in AI isn't that large models keep getting better, it's that small ones are catching up quickly. In 2022, reaching 60% on MMLU, a test of academic reasoning, required a 540-billion-parameter model. By 2024, a 3.8-billion-parameter model hit the same threshold — a 142-fold parameter reduction in two years. The direction is clear, and the implications are wide-ranging. Edge compute becomes real. Models that fit on a phone or a sensor don't need a data center. Inference shifts from hyperscaler to device, which is a different supply chain and a different competitive moat than anyone has built for. The data center API business model has a half-life. If a fine-tuned 7B model running locally can handle 80% of enterprise use cases, the case for paying per-token to a cloud provider shrinks fast. The most exposed providers are those whose moats are model quality rather than distribution, i.e., the capex-heavy frontier models. Compute becomes less scarce.

