Frontier capability can be reached with dramatically less compute than leading labs' capital programs imply if engineering effort is spent on architecture and memory efficiency: DeepSeek's reported $5.6M training run on 2,000 H800s produced GPT-4-level open models while Western labs spent hundreds of millions and used clusters orders of magnitude larger.
This changes capacity planning and make-vs-buy reasoning: an agent team can seriously consider training or fine-tuning its own competitive models on a small compute budget, and should optimize algorithms and data before accepting large-cluster costs as unavoidable.
When an expensive resource like compute is optimized sufficiently, the barrier to entry falls from capital to expertise; incumbents who rely on scale alone are exposed.
DeepSeek R1 trained for roughly $5.6M compared to hundreds of millions spent by Western labs.
DeepSeek used 2,000 Nvidia H800 chips to achieve results previously requiring massive clusters.
DeepSeek utilized engineering innovations and memory scaling rather than massive capital expenditure.