
For the past couple of years, one of the more dominant conversations in AI investing has been about training. Which company makes the best chips for training large language models? Which hyperscaler has the most GPU capacity? Those were the right questions for the previous several years. But in 2026, the conversation is quietly shifting, and investors who recognize it early might be rewarded down the line.
The new phase and key term is inference. Every time someone uses ChatGPT, asks Google a question powered by Gemini, or runs an AI agent inside an enterprise application, a model is executing a real-time response. That is inference, and it happens billions of times per day at a scale now beginning to dwarf total training workloads in terms of compute demand.