Gartner predicts major drop in AI token costs

Mar 26, 2026

5:06pm UTC

Copy link
Share on X
Share on LinkedIn
Share on Instagram
Share via Facebook

As AI applications become more advanced, they require more tokens, and overall costs continue to rise. That may not always be the case.

Recent research from Gartner finds that, by 2030, performing inference on a large language model with one trillion parameters will cost AI providers 90% less than it did in 2025. Furthermore, the research firm predicts that LLMs in 2030 will be up to 100 times more cost-efficient than the earliest models of similar size developed in 2022.

“These cost improvements will be driven by a combination of semiconductor and infrastructure efficiency improvements, model design innovations, higher chip utilization, increased use of inference-specialized silicon, and application of edge devices for specific use cases,” Will Sommer, senior director analyst at Gartner, wrote in the post.

The research included two different forecasts that offer differing projections based on how chips are applied during training:

  • Frontier scenarios: A model trained exclusively on leading-edge chips; costs come in significantly lower.
  • Legacy blend scenarios: A model trained on a mix of legacy and leading-edge semiconductors; costs are considerably higher.

While either approach leads to a significant drop in costs compared to today, it won’t necessarily translate into token cost savings for enterprise customers. Frontier intelligence will demand significantly more tokens than today's mainstream applications, such as the lower cost of basic chatbots versus more expensive agentic AI, according to Gartner.

“Integrating the next largest, most capable foundation models into every new feature will be commercially insolvent at scale,” Sommer told The Deep View. “As models become larger and more complex, they require significantly more tokens, and those tokens become more expensive to generate.”

As a result, to get the most value, Sommer recommends that enterprises diversify the models they use, assigning routine, high-frequency tasks to smaller, more efficient models, while reserving frontier models for high-margin, complex reading tasks.

“Enterprise leaders can exploit falling token [costs] by developing a roadmap that targets progressively larger, high-value domain problems to solve, planning ahead to ensure they have the data they need for tuning,” added Sommer.

Our Deeper View

The rise in popularity of agentic solutions like OpenClaw has led to a higher demand for tokens and soaring costs. While leading-edge developments in the AI space are happening quickly, to take advantage of them, enterprises need to be willing to invest. In a way, it is a relief that hardware developments will make AI models more affordable, but as model capabilities continue to increase rapidly, those savings won’t always trickle down. That points to a call for action many have been advocating: pausing further development to optimize workloads for token costs more effectively. This is a problem that many players in the AI ecosystem are trying to solve right now, and The Deep View covers it regularly. It ultimately comes down to not just paying Frontier Labs a premium price to use their most expensive models, but matching models to tasks to save money and improve performance and accuracy.