Companies going all-in on AI agents should brace for sharply higher costs. Research firm Gartner predicts that inference costs per “agentic workflow” will increase more than fivefold through 2028, even as AI models themselves keep getting cheaper.
That sounds contradictory, and it is. Gartner calls the phenomenon the Inference Paradox: the more efficient and cheaper AI models become per token, the more that saving gets swallowed up by increasingly complex applications. The result is that companies’ overall AI bill keeps climbing, while the return on that spending is anything but guaranteed.
From chatbot to autonomous agent
The distinction comes down to what AI is now being asked to do. A simple chatbot reads a query and quickly returns a probable answer. An AI agent does much more: it reasons, corrects itself, negotiates with other systems, and carries out multiple steps in sequence before a task is complete.
“Where a simple chatbot must read and interpret a query and quickly respond with a probabilistically reasonable answer, an AI agent must constantly reason, negotiate, and question itself,” said Gartner Senior Director Analyst Will Sommer. All those extra steps come at a cost: routing a task to a reasoning AI agent is at least five times more expensive than a simple chatbot interaction, according to Gartner, and often far more as task complexity grows.
Three forces driving costs up
Gartner points to three trends that together produce the paradox. First, the cost economics of foundational models are improving rapidly, with tokens becoming cheaper per unit. That same efficiency gain, however, makes it attractive to shift to more powerful, more expensive models. And to complete the picture, more sophisticated AI workflows consume far more tokens than a simple chat exchange, which drives up the total bill regardless.
In short: the pace of innovation in AI capability is outrunning the falling cost curve of tokens.
No quick fix
According to Gartner, there is no simple, low-cost solution that works across the board. Companies that want to keep their AI products competitive have little choice but to build and maintain complex, multi-model ecosystems.
Sommer specifically warns against reaching for AI agents by default: “Defaulting to generic autonomous intelligence will result in unbounded costs orders of magnitude higher than those of optimised product ecosystems.”
What companies can do
To keep costs in check, Gartner advises product leaders to adopt inference tiering: deliberately routing and orchestrating tasks to the most cost-effective model capable of handling them, rather than defaulting to the most powerful, most expensive AI for every job. That does require significant effort across complex workflows; there is no easy shortcut.
Source: Gartner, press release, August 17, 2026
