LLM Token Cost Drops, Agentic Work Expense Climbs

TL;DR. The cost of LLM inference per token has fallen dramatically, while multi-step agentic AI tasks are becoming significantly more expensive to run. - LLM inference costs dropped tenfold annually, faster than historical compute or bandwidth commoditization. - Agentic tasks consume 1000 times more tokens than simpler queries, leading to unpredictable and high operational costs. - Hidden token sequences in reasoning models further increase consumption, five to twenty times standard model use. - Enterprise-scale operations face token bills that outpace infrastructure spend due to explosive total consumption.

Sources

Back to QLANKR News