Getting a grip on shadow tokens and AI blowouts

And these specifics aren’t immediately apparent at pilot. Tools can appear inexpensive in controlled experiments yet unpredictably scale depending on session length, context window size, model selection and whether agents run in parallel. This is the fallacy of the $20-per-seat enterprise plan — tokens are charged separately at API rates with no ceiling. The final dollar value of any session is set by factors that finance can’t always model in advance, particularly when these decisions usually rest with the engineers themselves.

According to Deloitte, only 21% of organizations deploying agents have a mature governance model, a real concern because they’re token-eating machines. This is what was happening at Uber — Claude Code in agentic mode was autonomously reading codebases, planning changes across dozens of files and opening pull requests. Each step quickly adds up, with Anthropic’s own documentation noting that agents consume approximately seven times as many tokens as standard sessions.

This is shadow IT and shadow AI, evolved. This time, however, many leaders approved the tool in question without guardrails governing consumption. AI hype adds fuel to the fire and normalizes long sessions. Uber’s CTO, for example, described a company-wide shift toward “agentic software engineering” with employees “who are quietly experimenting, quietly shipping and quietly pushing things forward”. This is an exciting way to test the limits of what’s possible, certainly, but it’s also a position that goes a long way to explaining how the company spent its annual AI budget by April.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *