Per-token AI prices have fallen 41% since March 2026, according to the Silicon Data Token Expenditure Index reported this week. On paper, that should be the best news finance teams have had all year. In practice, enterprise AI bills are still climbing, and some are climbing faster than the price cuts are landing. JPMorgan's response, reported alongside the same story, was blunt: engineers are now capped at $2,000 a month of Claude usage, full stop. Uber's CTO said something similar back in August — the company blew through its entire 2026 AI budget and is done treating token volume as a productivity metric. The instinct behind both moves is understandable. The fix is still wrong, and the reason why is worth thirty seconds of attention before your team sets next quarter's AI budget the same way.
The Paradox: Cheaper Tokens, Bigger Bills
A falling per-token price only lowers your bill if usage stays flat. It doesn't. Every enterprise rolling out agents this year is also expanding what those agents are allowed to do — more tools, longer context windows, more autonomous multi-step runs instead of single Q&A calls. Volume grows faster than price falls, and a meaningful share of that volume is waste: re-derived context, retried steps after a bad tool call, verbose intermediate reasoning nobody reads. A 41% price cut on a workload that's grown 70% and gotten sloppier isn't a 41% saving. It's a bill that goes up while every line item on the invoice says the unit price went down — which is exactly the kind of number that makes a CFO stop trusting the AI team's projections.
Why a Hard Cap Is the Wrong Kind of Fix
JPMorgan's $2,000-a-month ceiling is a governance decision made by someone who doesn't yet have the visibility to make a better one. A flat per-engineer cap treats every token the same, regardless of whether it's running a production support agent generating revenue-relevant resolutions or a background script re-summarizing the same document for the fifth time that day. It will stop the worst offenders. It will also throttle the engineer whose agent workload is legitimately heavier because their team's use case is more valuable, and it does nothing to fix the underlying waste — it just makes people hit the ceiling and ask for an exception, which puts the same unmeasured spend back on someone's desk a month later.
The number behind the cap
New research led by Stanford's Erik Brynjolfsson found that frontier models misjudge their own token consumption by up to 30x, run to run, on functionally identical tasks. If the model itself can't reliably predict what a task will cost before running it, a flat monthly cap set by a human is a guess wearing a policy's clothing — and the fix isn't a bigger or smaller number, it's actually measuring the thing.
What Most Enterprises Don't Actually Know
Ask most AI teams for their cost breakdown and you'll get a total. Ask which workflows are driving it, which of those are producing measurable business value, and which are pure overhead, and the conversation usually stalls. That gap is the real problem a token price drop can't solve and a spending cap can't patch:
- Which specific agents, workflows, or teams are responsible for which share of the monthly bill — not a total, a breakdown
- Whether a given workflow's cost per successful outcome is trending down as the model gets cheaper, or up because the workflow itself has gotten less efficient
- How much of the spend is repeated context that prompt caching would make nearly free, versus genuinely novel work
- Whether tasks are being routed to a flagship model by default when a smaller, cheaper model would resolve them at the same accuracy
- What happens to spend the month after a cap is introduced — does waste actually drop, or does everyone just get better at requesting exceptions
None of these are exotic questions. They're the same breakdown a finance team would demand for any other line item on the P&L — cost center, driver, trend, and the value it's producing in return. The only reason AI spend gets a pass on this scrutiny is that most billing dashboards stop at a single monthly total by design, and most teams adopted agents faster than they built the instrumentation to answer questions the total doesn't. That's not a model-cost problem. It's an observability gap, and it's the same gap whether your per-token price fell 41% or rose 41% — the bill just makes the gap easier to ignore when prices are moving in your favor.
A Realistic Scenario
A Wizeb client — a logistics software provider running AI agents for shipment exception handling — saw their monthly AI bill rise for three straight months despite their model provider cutting per-token prices twice in that window. Their instinct, like JPMorgan's, was to set a hard monthly ceiling per team. Before doing that, we ran a workflow-level cost audit and found the real driver: one agent, responsible for re-checking shipment status against three separate carrier APIs, was re-fetching and re-summarizing the same carrier response inside its own reasoning loop up to four times per exception, because its context wasn't being cached between calls. That single workflow accounted for 38% of the entire AI bill. We added prompt caching on the repeated carrier-response context, routed the status-recheck step to a smaller model, and added per-workflow cost dashboards so the team could see cost-per-resolved-exception in real time instead of a lump monthly total. The bill dropped 52% in the first month — no cap, no exceptions process, no engineer told to use the tool less.
How Wizeb Approaches This
We don't start an AI cost engagement by asking a client what number they want the bill to hit. We start by building the visibility that JPMorgan's cap and most flat budgets are missing: per-workflow and per-agent cost breakdowns tied to the business outcome each workflow produces, not just a monthly total. From there, the actual fixes — prompt caching, model routing, and eliminating redundant calls — are well understood and fast to implement once you know where to point them. A price cut from your model provider should show up as savings, not get absorbed by waste nobody was measuring. That's the difference between a budget cap and a budget you actually understand. Start at wizeb.com/services/token-optimization.
Know where your AI budget actually goes
Wizeb's AI cost optimization engagement builds per-workflow spend visibility, then applies prompt caching, model routing, and redundant-call elimination to the workflows actually driving your bill — typically a 40–70% reduction, without a single hard cap. Visit wizeb.com/services/token-optimization to start the conversation.
Three Questions Before You Set Next Quarter's AI Budget
- 1Do you know which specific workflow or agent is responsible for the largest share of last month's AI bill — or only the total?
- 2If your model provider cut prices again next month, would your bill actually go down, or would usage simply absorb the difference?
- 3If you set a hard spending cap tomorrow, would it fix the waste underneath it — or just move the same unmeasured spend into an exceptions queue?
