News The Pentagon Says It Has Stopped Using Anthropic's AI Tools — Weeks After Claude Was Still Powering Military Intelligence Work Business a16z's Latest Consumer AI Report Finds a Split Market: 29 of the Top 50 AI Businesses by Consumer Spending Never Appear on Traffic Rankings Business OpenAI Hit With Trademark Lawsuit Over 'Astra' as TradeSun Claims Reverse Confusion Business SemiAnalysis Puts a Number on Subscription Value: Claude's $200 Plan Now Delivers Roughly 5x OpenAI's API Equivalent Opinion Only 11% of Companies Can Predict What AI Will Cost Them — and New Research Says the Cheaper Model Often Costs More Security Safeworld Raises $12.2 Million to Find Out What Generative-AI Robots Do Before You Deploy Them News Huawei Says It Has Mass-Produced 381 Chips Under Its 'Tau Scaling Law' — and Kirin 9050 Pro Is the Proof Business DeepSeek's Funding Round Doubles to $12 Billion as Tencent and CATL Anchor It — and Term Sheets Could Push It to 100 Billion Yuan News The Pentagon Says It Has Stopped Using Anthropic's AI Tools — Weeks After Claude Was Still Powering Military Intelligence Work Business a16z's Latest Consumer AI Report Finds a Split Market: 29 of the Top 50 AI Businesses by Consumer Spending Never Appear on Traffic Rankings Business OpenAI Hit With Trademark Lawsuit Over 'Astra' as TradeSun Claims Reverse Confusion Business SemiAnalysis Puts a Number on Subscription Value: Claude's $200 Plan Now Delivers Roughly 5x OpenAI's API Equivalent Opinion Only 11% of Companies Can Predict What AI Will Cost Them — and New Research Says the Cheaper Model Often Costs More Security Safeworld Raises $12.2 Million to Find Out What Generative-AI Robots Do Before You Deploy Them News Huawei Says It Has Mass-Produced 381 Chips Under Its 'Tau Scaling Law' — and Kirin 9050 Pro Is the Proof Business DeepSeek's Funding Round Doubles to $12 Billion as Tencent and CATL Anchor It — and Term Sheets Could Push It to 100 Billion Yuan

Only 11% of Companies Can Predict What AI Will Cost Them — and New Research Says the Cheaper Model Often Costs More

Only 11% of Companies Can Predict What AI Will Cost Them — and New Research Says the Cheaper Model Often Costs More

A Wall Street Journal-reported survey of nearly 400 businesses found only 11% could accurately forecast their AI spending. AI costs like a worker, not software. New analysis of more than 6,800 agent tasks found that in about 32% of model comparisons, the cheaper-per-token model ended up costing more — one task saw a budget model burn nearly 1,000 steps and $14 where a pricier model spent about 85 steps and $1.

American businesses spent the past two years encouraging their employees to use as much AI as possible — a practice that earned its own nickname, #tokenmaxxing. Then came the bill, the Wall Street Journal's Jillian Vordick and Stephanie Stamm reported on Monday. The hard part is that the bill is nearly impossible to predict: in a recent survey of nearly 400 businesses, only 11% said they were able to accurately forecast their AI spending.

The reason is that AI does not cost like software. A traditional application does roughly the same amount of work per user per month, which is why per-seat pricing works. An AI agent takes actions, makes decisions, and sometimes makes mistakes and starts over — so the cost of a "task" depends on which path the model takes, how many tools it calls, and how many times it retries. Companies budgeting AI like a SaaS line item are modeling the wrong thing.

The clearest evidence comes from an analysis by researchers at Stanford, Carnegie Mellon, UC Berkeley and Microsoft Research covering more than 6,800 agent tasks. It found that in roughly 32% of model comparisons, the model that was cheaper per token ended up costing more in total — because cheap models tend to take more steps, retry more often, or fail outright. In one case, a budget-tier Gemini Flash model ran nearly 1,000 steps without completing a task and consumed about $14 in tokens; the more expensive Gemini Pro finished the same task in about 85 steps at a total cost of about $1.

Variance is not just a cheap-model problem. Microsoft's own research found that the same model running the same task can differ by up to 30 times in token consumption between runs — and that spending more tokens does not necessarily improve the success rate. For finance teams used to variance measured in single-digit percentages, that is not noise; it is a different cost species.

Adoption makes the stakes concrete. The Ramp AI Index found that 43.5% of U.S. businesses were paying for Anthropic's subscriptions or tokens as of July, and that spending is extraordinarily concentrated: the top 1% of businesses spent a median of $7,400 per employee on AI tools in July, against $11.95 for the median company — a spread of more than 600 times between the heaviest users and everyone else.

Consultants are formalizing the playbook. McKinsey, which published its 2026 State of AI research alongside a briefing on agentic AI economics this week, points out the paradox: models matching GPT-4's capability now cost a fraction of the $60 per million output tokens that GPT-4's 8K-context API charged in 2023, yet enterprise AI spend keeps climbing, because the volume of work given to agents has exploded. Partner Tanguy Catlin's three levers: get visibility into spend by use case, team, agent, model and person; optimize workflows by routing tasks to appropriately sized models, caching reusable context and cutting unnecessary tool calls and agent loops; and tighten procurement by cleaning up idle licenses, managing quotas and renegotiating terms. "There is no single lever for cost control," he said. Colleague Lari Hämäläinen offered the working test for whether an agent is worth its bill: if a human needs an hour and verifying the agent's output takes six minutes, the agent creates value as long as its task success rate exceeds 10%.

What it means: the unit of AI procurement is quietly shifting from the seat to the task. As companies discover that sticker price per token is a poor predictor of total cost, expect model routing — sending complex work to frontier models and routine work to small ones — to become standard enterprise infrastructure rather than an optimization afterthought. The vendors that win in this regime will not be the cheapest per token; they will be the ones that finish the task in the fewest steps.

Comments (0)

Log in to join the discussion

Log In

No comments yet