AI subscription prices have stayed flat while the work inside them quietly shrank. That is the core finding of a new study from analyst firm SemiAnalysis, which measured the actual token limits behind major AI subscription plans and converted them into API-equivalent dollar values. Its headline conclusion: Anthropic's $200-per-month Claude plan delivers roughly 5x the API-equivalent value of OpenAI's competing $200 tier on the mid-tier models both companies market as daily drivers — Claude Opus 5.5 versus GPT-6.1 Sol. "Anthropic is an overwhelmingly better deal," the report states.
The gap widened sharply last week. On September 29, OpenAI reopened its $200 Pro plan to new subscribers and simultaneously cut token allowances across every model tier roughly in half — by its own description, the change would "net out at half the dollar in API spend compared to the old Pro $200 plan." Existing subscribers keep the higher limits until October 29; new sign-ups get the reduced quotas immediately. OpenAI also flattened its ladder: Pro 100, Pro 200 and the new Pro 500 tier now offer the same tokens-per-dollar, ending the old structure where each step up roughly doubled per-dollar efficiency.
The new $500 tier does not restore the value. SemiAnalysis measured only about 21% more GPT-6 Astra capacity than the old $200 plan offered — and because OpenAI cut cached-input pricing for GPT-6.1 Sol at the same time, the Sol-class API-equivalent value on the $500 tier actually went down. The real addition is speed: UltraFast mode, running at roughly 300 tokens per second. In other words, a meaningful share of what subscribers pay for at the top tier now buys velocity, not volume.
Anthropic moved in the opposite direction. With the Opus 5.5 API price cut — input and output down 20%, cache reads down 60% — the company raised subscription Opus allowances by roughly 20% on Max plans and about 50% on Pro. That still did not fully offset the API price reductions in subscription terms, but it left far more value on the table than OpenAI's approach, which the report calls "the nuclear option of just immediately cutting to Fable-level limits across the board."
The economics behind the squeeze are brutal, and SemiAnalysis quantifies them. Subscriptions represent only about 10% of overall lab revenue but can consume over 40% of inference compute; for Anthropic, they drag blended revenue per megawatt down by roughly $36 million. A user maxing out Opus 5.5 on the $200 plan runs at an estimated -369% gross margin at full utilization, versus about +1% for a user maxing the flagship Fable 5.1 — a deliberate design that makes the cheaper model the value leader. At a more realistic 20% utilization, those figures become 6% and 80%. These are SemiAnalysis's own models, not vendor disclosures.
The methodology is more careful than a screenshot of a usage meter. The team isolated one token type per experiment — using a passage of War and Peace because some models refuse gibberish — counted tokens in steps until results converged within ±5%, and subtracted fixed overheads. The workload mix matters: about 96.6% of tokens in its agentic scenario come from cache reads, mirroring agent-style context reuse, so heavy agent users capture the 5x gap far better than chat users would.
Chinese labs are still in the subsidizing phase the American incumbents are exiting. On SemiAnalysis's dashboard, MiniMax's plans at $22, $55 and $132 per month convert to roughly $427, $1,305 and $3,088 of API-equivalent value on MiniMax-M3 — a 19x-to-24x multiple of the fee, well above OpenAI's roughly 10x-15x. Z.ai's GLM-5.3 spans about 7.7x-11.6x depending on tier, and Moonshot's Kimi K3 about 2.5x-6.7x. Third-party wrappers like Cursor and Cognition deliver worse value than first-party subscriptions for the same underlying models.
The report is honest about its metric's limits. API-equivalent value flatters models with high API list prices — a cheap-API model with more tokens can look worse while completing more work — and OpenAI's Pro plans impose no 5-hour rolling window, letting time-shifted users drain more of their monthly allowance. "It doesn't make sense to say a plan is worth $X in isolation," the authors conclude; the real unit of subscription value is the (plan, model, workload) tuple. The full dataset sits behind SemiAnalysis's paid Tokenomics subscription, so the headline figures rest on one firm's measurement — but it is a third-party measurement, and both companies' own pricing moves are a matter of public record.
Comments (0)
Log in to join the discussion
Log InNo comments yet