ChatGPT GPT-6 Is Now the Default for All of ChatGPT's 1.2 Billion Weekly Users, With Answers That Render Interactive UIs News Healthleap Raises $38 Million for AI That Reads Hospital Charts Overnight to Catch Missed Malnutrition and Delirium Business Broadcom Seeks More Than $50 Billion to Finance OpenAI Custom Chips, With Apollo and Blackstone in Early Talks Claude Anthropic Cuts Haiku 5.5 API Prices by Up to 90% to Match GPT-6 Luna, and Gives It a 1M-Token Context Window Gemini Google Opens SynthID to Everyone: A Public Website That Tells You Whether an Image, Video or Audio Clip Is AI-Generated Meta Meta's AI Caught 97% of 33.2 Million Child-Exploitation Items Before Users Reported Them — Now an LLM Hunts "Signposting" Ads Business CoreWeave Enters India With 240 MW in Mumbai: Three Buildings at Adani's Taloja Campus Will Run NVIDIA Vera Rubin From 2028 Microsoft Copilot Microsoft Bets Windows on Hybrid Intelligence: MAI Code 1.1 Flash Comes to Win11 as the $5,999 Surface RTX Spark Dev Box Opens for Preorder ChatGPT GPT-6 Is Now the Default for All of ChatGPT's 1.2 Billion Weekly Users, With Answers That Render Interactive UIs News Healthleap Raises $38 Million for AI That Reads Hospital Charts Overnight to Catch Missed Malnutrition and Delirium Business Broadcom Seeks More Than $50 Billion to Finance OpenAI Custom Chips, With Apollo and Blackstone in Early Talks Claude Anthropic Cuts Haiku 5.5 API Prices by Up to 90% to Match GPT-6 Luna, and Gives It a 1M-Token Context Window Gemini Google Opens SynthID to Everyone: A Public Website That Tells You Whether an Image, Video or Audio Clip Is AI-Generated Meta Meta's AI Caught 97% of 33.2 Million Child-Exploitation Items Before Users Reported Them — Now an LLM Hunts "Signposting" Ads Business CoreWeave Enters India With 240 MW in Mumbai: Three Buildings at Adani's Taloja Campus Will Run NVIDIA Vera Rubin From 2028 Microsoft Copilot Microsoft Bets Windows on Hybrid Intelligence: MAI Code 1.1 Flash Comes to Win11 as the $5,999 Surface RTX Spark Dev Box Opens for Preorder

Anthropic Cuts Haiku 5.5 API Prices by Up to 90% to Match GPT-6 Luna, and Gives It a 1M-Token Context Window

Anthropic Cuts Haiku 5.5 API Prices by Up to 90% to Match GPT-6 Luna, and Gives It a 1M-Token Context Window

Anthropic released Claude Haiku 5.5, its fastest and cheapest small model: $0.10/$0.50 per million tokens on prompts under 100K (a 90% cut that matches GPT-6 Luna), a 1M-token context window, and the first Haiku-class effort controls. Sonnet 5.5 cache reads were halved, and Max and Team subscribers get monthly API credits.

Anthropic released Claude Haiku 5.5 on Wednesday, the third member of its Claude 5.5 family to ship in roughly a month after Opus 5.5 and Sonnet 5.5. The company is positioning it as its fastest and cheapest small model ever, and the pricing move is aggressive enough to redraw the low end of the API market: for prompts under 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens — a roughly 90% reduction from Haiku 4.5, which charged $1 and $5.

That price exactly matches OpenAI's GPT-6 Luna, the small model OpenAI launched on September 22 at the same $0.10/$0.50 rates. Above the 100,000-token threshold, Haiku 5.5 steps up to $0.50 input and $2.50 output, still a 50% cut from its predecessor. Cache reads drop to $0.01 per million tokens on short prompts. Anthropic says about 90% of requests to the previous Haiku fell under the 100K line, which is why its headline figure — around 75% cheaper on average — sits below the per-token reduction once real workloads and tokenization are factored in.

The capability ceiling moved too. Haiku 5.5 carries a one-million-token context window, five times the 200,000 tokens Haiku 4.5 shipped with a year ago, with a 128,000-token standard output limit (300,000 available in beta) and text-and-image input. It is also the first Haiku-class model with an adjustable effort setting, letting developers dial a single request toward lower cost or higher intelligence instead of switching model tiers. The model is live now as claude-haiku-5-5 on the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure.

Anthropic's own benchmarks — which have not been independently verified — put Haiku 5.5 at 72.4% on the offline subset of OSWorld 2.1, ahead of GPT-6 Luna at 48.9% and even Sonnet 5.5 at 70.6%, and at 39.2% versus 16.4% on Terminal-Bench 4.0. The company is upfront that Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding; Haiku is aimed at the high-volume, cost-sensitive work that quietly consumes API budgets: summarization, compaction, classification, database queries and subagent calls inside coding loops.

The launch came packaged with two adjacent sweeteners. Sonnet 5.5 cache reads were cut in half, from $0.20 to $0.10 per million tokens, which Anthropic says lowers the cost of most agentic workloads on that model by around 20% — cache reads are a large share of token consumption when an agent re-reads context on every step. And Claude Max and Team subscribers will start receiving monthly API credits this week for building on the Claude Platform: $100 a month for Max 5x, $200 for Max 20x, and up to $500 pooled across users on Team plans.

The timing reads as deliberate. Anthropic confidentially filed for an IPO on June 1 and is reportedly targeting a Nasdaq listing at a valuation that could approach $2 trillion, up from the $965 billion post-money mark set in its May Series H round. Three model launches in a single month, each hitting a different price and capability tier, is the kind of usage-growing cadence a company wants on the record before a listing that size. For developers, though, the practical takeaway is simpler: at a tenth of a cent per thousand input tokens, workloads that were once ruled out on cost — quality-checking every customer message rather than a sample — now pencil out.

Comments (0)

Log in to join the discussion

Log In

No comments yet