Business Sea Limited Becomes the First ASEAN Enterprise to Adopt Nvidia's Vera Rubin Platform as AI Day Singapore Showcases Regional Builds Claude Claude Opus 5.5 and Sonnet 5.5 Now Run Inside AWS GovCloud, Clearing the Way for ITAR-Regulated Workloads News Reflection AI Ships Beam: a 501B-Parameter Open-Weight Model That Claims GLM-5.2 Reasoning at 3-4x Less Inference Compute AI Agents TikTok Puts an AI Shopping Agent Inside Its Feed — With One-Click Checkout From Salesforce, Shopify and Stripe Partners Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sea Limited Becomes the First ASEAN Enterprise to Adopt Nvidia's Vera Rubin Platform as AI Day Singapore Showcases Regional Builds Claude Claude Opus 5.5 and Sonnet 5.5 Now Run Inside AWS GovCloud, Clearing the Way for ITAR-Regulated Workloads News Reflection AI Ships Beam: a 501B-Parameter Open-Weight Model That Claims GLM-5.2 Reasoning at 3-4x Less Inference Compute AI Agents TikTok Puts an AI Shopping Agent Inside Its Feed — With One-Click Checkout From Salesforce, Shopify and Stripe Partners Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third

Reflection AI Ships Beam: a 501B-Parameter Open-Weight Model That Claims GLM-5.2 Reasoning at 3-4x Less Inference Compute

Reflection AI Ships Beam: a 501B-Parameter Open-Weight Model That Claims GLM-5.2 Reasoning at 3-4x Less Inference Compute

Brooklyn startup Reflection AI officially launched Beam, its first frontier open-weight model: a 501B-parameter MoE with 23B active parameters, trained on 23.8 trillion tokens using 6,144 Nvidia GB300 chips. The company claims Beam matches Z.ai's GLM-5.2 on advanced reasoning while using 3-4x less inference compute; full Apache 2.0 weights ship later this month.

Reflection AI on Monday officially unveiled Beam, the first frontier-scale open-weight model from the Brooklyn startup founded by two former Google DeepMind researchers — confirming weekend reporting that a launch was imminent. The company calls Beam a "workhorse model" aimed at enterprises, the public sector and developers looking for a Western alternative to open models from Chinese labs.

The specs are the pitch. Beam is a sparse mixture-of-experts model with 501 billion total parameters but only 23 billion active per token, text-only, with a 1 million-token context window. It was pretrained on 23.8 trillion tokens using 6,144 Nvidia GB300 NVL72 GPUs, finishing the run in under four weeks with what Reflection reports as 92.3% goodput — a measure of how much GPU time actually went to useful computation. The company then ran more than 100 million reinforcement-learning rollouts across roughly a million synthetic coding, agentic and STEM environments, using another 10,500 GB300 chips.

On its own benchmarks — which have not been independently verified — Reflection says Beam matches Z.ai's GLM-5.2 on advanced reasoning while using 3-4x less inference compute, scores 80.9 on SWE-Bench Verified, 80.1 on Terminal-Bench v2.1 and 97.8 on AIME 2026, and outscores Thinking Machines' Inkling on four coding benchmarks where both report results. The company concedes Beam trails Moonshot's Kimi K3 on raw capability, and notes Inkling is multimodal while Beam is text-only.

Availability starts narrow: an early version is accessible through a waitlist while red-teaming finishes. Full weights are promised under an Apache 2.0 license later this month, along with a model card, documentation and the software stack needed to run and fine-tune the model, with distribution through hyperscalers and neoclouds.

The launch ends months of "has yet to ship" descriptions of a company that has raised close to $4.7 billion at a $25 billion pre-money valuation from Nvidia, Sequoia Capital and Lightspeed — Nvidia itself put $800 million into the last round. Reflection has also locked up compute: deals worth more than $7 billion with SpaceX and Nebius secure GB300 access through 2029, and it is building a 250-megawatt sovereign AI cloud in South Korea with retail conglomerate Shinsegae. It says Beam has been validated on Dell's AI Factory platform and that it is a certified participant in the Department of Energy's Genesis Mission.

The strategic logic is explicit. Reflection is positioning against closed labs like Anthropic and OpenAI, against Chinese open-weight leaders like DeepSeek, Qwen and Z.ai, and against Western open players like Mistral, Meta and Cohere. Its product concept is the "AI factory": institutions train Beam on their own proprietary data to build customized, local AI systems. Axios has reported that hedge funds and trading firms are among the most eager potential customers, and Jensen Huang has long championed the same idea — one that conveniently sells more Nvidia GPUs.

The caveat matters: every performance number published so far comes from Reflection's own technical report. The real test begins when the Apache 2.0 weights land and outside labs can run the same evaluations. If the 3-4x compute-efficiency claim survives independent scrutiny, Beam becomes the first serious American answer to the open-weight models that have set the pace; if it does not, Reflection joins the list of heavily funded labs whose benchmarks looked better than their product.

Comments (0)

Log in to join the discussion

Log In

No comments yet