Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

A Nine-Person Startup Sells AI Training Environments to Four of the Five Big US Labs. Now It Has $30M

A Nine-Person Startup Sells AI Training Environments to Four of the Five Big US Labs. Now It Has $30M

Halluminate, a nine-person San Francisco lab, raised $30 million in a Series A led by Oak HC/FT, taking its total funding to $38.5 million. It builds benchmarks and reinforcement-learning environments that teach models complex financial work, and says four of the top five closed-source US AI labs are paying customers. On its own due-diligence benchmark, seven frontier models averaged no better than 51%.

Halluminate, a nine-person startup in San Francisco, has raised $30 million in a Series A led by Oak HC/FT, with participation from existing investors Y Combinator, Orange Collective, FT Partners and Heavybit, plus individual researchers from Anthropic, OpenAI and Meta. The round, announced October 1, takes total funding to $38.5 million. The company did not disclose a valuation.

The business is unusual for its size. Halluminate builds benchmarks and reinforcement-learning environments for knowledge work, starting with financial services. The method is to find where a frontier model fails at a real task — building a financial model, redlining a contract, producing a client-ready deck — and then package that failure into a simulated environment the lab can train against. CEO Jerry Wu calls the company a “verticalized data research lab.” It was founded in 2024 by Wu and CTO Wyatt Marshall.

The commercial numbers are the striking part. In roughly nine months, five people took Halluminate from zero to a mid-eight-figure contracted annualized revenue run rate at positive gross margins, and the company expects to cross nine figures by the end of the year. Wu says four of the top five closed-source US AI labs are paying customers. He also says the team now builds five times as many environments a week as it did in February and gets ten times the effective output from them.

The public evidence of the problem it is selling arrived in August, when Halluminate released Westworld Due Diligence, a benchmark built from 88 tasks drawn from anonymized private-equity transactions and written and reviewed by practising deal professionals. Seven frontier models ran it; the highest average score was 51%. One task asked an agent to redline a statement of work while working through a 160-file data room, 21 emails spread across nine threads and four sets of meeting notes — and to track terms that changed over weeks while leaving provisions meant to stay untouched alone. Across the benchmark, models repeatedly failed to carry instructions through to the final deliverable, dropping required changes, using the wrong analytical method, or leaning on information that had since been superseded. The scores are Halluminate's own and have not been independently replicated.

The investment case rests on where model training goes next. Oak HC/FT general partner Matt Streisfeld argues finance offers an unusually broad span of complex knowledge work, from banking and private equity to consulting and accounting, and that as agents take on work stretching from hours into days, the quality of specialized training environments matters more than raw model capability. “When the agent starts getting into long horizon work,” he said, “testing work and specialization will really be key.” The firm's own thesis note puts lab spending on training data at roughly $7 billion a year today, on a path to several times that by 2030 as budgets shift from static annotation to agentic reinforcement learning. Scale AI wrote in February that nearly half its new data-training projects involve reinforcement-learning environments.

The field is already consolidating around that premise. Deeptune, which built simulated work environments for training agents, raised a $43 million Series A led by Andreessen Horowitz in March and agreed to be acquired by Mercor four months later. Scale AI sells its own environments product. Halluminate's counter is depth in one vertical rather than breadth, and it is deliberately concentrating on a handful of frontier labs instead of chasing enterprise customers.

That focus has a cost. Four buyers account for the bulk of the revenue in a market where each of them can build environments in-house, switch vendors, or fund a competitor. And as models improve, the environments have to get harder to stay useful. Wu has a name for the pressure: the “Moore's law of environments.” He estimates the complexity of his company's environments needs to roughly double every six to eight months to keep pushing frontier models, which means the IP is less any single environment than the machinery for producing the next generation of them. For a nine-person team, that is both the pitch and the treadmill.

Comments (0)

Log in to join the discussion

Log In

No comments yet