Halluminate, a nine-person startup in San Francisco, has raised $30 million in a Series A led by Oak HC/FT, with participation from existing investors Y Combinator, Orange Collective, FT Partners and Heavybit, plus individual researchers from Anthropic, OpenAI and Meta. The round, announced October 1, takes total funding to $38.5 million. The company did not disclose a valuation.
The business is unusual for its size. Halluminate builds benchmarks and reinforcement-learning environments for knowledge work, starting with financial services. The method is to find where a frontier model fails at a real task — building a financial model, redlining a contract, producing a client-ready deck — and then package that failure into a simulated environment the lab can train against. CEO Jerry Wu calls the company a “verticalized data research lab.” It was founded in 2024 by Wu and CTO Wyatt Marshall.
The commercial numbers are the striking part. In roughly nine months, five people took Halluminate from zero to a mid-eight-figure contracted annualized revenue run rate at positive gross margins, and the company expects to cross nine figures by the end of the year. Wu says four of the top five closed-source US AI labs are paying customers. He also says the team now builds five times as many environments a week as it did in February and gets ten times the effective output from them.
The public evidence of the problem it is selling arrived in August, when Halluminate released Westworld Due Diligence, a benchmark built from 88 tasks drawn from anonymized private-equity transactions and written and reviewed by practising deal professionals. Seven frontier models ran it; the highest average score was 51%. One task asked an agent to redline a statement of work while working through a 160-file data room, 21 emails spread across nine threads and four sets of meeting notes — and to track terms that changed over weeks while leaving provisions meant to stay untouched alone. Across the benchmark, models repeatedly failed to carry instructions through to the final deliverable, dropping required changes, using the wrong analytical method, or leaning on information that had since been superseded. The scores are Halluminate's own and have not been independently replicated.
The investment case rests on where model training goes next. Oak HC/FT general partner Matt Streisfeld argues finance offers an unusually broad span of complex knowledge work, from banking and private equity to consulting and accounting, and that as agents take on work stretching from hours into days, the quality of specialized training environments matters more than raw model capability. “When the agent starts getting into long horizon work,” he said, “testing work and specialization will really be key.” The firm's own thesis note puts lab spending on training data at roughly $7 billion a year today, on a path to several times that by 2030 as budgets shift from static annotation to agentic reinforcement learning. Scale AI wrote in February that nearly half its new data-training projects involve reinforcement-learning environments.
The field is already consolidating around that premise. Deeptune, which built simulated work environments for training agents, raised a $43 million Series A led by Andreessen Horowitz in March and agreed to be acquired by Mercor four months later. Scale AI sells its own environments product. Halluminate's counter is depth in one vertical rather than breadth, and it is deliberately concentrating on a handful of frontier labs instead of chasing enterprise customers.
That focus has a cost. Four buyers account for the bulk of the revenue in a market where each of them can build environments in-house, switch vendors, or fund a competitor. And as models improve, the environments have to get harder to stay useful. Wu has a name for the pressure: the “Moore's law of environments.” He estimates the complexity of his company's environments needs to roughly double every six to eight months to keep pushing frontier models, which means the IP is less any single environment than the machinery for producing the next generation of them. For a nine-person team, that is both the pitch and the treadmill.
Comments (0)
Log in to join the discussion
Log InNo comments yet