The AI industry's most expensive assumption is that better models require more compute. DatologyAI is productizing an attack on that assumption. On October 9 the data curation startup launched Curation Studio, a guided platform that turns its internal frontier-data tooling into a product any model team can run — inside their own cloud, on their own data, with no bytes leaving their environment. The company describes it as the output of three years of research, hundreds of thousands of experiments, and infrastructure built to handle petabyte-scale corpora.
The pitch is that data quality is a compute multiplier. Datology points to repeated findings from its research and customer work that models trained on well-curated data match or beat models trained with 10 to 100 times more compute on uncurated data, while being smaller and cheaper to serve. The commercial logic follows: teams that curate well can build domain models that rival frontier offerings and cut inference costs, rather than renting intelligence by the token forever.
The company backed the launch with benchmark comparisons it ran itself. The baseline was deliberately strong: a hand-mixed corpus of 39 of the best publicly curated datasets — including NVIDIA's latest Nemotron releases, The Stack v3, FineWeb and DCLM, mixed following the recipes behind Nemotron, OLMo and Granite — with the curated set drawn from exactly the same sources and trained the same way. Fully post-trained, a 30B-A3B model built on curated data outscored that baseline 46.8% to 37.7%, with both models pretrained on 1 trillion tokens. A 12B-A1.4B curated model pretrained on just 0.44 trillion tokens beat the 30B baseline outright, using a fifth of the pretraining compute, and another curated configuration matched the largest baseline at 6 times less training compute. These are vendor-reported numbers and have not been independently verified.
The product walks through four stages. Clean fixes ingestion problems, applies heuristic filters and decontaminates against benchmarks so teams never train on their own test set. Calibrate scores every document, reweights the corpus toward the tasks a team cares about and drops redundant copies, balancing task-optimized data against general-purpose data to avoid overfitting. Create generates new synthetic training data by grounded rephrasing of real, high-quality documents — the approach Datology pioneered in its BeyondWeb research — rather than free-form generation. Compose mixes the resulting subsets, with mixtures varying by model size and training stage. Teams can accept the presets or fork every stage.
The flagship customer example is Thomson Reuters. Using Curation Studio's pipeline to mid- and post-train a Qwen model, the publisher produced Thomson-1, which the companies say outperformed frontier models including GPT-5.6 Sol in blind head-to-head comparisons on complex legal workflows, at 100 times lower cost per task. That figure, like the benchmarks above, comes from Datology's own reporting.
The scale behind the product is considerable. Datology says it has processed more than 23 pebibytes of data over three years, with experiments consuming over 60 PiB; a single recent run pushed 355 TiB of source data through roughly 670 TiB of intermediaries to produce a final 26 TiB corpus, orchestrated across more than 50 custom operators, 540 operator configurations and 2,460 operations. Synthetic data generation bursts onto on-demand GPU compute from Modal, scaling to hundreds of NVIDIA B300s.
Deployment is the differentiator enterprises will notice first: the curation engine runs inside the customer's VPC, so proprietary data, lineage and outputs stay in the customer's account, and heavy model work runs on whatever GPU or inference provider the customer chooses. Datology counts Thomson Reuters, hedge fund HRT, several of the Mag 7, Deepgram, Arcee and Unconventional AI among its customers. The initial release covers LLM training across web, math, code and multilingual data, with a public deep-dive session scheduled for October 27.
What the launch really signals is the commoditization of a discipline that until now only frontier labs could practice. If curation tooling of this quality becomes broadly available, the moat shifts from who can afford the most tokens to who understands their data best — a shift that favors domain owners with proprietary corpora over pure-scale players. The open question, as with all vendor benchmarks, is replication; Datology has published its methodology and promises deeper engineering write-ups, which gives the community something concrete to test.
Comments (0)
Log in to join the discussion
Log InNo comments yet