NASA and IBM Research, working with several academic partners, have released the NASA-IBM Lunar Foundation Model — an open-source foundation model designed to make decades of lunar observation data usable for machine learning. The model, its code, its machine-learning-ready pretraining datasets and its benchmark collections are all public: hosted on Hugging Face, with code on GitHub and integration into the open-source TerraTorch toolkit.
The training corpus, called SomBench, is described by the team as the largest co-registered multimodal lunar corpus assembled to date: nearly 2 million tile bundles spanning 11 modalities and two spatial scales. Roughly 1 million of those are high-resolution images from the Lunar Reconnaissance Orbiter's Narrow Angle Camera at about one meter per pixel, alongside nearly 964,000 multispectral images from the Wide Angle Camera at 100 meters per pixel. NASA notes that LRO's data volume alone exceeds that of all its other planetary missions combined.
Seventeen years of continuous LRO observations form the backbone of the dataset, supplemented by data from NASA's GRAIL and Lunar Prospector missions and Japan's Kaguya orbiter — more than 30 spatially aligned data layers drawn from nine instruments across four missions. Rather than fine-tuning an existing model, the team trained from scratch on an architecture related to TerraMind, IBM's multimodal Earth-observation model, feeding imaging geometry such as illumination angles and sun position as explicit input alongside each tile.
In benchmarks across crater detection at 100-meter and 1-meter scales, polar ice deposit prediction and segmenting Irregular Mare Patches, the pretrained model matched or beat common baselines and an architecturally identical control model with random initialization. The largest gains came in ice deposit prediction, where the model cut prediction error by up to 22% compared with the best baseline, SwinV2-B, according to IBM. On coarse-scale crater detection it beat the same baseline by nearly 19% while using only half the training data.
"NASA has spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job," said Kevin Murphy, NASA's chief science data officer and acting chief data and AI officer. The argument is the familiar foundation-model pitch applied to planetary science: observation data is plentiful, labels are scarce, and a pretrained model that adapts to new tasks with a handful of labeled examples is worth more than a shelf of single-purpose algorithms.
The team is candid about the boundaries. The model is not suited to absolute geodetic positioning — in generation tests, latitude and longitude were off by dozens of degrees in some cases, and reconstructed elevation structures had the right shape but shifted absolute heights. The authors position the release as a reusable foundation for downstream tasks, not a replacement for physical measurement instruments.
The release extends a NASA-IBM collaboration governed by a Space Act Agreement that has already produced the Prithvi family of Earth-observation foundation models and Surya, a model for space-weather prediction. The pattern is deliberate: point foundation-model techniques at domain-specific archives where governments sit on mountains of unlabeled, multi-instrument data that is expensive to curate and harder to make ML-ready.
For lunar science the practical payoff is concrete — better maps of where polar ice might actually be, which shapes where landers go and what future crews might dig. For everyone else building AI on scientific data, it is a template: NASA did not publish a paper and stop there; it released the model, the corpus and the benchmarks together, which is what turns a broken archive into infrastructure.
Comments (0)
Log in to join the discussion
Log InNo comments yet