A Planet-Sized Data Problem
Every day, NASA's fleet of Earth-orbiting satellites beams down a staggering amount of information. These observations capture everything from the health of rainforests and the extent of polar ice to the moisture in farmland soil. This data is crucial
for understanding how our world is changing. By 2030, NASA expects its Earth science data archive to grow to 600 petabytes. To put that in perspective, a single petabyte is one million gigabytes. This deluge of information is far too vast for scientists to manually sift through. Traditional analysis methods are slow, meaning critical insights into phenomena like wildfires, floods, and droughts can be delayed. The challenge isn't a lack of information, but the overwhelming scale of it.
AI as a Powerful New Analyst
This is where artificial intelligence comes in. AI, particularly machine learning, excels at finding patterns in enormous datasets that would be invisible to the human eye. Think of it like a tireless research assistant who can read millions of documents or scan billions of images in a fraction of the time it would take a person. By training AI models on this satellite data, scientists can automate the process of identifying key environmental changes. This allows researchers to move from the painstaking task of data processing to the more critical work of discovery and interpretation.
Introducing Foundation Models
To supercharge this effort, NASA has partnered with tech giant IBM to build a new type of AI known as a "foundation model". These are massive, versatile AI systems trained on broad sets of unlabeled data. The collaboration has produced a family of models known as Prithvi, the Sanskrit word for Earth. One model was trained on years of NASA's Harmonized Landsat and Sentinel-2 (HLS) satellite imagery, creating a powerful geospatial tool. Another was trained on 40 years of climate and weather data from NASA's MERRA-2 dataset. The goal is to create flexible, reusable AI systems that can be easily adapted for a wide range of scientific tasks without starting from scratch each time.
From Data to Real-World Action
So, what can these AI models actually do? The applications are vast and vital. Scientists have already fine-tuned the Prithvi models for specific, high-impact tasks. These include mapping floodwaters after a storm, identifying burn scars left by wildfires, and classifying different types of land use, such as distinguishing between crops and forests. Future applications could include predicting severe weather patterns, tracking changes in wildlife habitats, and even providing early warnings for locust breeding grounds, as one research group has already demonstrated. This allows for a much quicker response to natural disasters and more effective management of natural resources.
An Open-Source Future
Crucially, NASA and IBM are making these powerful tools open source, meaning the models and the data they were trained on are freely available to the global research community. This is part of NASA's broader Open-Source Science Initiative, which aims to make scientific knowledge more accessible, transparent, and collaborative. By releasing these tools on public platforms like Hugging Face, the agencies are empowering a global community of scientists, startups, and public entities to build their own applications. This democratic approach not only accelerates discovery but also fosters innovation in ways the original creators might not have even imagined.














