What Data Did Google Buy and Why?
The data acquired by Google isn't about passenger names or credit card numbers. Instead, it’s a deep dive into the operational soul of an airline. The package includes a staggering 100 million internal emails, 500 million Microsoft Teams chats, 30 million lines
of software code, and billions of records related to pricing, competitor flight data, and employee productivity dating back decades. For Google, this is a goldmine. While the company has scraped much of the public internet, this kind of proprietary corporate data is unique. It provides a real-world, complex look at how a large organization functions, communicates, and solves problems—invaluable information for training more sophisticated AI models.
The Great AI Data Rush
This deal is a prime example of the tech industry’s escalating hunt for high-quality training data. As AI companies push their models to perform more complex reasoning and tasks, they are quickly exhausting publicly available information. This has led to a new frontier: the acquisition of private corporate datasets. Companies are realizing that decades of their internal records—from emails to project management files—are now incredibly valuable assets. A Mercor spokesperson noted that these records show “how real work gets done” and are now some of the most valuable materials for training the next generation of AI. This rush for unique data is also why tech companies have reportedly turned to buying and scanning vast quantities of physical books to feed their models.
Google’s Strategic Play
For a relatively modest $10 million, Google has acquired a dataset that could supercharge its ambitions in several areas. The most obvious application is in the travel sector. This data could enhance Google's existing products like Google Flights with better predictive models for pricing and logistics. More broadly, the data on internal communications and workflows could be used to train Google's enterprise AI offerings, making tools within Google Workspace smarter and more attuned to corporate environments. This isn't just about understanding an airline; it's about learning the language and logic of a complex modern business, a skill Google can then sell to other industries.
A Blow to the Underdog
The auction wasn't a one-horse race. Google's final $10 million bid beat a competitive $7.5 million offer from Mercor, an AI startup that connects experts with AI training projects. The bidding war, which started around $5 million, shows just how coveted this new class of data has become. While Mercor is a significant player in the AI space, its loss to Google highlights the immense power of Big Tech's deep pockets. As the race for AI dominance intensifies, smaller firms may find it increasingly difficult to compete for the essential data resources needed to build and refine cutting-edge models.
What About Privacy?
The sale of such a vast amount of internal data naturally raises privacy concerns, especially for former Spirit employees. According to court filings and statements from Google, the deal explicitly excludes sensitive customer information like passenger profiles and loyalty program records. Furthermore, all the data will be “rigorously scrubbed of any personally identifiable information by a third party” before Google receives it. However, privacy advocates argue that even de-identified data can sometimes be re-identified, especially within a rich dataset of interconnected communications. The sale sets a new precedent for what happens to a company's digital history after it goes bankrupt, raising new ethical questions about the afterlife of workplace communications.














