A $10 Million Data Deal
Google has successfully bid $10 million for a vast collection of corporate data from Spirit Airlines, which ceased operations in May 2026 and is liquidating its assets in bankruptcy court. The tech giant outbid Mercor, an AI-focused firm, which offered
$7.5 million, underscoring the growing demand for large-scale enterprise datasets. A Google spokesperson stated that the acquisition of this "enterprise dataset" can be "helpful in improving our products and AI models." The sale is part of a trend where AI companies are looking beyond the public internet for training material, turning to the internal operational data of companies—both living and defunct—to teach AI how to handle complex, real-world tasks.
What Kind of Data Is Included?
The dataset is enormous and diverse. It includes billions of passenger transaction records, competitor flight pricing data, and internal company communications like hundreds of millions of Microsoft Teams messages and emails. Also included are operational records related to aircraft operations, employee productivity, marketing campaigns, and even audits and fraud detection logs. Google has clarified that it is not purchasing sensitive customer lists or credit card information. However, the sheer breadth of the data, which captures nearly two decades of a major airline's operations, provides an intricate look at how a complex business functions, from internal problem-solving to external customer interactions.
The Promise of 'Anonymization'
Both Google and the court filings emphasize a critical point: the data will be 'de-identified' or 'anonymized' before Google receives it. A court-appointed third party will oversee this process, which involves scrubbing the data of personally identifiable information (PII) like names and email addresses. However, reports based on the court filings reveal some nuances. Google will choose and pay for the firm that performs the anonymization, and the process must be "reasonably satisfactory" to Google. Crucially, the agreement requires the process to maintain "referential integrity," meaning the links between different data points for a single, now-anonymous individual must be preserved. This allows the AI to trace a complete journey—from a customer complaint email to an internal maintenance log, for example—without knowing the person's actual identity.
Why This Data Is So Valuable for AI
AI models, particularly large language models (LLMs), become more capable and nuanced when trained on diverse, real-world data. While the internet provides a vast amount of text, corporate data shows how work actually gets done. Internal communications, project management records, and customer service logs contain the messy, context-rich details of human collaboration and problem-solving. This kind of information is invaluable for training AI to perform sophisticated tasks like managing customer service, debugging software, or handling complex logistics. By analyzing how Spirit's employees communicated and what actions they took, Google's AI can learn patterns that are impossible to find in public web pages or books.
The Blurring Line Between Privacy and Progress
This deal sits at the intersection of technological advancement and data privacy concerns. On one hand, using such datasets could lead to significant improvements in AI-powered services. On the other, the practice raises important questions about the fate of our data when companies go bankrupt. The Association of Flight Attendants-CWA, representing thousands of former Spirit employees, called the plan "outrageous" and filed a court objection. Privacy experts caution that even with names removed, the risk of 're-identification' is not zero. Studies have shown that by cross-referencing anonymized datasets with other available information, it can be possible to trace data back to a specific individual. As AI's hunger for data grows, this acquisition highlights a new frontier where the definition of 'personal information' and the protections around it will be continually tested.














