Google won a bankruptcy auction for part of Spirit Airlines’ enterprise dataset and software code for $10 million, per a court filing dated 14 August. It outbid Mercor.io, an AI-focused recruitment firm, which offered $7.5 million.

The stated purpose is training and product improvement — revenue management data, pricing models, booking curves, and flight behaviour data to feed Google’s AI models and travel products.

What is in it

  • 100 million emails
  • 500 million Microsoft Teams chats
  • 7.2 billion records for competitors’ flights
  • 7.5 billion passenger transaction records dating to 2008
  • more than 175,000 employee records dating to 1986
  • over 30 million recorded customer service calls
  • more than 15 million customer service chat records

Explicitly excluded: customer and loyalty data covering 97.5 million passengers, 52.4 million loyalty members, and 740,000 co-branded cardholders.

Google’s position is that it “will not receive any personal information from this dataset,” and that “any data we receive will be rigorously scrubbed of any personally identifiable information by a third party before receipt.”

The exclusion is narrower than it sounds

The excluded categories are the loyalty database and the customer database — the structured tables of names, addresses, and card numbers. Those are the obvious personal data.

The included categories are where personal information actually lives in an operating company.

100 million emails are corporate email. Corporate email contains employee names in every header, customer names and booking references in service threads, complaint correspondence, legal and HR matters, medical accommodation requests, disciplinary discussions, and every escalated passenger incident the airline ever handled. Email is the least structured and most personally revealing corpus any company holds.

500 million Teams chats are worse. Internal chat is where people speak informally about colleagues, customers, and themselves. It is the corpus most likely to contain candid, damaging, personal statements, written by people with no expectation it would be sold.

30 million recorded customer service calls are voice recordings. A voice is a biometric identifier under Illinois BIPA, Texas CUBI, and GDPR Article 9. Beyond the biometric, calls contain names, booking references, partial payment details, reasons for travel, medical assistance requests, and bereavement fare conversations, spoken aloud by identifiable people who were told the call was recorded for quality and training purposes — not for permanent transfer to a third party for AI training after their airline collapsed.

175,000 employee records back to 1986. Forty years of HR data belonging to people who mostly no longer work there, many of whom are dead, none of whom had any relationship with Google.

7.5 billion passenger transaction records since 2008. Transaction data is who flew where when, and it is the canonical example of data that is trivially re-identifiable from a handful of known trips.

The scrubbing problem

The safeguard is that a third party will strip personal information before Google receives anything. Two problems.

Google picked and paid the anonymiser. The party certifying that the data is clean was selected and compensated by the party that wants the data. This is the same structural conflict as an issuer-paid credit rating, and it produced the same known result there.

De-identification of unstructured data at this scale is not a solved problem. Removing names from a database column is tractable. Removing personal information from 100 million free-text emails and 500 million chat messages requires a system that understands context — that “the passenger in 14C who had the seizure on the Fort Lauderdale run” identifies a person, that an internal joke about a named colleague is personal data, that a booking reference is an identifier. Automated redaction over a corpus that size will have an error rate, and an error rate applied to 600 million documents is a large absolute number of leaked identities.

For audio, the difficulty is categorical rather than incremental. The voiceprint is the identifier. You cannot redact a voice from a recording of a voice and still have a recording. Transcribe-and-discard would work; retaining audio for model training does not.

And the re-identification literature is unambiguous on the structured half. Sweeney showed 87% of Americans are uniquely identified by ZIP, birthdate, and sex. De Montjoye showed four spatiotemporal points uniquely identify 95% of people in a mobility dataset. Airline itineraries are spatiotemporal points with names attached at origin. “Anonymised travel records” is a phrase that has failed every serious test applied to it.

The bankruptcy channel

The broader issue is the mechanism, not this transaction.

Bankruptcy is where privacy commitments go to die. Spirit’s passengers, employees, and correspondents gave data to Spirit Airlines under Spirit’s privacy policy. That policy — like nearly every privacy policy — permits transfer of data as part of a sale of assets. When the company fails, the estate’s obligation is to maximise recovery for creditors, and the data is an asset with a bid on it.

There is a partial safeguard. The FTC has intervened in bankruptcy data sales before, and successfully: the Toysmart case in 2000 established that selling customer data in contravention of the privacy promises made when it was collected is an unfair practice, and RadioShack in 2015 produced restrictions on what could transfer. A consumer privacy ombudsman can be appointed under the Bankruptcy Code to review such sales.

The exclusion of the loyalty and customer databases here is very likely the product of exactly that pressure. Which is worth crediting — and also worth noting that the pressure protected the structured tables and left the emails, the chats, and thirty million voice recordings on the table.

And the price is the story. $10 million. For 8 billion transaction records, 600 million messages, and 30 million voice recordings. The going rate for the complete operational history of a company that carried nearly 100 million passengers is roughly the cost of a mid-sized office building.

What it means in practice

Every privacy policy you have accepted contains a bankruptcy clause. Go and look. It will say your data may be transferred in connection with a merger, acquisition, or sale of assets. That clause is the one that governs when the company dies.

AI training demand has created a market for corporate exhaust. Emails and internal chat used to be a liability — discovery risk, storage cost. They are now inventory, because that is what language models are hungry for. Every company’s retention calculus just changed, in the wrong direction.

“Scrubbed by a third party” is an assertion, not an audit. Ask who the third party is, who chose them, who pays them, what standard they applied, and whether anyone independent verified the result. Absent answers, it is a press line.

Recorded calls are biometric data and are being treated as text. This is the least-examined part of the transaction and probably the most legally exposed.

What you can do

  1. If you flew Spirit or worked there, understand that the loyalty and customer databases were excluded — but your emails to and about the company were not. Nor were your recorded calls.

  2. Search any privacy policy for “bankruptcy,” “merger,” or “sale of assets” before you hand over sensitive data. It is the clause that survives the company.

  3. Assume every recorded customer service call is permanent and transferable. “This call may be recorded for quality and training purposes” does not say whose training.

  4. Exercise deletion rights while the company is still alive. Once an estate is in bankruptcy, the trustee’s duty runs to creditors, and a deletion request is a claim against an asset. If you live in a state with deletion rights, use them before you need them — California’s DROP is the most efficient route.

  5. Watch for a consumer privacy ombudsman appointment in bankruptcies affecting you. It is a real statutory mechanism, it is discretionary, and it works better when someone is paying attention.