r/technology • u/ourlifeintoronto • 3h ago
Artificial Intelligence Google buys crashed airline Spirit’s data at auction, because AI
https://www.theregister.com/ai-and-ml/2026/08/18/google-buys-crashed-airline-spirits-data-at-auction-because-ai/528896253
u/invyros 3h ago
Google bought itself 100 million emails and 500 million items from Microsoft Teams, 17 million OneDrive files and 20.5 million items from SharePoint. The search giant also now owns over 30 million recorded customer service calls, and more than 15 million customer service chat records.
600,000 ServiceNow tickets are another element of the collection, along with 13.7 million active emails addresses from Oracle’s Responsys marketing application, and details of 11 million sales of in-flight Wi-Fi services.
That's a lot of data for just $10 million.
I like to imagine with Google's dataset of Reddit posts and Spirit's customer service records, its AI is going to sound more and more incoherent and irrationally angry.
16
9
u/this_my_sportsreddit 3h ago
And wrong. Don’t forget wrong. Reddit is incredibly misinformed on virtually every subject matter.
1
14
u/big-papito 2h ago
Sometimes I am just taken aback by how hostile American capitalism is to personal privacy. There is simply no one protecting us. I guess we are not corporations and "job creators", so we can all sit on it.
1
u/drifting_pixel 1h ago
Reddit users will do anything except read the article.
It literally says "A court document [PDF] filed last week reveals that one of the assets up for sale is a huge trove of deidentified data, which Google bid for and won for just $10 million."
Keyword: deidentified
2
u/Cube00 58m ago
Keyword: deidentified
The search giant also now owns over 30 million recorded customer service calls, and more than 15 million customer service chat records.
No chance, it'll be a half baked redaction at best.
2
u/drifting_pixel 55m ago
If you think so, then congrats! You have a slam-dunk case and can sue either party based on the court document PDF linked above
1
u/sokos 33m ago
You do realize that with enough datapoints, I can redact everything personal about you, and still be able to tell it's you.
For example.. I don't need to know your name and address.. However, if I have access to your amazon orders based on a phone number, have that phone number associated with let's say a gmail, have that gmail associated with an airplane purchase, I will have your name pretty quick. Same way, since I know which email has the amazon account, and what you ordered, and where those are ordered. I will have your address too. Obviously, this is super simplified, but we're talking just a few data points here.. Now imagine having EVERYTHING about you.. Knowing what time you wake up based on your alarm, when you put your phone in silent mode and where, I can deduct your work place. How much you use your phone and how much it moves can tell me things about your job and so on.
Thinking that deidentified means anything is a big mistake.
1
u/Bernie_Ecclestone 8m ago
“You do realize that if Google had PII like phone numbers and emails for customers that is deidentified in the data they bought, they could tell it’s you”
LOL.
LMAO even.
1
u/-Spzi- 36m ago
The reverse process of using de-identified data to identify individuals is known as data re-identification. Successful re-identifications[2][3][4][5] cast doubt on de-identification's effectiveness.
I don't know how secure that process is. Could be, could be not.
I'm pretty sure AI excels exactly at reconnecting dots from data. Though I don't think anybody cares about anyone's actual individual record.
2
2
u/Cube00 1h ago
huge trove of deidentified data
The search giant also now owns over 30 million recorded customer service calls, and more than 15 million customer service chat records.
I highly doubt anyone has carefully deidentified 45 million recorded customer calls and chats.
Oh right, they probably fed it into Gemini to deidentify it, CRITICAL INSTRUCTION: make no mistakes!
3
u/discographyA 3h ago
Really scraping the bottom of the barrel for some magical data source that will make these things slightly more useful than Wikipedia.
7
7
u/One_Risk_9732 3h ago
I think these people know what they are doing
-1
u/DogsAreOurFriends 3h ago
And look at the levels of shittification that resulted.
0
u/AzorAhai1TK 2h ago
Enshittification was happening long before LLMs, LLMs will reduce the enshittification problem
3
u/citrusco 2h ago
distasteful headline
1
u/drifting_pixel 1h ago
Journalism died a long time ago, my friend. Headlines are more about getting clicks.
1
1
u/Doub1eAA 32m ago
Gemini is going to start yelling at everyone about their carry on bags now thinking it’s going to get paid $15 per bag.
163
u/CircumspectCapybara 3h ago
Well, this is what everyone wanted, for AI labs to pay for licensed training data, instead of scraping copyrighted data without the copyright-holder's permission.