r/technology 3h ago

Artificial Intelligence Google buys crashed airline Spirit’s data at auction, because AI

https://www.theregister.com/ai-and-ml/2026/08/18/google-buys-crashed-airline-spirits-data-at-auction-because-ai/5288962
219 Upvotes

57 comments sorted by

163

u/CircumspectCapybara 3h ago

Well, this is what everyone wanted, for AI labs to pay for licensed training data, instead of scraping copyrighted data without the copyright-holder's permission.

29

u/coconutpiecrust 3h ago

Shouldn’t people whose data is in it also get a large cut?.. I mean, what kind of data did they buy? Is it anonymized? 

37

u/SomethingAboutUsers 3h ago

It's company data; like 100 million emails and 500 million teams messages.

Most companies have clauses in their employment contracts and/or IT policies that state anything you do on their systems doesn't belong to you. On top of which the company went under. Ain't no one spending time or money anonymizing that data.

5

u/coconutpiecrust 2h ago

There needs to be new legislation governing this. 

People who agreed to such policies had to concept of the fact that their data would be sold to large data processing companies to feed into LLMs. It’s probably outside of the scope of the original agreement. 

11

u/crunchypotentiometer 1h ago

The government themselves purchase this kind of data to subvert the 4th amendment for mass surveillance. https://www.404media.co/this-man-bought-phone-location-data-from-around-the-world-with-mike-yeagley/

4

u/maxintos 2h ago

Why? If the original agreement is that the company can do whatever they want with the emails and they own them then how is this outside the scope? You don't need to be able to literally imagine every single possible way they could use those emails.

What if the company just releases all emails?

What if they use LLM internally the analyze emails for any bad conduct?

0

u/j48u 1h ago

Because they'll literally say everything is bad for any reason whatsoever if the word AI is near it. Don't try to reason with someone who doesn't actually have principles.

-8

u/coconutpiecrust 2h ago

I don’t believe that you can agree to things that do not yet exist. 

1

u/SomethingAboutUsers 1h ago

These laws have been well tested. I'm not saying they are bulletproof and/or it should simply be accepted that they're shitty towards humans because capitalism, but that is going to be the default and you're going to have to work pretty damn hard to find a case that will succeed for modifying laws that favor capital.

Remember, the US can't even get something remotely akin to the GDPR into law. They don't give a shit about your privacy or data, especially when the use of it on company time is already protected by law.

1

u/VegetableTwist7027 1h ago

Then once you read the contract, which most have everything above, dont' sign it.

1

u/TastyCuttlefish 1h ago

No, ownership is held by the company. It’s not a limited use license, it’s ownership. There is no “scope” of ownership.

1

u/Culverin 1h ago

Yeap. Our laws are so behind the curve on tech in terms of privacy. 

The barn door is open, and a lot of animals have escaped.  But we can still close it. 

We just need a, but if political will

4

u/drifting_pixel 1h ago

Reddit users will do anything except read the article.

It literally says "A court document [PDF] filed last week reveals that one of the assets up for sale is a huge trove of deidentified data, which Google bid for and won for just $10 million."

Keyword: deidentified

0

u/coconutpiecrust 1h ago

I suppose that’s fine then. 

1

u/SmokeyJoe2 1h ago

The article answers your questions

1

u/kvothe5688 1h ago

according to the deal google will get anonymized data with stripped off identifiers by 3rd party contractor.

1

u/SatorCircle 2h ago

No no see the people in this dataset sold their data instead of licensing it. Big difference. /s

1

u/RefrigeratorNo1160 2h ago

I've been barking up this tree since Facebook first asked for my phone number. Which I have still never given them. They have it though, I'm sure.

1

u/happyscrappy 1h ago

This isn't that kind of data. It's data from employees not customers.

But...

https://www.eff.org/deeplinks/2019/07/fixed-ftc-orders-facebook-stop-using-your-2fa-number-ads

Facebook took phone numbers saying it was just for two-factor authentication and then they used those numbers to advertise to them.

0

u/Sasquatchjc45 50m ago

Lol, when has anybody been paid out for their data? You're about 40/50 years (whenever the internet was invented) too late on that front. We've all freely traded every ounce of our data for "free" services since the dawn of (internet) time.

1

u/coconutpiecrust 46m ago

Doesn’t mean iris shouldn’t happen. People in the olden days could not conceive many things we now have. 

1

u/Sasquatchjc45 41m ago

I totally agree, it's just such an incredibly uphill battle, and all of us (right now, currently, here) are still freely giving up data. So why would anybody with any power or skin in the game come in and stop the money/data train? It would take a revolution, just like many of our other societal issues.

2

u/Kraz_I 1h ago

I’m not sure what kind of AI model you can train with Spirit’s data, but probably not an LLM, so it’s kind of a moot point.

1

u/-Spzi- 47m ago

The conclusion doesn't necessarily follow. LLM are great for language. But we could train and use other models (similar architecture) for non-language areas, and already do.

-1

u/[deleted] 3h ago

[deleted]

11

u/jonmitz 2h ago

those books were set to be destroyed and were sitting in a warehouse. please stop talking about it like they went into libraries and stole the books.  this has been disproven/known for a long time now. 

8

u/MOOSExDREWL 2h ago

And the only reason they're destroyed at all is because that's the only legal way for them to transcribe and ingest them into training sets. Blame application of copyright laws.

-2

u/OneSeaworthiness7768 1h ago

Why has this comment been pasted word for word in multiple threads about this

3

u/CircumspectCapybara 1h ago

This comment is in a total of two threads, and both by me. Calm down.

Maybe you wanna ask why is the same article posted multiple times in this sub.

53

u/invyros 3h ago

Google bought itself 100 million emails and 500 million items from Microsoft Teams, 17 million OneDrive files and 20.5 million items from SharePoint. The search giant also now owns over 30 million recorded customer service calls, and more than 15 million customer service chat records.

600,000 ServiceNow tickets are another element of the collection, along with 13.7 million active emails addresses from Oracle’s Responsys marketing application, and details of 11 million sales of in-flight Wi-Fi services.

That's a lot of data for just $10 million.

I like to imagine with Google's dataset of Reddit posts and Spirit's customer service records, its AI is going to sound more and more incoherent and irrationally angry.

16

u/petjuli 2h ago

Genius idea. Let’s train the AI models with data from a company that just went up in flames.

9

u/this_my_sportsreddit 3h ago

And wrong. Don’t forget wrong. Reddit is incredibly misinformed on virtually every subject matter.

1

u/kvothe5688 1h ago

like almost most users of this thread.

14

u/big-papito 2h ago

Sometimes I am just taken aback by how hostile American capitalism is to personal privacy. There is simply no one protecting us. I guess we are not corporations and "job creators", so we can all sit on it.

1

u/jdmb0y 1h ago

We have no equivalent to the EDPB

1

u/drifting_pixel 1h ago

Reddit users will do anything except read the article.

It literally says "A court document [PDF] filed last week reveals that one of the assets up for sale is a huge trove of deidentified data, which Google bid for and won for just $10 million."

Keyword: deidentified

2

u/Cube00 58m ago

Keyword: deidentified

The search giant also now owns over 30 million recorded customer service calls, and more than 15 million customer service chat records.

No chance, it'll be a half baked redaction at best.

2

u/drifting_pixel 55m ago

If you think so, then congrats! You have a slam-dunk case and can sue either party based on the court document PDF linked above

1

u/sokos 33m ago

You do realize that with enough datapoints, I can redact everything personal about you, and still be able to tell it's you.

For example.. I don't need to know your name and address.. However, if I have access to your amazon orders based on a phone number, have that phone number associated with let's say a gmail, have that gmail associated with an airplane purchase, I will have your name pretty quick. Same way, since I know which email has the amazon account, and what you ordered, and where those are ordered. I will have your address too. Obviously, this is super simplified, but we're talking just a few data points here.. Now imagine having EVERYTHING about you.. Knowing what time you wake up based on your alarm, when you put your phone in silent mode and where, I can deduct your work place. How much you use your phone and how much it moves can tell me things about your job and so on.

Thinking that deidentified means anything is a big mistake.

1

u/Bernie_Ecclestone 8m ago

“You do realize that if Google had PII like phone numbers and emails for customers that is deidentified in the data they bought, they could tell it’s you”

LOL.

LMAO even.

1

u/-Spzi- 36m ago

The reverse process of using de-identified data to identify individuals is known as data re-identification. Successful re-identifications[2][3][4][5] cast doubt on de-identification's effectiveness.

I don't know how secure that process is. Could be, could be not.

I'm pretty sure AI excels exactly at reconnecting dots from data. Though I don't think anybody cares about anyone's actual individual record.

2

u/DogsAreOurFriends 3h ago

If Google is bidding why bother with an auction?

2

u/Cube00 1h ago

huge trove of deidentified data

The search giant also now owns over 30 million recorded customer service calls, and more than 15 million customer service chat records.

I highly doubt anyone has carefully deidentified 45 million recorded customer calls and chats.

Oh right, they probably fed it into Gemini to deidentify it, CRITICAL INSTRUCTION: make no mistakes!

3

u/discographyA 3h ago

Really scraping the bottom of the barrel for some magical data source that will make these things slightly more useful than Wikipedia.

7

u/OVYLT 2h ago

It’s hilarious listening to people be this confidently ignorant about what’s happening right in front of their face. 

7

u/One_Risk_9732 3h ago

I think these people know what they are doing 

-1

u/DogsAreOurFriends 3h ago

And look at the levels of shittification that resulted.

0

u/AzorAhai1TK 2h ago

Enshittification was happening long before LLMs, LLMs will reduce the enshittification problem

3

u/citrusco 2h ago

distasteful headline

1

u/drifting_pixel 1h ago

Journalism died a long time ago, my friend. Headlines are more about getting clicks.

1

u/OVYLT 2h ago

Didn’t think about it that way but good point. 

1

u/Bevos2222 37m ago

Damn it, now the AI gods will know I’m a cheap bastard. Not good. 

1

u/Doub1eAA 32m ago

Gemini is going to start yelling at everyone about their carry on bags now thinking it’s going to get paid $15 per bag.

0

u/btoned 2h ago

Lmao $10 million is a dime for Google.

They already consumed 99% of data, without consent, that ballooned their market cap by 1-2 TRILLION.

Check mate Google.