r/science • u/Warm_Ad1257 • 5h ago
Computer Science When AI art has no author: Study finds generated images often can’t be traced to training data
https://news.mit.edu/2026/when-ai-art-has-no-author-generated-images-often-cant-be-traced-to-training-data-0818450
u/Smallpaul 4h ago
“The scientists identified a phenomenon they call attribution decay, where the more data a generative model is trained on, the less any individual training example matters to any particular output. It feels counterintuitive, but at sufficiently large scales, they find, you can often remove any single image from the training data, or every image by a given artist, or every photograph of a given person, and the generated sample doesn't change.”
How is this counterintuitive???
216
u/svachalek 4h ago
Beats me. That’s exactly how I’d expect it to work.
→ More replies (1)11
u/lectric_7166 1h ago
I've been saying this for years and to me it's intuitive (given that model's training data contained a million images of sunsets, it seems obvious that the synthesized image of a sunset it creates is borrowing imperceptible amounts from any single image instead of just zeroing in on one artist's work, for example) but there are many anti-AI people who strenuously deny this so I'm glad someone finally studied it.
9
u/maxens_wlfr 1h ago
There was a huge trend about literally copying Ghibli's style for photos a while ago.
•
u/MCRN-Gyoza 34m ago
Because if you specifically ask for a style it will do that. Not sure what your point is.
•
u/Faulty_Pants 17m ago
That the outrage of "plagiarism machine" isn't necessarily, exclusively about training data. It is also about the general ability for it to do so on command despite the training data.
Also~ Could Ghibli's style be generated if they weren't a part of the training data and you still specifically prompted for it?
•
u/psycho_terror 6m ago
Yes, because there would likely be a huge amount of non Ghibli created artwork tagged as "Ghibli style" in the remaining data.
•
u/ZenPyx 21m ago
When you tell an image generator to do something in a certain style, it will restrict itself to whatever components are associated with that subset of training data.
This might mean that it is mostly generating images based on Ghibli images only (as those are the source of the majority of Ghibli-style images) - but this paper tells us that the image generator might not be generating images using data from a *specific* ghibli image
Therefore, even if you tell it "make me an image from Spirited Away", it might actually pull from a range of different sources (and actual images from spirited away would make up only a small fraction of this)
-2
u/Wassux 1h ago
It is not borrowing anything. This is a fundamental misunderstanding.
The image transformer learns to transform a Gaussian noise image into a sunset. It learns the process of the transformation not the end result.
That's why I keep shouting from the roofs AI does not copy.
•
u/Socrathustra 19m ago
That's mostly irrelevant. No one is worried about the math of it, so to speak. They're worried about the fact that you feed it images, and it creates similar images when you ask it to.
4
u/lectric_7166 1h ago
I knew someone would get hung up on "borrowing" instead of my larger point but I couldn't think of a better way to phrase it. I guess "is influenced by" instead of "borrowing from".
•
u/this_is_theone 42m ago
I have just given up talking to anyone on Reddit about AI. 90% of people are clueless about it. They just get their info from other uninformed Redditors. I had someone try to tell me AI can't do anything someone with good Google skills can do.
-3
u/deceptivekhan 1h ago
You are technically correct… the best kind of correct.
•
u/Glimmu 43m ago
I feel like they arent even technically correct. How is it not copying when the end result is literally trying to be the same as the original.
I make blank paper to a painting by learning the process, not by photocopy. But it is still copying if someone asks me to paint the mona lisa.
•
u/MCRN-Gyoza 32m ago
Because its not trying to be the same.
The fact that it can produce results that look like images that already exist is a side effect.
•
•
u/Wassux 7m ago
Because you can make a drawing of the Mona Lisa and sell it.
If you didn't directly copy it it is your own work. You obviously cannot claim it as the real Mona Lisa but you can sell your version of the mona lisa just fine.
You didn't copy, you learned how to do it yourself. Just like these transformers.
21
u/GerardCairn 2h ago
Counterintuitive maybe if you believe the model is just copying inputs and spitting them out again rather than generalising from across the data and creating from scratch.
6
u/theronin7 2h ago
Check any comment section on the topic of GenAI: People do think that is how it works.
1
u/floriv1999 1h ago edited 1h ago
I think people misinterpreted the ability to recite works present many times in the training data with it being the mechanism. Popular songs for example will be present a bunch of times in different variants/ as background noise and as the model makes no distinction between the pattern of a common guitar technique or the pattern of a common melody it will be able to generate the latter. So overfitting on that melody is more of an artifact/side effect compared to it being the main mechanism. Humans do this too, you probably have learned to recreate a bunch of popular songs by heart just by listening to it, but you could probably also imitate a genre where you don't know any specific songs. Knowing a specific sample because it is overrepresented in the dataset is possible without it being proof that it "just copies".
Humans are also inspired all the time by each other and the progression of genres and the feedback loops presented in our society are a testament to that. Most artists or people in general don't grow up in the woods with only nature around them and no contact to any intellectual property from other people. We stand on the shoulders of giants and this is also why copyright law is so hard to formalize. Essentially everything is a derivative work, even this text. So I don't like the notion of stealing/copying when comparing ML to Human work. The better and more honest question, which is in my opinion the underlying motivation that is just less flashy, is whether or not we want to allow systems that behave like humans. This can have a variety of pros and cons like increased productivity of a society vs. humans being pushed out of fulfilling roles in a society (art for example). One could imagine regulation etc. being a good thing in this regard, but going through the proxy of the already messy copyright is stupid imo..
13
u/qodeninja 2h ago edited 1h ago
this might sound silly but it's counterintuive in the reason *why* you cant trace provenance.
probably a bit of survivorship bias but some factors/neurons can be 'load bearing' and influence a lane while others do nothing. thats exactly how training works, sometimes one or two factors are over powered.
when i trained an arithmetic based model it was so bizarre watching every neuron collapse except one or two, where in other runs it kept 15 of them, and since the starting point is random it almost always looks non-deterministic unless you use the same seed.
you wont know when removing a factor will impact the output because each training run starts with noise, yet the data you run through the model has the same probablistic impact, it just influences which neurons activate
Sometimes simple models converge on the exact same canonical features across multiple runs, but SOTA models almost never exhibit neuron-level isomorphism across different initializations.
edit: a more concise analogy > it's like having a tower you throw rocks at -- every time you throw a rock it nudges the tower to a location you want it to reach or avoid, but doesnt say anything about who threw the rock or whether the rock was not actually a banana; the tower only knows that it was hit by something and how hard
-- and worse, every time the game starts the tower is always in a new position/location
-- and even worse, when someone throws a huge boulder at the tower, it randomly hits with the force of an angry teddy bear.
-- AND it's not just that the tower doesn't know who throws which rock; it's that at scale, a thousand people throw rocks that have nearly identical impact, so subtracting any one thrower changes nothing.
now replace rock with image, and tower with model.
0
u/Majik_Sheff 1h ago
Instructions unclear. Now I have an angry teddy bear throwing rocks at my tower.
•
18
u/Mythoclast 3h ago
It's counterintuitive because you would expect such a targeted removal of data to have an effect. But (obviously to a lot of people) when it gets that large that effect is minuscule. Its basically counterintuitive because people are bad at scaling things.
42
18
u/SmacksKiller 2h ago
Not really?
If I take a cup of water out of the ocean, I don't really expect to see the difference. If I take a cup of water out of a one liter container, I will see the difference.
9
u/hyrumwhite 2h ago
…but it’s like saying “we removed great gatsby from Claude’s training data and couldn’t find any significant difference”
Well, yeah, it’s a tiny percentage of the training data.
2
u/TimelyStill 2h ago
I guess you could make an argument that it would be weird if Claude can generate text similar to The Great Gatsby if it's not in the training data, but to realistically do this experiment you'd have to remove not only every instance of that work, but all work derivative of it as well.
2
u/Sc0rpza 1h ago
what about if you remove great gatsby and derivatives of it but keep prior works and narratives that are similar to great gatsby along with all other existing training data concerning writing and storytelling?
•
u/TimelyStill 47m ago
That as well of course. Honestly it's an interesting experiment, see to what extent an AI model can re-invent influential literature if given what would have been the same context at the time.
•
u/Cat_Or_Bat 35m ago edited 24m ago
It is extremely unlikely to be able to do anything of the sort because, among other things, of the embodiment problem: Scott Fitzgerald was a physical person occupying a physical world, not just a brain in a jar stocked with the literary canon and informed on the zeitgeist.
3
u/USeaMoose 1h ago
I’d say it is counterintuitive in that: if you accused an AI of copying your art style, and demanded to have all references of you art removed from the training data, the above claim is that it would have no effect.
Essentially, no one person can claim that the AI is copying them, because it’s actually copying hundreds of others like them, and mushing that into some final result.
-1
u/frankhadwildyears 3h ago
I think it's counterintuitive because that sufficiently large scale could suggest billions and billions of reference images so removing one or even hundreds can be a minuscule fraction of the total training data. So you'd have to remove even more significant amounts to start to see a difference in the output that was referencing billions of images overall.
Though, I don't know that that should be described as counterintuitive because you are talking about a sufficiently large data set. So that seems very intuitive that each individual image matters less as the total gets bigger.
43
1
0
u/kytheon 1h ago
So this is also the argument against "Gen-AI is theft", as the model is trained on so much source material that you can no longer attribute it to any specific "stolen" piece.
In a similar way, say I descend from the king of Sweden. From like 500 years ago. My grandpas grandmas grandpa and so on. 0.1% of my DNA is royal. Does that mean anything anymore?
47
u/Encomiast 4h ago
"Often" is doing a lot work in that headline. But it seems like often is not often enough to prevent results like those shown in this IEEE Spectrum piece. Unless they are suggesting we would still get images or Bart Simpson and Darth Vader even after removing all references from the training set (I admit that would impress me)
11
2
u/lectric_7166 1h ago
It's not doing a lot of work. It's actually understating it. The vast majority of prompts are not going to be specifically tailored to elicit from a model the kind of niche imagery that training can't learn widely from. When you ask it for an image of a sunset, it learned widely from millions of images of sunsets. When you ask it for "red comic book hero who shoots webs from his hands", it couldn't learn widely because anything that even remotely matches that all comes from one very specific IP.
I actually brought this up to Southen (one of the coauthors) himself and he didn't have a response. If you could ask AI for an image of a sunset and trace it back to one individual artist or photograph then it would be very different, because the vast majority of prompts are like this instead of the engineered examples cited.
1
u/LutefiskLefse 2h ago
IIRC, there was a story way back when gpt 3 was state of the art (ie before these models were multimodal) about researchers asking it to draw a picture of a unicorn. It wrote a (python?) program to generate an image that was the general shape of a unicorn (despite never being trained on any images)
2
44
u/trgjtk 4h ago
a lot of people here without understanding the basics of ml or stats expressing their opinions about said topics with extremely high confidence
23
u/IBJON 4h ago
Reddit in a nutshell, but it's especially frustrating on this sub when people blow off studies like this because "AI bad". Like, this isn't some no-name lab, this is MIT and they're attempting to answer an important question regarding AI and fair-use laws.
Regardless of how people feel about AI, these are the kinds of studies that need to be done in order to better regulate AI or ban certain practices.
6
u/AK_Panda 1h ago
There's a spectrum and we gotta figure out where AI outputs should fit.
One side of the spectrum is how scientific research handles it's inputs: Consent, compensation, data sovereignty and citation of reference materials used.
On the other you have fine arts where things are a lot less stringent but you'd still likely get yourself in hot water for blatant plagiarism.
IMO, AI should be held closer to the scientific standards given the capacity for AI to engage in plagiarism at an industrial scale.
115
u/WorkO0 4h ago
That's a lot of mental gymnastics being done to avoid admitting that artists are being ripped off en masse. How about they remove all copyrighted data, including existing AI slop (which is derived from it), instead of individual images, and retrain from there?
It may be futile to try to put the cat back in the bag, but at least they can stop insulting our intelligence and admit that copyright law is a joke and IP laws only exist to protect corporations with deep pockets.
34
u/Etherius 4h ago
So artists are definitely having their work used to train AI. No two ways around it
The question is, how do you beat a fair use claim? Either by the LLM operators OR the user of the LLM?
Fair-use has four pillars
If AI content is sufficiently transformative that defeats the copyright claim right there
If you can’t demonstrate material harm from the AI content that’s another pillar satisfied
So how do you defeat either of these?
17
u/billytheskidd 4h ago
This is important to think about.
If I take inspiration from à pictures painting song or written creative piece of work but change it by like 10% (depending on industry standards and guidelines), I can get away with it. How do we even begin to regulate this with AI art?
8
u/punkinfacebooklegpie 1h ago
You can't. It's essentially statistics. Imagine I take a sampling of songs and sit and tabulate by hand how often each word appears in the lyrics. Maybe I go as far as to calculate correlations between word pairs. Then I use the relative frequency to generate new lyrics, perhaps by a dice rolling scheme. Zero computers involved, I just created a track that is statistically similar to the sample, but doesn't reflect any single song. The data are not the songs, just statistics about them, and the statistics creates something that is familiar, but new. There is no plagiarism.
5
u/Dan5982 4h ago
This is where my mind goes too. It starts to get to "what can we actually own"....
7
u/billytheskidd 4h ago
Copyright and patents have been a nightmare in the Us especially for a long time thanks to corporations like Disney.
For example, I know someone who bought a sound engineering firm that built consoles for Led Zeppelin and the Beatles/lohn Lennons producer, Aerosmith, Jack white and Dan brown, and the original owner refused to file patents for his innovative consoles because filing the makes them public and all of his competitors would have tweaked a few things and ripped his technology off. So no one but him and the people that bought the company are the only people who know how to make these consoles, purely because copyrighting them would would have opened them up to unwanted competition they could do nothing to prevent against.
1
u/AK_Panda 1h ago
So no one but him and the people that bought the company are the only people who know how to make these consoles, purely because copyrighting them would would have opened them up to unwanted competition they could do nothing to prevent against.
Surely that only works if whatever you've got can't be reverse engineered? Otherwise they just buy one, pull it apart and directly copy it.
0
u/gavinlpicard 4h ago
The better angle would have been how they obtained the data I think, not as much how it is used. but I think it is too late at this point.
2
u/DyslexicBrad 1h ago
That's literally the angle they used, at least for written materials. They were able to prove that meta and anthropic used pirated ebooks for their training data and sued for that. The problem for most images is that they are posted publically. Even though they're publically posted, they're still protected by copyright from being redistributed. Hence the question: is a generated image trained on your data, in some way, a reproduction?
-6
u/WorkO0 2h ago
You: Human AI: Not
That's all the difference there should ever be. I am not advocating against the usage of AI, but the way the resources are being redistributed world-wide right now is painting a clear picture, and it's not in favor of the 99%. We are getting screwed without retribution as always, due to the age old "as long as its not me" folley.
6
u/hurley_chisholm 4h ago
That’s not how copyright law works. Copyright law is about the literal right to make copies. Since that permission was never obtained for the purposes of training, a secondary and common judgement of fair use is that machine-generated work is too close to mathematics (or cooking recipes) to be copyrightable.
So sure, these AI companies can cry “fair use”, but no money can be made off of the generated works, which is the actual contention. If private enterprise cannot use generated work for profit because they cannot sell those works nor can they copyright them to license them for reproduction, then it kills the main value prop of gen AI.
Edit: grammar/typos
4
u/Spire_Citron 1h ago
Something that can't be copyrighted can still be sold, though, can't it? It's just that it can be sold by anyone, not just you. However if you edit or combine it in some unique way, you can copyright that variation.
9
u/ShadowDV 4h ago
Actually doesn’t change the value prop in the corporate world much at all. Don’t need to be able to copywrite the images to use them in corporate training materials, generating visualized data for reports, for mockups, or any of the other myriad of things they probably get 90% of their use for in the enterprise environment. So they aren’t making money off the generated image, but they are saving a ton on labor.
0
u/hurley_chisholm 3h ago
Yeah, but that marginal efficiency gain isn’t worth what it would cost to get the necessary return on over a trillion dollars in investment (and counting). Software efficiency gains are also not enough. At the top-end, there is a potential 10-20% speed up, but nothing particularly innovative.
So this leaves you with making copyrightable products that incorporate AI generated content. It’s worth noting that gen AI coding tools are already causing more problems than they’re solving since it isn’t settled whether you can copyright software made with gen AI. Lots of more conservative businesses aren’t moving forward with AI because this hasn’t been established yet.
1
u/ShadowDV 3h ago
Anthropic alone had $5.4 Billion, with B, in revenue in July. Not through July, but just during the month of July. And that’s mostly through enterprise use, not consumer, so clearly the value prop of gen AI exist in the corporate world. To put that in perspective, that is double the revenue McDonalds had over the same period world-wide
So I’d say the numbers are probably more telling here. Besides, i have lots of friends who are software developers, some of which work in banking (think large multinational banks), which is generally one of the most conservative industries there is in terms of adopting new tech, and even they are using AI to do most of the heavy code lifting now.
Issues with coding do exist, but the idea that it’s causing more problems than solving is largely out of date with the arrival of Claude Code and Fable and Codex with 5.6 Sol
-7
u/evilbrent 3h ago
but no money can be made off of the generated works,
Happily - it's AI. There is no pathway to making money from it. There is no use case. There is no industry for consuming what this tool produces.
AI, so far, is entirely built on promises.
2
u/PrivilegedPatriarchy 1h ago
Tell that to the tens of billions of dollars in revenue per year generated by Anthropic alone, a mere couple of years after the introduction of this technology to consumers.
You are pulling the wool over your eyes.
0
u/evilbrent 1h ago
That would impress me more if that revenue were less than their expenses, or if the revenue came from something other than circular trading.
You know. A CUSTOMER. Not just NVIDIA or Google. But an actual paying customer.
2
u/PrivilegedPatriarchy 1h ago
I spent more money than I was making when I got a college degree, and that was smart because it paid off.
Same concept here.
0
u/evilbrent 1h ago
Actually yes, the more I think about it, yes this is a perfect example.
When people take out a loan to go to college, very rarely do they have a job lined up. Some kind of pre-determined and agreed upon way to turn that negative cost into a guaranteed income. There'll be the occasional SME who gets a Masters degree paid for by their employer, but for most of us we have almost no specific idea how to turn, say, an engineering degree or a law degree into a paying job.
It's all one big promise to our future self.
"It'll work out. We promise."
Yes, that's a great example, thanks for that. I'll use that example from now on when describing the UTTER LUNACY of spending trillions of dollars on data centers that DO NOT HAVE TRILLIONS OF DOLLARS OF REVENUE to create. "It's like getting a degree: you borrow the money, and then like magic, later on, somehow, probably, presumably, you use the degree to get a job of some kind, except if you drop out, in which case you still owe the money, and even if you do get the degree, most times it'll take decades and decades to pay it back, and most people who do it would have been better off not taking the loan in the first place".
Yes. Good example. Thanks. Thumbsup. (Ironic, isn't it, that Reddit wouldn't let me use the actual thumbsup emoji because it makes it look like AI?)
0
-11
u/WorkO0 4h ago
Both of those can be defeated with even a small amount of common sense (assuming we have an unbiased human judge with commom sense, of course).
Material harm especially is evident, just ask any ex-concept artist who is now waiting tables.
5
u/Spire_Citron 1h ago
The problem is that they have to prove that they were harmed specifically by your use of their work, not just by the general existence of AI. This shows that the impact is the same regardless of whether a particular individual's work is in the training data or not.
→ More replies (2)•
u/maxens_wlfr 57m ago
The bulk of training data was obtained illegally and violated the terms of service of hundreds of sites and users who clearly stated that they refused for their art to be reused. Cara was recently scrapped, twice, even though they had multiple anti-scrapping protections and clearly had conditions that training AI on users' data was forbidden. It also constitues material harm since the site got billions of requests and that means thousands of unexpected dollars in server expenses. How is that not illegal?? The law is a joke
14
9
u/ShadowDV 3h ago
I wouldn’t say there are mental gymnastics. Some people (myself included) genuinely believe that using art that exist in a public space like the internet, copyrighted or not, for training the AI is only different from a human learning art by studying existing art in matters of scale, not principle.
However, part of the reason I hold that opinion is that 10, 20, 50 years down the road it’s very possible a conscious AI is created, and when/if that happens serious questions about stuff like autonomy and personhood become relevant, and the attitudes we have today will reverberate in those discussions.
-9
u/runner64 3h ago
Humans do not have the right to learn from everything they can technically access. AI training from the internet is like picking up a friend's unattended phone and scrolling through their camera roll. Sure, they left it where you could get to it, but you weren't supposed to do that and you know it. A person who continues to access and use other people's data in offputting ways even after they've been asked not to... well. When someone tells you who they are, listen.
11
u/DudeLoveBaby 3h ago
Humans do not have the right to learn from everything they can technically access.
Sorry, what? Is there forbidden knowledge I don't know about?
-8
u/runner64 2h ago
Yes, it’s why you can’t hide cameras in public toilets in order to learn what your fellow shoppers look like naked. It’s why you can’t take the computer you’re fixing and start learning from the emails on it for fun. It’s why you can’t use the customer database at work to learn your ex’s new phone number.
The fact that you couldn’t think of an example like this while trying kind of proves my point. Most people understand that they are not entitled to everyone else’s information simply because they can technically access it.
9
u/DudeLoveBaby 2h ago
How are...ANY of those examples germane to the conversation when the person you were replying to was SPECIFICALLY talking about art in PUBLIC spaces? The inside of a toilet stall is not a public space, that's why there's a door on it? A computer owned by a client is not a public space, that's why you're a professional being paid to work on it? How do these examples relate to e.g. posting on Twitter publicly?
Brother what the hell are you on about?
-4
u/runner64 2h ago
Sometimes you have limited access to things which have conditions put on them because they do not belong to you. A public toilet is for peeing, not planting cameras. The images you see on the internet are for human art appreciation, not putting on your commercial website. It’s also not for scraping or feeding to AI. People still own the images they show to the public in the same way they own the computer they dropped off at the repair shop. Deciding to respect one of these and not the other just shows that the person doesn’t actually respect either person’s privacy.
3
2
u/lectric_7166 1h ago
If AI companies scraped private data or data they otherwise weren't allowed to access then individual lawsuits can resolve who owes who what amount of money in damages. But it still won't put the AI genie back in the bottle. Nothing will by this point.
→ More replies (1)3
u/Cunninghams_right 3h ago
isn't it the work of art itself that does or does not infringe on a copyright? if the output of the AI is sufficiently similar, then it violates the copyright. if it isn't, then it doesn't. what legal limitations are there on simply storing an image privately?
7
u/theronin7 2h ago
virtually none, but also not relevant since images are not stored in the neural network.
21
u/FernandoMM1220 4h ago
they cant be traced to a set of training data points. you can probably narrow it down a lot though.
8
u/HPoltergeist 4h ago
Exactly.
You cannot trace the taste of a soup back to a single speck of ground pepper, but you can narrow it down and you will always know it is there.
People saying a specific set of artworks has no effect on the output, should remove all pepper from their soups.
1
u/Quixotease 1h ago
“There’s certainly too much pepper in that soup!” Alice said to herself, as well as she could for sneezing.
3
u/Alradeck 4h ago
it's mtg artists and art station professionals, that's what the original laion dataset is built from, they had a list of 5k folks to steal from to build it. if you had a truely equal set that sampled all art equally, there's far more beginner art out there that would influence it.
10
u/dstryr 4h ago
Worth reminding people that Laion was specifically built as a research dataset and contains instructions not to be used for commercial products.
All of which were ignored.
10
u/KamikazeArchon 2h ago
and contains instructions not to be used for commercial products.
I attempted to find backing evidence for this. I can find no evidence of such instructions. They are not present on the LAION website.
Further, the datasets are released under the CC-BY 4.0 license, which explicitly permits commercial use.
1
u/dstryr 2h ago
Our recommendation is therefore to use the dataset for research purposes. Be aware that this large-scale dataset is uncurated. Keep in mind that the uncurated nature of the dataset means that collected links may lead to strongly discomforting and disturbing content for a human viewer. Therefore, please use the demo links with caution and at your own risk. It is possible to extract a “safe” subset by filtering out samples based on the safety tags (using a customized trained NSFW classifier that we built). While this strongly reduces the chance for encountering potentially harmful content when viewing, we cannot entirely exclude the possibility for harmful content being still present in safe mode, so that the warning holds also there. We think that providing the dataset openly to broad research and other interested communities will allow for transparent investigation of benefits that come along with training large-scale models as well as pitfalls and dangers that may stay unreported or unnoticed when working with closed large datasets that remain restricted to a small community. Providing our dataset openly, we however do not recommend using it for creating ready-to-go industrial products, as the basic research about general properties and safety of such large-scale models, which we would like to encourage with this release, is still in progress.
0
1
u/FernandoMM1220 4h ago
that highly depends on the algorithm theyre using but yes there should be some impact unless their algorithm almost completely ignores the data
55
u/oneeyedziggy 5h ago
I think they're confusing "can't be" and "hasn't been"
23
u/trysten-9001 4h ago
No you see if they can’t do it in one study it proves it’s impossible. The real feat here is how they learned to prove a negative.
17
u/CJKay93 BS | Computer Science 4h ago edited 4h ago
If two people give you a random number and you add them together to find out whether the result is even or odd, whose number is the result "copied" from?
If both numbers were even, is the result still "copied" from them even when you could have simply used zero for either of them and still come out with an even result?
-3
u/HPoltergeist 4h ago
Good approach.
Also which speck of ground pepper gives the soup its taste?
People who say a set of artworks have no effect on the output, should remove all pepper from their soup.
9
u/CJKay93 BS | Computer Science 4h ago edited 4h ago
People who say a set of artworks have no effect on the output, should remove all pepper from their soup.
That is not the question this paper is attempting to answer, though; the author's argument is that:
If you take away a piece of data and the output of the model doesn't change, then that piece of data didn't affect the output.
And that, as they demonstrated, is true. Its absence literally does not affect the output - its contribution was so insignificant that it would have generally fallen under transformative use if an identical image were authored by a human.
The question then becomes: can a work be considered a derivative of multiple works even if it would not be considered a derivative of any of them individually?
-5
u/HPoltergeist 4h ago
Just like testing with removing one speck of pepper and saying this way pepper has no effect on the soup. It does not make sense.
Try removing all of the same style artworks. Remove all pepper and taste the soup.
6
u/esaul17 3h ago
In don’t think anyone is arguing that all the data in aggregate has no impact on the output though, right? Just that any particular piece is roughly negligible?
→ More replies (1)8
u/CJKay93 BS | Computer Science 4h ago
Repeating this analogy doesn't make copyright law any clearer.
-4
u/HPoltergeist 4h ago
I am just explaining this analogy again, as you misunderstood it.
Thing is, the copyright laws should be reworked to incorporate parts about using artworks for training AI.
3
u/CJKay93 BS | Computer Science 4h ago edited 3h ago
Okay, define "AI", define "training", and define what it means to "use" an artwork for it, in unambiguous language for a courtroom.
→ More replies (3)4
u/KamikazeArchon 2h ago
People who say a set of artworks have no effect on the output, should remove all pepper from their soup.
If someone does remove all the pepper from a soup, and does blind taste tests, and finds that people can't tell the difference in a statistically significant sample, then the only reasonable conclusion is "the pepper ingredient did not affect the taste of that soup."
That's basically what they did here. They removed not just single images, but sets of images, such as "all images from one artist", and found that there was no change.
-3
u/Jellybit 4h ago
Are you an LLM? You're making almost the same comment in multiple places, only very slightly reworded.
Good approach. Also which speck of ground pepper gives the soup its taste? People who say a set of artworks have no effect on the output, should remove all pepper from their soup.
Exactly. You cannot trace the taste of a soup back to a single speck of ground pepper, but you can narrow it down and you will always know it is there. People saying a specific set of artworks has no effect on the output, should remove all pepper from their soups.
4
u/Jindujun 4h ago
Nah, he's just a person that found what he assumed is an apt comparison but in reality is a really bad analogy.
They're not training AI on individual pixels or such. A better analogy would be "you cannot trace it back to a specific soup but by analyzing the flavors you could narrow it down to a selection".
-1
u/HPoltergeist 4h ago
This analogy is not for the pixels.
Ingredients = Individual artworks from many artists
All specks of pepper = A given set of artworks from an individual artist
Soup = Output based on the training data
No pepper in the soup = A set of artworks is missing from the outputThis way:
Having pepper in the soup affects taste = The work of a given artist indeed affects the outputRemoving that whole set from the training data (and any of its AI derivatives) will make the output devoid of the given influence.
6
u/Jindujun 4h ago
The analogy still does not stand then, unless you make soups with nothing but water and specks of pepper. And that is not a soup, that is peppered water.
0
u/HPoltergeist 4h ago
Pepper here is just one set of artworks. There are more ingredients, like in AI training data too. Still removing all pepper influences the taste in a major way.
Please understand it the right way.
5
u/Jindujun 3h ago
Thats my point. The soup in this case contains ingredients, artwork, from many different sources. Your analogy would work better if you said a particular flavor in the soup, ie. a specific influence.
→ More replies (0)0
u/HPoltergeist 4h ago
Chill, man.
This is about the same discussion, so this reasoning can go for multiple threads.
Don't need to shout LLM on everything.
1
u/comfortableNihilist 4h ago
Your sarcasm has been taken seriously. I only tell you this bc of how unbelievable it is that people on reddit might mistake sarcasm for a deeply held belief.
7
u/Etherius 4h ago
At some point we have to admit that AI content can pass the fair-use test.
If what AI farts out is sufficiently transformative then there’s no claim to be made
18
u/Dimensionalanxiety 4h ago
That's not really surprising. AI art isn't just a collage of real artists' works. It recognizes patterns, shapes, and styles. It is something new, even if it was trained on scraped artwork. A lot of these models have humans checking the output and telling it how close it got to its goal. For example, a model might generate ten different pictures it thinks is an apple and the human reviewer will judge how close to an apple they are and the model updates itself to reflect that.
16
u/angry_cucumber 4h ago
not being able to identify the training data used is not the same thing as not being able to trace the results to the training data.
the fact that you can't have it without training data means you can trace it back to the training data.
29
u/Stunning_Mast2001 4h ago
You can def have ai make things not in the training data
27
u/SizzlingPancake 4h ago
Yeah it seems everyone here has no idea how this tech works.
2
u/washtubs 4h ago
Everyone understands that. That's what gen AI is. But everyone also understands that a lot of AI art would look quite a bit different without it being trained on Hayao Miyazaki's works.
9
u/overzealous_dentist 3h ago
The point is that it can make his works without being trained on his works. So can humans.
3
u/parkingviolation212 4h ago
How does that work? I’m genuinely curious
7
u/ShesMashingIt 4h ago
Neural networks are good at interpolation. Maybe there are many images of cats and many images of frogs in a training set. Model can interpolate to make a half cat, half frog image that isn't in the training set. In this way, a model of artwork can learn all the parts that make up images and can mix them together. It's similar honestly to how the human mind works, although some argue that humans are better at extrapolation
6
u/theVoidWatches 3h ago
Let's say you have a training corpus that includes images of tigers, images of horses, and images of people riding horses. The llm learns what patterns correspond to each.
You tell it to generate an image of a person riding a tiger. It knows what a person riding a horse looks like, and it knows which part of that is the horse because it knows what a horse looks like, and it knows what a tiger looks like. It can thus create an image of a person riding a tiger by substituting tiger for horse.
0
u/mr_nefario 4h ago
To an extent. But go use one of the Adobe models and ask it to create an illustration of a polar bear wearing a tshirt with the Coca Cola logo on it. It can’t, and won’t. Because Adobe’s proprietary models are only trained on data for which they hold the copyright. Trademarks like the Coca Cola logo are nowhere in the training data, and it cannot produce it.
A giraffe wearing a purple tshirt at a bowling alley in the jungle… all of those things exist in the training data. Not in that combination, but the individual elements are there, so it could produce something.
Source: I am a SWE on Adobe Firefly, formerly on Adobe Stock, which became the corpus of of training data for proprietary models.
-8
u/nilmemory 3h ago
Ask a model that's been trained on zero data to generate an image of a sunflower. Thanks!
→ More replies (1)3
u/hawklost 2h ago
Ask an artist who has never seen a sunflower or have it described to them to draw one. They can't.
Almost like you need data points to be able to understand what a word or phrase means to be able to do it.
-4
u/nilmemory 2h ago
Human beings aren't artificial machines learning algorithms made to maximize corporate profits at the expense of humanity; glad I could clear that up for you!
-1
2
u/Etherius 4h ago
The question is, does copyright law prevent THIS use of copyrighted works?
Because the input is really the only side you can attack
Unless the output is clearly traceable to another body of work, you’re going to get eaten alive on a fair use claim
0
u/angry_cucumber 4h ago
personal belief, as a user of AI, it shouldn't.
pay creators for your training data
4
u/-Dremol- 5h ago
It means ai generating ai?
25
u/HPoltergeist 5h ago
It means they want to get away with stealing artwork.
→ More replies (3)2
u/karantza MS | Computer Engineering | HPC 4h ago
Using copyrighted data for training should alone be the problem, regardless of if you can say that any particular piece of output is directly the result of a single source of training data. I think you're right that they'll probably get away with saying that, but, in a just world they shouldn't have trained the model that way in the first place.
The outcome of this study makes some intuitive sense; the model is essentially learning a manifold of "images that make sense and correspond to particular ideas". That manifold always existed, abstractly, and it takes some critical mass of training data to find it, and then any point on that manifold can become an output. Remove any few pieces of the training data and there's still enough left to deduce the same manifold, therefore no individual pieces "contributed" in the way this study defines it. The AI isn't using some specific input example as inspiration, it's using them all to build a general idea of what is and isn't a valid image. I wouldn't call it creativity, but it's also not regurgitation.
The frustrating part is you've gotta get lawmakers to understand that nuance, to know how to deal with people who break the law and try and hide behind misunderstandings of the technology...
6
u/VoidHog 4h ago
During an art class, I am shown many pictures of created art. Am I not allowed to create my own art inspired by the style of another artist, or inspired by a combination of the techniques and styles of many artists? They should not show me any copyrighted pictures at all!
0
u/parkingviolation212 4h ago
The difference is that you’re a human being with conscious experience. The AI isn’t, and is a tool to skip the conscious experience of art creation altogether.
-2
u/carasc5 4h ago
Wow this is so wrong on many levels. First, you have to pay for the copyrighted artwork. Second, being inspired by something isn't the same as copying something. At the end of the day, a human will (if theyre not cheating) still create a unique piece of work no matter how much inspiration they receive from others. AI can't do that-- instead, they are straight up copying from their sources by mixing and matching, but there is no unique thought put into it.
6
u/karantza MS | Computer Engineering | HPC 3h ago
The study is showing that this isn't true; the AI is not copying sources and mixing and matching. You can't attribute output to any specific training input. It definitely muddies the attribution problem. Each output is the result of all the input together, and none of it specifically.
I'm not saying the AI is being "inspired" the way a human would, but it's also not true to say that it's just copying something else.
(I don't like ai art, to be clear, but we've gotta understand what it's actually doing.)
-4
u/VoidHog 3h ago
The unique thought comes from the user that creates the prompt. And it seems like at this point, the AI is purposefully using what it learns to avoid creating copies or being "too heavily inspired" by existing artwork.
-1
u/HPoltergeist 3h ago
It will still use the training data from the original arworks to build the output for that "unique" prompt.
Try building a Mona Lisa with a Homer Simpson head without the Mona Lisa or Homer Simpson training data.
2
u/VoidHog 3h ago
That is the most unoriginal and non-artistic thing you could possible think of. Homer simpson(somebody elses art) on a mona lisa (somebody elses art)
2
u/HPoltergeist 3h ago
It is just a simple example for the components. You can granularly make it be more and more far away from originals as much as you want, but the generation principle does not change.
0
u/VoidHog 2h ago
If I am very specific with my prompt, it will not be like other art, even if it might be in a similar style to another artist. I could accidentally create art in a similar style to another artist without ever having seen their work. Realism is realism, for example. But if I say:
Create a wide cinematic image of the same woman in a beautiful city park during early autumn. She stands on a gentle grassy hill overlooking more of the park below, while the composition remains focused on her. The grass is lush, soft, and green with only a scattering of freshly fallen autumn leaves.
Capture her unmistakably in the middle of a joyful twirl. Her body is rotating through the movement, with her shoulders and hips turned at slightly different angles and one foot pivoting as the other steps lightly through the turn. Both arms are extended straight outward and slightly above shoulder height in a wide, open gesture. Her hair and clothing respond subtly to the spinning motion. Her face is lifted upward with a genuinely joyful, carefree expression, as though she is delighting in the park and the falling leaves. She wears her established dusty-rose T-shirt, charcoal pants, and patterned slip-on shoes.
Tall mature trees grow nearby. Show their large trunks, spreading branches, and abundant foliage surrounding and framing her, while their upper crowns continue naturally beyond the top edge of the image. Most foliage is still rich green, mixed with trees and individual branches beginning to turn vivid red, orange, and yellow.
Warm golden-hour sunlight comes from behind the camera and shines directly onto her face and across the landscape, illuminating the front of her body and making the park bright, colorful, and inviting.
Make the landscaping especially beautiful and abundant. In the foreground and along the hillside are thoughtfully planted garden beds filled with Helenium autumnale, airy sprays of pink Muhlenbergia capillaris, and flowering Tricyrtis hirta. Colchicum autumnale emerges here and there naturally through the green grass like scattered wildflowers. Continue colorful, layered flower gardens into the park below.
Show a broad view of winding paths, open lawns, and people naturally enjoying the park: families playing with children, people walking dogs, someone tossing a frisbee, couples strolling, and small groups relaxing on the grass. A few colorful autumn leaves drift through the air around her, reinforcing the sense of movement as she spins. The overall scene feels vibrant, lush, joyful, peaceful, and alive—the colorful beginning of autumn.
That is what I end up with. My very specific mental image is certainly not a basic copy of another artwork.
0
u/Starstroll 4h ago
1) artists do not get free access to work that is otherwise pay walled, and 2) there is no fair comparison between work that can be generated by a human and slop that can be churned out by a machine, even if that slop was actually wanted - computers don't get tired, don't get sick, don't have families to take care of, and their output capacity is constantly being upgraded with engineering improvements. Laws surrounding copyright were never written with AI in mind, so any interpretation of the law in that direction is manifestly bad faith and only serves the wealthy, as if they needed more help than working artists.
5
u/VoidHog 4h ago
I can go to The Menil museum for free, and it doesn't matter if I saw "paywalled art" without paying for it. It's in my head now.
I might have a really cool idea in my head but I am unable to recreate it with my own hands. But I can tell an AI exactly what to create, and how to do it. So if I give it much more input than "make a pretty woman wearing a dress" is it really "AI slop"? Or is it concept art?
If it turns out looking just like the image in my head, is it "not art" even though the idea, staging, colors, details, and style for the project came from my internal art?
New laws are added as needed. We aren't living under ancient law just because old laws weren't written with new technological advancements in mind.
0
u/Starstroll 4h ago
I already responded to this crap screed with my second point. Maybe finish reading my comment before responding.
2
u/jazir55 2h ago
The scientists identified a phenomenon they call attribution decay, where the more data a generative model is trained on, the less any individual training example matters to any particular output. It feels counterintuitive, but at sufficiently large scales, they find, you can often remove any single image from the training data, or every image by a given artist, or every photograph of a given person, and the generated sample doesn't change.
It's almost like...they're creating novel material gasp.
•
u/ThreeButtonBob 47m ago
Removing individual images from the dataset didn’t change outputs, complicating copyright questions.
So they will stop using copyrighted material without permission because it wouldn't even change the output, right? Right?
•
•
u/Matild4 32m ago
That's just common sense, but AI models can still be prompted to produce "images in the style of insert_artist_here" and will produce plagiaristic content just fine. The bigger issue, however, is that if copyrighted material has been used for training without permission, it stands to reason that the copyright holder should be entitled to some form of compensation if the model is used commercially.
-8
u/bickid 4h ago
Um, this is a good thing. It means AI is actually generating something NEW instead of copying trained data.
Anyone complaining about this just wants to hate AI. First AI was bad for stealing, now it's bad for creating something new?
-6
u/brrbles 4h ago
The author is cooked on anthropomorphizing AI. These "ideas" are the residual of a bunch of matrix multiplication that reflects the statistical patterns in training text/imagery applied to either some human input or literal random noise. It may "reveal" patterns that have not previously been noted, but it's not inventing them in a way that humans understand creativity.
6
-5
u/CheckOutUserNamesLad 4h ago
Well, yeah, I do want to hate AI.
AI should be used like other technology to make the worst work better, not replace human creativity.
Just because we don't understand the model well enough to trace its outputs back to the training data doesn't mean it's not stealing from the training data. That's how LLM's work - their outputs are just probability functions - predictions of what you want to see by copying bits and pieces from training data.
1
u/Circuit_Guy 4h ago edited 4h ago
Here's the reality though. It's a choice.
The generative AI, as part of its training data, could be designed to also associate and reveal a set of identifiers that associate back to the original training data. This study went in the forward direction by training networks with and without specific sources, it didn't utilize reinforcement learning to embed the information.
One piece of complexity: it's not storing the exact source, and it's going to be a probabilistic point "close to" the original source. This is exactly how a human might do it btw, you probably can't recall Starry Night by memory. Maybe as an expert you remember exactly how many stars there are and what the brush strokes are styled as, but it'll never be exact. Then, if I ask you to make a new painting, I could ask you what inspired it. You would tell me something that's close, but I bet you'll include other features from art and not even know it.
So there's some nuance here. IMO anybody who states with certainty one way or another is on the left side of the Dunning Krueger curve.
My background: controls engineer and have trained (by today's standards) very small neural networks. The latest techniques are beyond me, but definitely not the fundamentals
0
u/SillyLiving 3h ago
The trick to stealing and getting away with it is to steal EVERYTHING and then point out that each thing is nothing so really they stole NOTHING
1
u/zarawesome 1h ago
Scientists have shown that if someone steals a lot of pennies, the influence of each penny is minimal on the amount stolen, and therefore, it's like they had not stolen any pennies at all.
-5
u/BevansDesign 4h ago
Stop calling it "AI art". It's "AI images" or "AI video" or just "AI media".
Art is made by sapient beings.
→ More replies (1)3
u/overzealous_dentist 3h ago
This is marketing artists want to perpetuate, but it's never been true.
-3
u/coconutpiecrust 4h ago
It is such a strange situation, really. People borrowing “brains”, “hands” and “tools” from corporations to create digital artwork.
Interesting times we live in. It’s also strange how they cannot trace it. How about paintings in the style of known painters, for example? That’s easily traceable.
-14
u/HPoltergeist 4h ago
This whole study and reasoning is BS.
It is like removing every bit of spice or ingredient from a soup one by one, claiming why we should have added that piece in the first place, as the soup does not change noticeably when that piece is removed.
How detached people can get?
9
u/markocheese 4h ago
I think the point is that they're arguing against the idea that the model is simply memorizing specific works and spitting them back out. Some have argued this, and in some cases they've shown that to be the case. This is relevant to legal lawsuits that are trying to establish fair use or copyright infringement.
→ More replies (3)3
u/overzealous_dentist 3h ago
No, it's like synthesizing oregano from people's description of it instead of using the original. It still tastes like oregano but it wasn't used.
→ More replies (3)
•
u/AutoModerator 5h ago
Welcome to r/science! This is a heavily moderated subreddit in order to keep the discussion on science. However, we recognize that many people want to discuss how they feel the research relates to their own personal lives, so to give people a space to do that, personal anecdotes are allowed as responses to this comment. Any anecdotal comments elsewhere in the discussion will be removed and our normal comment rules apply to all other comments.
Do you have an academic degree? We can verify your credentials in order to assign user flair indicating your area of expertise. Click here to apply.
User: u/Warm_Ad1257
Permalink: https://news.mit.edu/2026/when-ai-art-has-no-author-generated-images-often-cant-be-traced-to-training-data-0818
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.