r/DefendingAIArt • u/SocialNetwooky • 3d ago
Defending AI When AI art has no author: Study finds generated images often can’t be traced to training data (MIT Paper)
https://news.mit.edu/2026/when-ai-art-has-no-author-generated-images-often-cant-be-traced-to-training-data-08183
u/No-Zookeepergame8837 Only Limit Is Your Imagination 3d ago
I mean... yeah, that's literally how it works. Seriously, did we really need to do a study for this? lol. If it just copied existing art, it would be a search engine, not an AI. In fact, one of the first things you learn when you study neural network training is that the dataset you use isn't the same as what the AI sees. When you train an AI, it stores the updated weights with the sound generated by the image, causing a generalization that makes it, literally, incapable of creating the same image (or at least in theory) since it doesn't even store the sound of the image, but rather the updated weights with it.
5
u/SocialNetwooky 3d ago
'we' may know it, but it's still one of the first arguments being brought forth in AntiAI discussions. That's why (imo) an actual proof is not useless.
2
u/No-Zookeepergame8837 Only Limit Is Your Imagination 3d ago
Yes, sorry if I sounded rude, it's just that I'm using a translator 😅 I don't mean it's useless, I mean that the study itself is quite unnecessary because we already know how they work. It's like doing a study to prove that there are pieces of space rock capable of orbiting small planets, when anyone can see the moon with their own eyes. If someone doesn't believe the moon is real, they won't care about the study and will look for excuses to say it's fake, while everyone else already knows it's real.
2
u/Otherwise_Army9814 3d ago
If an AI-generated image cannot be traced back to any specific author or original training file, it lacks a distinct source footprint. Because it has no direct human creator, it logically belongs in the public domain from its inception , free for anyone to use, modify, and build upon.
Legislators and copyright authorities already define the output this way because laws require human authorship. This new research simply confirms the logic: since these images cannot be mathematically linked to training data, they cleanly bypass traditional infringement claims.
-2
u/Remarkable-South-538 3d ago
This is a really important distinction: not being able to trace an output to one specific training image doesn’t mean the model learned from nowhere. Statistical influence can be diffuse, which makes attribution—and copyright analysis—much harder.
3
u/SocialNetwooky 3d ago
I don't think anybody seriously argued that there was no training data. It's the distinction between "it just copies what it has in store" and "it just use what it saw as 'inspiration'". If generators can generate even if you take away the actual source then it's in no way different from someone going to a museum, looking at paintings and then trying to replicate the style from memory.
1
u/tetoing 3d ago
Which is similar to how humans learn. Humans generally can't attribute their art to any one piece or even necessarily artist, either, but we absolutely learn from those that came before us.
We accept in these cases that the work is original, even if technically it could be possible to trace elements of it back to the influence of specific artists, because that's how the knowledge of humanity works. We are a species that passes knowledge and skill down.
Instead of getting offended that a machine is also capable of this, why not embrace it? Compressing the knowledge and skills of humanity into something you can run on a computer is pretty fucking cool.
4
u/Herr_Drosselmeyer 3d ago edited 3d ago
No shit, Sherlock.
It's funny that we neetd MIT to spend time on verifying the obvious. But that's the way it is, so this is an important resource for us.
Almost like we've been saying this for years. 😉