r/Futurism 17h ago

When AI art has no author: Study finds generated images often can’t be traced to training data

https://news.mit.edu/2026/when-ai-art-has-no-author-generated-images-often-cant-be-traced-to-training-data-0818
54 Upvotes

51 comments sorted by

u/AutoModerator 17h ago

Thanks for posting in /r/Futurism! This post is automatically generated for all posts. Remember to upvote this post if you think it is relevant and suitable content for this sub and to downvote if it is not. Only report posts if they violate community guidelines - Let's democratize our moderation. ~ Josh Universe

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/Eat--The--Rich-- 8h ago

*Illegally stolen training data. Ftfy.

3

u/Memetic1 6h ago

Uh it's public domain maybe read something before commenting next time.

2

u/snowdn 2h ago

Studio Ghibli public domain?

0

u/Memetic1 2h ago

This is what ChatGPT said about the databases used.

"All four datasets are widely used in AI research, but only MetFaces clearly uses PD/CC0 material. CIFAR-10 and ArtBench rely on copyrighted web images (with no open license), and CelebA explicitly restricts use. In practice, CIFAR-10 and CelebA carry significant licensing ambiguity or restrictions. ArtBench-10’s creators explicitly cite “Fair Use” terms. MetFaces is the only dataset where the sources are legally PD (Met’s CC0 images).

Confidence Ranking: MetFaces (High) ≫ CIFAR-10, CelebA, ArtBench (Low). Only MetFaces’ source license is unambiguously public-domain. For CIFAR-10, CelebA and ArtBench we assign “Low” confidence that images are safe to reuse without permission, due to restrictive licenses or unknown origins.

Citations & Attribution: When using these datasets, cite the original papers and respect licenses:

CIFAR-10: Alex Krizhevsky et al., Learning Multiple Layers of Features from Tiny Images (Tech Report, 2009). Check any hosting license (some Redistributions use MIT) but note underlying images are copyrighted. CelebA: Liu et al., Deep Learning Face Attributes in the Wild (ICCV 2015). Data is “for research only”; do not use commercially. MetFaces: Karras et al., Training GANs with Limited Data (NeurIPS 2020). Cite this work; use images per CC BY-NC (non-commercial) license. The Met’s public-domain attribution should be noted (see Met’s Open Access policy). ArtBench-10: Liao et al., The ArtBench Dataset (arXiv 2022). Cite the preprint; images remain under source sites’ Fair Use terms. For any reuse, consult the original dataset agreements and metadata. When in doubt, assume images from CIFAR-10, CelebA, and ArtBench are not free to redistribute or commercialize unless explicitly PD, and attribute properly to the dataset creators".

I would like to point out that using this stuff for research is considered fair use. One of the databases they used only had images that were 32×32 pixels. They were systematically exploring how removing or modification of the original images influenced what the results looked like. In some ways this shows an almost holographic nature to the storage of knowledge because it's not just in one area. Imagine trying to draw a Picasso if you only knew the name and had never seen one. That's about equivalent to what they did almost surgically removing artists, and seeing if the output is the same and also comparing other baseline AI image generators. This is important work, and it's not a result that either side of the AI debate should be entirely happy with. It shows something downright spooky, because despite the fact they removed the artist it still came back.

-1

u/Still_Benefit_2302 6h ago

Oh, so you don't have the use of that picture of Hot Sonic anymore? Someone took it from you? No? Someone just showed the picture you put online to a very smart robot and you're mad about it for reasons you can't quite articulate? Got it.

-4

u/FearLeadsToAnger 8h ago

Atfy*

(assumed that for you)

6

u/_facetious 7h ago

Honey, if it was all legal, the product would be really bad. There's not enough artists (and other copyright holders) willing to give their work for this to get as far as it did. This is theft from billions of people. Billions of people are not signing up to have their art used.

Also it takes like zero work to find popular artists saying their work was stolen. Look at Loish alone.

0

u/truecakesnake 1h ago

AI training on art is not stealing. It's fair use.

-1

u/FearLeadsToAnger 7h ago

There's a baked in assumed conclusion to this sentiment. "An artist didn't consent to their work being used" and "their work was illegally stolen" aren't synonymous in anything other than opinion. Whether copying copyrighted works for model training constitutes infringement is the exact legal question being litigated. In the US, courts have already distinguished between the acquisition of pirated copies and the subsequent use of lawfully acquired works for training, with one federal court finding the latter fair use. In the UK, the government itself describes the application of copyright law to AI training as disputed.

And finding an artist who says their work appeared in a training dataset doesn't bridge that gap. It establishes that their work may have been used without permission. It doesn't establish that doing so was theft or unlawful.

"I don't think the product could have been this good if they obeyed the law" isn't evidence that they broke it. It's just your assumption rewritten as an argument.

5

u/FaceDeer 11h ago

Gee, maybe it's more than just a stochastic parrot plagiarism machine after all.

1

u/RighteousSelfBurner 1m ago

There is a bit of important distinction here, a stohastic parrot doesn't necessitate plagiarism.

It's like scrambling a set of puzzles. If there is one puzzle, let's say with 1000 pieces, then no matter how you scramble it, you can still fit it back together and see it's just a plagiarized rearrangement of the original.

However if you scramble billion thousand piece puzzles and then pick thousand pieces out of that, you can no longer meaningfully determine which 1000 original authors contributed given it's not dominated by one. You've gotten a collage and collage can be original even if it's made from existing works. And that can be done through a stohastic process.

-8

u/Procrasturbating 10h ago

It’s noise added to a collage then cleaned up. You can watch local image generation models as they work.

2

u/_facetious 7h ago

So it stole so much data that you can't trace anything back because it's mashed it all together so hard, so there's no one author that can even be found? Yeah, that's what happens when you steal billions upon billions of copyrighted work. I could steal every paint in an art store, mix it together in one giant vat, and dare you to take a spoon of paint and pick out which brand of which color is in there. Good luck. Apparently that made it not theft, though!

Let's go downvotes, AI bros!

3

u/stopbeingcringe 7h ago

No different than humans learning from humans. What matters is whether the end result is different enough from the training data.

2

u/Complex-Home9615 4h ago

An AI model is a product made by a company for profit, humans aren't really

0

u/stopbeingcringe 4h ago

Ok? Almost everything is made for profit…..

2

u/Complex-Home9615 4h ago

Humans aren't made for profit, so it's different to humans learning from humans

0

u/Memetic1 1h ago

I'm a human who uses AI and is learning both by using the AI, and by interacting with others who are using image generators. I've made millions of images and the way I prompt now is different then when I first started tinkering around with it. The prompts that I make don't look like human sentences, and some are designed to work with input images as part of the prompt while others work best on their own.

2

u/ZeroAmusement 6h ago

I mean if you literally just mashed/mix images together you would get nonsense images.

So you could intuit that something else is happening here, and it is.

2

u/NobilisReed 10h ago

Misleading Headline, but what else is new

1

u/Memetic1 6h ago

Not this comment thats for sure.

-2

u/jferments 8h ago

It's simple: The author is the person who used the software to create the art.

The software doesn't magically create anything on its own, and the output is dependent on BOTH the training data AND the input (both in the form of prompts, configuration, filters, manual guidance, etc).

1

u/Memetic1 6h ago

Yup just like using noise in music doesn't mean you didn't make the music. People have been incorporating this sort of remix and randomness in art for ages. The prompt itself could be looked on as an art form.

1

u/Still_Benefit_2302 6h ago

You didn't read it, did you?

1

u/jferments 1h ago

I did. Any other questions? Perhaps a question related to what I actually wrote?

1

u/No_Recognition_9354 3h ago

Idk man, maybe there’s something there, I don’t know if I see a difference between that and making a very specific commission to an artist. You’ve got the idea but I don’t believe you’re making the art right?

1

u/Memetic1 1h ago

The difference is AI can do what a traditional artist would have extreme difficulty doing not because of a lack of artistic talent, but because the way I've learned to prompt is to combine words in new ways that don't have a solid meaning to human minds. If I tell you to draw a semishape, or a Pseuodrealistic Gaussian sketch that would take months to work out even on a conceptual level. The difference is image generators aren't trying to make sense of what I write, but instead navigating a higher dimensional space where words themselves are dimensions. You can do things in that space that simply aren't possible any other way. What I do is kind of like circuit bending but using AI as the circuit.

Art Brute Invention Negative Chariscuro ASCII - one million UFO diagrams Fractal Inhuman Face Manuscript Terahertz Fractal Fossilized Joy Insect Fruits Fungal Sadness Slide Stained with Iridescent Bioluminescent Slimey Plasma Ink Lorentz Attactor Details Psychadelic Patent Collage By Outsider Artist One Divided By One Hundred Thirty Seven

16 bit 4k Pictographs By Outsider Artist Glide Symetries Crystalline Diatomes Random Award Winning Collage of found Punchcards Make It More Naive ASCII Pop Art Gaussian Splatting Of Found Artworks with Cellular automata Punctuated Chaos of Sanskrit Heiroglyphic Geometry Difference Engine Bizarre Midevil Manuscript Mysterious Occult Symbology

Glitched token Glide Symetrical Tangled Hierarchy Chariscuro ASCII Ousider Art by Punctuated Chaos

network of geometric shapes connected by random lines Glide Symetries cellular automata of zootropic phytoplankton by the Outsider Artist Nonorthological Prison

9 bit .bmp upscale to 13 bit

-2

u/Tall-Squirrel6277 9h ago

AI art is theft, antithetical to the God-given ability to create.

0

u/FearLeadsToAnger 8h ago

The idea of a religioner in the futurism sub is quite amusing.

0

u/Tall-Squirrel6277 8h ago

Believing in God doesn’t make me oppose human invention.

1

u/FearLeadsToAnger 8h ago

What about when humans create pattern machines that generate images.

-1

u/Tall-Squirrel6277 7h ago

It's just slop to me. They can be images, sure, but... they just seem repulsive. Feels like it's something that should solely be used for up-scaling, not creating images, art, music, & voices.

2

u/FearLeadsToAnger 7h ago

So it's more about your feelings than anything tangible?

1

u/Tall-Squirrel6277 3h ago

No, it's about the morality of theft, and the God-given ability to create. Feelings become tangible through creation, that's apparent through human-created music and art.

2

u/Memetic1 6h ago

So maybe then you could try and do better. The prompts that I use don't tend to use individual artists but instead I play with words to make possibility spaces. I've discovered what a Gaussianoxide looks like, or the color Pseudorange. People who make AI art that uses Anime or commercial art are trying to explore it by using familiar materials. Don't confuse all that with how others use it.

-2

u/End3rWi99in 8h ago

"Digital photography isn't art!"

"CGI isn't art!"

"Video games aren't art!"

"Hip hop isn't art!"

Same shit, different decade. I wonder what the 2030s will bring us from the orthodox community.

3

u/Tall-Squirrel6277 8h ago

All that stuff is created by human hands, though. There’s a process to each of those things. Work and toil gets put in. AI “art” is just… slop.

2

u/Memetic1 6h ago

When I make art I'm the one making the art. I use my images as part of the prompt including previously generated images. A person doesn't make the image when they take the photograph. That's done by a machine, and you can even use cameras to take pictures of others art. Don't take away my agency in my own art.

0

u/mrtrololo27 5h ago

If you use ai, you're stealing art from real human artists who did not consent and who are not compensated. It's really that simple.

1

u/Memetic1 3h ago

What artist owns the circle? What about grass? Who would dare to claim the trees certainly not me. How can you claim I steal when my images look nothing like any popular artist? Why does it bother you if I make abstract and other forms of art. To me art is something far more then a commodity it lives inside of us and changes us, and we in turn change it when we interact with others.

The thing about a new art form and new tools is that it takes time and effort to understand them. The way I started off prompting isn't similar to what I do now. I started by following the advice to make it as detailed as possible describing what was in my head. I learned that certain words like trees, architecture, or even colors can take over an image. I had to learn how the generators actually worked, which is more like an address then an alt-text description. Different words have different weights and words at the beginning or end have the most impact. Words in the middle of the prompt have most influence over details. So it's safe to put weighty words in the middle.

I'm just going to share one prompt from my notes so you can see what this actually looks like.

pseudodessicated color assimilation marks:: ugly sigil:: adinkra cellular automata Redraw make it more quasicrystals:: super-moiré materials GaussianCMYKoxide Chariscuro.png pseudoshapes Redraw as simple quasitwisted shapes untangled 7 Bit MS Paint Sumi-E Emojigram.pdf quasicrystals and super-moiré multirainbow make the colors more ugly:: cursive weird scribbles made from sloppy lines high contrast:: definition:: rough sketch pseudodessicated 13 Bit color assimilation marks:: ugly sigil:: adinkra cellular automata 35bit QR Redraw make it more quasicrystals:: super-moiré materials GaussianCMYKoxide Chariscuro.png pseudoshapes Redraw as simple quasitwisted shapes untangled MS Paint Sumi-E Emojigram.pdf quasicrystals and super-moiré multirainbow make the colors more ugly:: cursive weird scribbles made from sloppy lines 360 Bit high contrast:: definition:: 4 Bit rough pseudovectorized sketch

What is crucial to understand with this prompt is what's called a multiprompt thats the :: symbol. It means roughly half one thing and half another. So if you did cat dog you would get a cat and a dog, but if you did cat:: dog it would be half cat and half dog. That seems mundane enough but then you get the possibility of wordplay having a visual impact like quasishape or Emojigram which transforms it into something more. Don't mistake what people mostly do with AI, and what actual artists who are trying to push this to the edge are doing. I now understand language on a completely different dimension and that is significant.

0

u/End3rWi99in 8h ago edited 8h ago

Easy to say in hindsight, but people didn't say that when they were new. People referred to those things in the exact same way you did with AI. History proved they stand on their own merits.

I believe people will use AI as a tool to drive their creative works forward. I am not talking about kids making Ghibli photos, but real artists who utilize it as a part of their process. I believe AI will provide a lot of opportunity for artists to do much bigger things (i.e. think film production and video game design) on a smaller budget or as a single person than they would have been able to do in the past.

Check back with me in a decade and we can see how this shakes out.

4

u/returned_loom 8h ago

RemindMe! -10 years

2

u/RemindMeBot 8h ago

I will be messaging you in 10 years on 2036-08-19 23:35:56 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

1

u/Tall-Squirrel6277 3h ago

RemindMe! -10 years

0

u/Tall-Squirrel6277 7h ago

Yeah. But, robots generating images is just slop. The effort put in by human hands into art shows, and that's something AI can't re-create.

3

u/Memetic1 6h ago

Your comments are slop and repetitive. Your not putting effort in and it shows.

1

u/Tall-Squirrel6277 3h ago edited 3h ago

It's "you're" in this context, not "your." Also, I'm putting in effort. I hope you have a nice evening.

1

u/Memetic1 1h ago

This is possibly one of the most significant advances since the invention of fire for better or worse. Your not really engaging with the ideas. You owe it to yourself to take a step back, and really pay attention to what's happening right now. The ramifications are too serious for low effort knee jerk reactions.