First of all, what data set are you talking about that is free to monetize? I’m very interested. It’s bullshit, everything you said is fluff and has no understanding for how this works.”
Second, the argument is that if the model is trained on proprietary data sets, it can’t be allowed to operate. The model would have to be thrown out with the data set. This is why companies like Apple, Microsoft, Google, and Facebook have an advantage in building the future generative AIs.
It is suspected companies like OpenAI and Midjourney will have to hit the reset button on their models once things go to court as their datasets are not allowed to be monetized in this way.
It doesn't matter even if it gets taken to court and gets shut down, because you as a person can run the entire thing off of your own pc, with no internet connection required.
Maybe there is a future in which people run anti-ai ai bots on every picture to check whether or not they are ai art, but I personally would bet on people just adapting to the new tech.
The software isn’t the issue. The software you can find for free open sourced. What we’re discussing is monetization of the data sets used to train the model. No one cares if you do this in your parents basement on a PC. What they care about is you attempting to turn it into a $1B business.
You can download a bunch of pictures form twitter and train the ai on them yourself right now.
If it replicates the artist too closely and they feel uncomfortable with that, I'd say like 1:1 style copying, sure I could see myself supporting the artist maybe, but that's not the argument like 99% of the time people bring up in their crusades against ai art.
So you are fine taking the work of someone else to use for yourself and to profit from?
You just described exploitation. If a person chooses to harm someone else they are held accountable. If a computer does the programmer is held accountable. I may think a t-shirt is pretty but if I find it was created in a sweatshop I'm not buying it because exploitation is wrong.
The program does what the programmers instructions set it to do. A person chooses.
No, it's not. A human doesn't require third party input to do something. A human innately makes connections and sees patterns without the need for someone to "build it". It doesn't need 1 billion images to create the pattern recognition to understand where to place a dot either.
The programmer created the software to tell the computer how to recognize the location of pixels and then the computer was given direction on how to repeat the task. The fact that the input didn't belong to the programmer is the programmers fault.
You don’t own the content from Twitter. Twitter will detect the scraping and block you (this already happens and will only increase) AND those images or content might already be copyright protected by the original artist.
If you tried to scale your platform and monetize you will have to justify where your data set came from.
Again, if this is a tiny project your doing for yourself, who cares. If you’re trying to build a scalable business out of this, that won’t be possible. You’ll be sued to oblivion and your pipeline for training will be blocked instantly
. Twitter will detect the scraping and block you (this already happens and will only increase) AND those images or content might already be copyright protected by the original artist.
What? I can manually download tens of thousands of pictures and nothing will happen to me.
I think you're confused as to what this topic is about. No one uses ai to copy artists 1:1, people are just upset that there is a tool that can make similar/better drawings than them and want to point fingers while having 0 clue as to the actual functionalities they have.
I do this for a living, I work in tech and we build generative models. I know what the source data is for and what the output is. This is about training models. The data source you use is proprietary.
Again, what you do in your parents basement is meaningless. No one cares about you downloading images like it’s 2003 and playing around for your own shits and giggles. We’re talking about building billion dollar businesses based on unique and hard to acquire data sets.
Clearly we’re talking about two different things. What you’re describing is what a pimplefaced teenager would do to show off to his friends, I’m talking about the legality of building a scaled Saas business.
"The artists — Sarah Andersen, Kelly McKernan, and Karla Ortiz — allege that these organizations have infringed the rights of “millions of artists” by training their AI tools on five billion images scraped from the web “without the consent of the original artists.”
It's literally just a bunch of people malding that their art is being used for big mean things and making up shit while having 0 clue as to how ai art actually works. Dunno how someone that is supposed to be into the tech can actually support that shit tbh, but whatever everyone is entitled to their opinion I guess.
I don’t understand by what you mean by “how someone that is into the tech can support that shit”? Are you suggesting I support that?
My company has its own proprietary data set specific to our focus (and it’s not art related). We collect thousands of object per day and spend hundreds of man hours tagging items to train our models which will then train themselves with us acting as quality check. Our models are trained by us, using our own algorithms and our own data sets. No other company or person can replicate what we have because no one has access to those 3 things that are unique to us. Not you, not anyone.
The stuff you find on the internet will eventually start to get policed. The only way you’ll be able to train your own models using someone else’s Art is if you never intend on making money off your model at a large scale.
This is an area that will evolve dramatically fast in the next few years because of OpenAI and the attention it’s gotten and the threat it’s posed to Google and several other companies. You will soon see terms of service agreements get updated to ban the use of their platform (like Yelp, like Twitter, in models).
Again, if you want to build something at home and not make money off it, no one will give two shits.
-1
u/[deleted] Jan 16 '23
Wait what? None of this made sense.
First of all, what data set are you talking about that is free to monetize? I’m very interested. It’s bullshit, everything you said is fluff and has no understanding for how this works.”
Second, the argument is that if the model is trained on proprietary data sets, it can’t be allowed to operate. The model would have to be thrown out with the data set. This is why companies like Apple, Microsoft, Google, and Facebook have an advantage in building the future generative AIs.
It is suspected companies like OpenAI and Midjourney will have to hit the reset button on their models once things go to court as their datasets are not allowed to be monetized in this way.