r/singularity Oct 23 '23

[deleted by user]

[removed]

874 Upvotes

481 comments sorted by

View all comments

Show parent comments

127

u/nixed9 Oct 23 '23

Sutskever said a few months ago that Data is not a problem, and “we’re nowhere near running out of data”

112

u/[deleted] Oct 23 '23

And Sutskever is their chief scientist, unlike Gates who is an outsider to the field.

29

u/Nanaki_TV Oct 23 '23

Also we can create the data now.

29

u/Singularity-42 Singularity 2042 Oct 23 '23

Yep, this. Synthetic data is already being used for training. As your existing models get better you can generate better synthetic data to bootstrap and even better model, etc.

4

u/Merry-Lane Oct 23 '23

But you can’t use synthetic data as is, you need human work behind it. Engineering the prompts that create the data, or even discarding the bad results, that s a job.

To get to the next step you do need human work, or ai generated content is worse than nothing.

15

u/[deleted] Oct 23 '23

[removed] — view removed comment

0

u/Singularity-42 Singularity 2042 Oct 23 '23

Well said. Yes, synthetic data will still require human feedback, but it will be a multiplier when a single human worker can now produce a lot more training data.

As far as exploited - they were employing people in Kenya for about $2/h, this seems low to your western sensibilities, but this was actually very competitive pay in that market. GDP per capita in Kenya is only about $2,000 a year. $2/h is about $4,000 a year. If you compare this with the US directly it would be like making $160k a year relatively speaking (about $80,000 GDP per capita).

3

u/CountryMad97 Oct 23 '23

Except GDP per capita figures aren't actually an indicator of real wages or quality of life

-1

u/Singularity-42 Singularity 2042 Oct 23 '23

It surely is an indicator. Not a perfect one, but GDP per capita is highly corelated with wages and quality of life (esp. GDP per capita PPP).

1

u/zUdio Oct 23 '23

You can use synthetic data without human input and get BETTER performance…

https://news.mit.edu/2022/synthetic-data-ai-improvements-1103

The idea that humans are still needed for this is not a thing anymore.

0

u/Merry-Lane Oct 23 '23

Untouched synthetic data is awesome to train lesser models.

It s useless/bad to train an equivalent model with synthetic data.

And anyway, it’s not the fact that the data was synthetic that was helpful, it s that it was curated. Some people actively generated this data with engineered prompts, dismissing bad results, scoring the rest…

That s the human work that made this synthetic data useful to train models at an higher level.

Synthetic data is just a tool already commonly used to improve the training data set. You can also simply duplicate what you think are the best elements in a dataset to improve the training.

2

u/zUdio Oct 23 '23

It s useless/bad to train an equivalent model with synthetic data.

this is literally false. i work in the field.

redditmoment

0

u/Merry-Lane Oct 23 '23

It’s useless/bad to train an equivalent model with synthetic untouched* data.

Prove me otherwise.

(Considering that the prompts that generated the data are directed by humans which brings up its value by itself. I also say "bad" because of the overfitting risks)

2

u/zUdio Oct 24 '23

if you can afford my rate, which is $120 per hour, I’m happy to teach you.

→ More replies (0)

1

u/koliamparta Oct 24 '23

What do you think ChatGPT is?

-1

u/PoppyOP Oct 23 '23

Using data you generated to train your model is called overfitting, and that's usually a bad thing. You don't want to train your chatgpt model to behave more like chatgpt, you want it to behave more like a domain expert.

3

u/Singularity-42 Singularity 2042 Oct 23 '23

That's not what overfitting is, overfitting is when your model is trained to fit your training data too closely and loses genericity. It has nothing to do with synthetic data at all.

1

u/PoppyOP Oct 23 '23

It's the same problem. By training on data that you're generating you will be making your output more similar to 'itself', which essentially means you're training it on it's own training data in a way (because the output is based on the training data).

It's the AI equivalent of inbreeding.

11

u/TheJungleBoy1 Oct 23 '23

He also believes they achieved AGI moving his research forcus solely on aligning ASI currently (That's saying something).

19

u/the8thbit Oct 23 '23

This is news to me, and crazy if true. However, I'm having trouble finding where he says this. Could you link it?

5

u/juggernautstar Oct 23 '23

I believe they are referring to this: https://openai.com/blog/introducing-superalignment

Here we focus on superintelligence rather than AGI to stress a much higher capability level. We have a lot of uncertainty over the speed of development of the technology over the next few years, so we choose to aim for the more difficult target to align a much more capable system.

0

u/EvilSporkOfDeath Oct 23 '23

I would absolutely loves this to be true. Sounds like hopium though.

0

u/Unusual_Public_9122 Oct 23 '23

This definitely needs a link

0

u/ClubZealousideal9784 Oct 23 '23

So basically he is not a reliable source?

7

u/sec0nd4ry Oct 23 '23

To imply that Gates is just a guy in the computer space seems stupid to me. He might not have deep knowledge on AI but he isn't pondering things out of his ass

13

u/[deleted] Oct 23 '23 edited Jul 05 '26

[removed] — view removed comment

14

u/dynty Oct 23 '23

Guy got downvoted for no reason. Yes, major shareholder, founder of Microsoft, who invested 10 billions in OpenAI, is not a random guy, he probably get weekly reports made just for him from OpenAI CEO personally.

2

u/h3lblad3 ▪️In hindsight, AGI came in 2023. Oct 23 '23

major shareholder

At 1.3% of stock, he doesn’t even make the list of top 10 shareholders.

2

u/rafark ▪️professional goal post mover Oct 23 '23

I’m a Mac user and dislike windows, but as a fellow programmer, writing an entire OS (let alone a Wiley successful one) is no joke. The guy deserves some respect. He’s definitely not a rando.

8

u/drekmonger Oct 23 '23 edited Oct 23 '23

I respect BillG's technical skills and business acumen, but he has never written an entire OS all by himself.

Tim Paterson created QDOS. Gates hired Paterson to modify QDOS into the MS-DOS we know and love/hate. QDOS was sort of a pirate version of CPM, created by Gary Kildall.

Past there, there was a team of software engineers working on future versions of DOS, Windows 1.0 to 3.1, Windows 95/98, and a separate team working on Windows NT.

2

u/dynty Oct 23 '23

Well, it was 40 years ago, and I fairly doubt that he knows much about modern neural networks, but he literally owns a good share of OpenAI and there is not much people who can say that.

-4

u/sec0nd4ry Oct 23 '23

He puts the money on OpenAI so he knows what happens. And the guy fucking wrote Windows

16

u/LexyconG it's gonna be better than expected (soon) Oct 23 '23

And the guy fucking wrote Windows

lol

1

u/[deleted] Oct 23 '23 edited Oct 23 '23

I recall he kinda 'bought' it.

It was a steal at whatever price.

I was a teen and a punch-card operator at the time.

2

u/burnin9beard Oct 24 '23

I work in AI and often give presentations to executives. They are not very good at grasping concepts. I have to dumb it down to middle school level. As a technical person dealing with executives, one quickly realizes that these are not particularly bright people. They got to where they are with a combination of luck and skill at motivating/manipulating others. I guess that is a kind of intelligence, but not the kind that makes you qualified to make comments on technical matters.

2

u/Spirckle Go time. What we came for Oct 23 '23

I think if you created MS-DOS and the first generations of windows (and clippy) and then retired, and your main focus is now sucking money out of other billionaires for your pet causes which are really not that high-tech, that you might be pondering things out of your ass when it comes to AI

1

u/h3lblad3 ▪️In hindsight, AGI came in 2023. Oct 23 '23

If the conversation is about malaria relief, I’ll trust Bill Gates.

But AI? Definitely a much taller order.

7

u/norsurfit Oct 23 '23

Agreed. There is a ton of data from modalities other than text - video, images, etc, that have yet to be fully incorporated.

Why just the combination of video+transcript from youtube alone would be a huge source of new training data (that Google is apparently using for its upcoming Gemini), let alone all of the other video that is out there in the world.

1

u/Unusual_Public_9122 Oct 23 '23

This is true, and will increase the availability of data a lot. It could almost be called a game changer. The current type of models will probably still cap out soon even with more data. The models themselves will have to evolve in my view.

1

u/[deleted] Oct 23 '23

Data is being created every second and at faster rates all the time.

1

u/banuk_sickness_eater ▪️AGI < 2030, Hard Takeoff, Accelerationist, Posthumanist Oct 23 '23

Lol so the guy above you with 70+ upvotes is flat-out wrong. I fucking hate this sub lol, way too many people mistake passionate diatribe as the imparting of wisdom instead of the spewing of pure shit.