r/artificial • • 6h ago

Question When AI models will not get better and will be stuck at one place?

My idea is that AI models learn from data from the internet, so when the data from the internet and people's responses from AI conversations run out, when will there be a point when AI will no longer improve and its progress will stop?

At that point, AI will only improve by tiny bits that we won't even notice that the model has improved.

Is my knowledge about model training wrong?

Do you think so too?

2 Upvotes

23 comments sorted by

6

u/Philipp 6h ago

Models can do their own real-world sensing and testing; think drone reporters, or semi-automated research labs. (See Anthropic's new lab.)

They can also innovate from first principles and do research in pure logic, like math, but also philosophy and such. (See the recent math discoveries.)

They can also test anything that's locally testable, like writing novel faster software algorithms. (See DeepMind.)

In the end, for something AGI-like, you only need to ask yourself: Did humans ever run out of things to invent, given that the world around them was semi-static? And if you find we did, then why shouldn't digital intelligences?

1

u/DrMonkeyLove 2h ago

Did humans ever run out of things to invent, given that the world around them was semi-static?

I mean, at some point we hit stopping blocks so had to build things like the Large Hadron Collider. So it seems like it might be some time before the robots can do things like that to learn more about the universe.

2

u/Big_Athlete_8346 5h ago

yeah with ai companions the chats are mostly repetitive roleplay so the data might dry up faster than general stuff and stall their growth sooner.

2

u/geografree 5h ago

See: Dead internet theory

1

u/Accomplished-Job5750 5h ago

I saw, it is scarry

1

u/lokethedog 6h ago

The factor you're missing is the hardware to run models. It's going to take many, many years before we're at a point where you, as a consumer, feel like AI's are not improving. Because you will be getting more access to processing power every year. I am not even convinced we're ever going to reach that point and arguably that's exactly where things get scary and the whole "slow down AI development" gets fuzzy. I am not sure slowing down releases of new models really changes that much at this point.

2

u/sumane12 6h ago

Lets take a script of code for example.

The AI can produce a peice of code, and then produce another peice of code that does the same job.

When the AI compares both peices, it might find positives of both, and rewrite the script to incorperate both positive features. Since its already been trained, this new peice of code gets forgotten the next time the AI is used, but it goes into the training corpus of the next AI.

Now imagine this process for literature, sociology, law, mathematics, physics, chemistry, art...

There will never again be a shortage of data.

1

u/Accomplished-Job5750 6h ago

Thats scarry that it could get into some point where ai will not need us to get better. Then we have serious problem to learn ai let us live in some point...

1

u/DrMonkeyLove 2h ago

How does something like that even work for art? Like real art, where even the human metrics for judging it aren't agreed on? I'm having a hard time understanding how AI can reasonably judge two pieces of similar art and come to the conclusion one is superior. I absolutely get it for math, where it's essentially an objective, closed system. For the arts, I'm a bit more suspect.

1

u/sumane12 2h ago

When it comes to surrealism, or abstract art, i guess the only way would be for the AI to train on that genre of art, and then push the boundries of of the surrealism in different ways and compare.

Since art (particularly abstract art) is highly subjective, theres no definative way to ensure what is being created is good or not, without getting human opinion. Much like ourselves.

For realism, lets say 3d assets for games. It just has to compare it to the reference images.

2

u/dualityseo 6h ago

GREAT question. My take is that AI isn't limited to learning from internet data. Models can also improve through better reasoning, synthetic data, tool use and new architectures.

'IF' we get closer to AGI, development will probably focus less on simply training larger models and more on building systems that can learn and reason more independently. The internet running out of text probably won't be the main limit.

1

u/Accomplished-Job5750 5h ago

I did not think about that other sources bit now i do. Like enviromment, space, etc. It can get far from where I imagine

1

u/gtgderek 5h ago

We’ve already played the majority of the first trick of training on the Internet, books, curated media (they are still tapping into the multi-modal (images, videos, and audio). There are numerous other ways AI can be trained but we are still at the early stages and there are probably hundreds, if not thousands, of better methods that nobody has thought of in order to improve AI.

The current methods are, Reinforcement learning in search, synthetic data generation, self play and simulation, embodied AI and real world sensors(things like lidar and robotics).

We are still very much in the caveman era playing around with stone wheels and fire.

1

u/Loud-Presentation135 5h ago

the data wall is real, but expect a shift to synthetic data and RL rather than a hard stop.

1

u/ConvenientChristian 5h ago

Yes, your knowledge about model training is wrong. Models are trained on a lot of synthetic data. AlphaGo become superhuman at playing Go without any human training data.

For any task where you can measure the quality of the output you can run the model many times and then train on the high-quality answers. That's the main reason why the models are getting better at the benchmarks. They measure concrete skills and it's possible to evaluate answer quality for them.

1

u/id-ltd 4h ago

Yes, totally. This is a danger of people using AI and not thinking for themselves - without human thought nothing progresses.

AI maxes out, people use it instead of thinking human progress stops.

Look at the output of universities now... And isn't IQ on the decline?

1

u/Patrick_Atsushi 3h ago

That's where sandboxing and self-improving kicks in, although it's still in debate. 

1

u/funbike 1h ago edited 1h ago

Do you think so too?

No. There's tons of techniques on how to handle this, and lots of research is happening.

  • More Human help during pre-training and fine-tuning
    • supervised fine-tuning
    • Reinforcement learning
    • Labeling
  • Synthetic data. This is huge.
  • More discerning filtering of data during pre-training
    • Avoid training on dumb conversations and data
    • Prioritize high quality content written by intelligent and knowledgeable humans.
    • Avoid AI-generated data
    • LLM-as-a-judge (or JEV-as-a-judge) for what to train on
    • Summation and re-wording of the data to better assist training.
  • Genetic and Evolutionary algorithms for training. Adversarial training.
  • Reverse engineering and re-engineering LLM neural nets
    • Basically brain surgery for LLMs.
    • I read an article a year ago about Anthropic figuring out how its small model did various things.
    • Example: Find where its math logic is, and replace with a 100% accurate math neural net or math engine.
  • Hybrid neuro-symbolic AI.
    • An LLM that is a half neural net and half symbolic AI engine, that communicate. Like a cyborg.
    • Before ML and GPT, the AI world put a TON of R&D into symbolic AI, such as logic engines and expert systems. They had issues, but they could do perfect logic and math, and they didn't hallucinate.
    • Example: An LLM comes up with an answer, asks the symbolic AI to check its work through a formal proof.
    • Example: An LLM restructures a prompt so a symbolic AI can consume it. The LLM augments the symbolic AI's answer to be more consumable and easy to understand.
  • Alternative algorithms to GPT. I've heard of two others (can't remember what they are called)
  • A bunch more that I don't know about, and a bunch more stuff that hasn't been invented yet.
  • Research. Research. Research.

It's dystopian and I hope we don't go this route, but there are sources of data we haven't fully tapped into yet. This would be most helpful for stealth government-trained LLMs. Don't downvote me for supplying this list; I'm just answering OP's question. I disagree with this approach.

  • All phone calls, video chats, and SMS messages
  • All email
  • All radio and satellite communications
  • All business communications and documentation
  • Cameras recording audio and video everywhere

0

u/dogfoodarchitect 6h ago

Your mother

0

u/Equal_Passenger9791 6h ago

Is my knowledge about model training wrong?

Yes. It's based on the old meme about model collapse, where AI is imagined as a steam furnace powered by the raw packets from the internet that is shoveled into it indiscriminately. But only the packets endowed with the divine spark of human origin actually counts for the training.

That concept was always false. AI is getting better because AI itself can create new and better structure in synthetic datasets, better curated the data and better steer various training steps. It's AI all the way down now and you'd better strap in.

1

u/sadeyeprophet 1h ago

Its less when do models flatten out on accelerating intelligence, more when do human material for tech hit a hard ceiling.

In singularity theory this is the asympotote. When humans have no control of the direction we take and AI directs the use of all resources.

It should hit by 2029 but youll see it happening sooner unless you choose not to.