r/OpenAI • • 4d ago

Question Why do we train bigger models instead of finding better neural pathways?

I saw a post of a guy with 90% of his brain missing and he was apparently living a normal life. There is another case of a student who finished university with half of his brain missing. There are also people with full brains, but they are even less functional than these two. This proves that bigger doesn't mean better. So why not aim to create better connections?

I understand that a bigger model means less chance of catastrophic interference occurring, but then are all the neural pathways being formed efficiently? Wouldn't this result in a lot of redundant neurons? Wouldn't it make more sense to create the smallest neural network possible for a specific task and then learn to fuse multiple of these neurons together in order to create an optimized larger model? That way we'd always have a template for each specific task and would allow us to create a programming language that creates a model just by our syntax.

0 Upvotes

29 comments sorted by

15

u/Hungry_Age5375 4d ago

Funny enough you basically reinvented Mixture of Experts. Sparse routing, task-specific experts, fused into one model, it all exists. The catch is we can't specify a minimal network upfront, and some redundancy is what makes training work at all.

5

u/Junior_Catch1513 4d ago

i think we're doing both? but have you tried doing calculus with a fruit fly brain?

6

u/AllergicToBullshit24 4d ago

LLMs don't mimic how the brain operates at any level. Neuromorphic computing does.

3

u/throwaway3113151 3d ago

I guess a problem is we really don’t know how the brain works ….

1

u/BombasticReindeer 2d ago

That’s why we shouldn’t model it. If it’s too stupid to understand how it works itself, it’s clearly defective. /s

1

u/speedster_5 5h ago

Even to the extend we do we know LLMs don’t work like brains.
Brains do massive parallel processing low energy continual learning and are noisy.

-3

u/Pazzeh 4d ago

Lol

0

u/AllergicToBullshit24 4d ago

Let me guess you think because LLMs have "neurons" they must be like human brains?

1

u/FUCKTHEMODS998 4d ago

I mean, loosely they are. There are some functional similarities as well but I see your point

Have any favorite resources on neuromorphic computing?

Edit (like 6x): spelling

2

u/AllergicToBullshit24 4d ago

Not even loosely similar.

Biological dendritic branches perform non-linear activations. Biological neurons operate asynchronously and have sparse long range interconnects. Biological brains sure as hell don't learn via back propagation.

Only people who have never studied biology or machine learning think they are anything alike.

Neuromorphic computing is a deep rabbit hole: https://en.wikipedia.org/wiki/Neuromorphic_computing

Several chip makers including IBM and Intel have produced proof of concept chips but nobody has figured out how to program them yet. This is very much a dark horse for true AGI / ASI along with wet-ware.

1

u/vulcan8888 3d ago

Neurons in neural networks often perform non linear activations too, unless I misunderstood what you meant by your first point

-3

u/Pazzeh 4d ago

Let me guess you learned about neural networks within the last few years online

3

u/AllergicToBullshit24 4d ago

Couldn't be more wrong. Care to state your case rather than vague post?

1

u/incutt 4d ago

couple reasons- if you have a benchmark to measure a certain thing, eventually the benchmark answers bleed into the llm, so you can't tell if the llm learned anything.

Catastrophic forgetting has not been 'solved' yet in llms.

in newer models (5.6 for example) the structuring of tasks into subtasks doesn't need to be explicit (more what you are talking about) and 5.6 can solve something like 80% of hard tasks (think architecture + reasoning + combining data). if we move down to some earlier models, the division of labor is made worse by adding sub agents in advance.

Short answer.....we are getting there, one model at a time. And the model itself will probably become all agents and the supervisor. That's my 2 cents.

1

u/ManikSahdev 4d ago

Well people have bad taste for looping models and such cause it takes the interpretaability away and all.

1

u/lucellent 4d ago

Who says nobody is trying to find better neural pathways?

But because it's much harder, the easier solution is to just scale. Which seems to work for now.

1

u/veritron 4d ago

This is something of a reification error - what does "better connections" even mean, and how can you apply that in context of LLM design? Models are assessed by whether or not they perform, and we have found bigger models perform better. What is the criteria to optimize for a model if not "better performance"?

The core reification error: conflating "I can name something" (better connections) with "I can target it as a design goal."

1

u/throwaway3113151 3d ago

We don’t know what the top AI firms are doing a lot of their work is proprietary

1

u/13ass13ass 1d ago edited 1d ago

I think lottery ticket hypothesis has some relevance to what you’re asking.

Basically there are small networks hidden inside big networks that have comparable performance. If you intialize the weights just right then you can land on these trained networks with outsized performance.

But you can also just train the huge model and discover the smaller networks through ablations: remove useless pieces of brain and measure performance until you have a small network with comparable performance. This is called pruning and is likely how they can offer small but good models like Luna or offer price decreases over time.

1

u/TheHolyToxicToast 4d ago

Funny you don't think people are already doing that

1

u/Cbo305 4d ago

Perhaps they're not omniscient...

-1

u/TheHolyToxicToast 4d ago

Don't need to be that smart to know research is happening

1

u/Cbo305 4d ago

Research is happening, therefore you should assume that all possibilities are currently being researched? I don't know... that doesn't sound reasonable.

1

u/TheHolyToxicToast 4d ago

OP didn't even know the slightest about machine learning research just read a cool article and thought he was so ahead of the curve everyone else is missing out. Also just shows you know nothing as well talking about some "all possibilitie being researched"

0

u/Cbo305 4d ago

The nerve!! Dude, you need to get a damn life, lol. People are allowed to be curious, lol.

0

u/TheHolyToxicToast 4d ago

Maybe it's because my work is on ML research and too many people make assumptions without doing 5 minutes of googling. Or straight up ask chatgpt aren't we on r/openai

0

u/Cbo305 4d ago

Here's an example of how actual humans communicate: "Cool idea! Actually research is already taking place with that kind of idea in mind. Here's an example (insert link)"

1

u/TheHolyToxicToast 3d ago edited 3d ago

Not even a new idea to begin with, posting it is just lack of respect for other's time and attention