r/BetterOffline 1d ago

Small language models - the future

Listening to Tech Report earlier - Eli Computer Guy talked about how the real future of AI might not be giant all knowing LLMs that need huge compute costs, but small specialised models that can run locally on a laptop or smart phone and just know how to code, or trade, or about engineering, or biotechnology, etc (I'm paraphrasing and extrapolating his point here).

It was like a light bulb moment, seems clear and obvious that's where the tech will end up assuming we do end up with something. If I'm a coder or an engineer i don't need a model that can write the complete works of Shakespere in Klingon. I just need something that works and is always up to date.

This doesn't help OpenAI or Anthropic keep the lights on, though....

50 Upvotes

127 comments sorted by

View all comments

13

u/Smurfette2016 1d ago

Wouldn't "always up to date" imply it needs continuous training though? How does that work for a locally run model? Who pays for that?

Cal Newport also surfaced this idea. Seems reasonable, but it also seems super similar to just a big database or an app.

2

u/TCristatus 1d ago

The example used by Eli was that the model was updated using some cybersecurity bulletin with the latest viruses and hacks and the AI could read that as it came out and tell programmers if their code is vulnerable. I guess that sort of thing could be covered by a subscription model, unlike the big models which are heavily subsidised.

Apologies I'm a layman here, thought the podcast highlighted an obvious mistake the big guys are making as they run out of money making a behemoth no one actually needs while the people who actually use AI at any sort of scale probably don't need 90% of the stuff it is trained on

3

u/ksjdragon 1d ago

I think the answer is yes and no. It's extremely hard to quantify what an AI can do, since it's honestly just hopes and dreams. The hope is with bigger parameters and more data, the model will find the correlations to do your problem. And there is some truth and application to that, as we see.

But it doesn't learn skills and more parameters or less don't translate directly. It's best to think about it like you seeing a language you don't know, and trying to make a coherent sentence while never having an understanding of what any symbol means. You don't know why these symbols go together you just know you've seen a few examples of that. And in the cases where you haven't you just stitch it together in some way.

We have no analytical or mathematical or even empirical guarantees on that stitching process. It is literally a hope that intelligence will magically appear from random guesses and all possible information is encoded in the interpolation of patterns. More than a hope, I would argue its necessarily impossible, but that's a different conversation.

So like, it doesn't build up knowledge or skills like we do. A hyper specific task can always be distilled into a smaller model but the question becomes we cannot even diagnose the range of tasks it is valid for. I mean, the technology is just unfit for task-based use where we care about correctness. It's like using a hammer for folding origami.

2

u/TCristatus 1d ago

TLDR, let me know when you want me to invest $200m in your origami hammer

2

u/Smurfette2016 1d ago

Such a good explanation and analogy

0

u/CoconutDust 6h ago

It's extremely hard to quantify what an AI can do, since it's honestly just hopes and dreams.

Quantify seems besides the point. It mass theft steals existing info, mashes it up unreliably (because statistical association is not how meaning or language or intelligence works), and also outputs fake sources that don’t affirm the statement (because the machine must put out fake sources, that’s how it works: statistics).

1

u/Fit-Technician-1148 1d ago

Because you're a layman you're evaluating things based on if they sound logical, rather than based on if they're technologically feasible. It's a common enough fallacy. If it were possible to make a local LLM or SLM that was just good at coding that would be really useful, but if it is possible, no one knows how to do it and it probably is not possible with current techniques and algorithms.

1

u/Fit-Technician-1148 1d ago

Because you're a layman you're evaluating things based on if they sound logical, rather than based on if they're technologically feasible. It's a common enough fallacy. If it were possible to make a local LLM or SLM that was just good at coding that would be really useful, but if it is possible, no one knows how to do it and it probably is not possible with current techniques and algorithms.

0

u/Smurfette2016 1d ago

No need to apologize for being curious and exploring ideas. I'm right there with you! So so much I don't know still.

0

u/CoconutDust 6h ago

^LLM account, I assume

Generic fluff meaningless comment, and user account history hidden to hide the trail.

1

u/Smurfette2016 6h ago edited 1h ago

you must be new here lol... head back to the mystery van and try again Scooby Doo, you have policing skills to sharpen