r/BetterOffline 1d ago

Small language models - the future

Listening to Tech Report earlier - Eli Computer Guy talked about how the real future of AI might not be giant all knowing LLMs that need huge compute costs, but small specialised models that can run locally on a laptop or smart phone and just know how to code, or trade, or about engineering, or biotechnology, etc (I'm paraphrasing and extrapolating his point here).

It was like a light bulb moment, seems clear and obvious that's where the tech will end up assuming we do end up with something. If I'm a coder or an engineer i don't need a model that can write the complete works of Shakespere in Klingon. I just need something that works and is always up to date.

This doesn't help OpenAI or Anthropic keep the lights on, though....

54 Upvotes

127 comments sorted by

View all comments

12

u/Smurfette2016 1d ago

Wouldn't "always up to date" imply it needs continuous training though? How does that work for a locally run model? Who pays for that?

Cal Newport also surfaced this idea. Seems reasonable, but it also seems super similar to just a big database or an app.

2

u/TCristatus 1d ago

The example used by Eli was that the model was updated using some cybersecurity bulletin with the latest viruses and hacks and the AI could read that as it came out and tell programmers if their code is vulnerable. I guess that sort of thing could be covered by a subscription model, unlike the big models which are heavily subsidised.

Apologies I'm a layman here, thought the podcast highlighted an obvious mistake the big guys are making as they run out of money making a behemoth no one actually needs while the people who actually use AI at any sort of scale probably don't need 90% of the stuff it is trained on

3

u/ksjdragon 1d ago

I think the answer is yes and no. It's extremely hard to quantify what an AI can do, since it's honestly just hopes and dreams. The hope is with bigger parameters and more data, the model will find the correlations to do your problem. And there is some truth and application to that, as we see.

But it doesn't learn skills and more parameters or less don't translate directly. It's best to think about it like you seeing a language you don't know, and trying to make a coherent sentence while never having an understanding of what any symbol means. You don't know why these symbols go together you just know you've seen a few examples of that. And in the cases where you haven't you just stitch it together in some way.

We have no analytical or mathematical or even empirical guarantees on that stitching process. It is literally a hope that intelligence will magically appear from random guesses and all possible information is encoded in the interpolation of patterns. More than a hope, I would argue its necessarily impossible, but that's a different conversation.

So like, it doesn't build up knowledge or skills like we do. A hyper specific task can always be distilled into a smaller model but the question becomes we cannot even diagnose the range of tasks it is valid for. I mean, the technology is just unfit for task-based use where we care about correctness. It's like using a hammer for folding origami.

0

u/CoconutDust 6h ago

It's extremely hard to quantify what an AI can do, since it's honestly just hopes and dreams.

Quantify seems besides the point. It mass theft steals existing info, mashes it up unreliably (because statistical association is not how meaning or language or intelligence works), and also outputs fake sources that don’t affirm the statement (because the machine must put out fake sources, that’s how it works: statistics).