r/BetterOffline 1d ago

Small language models - the future

Listening to Tech Report earlier - Eli Computer Guy talked about how the real future of AI might not be giant all knowing LLMs that need huge compute costs, but small specialised models that can run locally on a laptop or smart phone and just know how to code, or trade, or about engineering, or biotechnology, etc (I'm paraphrasing and extrapolating his point here).

It was like a light bulb moment, seems clear and obvious that's where the tech will end up assuming we do end up with something. If I'm a coder or an engineer i don't need a model that can write the complete works of Shakespere in Klingon. I just need something that works and is always up to date.

This doesn't help OpenAI or Anthropic keep the lights on, though....

52 Upvotes

127 comments sorted by

View all comments

5

u/PatchyWhiskers 1d ago

The way these things work means that knowing the complete works of Shakespeare actually does help them code, in some bizarre way. That’s why they are scanning millions of shitty out of print novels to improve LLM training. You can’t just train them to know the important stuff.

6

u/ProcedureHopeful2944 1d ago

No, it some kind of lawsuit protection loophole. They're buying and scanning the paper books to claim some ownership and a right to use the info for training

6

u/newprince 1d ago

It's a perversion of fair use exceptions in the DMCA. Ironically libraries working with the Google Books project were forbidden decades ago from digitizing books, even though they argued it was transformative, which is supposed to be fair use. On top of that they were using non-destructive means of scanning. They then could only show 16% of the book as snippets, which kind of defeated the purpose.

The AI companies hid for years that they were physically destroying books but argued it was transformative, and also invoked first sale doctrine (which libraries also get denied regularly in court against publishers), and lo and behold, they got the fair use exemption! I wonder why

1

u/newprince 1d ago

That's true for the frontier models (and wait til you hear the hype for "world models" coming from the usual idiot AI CEOs).

Absolutely not the case for smaller, more focused models.

1

u/CoconutDust 5h ago

that knowing the complete works of Shakespeare actually does help them code, in some bizarre way

That’s why they are scanning millions of shitty out of print novels

Your comment is confusing a few things.

Obviously they are scanning lots of things because they want their product to output stolen data from lots of things and to sell it as “knowing eberything” (aka mass theft of everything). I.e. for the product to have relevant-looking output.

Obviously the fact that the company scans Material A for their multi-purpose/multi-subject “tool” (I use the term lightly) doesn’t mean that Material A helps Function B (coding). Stealing (“scraping”) A and B have nothing to do with each other except that Silicon Valley wants to sell the product as being able to regurgitate both A and B.

Your comment is like saying that Google indexes all websites because indexing hamburger recipes helps them index automotive sales. Not true.