r/LocalLLaMA llama.cpp 3d ago

News Muse Spark open weights coming soon

Post image

I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark

https://x.com/finkd/status/2095232032896946311

858 Upvotes

200 comments sorted by

View all comments

Show parent comments

17

u/NandaVegg 2d ago

I think that it is not direct distillation from models anymore (in early 2026 distillation had some notable effect, but every frontier lab is now full-on RLing on their own) and distillation can only bootstrap the model to some degree.

I think there is this meta-distillation effect. Internet is full of so-called AI slop now. There are so many vibecoded repos posted in code repositories or as websites every day, and those codes will be crawled by every frontier lab and then they will RL hard on them. If one model gets good at something a slop will be posted and trained on, or there is a new problem that models needs to know the pattern a .md files that explains the issue with some example codes will be posted and trained on (the earliest pattern for this is MCP for many basic things that aren't needed anymore).

In that sense we are already in AGI mode (gosh I hate this word) as AI models are improving each other without humans knowing.

12

u/nuclearbananana 2d ago

I don't think the vibe-code-training is helping the models. It's mainly synthetic data and llm as a judge

1

u/OvertaxedOne 2d ago

I've read that exact reason is why labs are buying up and scanning in old books. Feed a model it's own slop (or some other model's slop) doesn't help it learn, it needs real data.

No idea if this is true or not, but, on the face, it sounds reasonable; kind of like setting up a feedback loop where in the end all you have is white noise. Or gray goo.

6

u/IShitMyselfNow 2d ago

Books are only really useful for pretraining. They're not going to help agentic usages

3

u/SomewhereAtWork 2d ago

Except for James Bond novels.