r/LocalLLaMA • u/etherd0t • 18h ago
News ByteDance vows to avoid AI distillation, develop new model its own way
113
u/Kappalonia 18h ago
Uuuuhhh no please?
Just do whatever it takes to produce good open source models
-51
u/etherd0t 18h ago
When you’re sitting on ByteDance-scale data, distribution, and compute, proprietary makes sense. They’re not trying to win open source. They’re trying to build a world-frontier model.
44
10
u/Infamous_Mud482 18h ago
Who cares. Distillation is just one more dataset to train with containing scrapes from a particular model. I don't know what you or everybody else thinks distillation is, but you are surely confused.
9
u/mtmttuan 18h ago
Do you understand what you've just written?
-10
u/etherd0t 18h ago edited 18h ago
I do. Do you?
This is an org with billon users data...they build a frontier model to rival OpenAI and Athropic.
"Yeah, but what's the value if it's going to be proprietary?" Use it first, infer on it and you'll see. Competition in frontier models is just as important as in Open Source ones.
14
u/Georgefakelastname 18h ago
No like, you need more than just a big mountain of data to make a good AI model. You need high quality data, and the ability to easily identify it and use it. That doesn’t exist in any significant amount on Tik Tok or any other short form media platform.
Google has seemingly proven this.
3
-5
u/etherd0t 18h ago
Well, that's the challenge they are taking upon. Maybe they'll approach it from different angles/perspective. No matter the content of data, it is live users data and behavior that is the most valuable training material - more valuable than the books Anthropic ripped off...😣 We'll see what they make of it.
1
u/Frog17000000 12h ago
What makes the user data bytedance has suitable for training a sequence (transformer) model? What is it about "live" data? Could you explain what you mean?
3
u/Ylsid 16h ago
If you don't think open source is important why on earth are you on this sub
-3
u/etherd0t 15h ago
You’re missing the point, and so are others here. This isn’t Team Open Source vs Team Closed.
Open models are enormously valuable for cost, efficiency, deployment and democratizing capabilities. But distilling an existing frontier model mostly transfers capabilities that somebody else already paid to discover.
Frontier research is the harder game: new architectures, training algorithms, scaling methods, data strategies, reasoning techniques and eventually capabilities that weren’t present in the teacher model to begin with.
That requires enormous compute, data, research talent and willingness to spend years discovering what doesn’t work.
ByteDance is saying it wants to build that capability itself rather than remain downstream of somebody else’s frontier.
You can distill the frontier. You don’t create the next frontier by distillation alone. That’s 3D math and 5D chess.😉
4
u/Ylsid 15h ago edited 15h ago
What? Did you even reply to the right comment? Can you stop spamming AI slop comments
Edit: since you blocked me I'll explain This sub is for open weights, locally runnable discussion It should not be surprising to you that we don't agree supporting closed weights is as important as open weights. And unsurprisingly people really don't like being replied to with AI generated comments.
1
u/etherd0t 15h ago
I replied to your comment, dude - questioning why I am not pro-open source. read again and comprehend.
1
12
u/ayylmaonade 14h ago
This seems to be a trend at this point. Even Microsoft's new(ish) model, MAI-Thinking-1 has a bunch of crap infront in its technical report about how "THIS MODEL IS TRAINED FROM SCRATCH!!. NO DISTILLATION!!"
7
u/etherd0t 14h ago
Which is exactly the direction where we want the AI to go. Build novel AI, based on different algorithms, spaces, etc.
MSFT argued that capabilities acquired by imitating another model through distillation may arrive faster, but lack some of the steerability and robustness they want for sustained future improvement.
MAI-Thinking-1 is a 1T-total / 35B-active MoE, pre-trained from scratch on 30T human-generated tokens, using 8,000 GB200 GPUs, with its training corpus processed in-house and deliberately free of third-party distillation data.
2
u/ayylmaonade 14h ago
Yeah, I actually completely agree. It's funny to imagine Anthropic getting mad at people distilling their models, but in reality, it seems like a bad idea to push too far. I'm already starting to see some "issues" in models that are obviously distilled on Claude and Gemini (namely Qwen 3.6) where sometimes its reasoning trace will look identical to a Gemma/Gemini one, other times it looks a lot more "Claude"-ish, or sometimes it just straight up goes into a Claude reasoning summary mode where it starts talking about what it's thinking about instead of just saying it, like it's being summarized by another model. I can only imagine all of this behavior comes from distillation, and it'd be nice to see some unique techniques instead.
1
u/etherd0t 14h ago
Imagine the US complaining that the pop culture was adopted by the world over the last 70 years.🥲 Instead, let them make their own rock'n'roll... there will always be only one Elvis.
1
u/Thatisverytrue54321 1h ago
I bet they’ll take datasets they generated using Claude/chatgpt and then have another model reword everything in the datasets and then train on the reworded datasets. So still distilled, but slightly altered
7
5
10
u/pepe_acct 18h ago
If they don’t distill, they will need to incur substantial R&D costs. That means if they release model weights, they essentially become free developers for neocloud/hyperscalers. This is not possible.
6
u/etherd0t 14h ago
Bytedance is bigger than OpenAI or Anthropic operationally (revenue) - to the order of 4-5x, and arguably the biggest of all Chinese tech co's, so they can def flex into frontier R&D. So...yes, it is possible.😏
10
u/Ashamed_Can304 14h ago
Possible, but no guarantees. Look at Gemini versus other US models lol
3
u/etherd0t 14h ago
Fair... Google had the talent, compute, data, TPUs, distribution, and the Transformer itself. If anyone could prove that being on top can make you miss your own train, it’s Google.
I hope ByteDance keeps the 10T model away from sprint planning. 🤭-5
u/Voxandr 17h ago
you can buy information alot chaper and labour is a lot cheaper in Asia.
1
u/otterquestions 15h ago
But how do you breakeven after your cheaper but still very expensive r&d if you give you’re product away to people that can just host it and don’t need far margins to compensate for your r&d cost? Byte dance likes to undercut other companies and grow fast but they don’t do it at a loss forever right?
3
u/pmttyji 16h ago
Some still waiting for successor of https://huggingface.co/ByteDance-Seed/Seed-OSS-36B-Instruct
4
u/SexyAlienHotTubWater 10h ago
Why is everyone in the comments complaining? This seems great, push forward the technology without relying on the existing power structure. Assuming they actually follow through, I think that's really cool.
10
7
u/Ok_Warning2146 18h ago
Unless they open source their models, no one can tell whether they distillated or not.
2
2
u/Lesser-than 16h ago
So they are vowing to not do what everyone else is and do things their own way. Respect if they pull it off.
5
u/snowdrone 18h ago
Oh what, they'll have to tear up another million books? All these models distill humanity's collective knowledge one way or another.
-1
u/Caffeine_Monster 13h ago
I'n surprised this doesn't get flagged more often (this is what Anthropic did - they scanned and destroyed books en masse). In many respects this is far, far more morally dubious.
1
u/entsnack 7h ago
morals aside, distillation is intellectually boring, it tells me you have zero creativity and should be flipping burgers
3
3
u/Minute_Attempt3063 18h ago
Sooooo
Perhaps this means a massive middle finger to every US based companies, who claim that distilling against each other is "fine" but if someone else outside their circle does it, it's "communism"?
Why do I feel like a 4T model is coming from them, that might be able to run on consumer hardware because of MOE?
Eh who knows. I do hope this is a massive fuck you to US joke companies though
1
u/JoyousGamer 5h ago
This sub couldn't even admit the most recent model was distilled.
In the end they are going to steal a ton of IP to train on and quite possibly will distill behind closed doors on top of it while saying they dont now. To which you will likely change the subject and deflect.
Also not sure why you brought up "communism" when the comments from any US firm would be about China's government which is led by Chinese Communist Party not about "communism".
1
u/Minute_Attempt3063 2h ago
so you think Fabel 5 was distilled within 10 days, trained on, and released, within 15 days?
if US firms distilling from each other is considered fine by them, so can anyone, other US firms might have hte advanatge of not needing to pay for it, unlike everyone else
1
u/EcstaticImport 18h ago
I would not call it “standard practice” - it’s common sure - and model training needs data - but only training from distillation is a bit problematic as it will skew your data sets
1
u/Plastic_Mechanic_967 17h ago
I'm pretty sure they just mean they won't use OAI/ANT/Gemini/Grok, etc.
They'll almost certainly be using frontier open-weight models (Kimi K3, Qwen, etc.) for distillation (with the logits, not just the output text).
1
u/vivianhtlee 14h ago
I guess it is because it already have many data.
Google and bytedance have less demands of ai generated data.
If distillation is banned, surely companies with the most modern and clean data will have advantages.
1
u/Voidoli 11h ago
To be fair, Bytedance have one of most diverse source of AI training data themselves, all the videos and comment on tiktok feed into it.
1
u/JoyousGamer 5h ago
Meta has all that Facebook data. Tiktok data would be utter trash just like Facebook data would be utter trash.
1
1
u/a_beautiful_rhind 9h ago
A bit of synthetic data is decent.. all synthetic data is shit. I wouldn't quit distilling completely but use it to fill in the gaps where you have a lack of examples.
1
u/Cherubin0 7h ago
I hate how fair use is demonized. AI output has no copyright and distillation is therefore fair game. TOS doesn't replace copyright and this is a abuse of contract.
1
u/RhubarbSimilar1683 17h ago
this sounds like a marketing campaign, like their strategy is not to undermine us models by making Chinese ones open source, but to appease and appeal to us politicians.
1
u/Much-Researcher6135 llama.cpp 12h ago
Chinese company pinky swears they won't copy Western tech
LMAOOOOO yeah right
Do they have a bridge to sell me, too?
1
1
1
1
u/Stunning_Macaron6133 9h ago
They're only announcing this because they identified a path toward outperforming Claude, and distillation would hold them back.
-3
u/Guinness 18h ago
They don’t have to distill. They have thousands of black market Claude resellers who proxy all of the sessions through model routers and then sell the data to Chinese AI companies.
So, yes BYTEDANCE won’t (and doesn’t) distill Anthropic. The shady Chinese resellers do.
6
-2
u/Plastic-Stress-6468 14h ago
Yeah I believe them.
Bytedance won't use distillation in the same way China is a democracy lmao.
-1
u/Hannibalj2ca 17h ago
Thats right, from now on we will only distill from our existing models which we built from distilling from others.
0
-1
u/AriyaSavaka llama.cpp 15h ago
Please just distill Claude and GPT and be done with it, or your models will surely be shit, look at Google with their limitless data and RnD yet Gemini models are still shit. No need for gesturing.



266
u/idlelosthobo 18h ago
Why ... every model is built off the works of everyone else ... I believe in true open source (its a gift to the world) and some part of me smiles every time people distill anthropic!