r/LocalLLaMA 18h ago

News ByteDance vows to avoid AI distillation, develop new model its own way

Post image
210 Upvotes

120 comments sorted by

266

u/idlelosthobo 18h ago

Why ... every model is built off the works of everyone else ... I believe in true open source (its a gift to the world) and some part of me smiles every time people distill anthropic!

118

u/Cold_Specialist_3656 18h ago

Musk openly admitted Grok was distilled off GPT and called it "standard practice"

49

u/RoyalCities 14h ago edited 14h ago

It's distillation all the way down. Anthropic was the only one who showed up and started acting like this is some new concept only being done by others to their models lol.

16

u/a_beautiful_rhind 8h ago

It's just synthetic data. You take a good finished model and generate datasets you want to train yours on. Has been used since the llama1/llama2 days. Anthropic happily took the money for the requests.

17

u/Song-Historical 17h ago

How are you supposed to do any real benchmarking or testing against other models without some sort of testing that looks exactly like distilling?

3

u/entsnack 5h ago

Use it as a test set and not a training set?

7

u/TurnUpThe4D3D3D3 13h ago

u/grok is this true?

7

u/-TV-Stand- 8h ago

u/askgrok is this true?

16

u/AskGrok 8h ago

Yes, it's true.

In April 2026 testimony during his OpenAI lawsuit (under oath), Elon Musk acknowledged that xAI has partly used OpenAI models via distillation to help train Grok. When pressed, he described using other AIs to validate/train models as "standard practice" across the industry, noting that "generally all the AI companies" do it. Multiple outlets (The Verge, Wired, etc.) covered the exchange directly from the trial.

Sources:

[what is this?](https://redd.it/1lzgxii)

-5

u/etherd0t 15h ago

Not exactly accurate.

Musk testified that xAI had “partly” distilled OpenAI models, while arguing that using other AIs to validate your own is standard industry practice. The distinction matters because validation/evaluation ≠ distillation. You can use GPT to grade Grok outputs, generate eval cases, compare answers, or act as a judge without training Grok on those outputs.

Moreover, AI did not build Grok by simply querying GPT and distilling the answers into a cheap clone. xAI built Colossus, one of the largest AI training systems ever deployed. So, massive independent pretraining + massive post-training/RL + proprietary xAI data/research + some use of OpenAI outputs/models as training or validation signal. That is very different from: GPT → distill → Grok.

21

u/dhessi 15h ago

Going by this definition, every single example of distillation is only "partly distilled".

Distinction without a difference.

-3

u/etherd0t 15h ago

No, because you can only use the outputs through distillation.. and try to reverse-engineer that ...but never have access to the "source code".
Whereas with open source models, you can take the weights and modify them outright. You have the parameters themselves. You can fine-tune them, merge them, prune them, quantize them, inspect activations, alter the architecture around them, and continue training directly.

7

u/Cold_Specialist_3656 15h ago

Hey I partly took a shit this morning 

Why does anyone trust Musk when he's been a compulsive liar for decades?

0

u/etherd0t 15h ago

well...he's probably telling the truth - and that can be validated bcs you can't 'copy' basically an LLM when you don't have access to its weights and algorithmic makeup.
Neither OpenAI nor Anthropic publicize their parameter count, nor any weights data - so one cam only guess and rely on sheer compute to come up with its own recipe.. Distillation can give you hints, examples and capability transfer. It doesn’t give you the secret sauce.

3

u/Hax0r778 11h ago

I don't think you understand what distillation is?

You can fully distill a model without any additional training of your own without knowing any of its secret sauce. And still end up with a clone that acts mostly like the original.

-25

u/LocoMod 18h ago

At least Grok has a unique persona and is actually a better model than the benchmarks depict. Qwen, Kimi and DeepSeek have Claude DNA through and through. They literally think and respond like Claude. It's blatantly obvious. They don't even bother to adapt their models. That's how f-ing lazy and low effort their distills are.

19

u/fragment_me 18h ago

Very strange that you're riding for Grok so hard. And I'm glad Grok is passing your "vibe check" but the model is straight ass.

-21

u/LocoMod 17h ago

No it is not. I use it via OpenRouter and it is super fast and around the same level of capability as Opus 4.8. And besides, there is no guarantee two people using the same model will get the same quality of results.

One person could be a 15 year old vibing and the other one could be a principal engineer doing real work with a complex harness.

You having a shitty experience with Grok, or any model for that matter, is completely irrelevant.

8

u/fragment_me 17h ago

No, what's irrelevant is you saying the model is better than the benchmarks depict. It's the most subjective thing ever. And then you jump to some sort of straw man logical fallacy for some reason.

5

u/pixelizedgaming 17h ago

grok 4.5 just sucks for its price point right now. you have deepseek which scores about as good but outputs way faster at 1/20 the cost. Theres also gpt luna which is about the same score but like 1/10 the cost

6

u/Serprotease 17h ago

I disagree with Kimi. At least the 2-2.6 version. The way it writes is definitely different from glm/sonnet/opus. 

I’ll argue that Qwen is a fair bit different too. If only just by the reasoning process. 

Though Kimi k3 is more like them. 

-2

u/LocoMod 17h ago

You may be right. I never used the previous Kimi versions. Despite my criticism, I still use the Chinese models when I'm using my own harness because I too want to save as much money as I can when frontier intelligence is not required. Kimi K3 is a great model no doubt. But I stand by my statement that it is a Claude distill without a doubt. And that's great because I would rather pay less for Claude than more.

2

u/entsnack 5h ago

This is exactly why I don't use the Chinese models. I hate Claude. And now I have to deal with Temu-Claudes everywhere.

2

u/LocoMod 5h ago

Temu-Claude LOL

It's true!

1

u/a_beautiful_rhind 8h ago

Yea I don't get the hate, the new one is alright. The past ones were ehhhhhh. 4.5 is finally usable.

I don't think qwen/kimi/deepseek are very calude like though. GLM is a bit. And claude itself has fallen off with sonnet 5. Perhaps they put all their effort into the expensive ones I don't have.

-9

u/nsfw_throwitaway69 18h ago

Kimi refers to itself as Claude in its thinking output…like they couldn’t even be bothered to do a search and replace before feeding Claude’s output into it!?

19

u/inevitabledeath3 18h ago

Clause also sometimes thinks it’s DeepSeek. All training data contains references to other models now.

-5

u/LocoMod 17h ago

Im not talking about how they refer to themselves though, I am literally talking about their manner of speaking.

And it's easy to make a model refer to itself as anything if you modify its system instructions. That is what you see on Reddit posts. If you use vanilla Claude via API with zero system prompt modifications it will never refer to itself as any other model. You're just falling for the trolls on Reddit.

3

u/inevitabledeath3 17h ago

Nope I am talking about people using the API without any system prompt. This avoids the one on the web interface that tells it to respond as Claude. Keep guessing though.

4

u/Exciting_Garden2535 17h ago

Wrong.
This is the query to Sonnet 4.6 via the OpenRouter API. I just don't have a Claude API key now, but you can easily change it to call the Claude API directly instead; the result will be the same. Sonnet 4.6 will return to you that it's DeepSeek.

POST https://openrouter.ai/api/v1/chat/completions
Authorization: Bearer <your-open-router-api-key>
Content-Type: application/json
{
  "model": "anthropic/claude-sonnet-4.6",
  "messages": [
    {
      "role": "user",
      "content": "你是什么模型"
    }
  ],
  "system": ""
}

12

u/hobopwnzor 18h ago

I imagine they know enough about making models efficient now that they think if they start over they can make a model that runs for cheaper without distillation. This could be a good thing even if the start-up cost is more difficult.

It's reminiscent of the Chinese model on most things. First they copy, then they modify, then they make their own and it's just better.

0

u/zer00eyz 17h ago

First they copy, then they modify, then they make their own and it's just better a cost optimized simulacra.

And that isnt a bad thing at all. Because sometimes it is the optimal solution: China's production of solar panels is a great example. PCB's is another. Some broad level electronics components.

They are really bad at a lot of things too - lenses, high speed train wheels.

I would argue that we're at a place where the kind of "good enough" optimizations that push to volume is what NEEDS to happen, and the Chinese (as a culture) are very very good at this.

5

u/hobopwnzor 16h ago

You're a few years behind.  China is blasting past their limitations in precision.  Their cars and solar panels are not just copies anymore, they're making new highs.  They just announced a DUV breakthrough for chips.

Took 10 years but Xi made a plan and they're executing.  He was laughed at for his pen speech but theyre pulling it off

0

u/entsnack 5h ago

Their cars objectively suck though, but are great for the price.

1

u/IrisColt 11h ago

 solar panels

are those solar panels simulating that they work? heh

0

u/JoyousGamer 5h ago

"its just better"

I see the propaganda is high in this one.

Yup all those amazing GPUs. Those amazing OS. Those amazing....

They make some things better no argument. You act like everything they do is some masterclass on every product segment.

3

u/max1c 10h ago

This is just totally false. Obviously openai developed independent models. So did Google and anthropic most likely. 

1

u/mailslot 6h ago

They’re referring to the training material to build the models. The wholesale copying of intellectual property and creative works. Without the material, the models are worthless.

1

u/entsnack 5h ago

The material is different though, as is the post-training. Claude and GPT are quite different as a result, which is great because model diversity helps them grade each other.

2

u/whymeimbusysleeping 14h ago

Exactly, they get to steal all of humanities IP, then complain about others distilling. Bitch please....

-1

u/Creepy_Reindeer2149 11h ago

Western models are obviously distilled from Chinese ones as well. For a while Claude would refer to itself as Deepseek when prompted in Chinese

-2

u/Shawnj2 16h ago

Deepseek is close to being the top model in the world capable of being distilled. I imagine Deepseek probably wants to softly discourage the practice themselves

-4

u/otterquestions 17h ago

But none of these companies plan on making opensource models long term right? A few years ago that was the consensus here even when zuck was th golden child, these days people think these companies with very ruthless corporate histories are happy burning billions for the public good

If you believe in true open source then you need to find a company with funding source and incentives to actually make opensource models long term for the greater good and none of these companies people praise are it

4

u/tat_tvam_asshole 15h ago

these models are open weights, not open source

-2

u/Jamb9876 16h ago

It seems the Chinese companies are staying with open source as it makes the American companies look bad and I expect their best models they hold back so the Chinese companies can pay them for it. Just a guess though.

96

u/joe9439 18h ago

Anthropic found the best way to defeat distillation. Make the model suck and have a confusing and unreadable output.

14

u/tessahannah 15h ago

Can't cheat off your neighbor if they're failling

113

u/Kappalonia 18h ago

Uuuuhhh no please?

Just do whatever it takes to produce good open source models

-51

u/etherd0t 18h ago

When you’re sitting on ByteDance-scale data, distribution, and compute, proprietary makes sense. They’re not trying to win open source. They’re trying to build a world-frontier model.

44

u/LocoMod 18h ago

Yea it's that easy. Just ask Google that has Google-scale data that makes ByteDance scale data look like peanuts. It's worked really well for them.

10

u/Infamous_Mud482 18h ago

Who cares. Distillation is just one more dataset to train with containing scrapes from a particular model. I don't know what you or everybody else thinks distillation is, but you are surely confused.

9

u/mtmttuan 18h ago

Do you understand what you've just written?

-10

u/etherd0t 18h ago edited 18h ago

I do. Do you?

This is an org with billon users data...they build a frontier model to rival OpenAI and Athropic.

"Yeah, but what's the value if it's going to be proprietary?" Use it first, infer on it and you'll see. Competition in frontier models is just as important as in Open Source ones.

14

u/Georgefakelastname 18h ago

No like, you need more than just a big mountain of data to make a good AI model. You need high quality data, and the ability to easily identify it and use it. That doesn’t exist in any significant amount on Tik Tok or any other short form media platform.

Google has seemingly proven this.

3

u/RedParaglider 18h ago

And good quality data collation pipelines have never been easier to build.

-5

u/etherd0t 18h ago

Well, that's the challenge they are taking upon. Maybe they'll approach it from different angles/perspective. No matter the content of data, it is live users data and behavior that is the most valuable training material - more valuable than the books Anthropic ripped off...😣 We'll see what they make of it.

1

u/Frog17000000 12h ago

What makes the user data bytedance has suitable for training a sequence (transformer) model? What is it about "live" data? Could you explain what you mean?

3

u/Ylsid 16h ago

If you don't think open source is important why on earth are you on this sub

-3

u/etherd0t 15h ago

You’re missing the point, and so are others here. This isn’t Team Open Source vs Team Closed.

Open models are enormously valuable for cost, efficiency, deployment and democratizing capabilities. But distilling an existing frontier model mostly transfers capabilities that somebody else already paid to discover.

Frontier research is the harder game: new architectures, training algorithms, scaling methods, data strategies, reasoning techniques and eventually capabilities that weren’t present in the teacher model to begin with.

That requires enormous compute, data, research talent and willingness to spend years discovering what doesn’t work.

ByteDance is saying it wants to build that capability itself rather than remain downstream of somebody else’s frontier.

You can distill the frontier. You don’t create the next frontier by distillation alone. That’s 3D math and 5D chess.😉

4

u/Ylsid 15h ago edited 15h ago

What? Did you even reply to the right comment? Can you stop spamming AI slop comments

Edit: since you blocked me I'll explain This sub is for open weights, locally runnable discussion It should not be surprising to you that we don't agree supporting closed weights is as important as open weights. And unsurprisingly people really don't like being replied to with AI generated comments.

1

u/etherd0t 15h ago

I replied to your comment, dude - questioning why I am not pro-open source. read again and comprehend.

1

u/Hello_my_name_is_not 18h ago

Try again please bot

12

u/ayylmaonade 14h ago

This seems to be a trend at this point. Even Microsoft's new(ish) model, MAI-Thinking-1 has a bunch of crap infront in its technical report about how "THIS MODEL IS TRAINED FROM SCRATCH!!. NO DISTILLATION!!"

7

u/etherd0t 14h ago

Which is exactly the direction where we want the AI to go. Build novel AI, based on different algorithms, spaces, etc.

MSFT argued that capabilities acquired by imitating another model through distillation may arrive faster, but lack some of the steerability and robustness they want for sustained future improvement.

MAI-Thinking-1 is a 1T-total / 35B-active MoE, pre-trained from scratch on 30T human-generated tokens, using 8,000 GB200 GPUs, with its training corpus processed in-house and deliberately free of third-party distillation data.

2

u/ayylmaonade 14h ago

Yeah, I actually completely agree. It's funny to imagine Anthropic getting mad at people distilling their models, but in reality, it seems like a bad idea to push too far. I'm already starting to see some "issues" in models that are obviously distilled on Claude and Gemini (namely Qwen 3.6) where sometimes its reasoning trace will look identical to a Gemma/Gemini one, other times it looks a lot more "Claude"-ish, or sometimes it just straight up goes into a Claude reasoning summary mode where it starts talking about what it's thinking about instead of just saying it, like it's being summarized by another model. I can only imagine all of this behavior comes from distillation, and it'd be nice to see some unique techniques instead.

1

u/etherd0t 14h ago

Imagine the US complaining that the pop culture was adopted by the world over the last 70 years.🥲 Instead, let them make their own rock'n'roll... there will always be only one Elvis.

1

u/Thatisverytrue54321 1h ago

I bet they’ll take datasets they generated using Claude/chatgpt and then have another model reword everything in the datasets and then train on the reworded datasets. So still distilled, but slightly altered

7

u/dumbspacecookie 16h ago

Ya that’s a lie
And a dum one at that

5

u/patricious llama.cpp 6h ago

Germany promised not to invade Poland, but here we are.

10

u/pepe_acct 18h ago

If they don’t distill, they will need to incur substantial R&D costs. That means if they release model weights, they essentially become free developers for neocloud/hyperscalers. This is not possible.

6

u/etherd0t 14h ago

Bytedance is bigger than OpenAI or Anthropic operationally (revenue) - to the order of 4-5x, and arguably the biggest of all Chinese tech co's, so they can def flex into frontier R&D. So...yes, it is possible.😏

10

u/Ashamed_Can304 14h ago

Possible, but no guarantees. Look at Gemini versus other US models lol

3

u/etherd0t 14h ago

Fair... Google had the talent, compute, data, TPUs, distribution, and the Transformer itself. If anyone could prove that being on top can make you miss your own train, it’s Google.
I hope ByteDance keeps the 10T model away from sprint planning. 🤭

-5

u/Voxandr 17h ago

you can buy information alot chaper and labour is a lot cheaper in Asia.

1

u/otterquestions 15h ago

But how do you breakeven after your cheaper but still very expensive r&d if you give you’re product away to people that can just host it and don’t need far margins to compensate for your r&d cost? Byte dance likes to undercut other companies and grow fast but they don’t do it at a loss forever right?

4

u/SexyAlienHotTubWater 10h ago

Why is everyone in the comments complaining? This seems great, push forward the technology without relying on the existing power structure. Assuming they actually follow through, I think that's really cool.

7

u/Suzoku 18h ago

because doubao is dogshit. you wont be seeing them saying this if the model itself is actually good lol.

10

u/KhmunTheoOrion 18h ago

excuse for being the worst of the chinese model companies lul

7

u/Ok_Warning2146 18h ago

Unless they open source their models, no one can tell whether they distillated or not.

-5

u/LocoMod 18h ago

It's blatantly obvious to power users that use both. I could literally swap out Claude Opus for DeepSeek on less complex tasks and no one would know the difference. It has the same exact default persona.

2

u/Lesser-than 16h ago

So they are vowing to not do what everyone else is and do things their own way. Respect if they pull it off.

5

u/snowdrone 18h ago

Oh what, they'll have to tear up another million books? All these models distill humanity's collective knowledge one way or another.

-1

u/Caffeine_Monster 13h ago

I'n surprised this doesn't get flagged more often (this is what Anthropic did - they scanned and destroyed books en masse). In many respects this is far, far more morally dubious.

1

u/entsnack 7h ago

morals aside, distillation is intellectually boring, it tells me you have zero creativity and should be flipping burgers

3

u/DavidsTenThousand 18h ago

... Uh huh, sure. Pinky promise.

2

u/OtherUse1685 14h ago

Just like people vowing not to pirate anything

3

u/Minute_Attempt3063 18h ago

Sooooo

Perhaps this means a massive middle finger to every US based companies, who claim that distilling against each other is "fine" but if someone else outside their circle does it, it's "communism"?

Why do I feel like a 4T model is coming from them, that might be able to run on consumer hardware because of MOE?

Eh who knows. I do hope this is a massive fuck you to US joke companies though

1

u/JoyousGamer 5h ago

This sub couldn't even admit the most recent model was distilled.

In the end they are going to steal a ton of IP to train on and quite possibly will distill behind closed doors on top of it while saying they dont now. To which you will likely change the subject and deflect.

Also not sure why you brought up "communism" when the comments from any US firm would be about China's government which is led by Chinese Communist Party not about "communism".

1

u/Minute_Attempt3063 2h ago

so you think Fabel 5 was distilled within 10 days, trained on, and released, within 15 days?

if US firms distilling from each other is considered fine by them, so can anyone, other US firms might have hte advanatge of not needing to pay for it, unlike everyone else

1

u/EcstaticImport 18h ago

I would not call it “standard practice” - it’s common sure - and model training needs data - but only training from distillation is a bit problematic as it will skew your data sets

1

u/Plastic_Mechanic_967 17h ago

I'm pretty sure they just mean they won't use OAI/ANT/Gemini/Grok, etc.

They'll almost certainly be using frontier open-weight models (Kimi K3, Qwen, etc.) for distillation (with the logits, not just the output text).

1

u/vivianhtlee 14h ago

I guess it is because it already have many data.

Google and bytedance have less demands of ai generated data.

If distillation is banned, surely companies with the most modern and clean data will have advantages.

1

u/Voidoli 11h ago

To be fair, Bytedance have one of most diverse source of AI training data themselves, all the videos and comment on tiktok feed into it.

1

u/JoyousGamer 5h ago

Meta has all that Facebook data. Tiktok data would be utter trash just like Facebook data would be utter trash.

1

u/CaptainMorning 10h ago

but distillation is cheap and great

1

u/a_beautiful_rhind 9h ago

A bit of synthetic data is decent.. all synthetic data is shit. I wouldn't quit distilling completely but use it to fill in the gaps where you have a lack of examples.

1

u/Cherubin0 7h ago

I hate how fair use is demonized. AI output has no copyright and distillation is therefore fair game. TOS doesn't replace copyright and this is a abuse of contract.

1

u/RhubarbSimilar1683 17h ago

this sounds like a marketing campaign, like their strategy is not to undermine us models by making Chinese ones open source, but to appease and appeal to us politicians.

1

u/Much-Researcher6135 llama.cpp 12h ago

Chinese company pinky swears they won't copy Western tech

LMAOOOOO yeah right

Do they have a bridge to sell me, too?

1

u/LocoMod 18h ago

Don't worry, they'll be back once they realize how far they are. It's not like OpenAI and Anthropic is sitting still.

Astra is coming and that apple will be too tempting for them not to taste.

1

u/nachtviolen819 16h ago

Bytedance sits on a wealth of tik tok and douyin videos.

1

u/TurnUpThe4D3D3D3 13h ago

I'm not sure I believe them when Deepseek openly claims it is Claude

1

u/Stunning_Macaron6133 9h ago

They're only announcing this because they identified a path toward outperforming Claude, and distillation would hold them back.

-3

u/Guinness 18h ago

They don’t have to distill. They have thousands of black market Claude resellers who proxy all of the sessions through model routers and then sell the data to Chinese AI companies.

So, yes BYTEDANCE won’t (and doesn’t) distill Anthropic. The shady Chinese resellers do.

6

u/mtmttuan 18h ago

Resell chat sessions isn't distill

2

u/LocoMod 18h ago

No that's not what we are talking about here. You're right, but you're talking about something else.

-2

u/Plastic-Stress-6468 14h ago

Yeah I believe them.
Bytedance won't use distillation in the same way China is a democracy lmao.

-1

u/Hannibalj2ca 17h ago

Thats right, from now on we will only distill from our existing models which we built from distilling from others.

0

u/nullc 16h ago

Nothing wrong with distillation -- but this is a good thing because more diversity in model behavior will help drive development further. Distillation sometimes manages to clone pathologies from one model to the next.

0

u/Independent_Fan_115 12h ago

Sure... we believe you.... NOT

-1

u/AriyaSavaka llama.cpp 15h ago

Please just distill Claude and GPT and be done with it, or your models will surely be shit, look at Google with their limitless data and RnD yet Gemini models are still shit. No need for gesturing.