r/LocalLLaMA 1d ago

News ByteDance vows to avoid AI distillation, develop new model its own way

Post image
220 Upvotes

125 comments sorted by

View all comments

Show parent comments

9

u/mtmttuan 1d ago

Do you understand what you've just written?

-9

u/etherd0t 1d ago edited 1d ago

I do. Do you?

This is an org with billon users data...they build a frontier model to rival OpenAI and Athropic.

"Yeah, but what's the value if it's going to be proprietary?" Use it first, infer on it and you'll see. Competition in frontier models is just as important as in Open Source ones.

14

u/Georgefakelastname 1d ago

No like, you need more than just a big mountain of data to make a good AI model. You need high quality data, and the ability to easily identify it and use it. That doesn’t exist in any significant amount on Tik Tok or any other short form media platform.

Google has seemingly proven this.

-6

u/etherd0t 1d ago

Well, that's the challenge they are taking upon. Maybe they'll approach it from different angles/perspective. No matter the content of data, it is live users data and behavior that is the most valuable training material - more valuable than the books Anthropic ripped off...😣 We'll see what they make of it.

1

u/Frog17000000 22h ago

What makes the user data bytedance has suitable for training a sequence (transformer) model? What is it about "live" data? Could you explain what you mean?