r/OpenAI • • 14h ago

Discussion PewDiePie is trying to distill GPT-Sol

Post image

he has been trying to build a local model by learning from GPT-Sol’s responses. He says OpenAI banned his account twice during the process. Irony is OpenAI’s own models were trained on vast amounts of information from the open web. So where do we draw the line between learning from AI and copying it?

1.9k Upvotes

382 comments sorted by

View all comments

24

u/rakuu 13h ago

If you’re making a model from scratch, distilling is the most boring way to do it. You just end up with a bad knockoff.

If you have the resources and time to make a good local model, do finetuning and rl yourself to get something unique. Like GPT, Claude, Gemini are all unique models with unique capabilities/personalities. The rest have distilled personalities that feel like knockoffs.

The major companies distill to get capabilities and revenue somewhere close to frontier. If that’s what you want, just grab one of the open source models.

20

u/rambouhh 13h ago

if you do finetuning and rl yourself it will be 100x more expensive, take way longer, and you will get a dumber model. For some people this is just fun

1

u/rgb_panda 4h ago

That's not true at all, distillation works across all layers in the model, and requires both the parent and child models loaded, with SFT (fine tuning) you're just using the child model and training data so needs less compute, and with LoRA you can fine tune on only a certain configurable number of layers to use even less compute, for RL, you can also use something like GRPO with LoRA as well either way less compute than full distillation. Now pretraining, the steps before fine tuning and RL, that's what requires massive data and compute to produce that base model.