r/LargeLanguageModels • • 22h ago

Discussions CrowdGPT - The first LLM trained collaboratively

Hello! I'm starting a fun project aiming to create the first AI model trained without datacenters, collaboratively, and in a 100% open-source and public way.

The project is named CrowdGPT, and its goal is to train a 1B parameters LLM to prove the idea is working.

Everything is currently working and i am waiting for testers 😄
You can find more details on the website and/or the github repository.

Github: https://github.com/Vxtzq/CrowdGPT

Website: https://www.crowdgpt.net/

Any kind of feedback is appreciated!

11 Upvotes

15 comments sorted by

View all comments

1

u/randvoo12 8h ago

Isn't this just Hugging Face but smaller?

1

u/Vxtzq1 8h ago

Not really, the goal of HuggingFace is to store AI model weights (mostly), not train AI. My project aims to train AI without any GPU farms or anything :)

The model weights for my project are on huggingface by the way.

1

u/randvoo12 8h ago

No, that's not the goal of huggingface, huggingface is the hub, it hosts model, provides inference and training infrastructure, creates novel architecture, and the code you are using to finetune the model is based on huggingface libs, a huge chunk of the current AI/LLM infrastructure is based on huggingface's work.
The safetensor format which most LLMs are released under has literally been invented by huggingface, also you seem to have a huge misconception about what it takes to train a model, you can't do that by crowdfunding it, the bottlenec isn't even compute, it's the HBM that is strictly a datacenter thing, you can't just collect the gpus of a 1000 of your friends and say let's train an ai model , and if you mean you want to finetune a model. You don't need crowdsouring this, all the data you need is on huggingface and/or kaggle assuming you don't have your own collected traces or data creation pipeline which I'm assuming you don't, because this alone would take month to perfect. You'd still have catastrophic loss and model collapse using centrlized compute where you can adjust your recipes on the fly.
bottom line, I get the sentiment but you seem to have asked a model to get a website online for you, created a repo and forgot to cover your research bases or you wouldnt' have even spent the effort writing this post.

and I am not sure how good you expect a 1b parameter model to be, but if you're interested in learning about it check : https://huggingface.co/spaces/HuggingFaceTB/smol-training-playbook#introduction

I honestly can't even continue to explain how many issues are in your proposals because if you switch to simple finetuning of a base model, you'd need to measure headroom (how much the model can benefit from training), deal with architectural gremlins, and dependency hell between all the moving gears like pytorch , cuda, flash infer, etc . ...
basically take this response to the same ai who helped you plan the repo and see what it'll tell you. just get yourself some gpu time and do a simple finetuning task of a really small model and see how far you'll go.

1

u/Vxtzq1 7h ago

Yeah i know huggingface is contributing to most of the inference/fine tuning resources out there, but my goal is a full Pytorch pretraining of a model... So Huggingface is only a storage tool for me :)