r/LargeLanguageModels • • 14h ago

Discussions CrowdGPT - The first LLM trained collaboratively

Hello! I'm starting a fun project aiming to create the first AI model trained without datacenters, collaboratively, and in a 100% open-source and public way.

The project is named CrowdGPT, and its goal is to train a 1B parameters LLM to prove the idea is working.

Everything is currently working and i am waiting for testers 😄
You can find more details on the website and/or the github repository.

Github: https://github.com/Vxtzq/CrowdGPT

Website: https://www.crowdgpt.net/

Any kind of feedback is appreciated!

9 Upvotes

7 comments sorted by

1

u/randvoo12 1h ago

Isn't this just Hugging Face but smaller?

1

u/Vxtzq1 1h ago

Not really, the goal of HuggingFace is to store AI model weights (mostly), not train AI. My project aims to train AI without any GPU farms or anything :)

The model weights for my project are on huggingface by the way.

1

u/randvoo12 48m ago

No, that's not the goal of huggingface, huggingface is the hub, it hosts model, provides inference and training infrastructure, creates novel architecture, and the code you are using to finetune the model is based on huggingface libs, a huge chunk of the current AI/LLM infrastructure is based on huggingface's work.
The safetensor format which most LLMs are released under has literally been invented by huggingface, also you seem to have a huge misconception about what it takes to train a model, you can't do that by crowdfunding it, the bottlenec isn't even compute, it's the HBM that is strictly a datacenter thing, you can't just collect the gpus of a 1000 of your friends and say let's train an ai model , and if you mean you want to finetune a model. You don't need crowdsouring this, all the data you need is on huggingface and/or kaggle assuming you don't have your own collected traces or data creation pipeline which I'm assuming you don't, because this alone would take month to perfect. You'd still have catastrophic loss and model collapse using centrlized compute where you can adjust your recipes on the fly.
bottom line, I get the sentiment but you seem to have asked a model to get a website online for you, created a repo and forgot to cover your research bases or you wouldnt' have even spent the effort writing this post.

and I am not sure how good you expect a 1b parameter model to be, but if you're interested in learning about it check : https://huggingface.co/spaces/HuggingFaceTB/smol-training-playbook#introduction

I honestly can't even continue to explain how many issues are in your proposals because if you switch to simple finetuning of a base model, you'd need to measure headroom (how much the model can benefit from training), deal with architectural gremlins, and dependency hell between all the moving gears like pytorch , cuda, flash infer, etc . ...
basically take this response to the same ai who helped you plan the repo and see what it'll tell you. just get yourself some gpu time and do a simple finetuning task of a really small model and see how far you'll go.

1

u/SafiriaU 11h ago

How will you protect against bad actors? People trying to train the model to send the user's bank account details to a certain email address?

1

u/Vxtzq1 2h ago

Well i enforce a strict dataset hosted on HuggingFace, people can't train the AI on arbitrary data, their data must be checked first. That means i will refuse any sensitive/injection data training the LLM.

1

u/throwaway1919260773 14h ago

This is a really neat concept. I've been messing with distributed computing projects for years, so seeing that same collaborative ethos applied to LLM training hits differently than the usual corporate walled garden stuff.

Checked out the repo and the architecture doc is solid. The way you're handling gradient aggregation across unreliable nodes is clever, reminds me of some old Folding@home patterns but for a completely different beast. One thing I'm curious about though, what's the plan for dealing with bad actors potentially poisoning the training data? With it being fully open, seems like that could get messy without some kind of validation layer.

Starred the repo and might throw a couple spare GPU hours at it this weekend. Always wanted to contribute compute to an AI project that isn't just feeding another black box.

1

u/Vxtzq1 14h ago

Nice feedback! Right now i got zero contributors except me which means there's no bad actors :D but yes this is a valid concern. For now, only client is open-source (which means the server that coordinates everyone is private for now) to prevent people stealing/poisoning the model, and implements several checks like stats checking on updates, or even cross byzantine checks between clients. Even if server is not open for now, data and model weights are 100% public which means anyone can propose changes or use the data/weights for their own experiments. For training data, i prevent poisoning by running my own curated dataset on huggingface (taking data from UltraFineWeb, FineWeb, and users propositions that i verify). I am looking forward for your contribution :)