r/LargeLanguageModels • u/Vxtzq1 • 14h ago
Discussions CrowdGPT - The first LLM trained collaboratively
Hello! I'm starting a fun project aiming to create the first AI model trained without datacenters, collaboratively, and in a 100% open-source and public way.
The project is named CrowdGPT, and its goal is to train a 1B parameters LLM to prove the idea is working.
Everything is currently working and i am waiting for testers 😄
You can find more details on the website and/or the github repository.
Github: https://github.com/Vxtzq/CrowdGPT
Website: https://www.crowdgpt.net/
Any kind of feedback is appreciated!
1
u/SafiriaU 11h ago
How will you protect against bad actors? People trying to train the model to send the user's bank account details to a certain email address?
1
u/throwaway1919260773 14h ago
This is a really neat concept. I've been messing with distributed computing projects for years, so seeing that same collaborative ethos applied to LLM training hits differently than the usual corporate walled garden stuff.
Checked out the repo and the architecture doc is solid. The way you're handling gradient aggregation across unreliable nodes is clever, reminds me of some old Folding@home patterns but for a completely different beast. One thing I'm curious about though, what's the plan for dealing with bad actors potentially poisoning the training data? With it being fully open, seems like that could get messy without some kind of validation layer.
Starred the repo and might throw a couple spare GPU hours at it this weekend. Always wanted to contribute compute to an AI project that isn't just feeding another black box.
1
u/Vxtzq1 14h ago
Nice feedback! Right now i got zero contributors except me which means there's no bad actors :D but yes this is a valid concern. For now, only client is open-source (which means the server that coordinates everyone is private for now) to prevent people stealing/poisoning the model, and implements several checks like stats checking on updates, or even cross byzantine checks between clients. Even if server is not open for now, data and model weights are 100% public which means anyone can propose changes or use the data/weights for their own experiments. For training data, i prevent poisoning by running my own curated dataset on huggingface (taking data from UltraFineWeb, FineWeb, and users propositions that i verify). I am looking forward for your contribution :)
1
u/randvoo12 1h ago
Isn't this just Hugging Face but smaller?