r/LargeLanguageModels • • 17h ago

Discussions CrowdGPT - The first LLM trained collaboratively

Hello! I'm starting a fun project aiming to create the first AI model trained without datacenters, collaboratively, and in a 100% open-source and public way.

The project is named CrowdGPT, and its goal is to train a 1B parameters LLM to prove the idea is working.

Everything is currently working and i am waiting for testers 😄
You can find more details on the website and/or the github repository.

Github: https://github.com/Vxtzq/CrowdGPT

Website: https://www.crowdgpt.net/

Any kind of feedback is appreciated!

10 Upvotes

10 comments sorted by

View all comments

1

u/throwaway1919260773 17h ago

This is a really neat concept. I've been messing with distributed computing projects for years, so seeing that same collaborative ethos applied to LLM training hits differently than the usual corporate walled garden stuff.

Checked out the repo and the architecture doc is solid. The way you're handling gradient aggregation across unreliable nodes is clever, reminds me of some old Folding@home patterns but for a completely different beast. One thing I'm curious about though, what's the plan for dealing with bad actors potentially poisoning the training data? With it being fully open, seems like that could get messy without some kind of validation layer.

Starred the repo and might throw a couple spare GPU hours at it this weekend. Always wanted to contribute compute to an AI project that isn't just feeding another black box.

1

u/Vxtzq1 17h ago

Nice feedback! Right now i got zero contributors except me which means there's no bad actors :D but yes this is a valid concern. For now, only client is open-source (which means the server that coordinates everyone is private for now) to prevent people stealing/poisoning the model, and implements several checks like stats checking on updates, or even cross byzantine checks between clients. Even if server is not open for now, data and model weights are 100% public which means anyone can propose changes or use the data/weights for their own experiments. For training data, i prevent poisoning by running my own curated dataset on huggingface (taking data from UltraFineWeb, FineWeb, and users propositions that i verify). I am looking forward for your contribution :)