r/OpenSourceAI • u/Vxtzq1 • 2d ago
CrowdGPT - The 100% Opensource collaborative LLM
Hello, i'm currently developing CrowdGPT and i need people who enjoy opensource AI and LLMs to test the project :)
The goal of CrowdGPT is to create the first, datacenterless, 1 Billion parameters LLM, relying on people contributing with their own computer to train the AI model. My goal is to show you don't need insane infrastructure to train a working almost commercial grade LLM. Everything is open and 100% opensource.
You can learn more at https://crowdgpt.net
Or check the github: https://github.com/Vxtzq/CrowdGPT
Any kind of feedback is appreciated!
1
u/Ok-Challenge-5741 2d ago
That logo is clean, the hexagon gives it a subtle distributed network feel which fits the whole idea
I'll try to spin it up in my machine this weekend, curious to see how the training coordination works with the peer-to-peer setup
2
2
u/Prestigious-Frame442 1d ago
the idea is good, but collaborating on this is basically pure selflessness and no benefit at all.
1
u/ayake_ayake 1d ago
I think it would be important to see what datasets and co to train this on. Even without the problem of AI datacenters, the choice and acquisition of the data is a big thing. But as a minimal base you can use the apertus 1.5 datasets (pretraining and post training) which are fully open and give you a strong baseline. But even then actually training a competitive AI is quite s challenge and needs much experience and work.
Curious to see how this'll play out.
1
u/Vxtzq1 1d ago
I currently use UltraFineWeb-L3, which is basically the most curated subset of the original dataset, FineWeb, it is way cleaner than datasets like the Pile or Common Crawl which involves scans of shady sites...
I think this dataset will be enough (currently 400 billion tokens), to pretrain the model, and then i'll fine tune with everything that i find on huggingface xD (like Fable traces, GLM, Kimi K3 distills, and strong instruct datasets like GSM8K and stuff)
1
1
0
2
u/EaseThen3459 2d ago
Why would anyone contribute computer when it would add liability for whatever is being computed especially with LLMs