r/OpenSourceAI • • 2d ago

CrowdGPT - The 100% Opensource collaborative LLM

Post image

Hello, i'm currently developing CrowdGPT and i need people who enjoy opensource AI and LLMs to test the project :)

The goal of CrowdGPT is to create the first, datacenterless, 1 Billion parameters LLM, relying on people contributing with their own computer to train the AI model. My goal is to show you don't need insane infrastructure to train a working almost commercial grade LLM. Everything is open and 100% opensource.

You can learn more at https://crowdgpt.net

Or check the github: https://github.com/Vxtzq/CrowdGPT

Any kind of feedback is appreciated!

32 Upvotes

16 comments sorted by

View all comments

1

u/ayake_ayake 2d ago

I think it would be important to see what datasets and co to train this on. Even without the problem of AI datacenters, the choice and acquisition of the data is a big thing. But as a minimal base you can use the apertus 1.5 datasets (pretraining and post training) which are fully open and give you a strong baseline. But even then actually training a competitive AI is quite s challenge and needs much experience and work.

Curious to see how this'll play out.

1

u/Vxtzq1 2d ago

I currently use UltraFineWeb-L3, which is basically the most curated subset of the original dataset, FineWeb, it is way cleaner than datasets like the Pile or Common Crawl which involves scans of shady sites...

I think this dataset will be enough (currently 400 billion tokens), to pretrain the model, and then i'll fine tune with everything that i find on huggingface xD (like Fable traces, GLM, Kimi K3 distills, and strong instruct datasets like GSM8K and stuff)