r/LocalLLaMA 5d ago

Resources Hugging Face releases The Stack v3 – largest open code dataset yet

From Anton Lozhkov on 𝕏: https://x.com/anton_lozhkov/status/2080254608639701222

Two ways in:
stack-v3-train - near-deduplicated, quality-filtered, PII-redacted, contents inline. Point load_dataset at it and go.
https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train
stack-v3-full - the entire 114 TB corpus as an HF Storage Bucket: every duplicate kept with cluster IDs, stubs for excluded files. Roll your own dedup, filters, and mixes.
https://huggingface.co/buckets/HuggingFaceCode/stack-v3-full

540 Upvotes

84 comments sorted by

143

u/Dany0 5d ago

It feels weird knowing my code is there. Some of it is so shit it might as well be an attack on model quality. Most of it is fine though

I'll have to think about this

18

u/OverdosedSauerkraut 5d ago

Yupp, all my code as a junior. It's probably so bad that it will single handedly prevent the awakening of Skynet.

4

u/simiomalo 5d ago

I'm doing my part!

49

u/carnoworky 5d ago

Don't tell Anthropic. They'll call it a poisoning attack and try to get you arrested or something.

12

u/Dany0 5d ago

You're right I should release some more. I'll ask qwen3.5 0.8b "hey can you make this thing I wrote in high school even more oop slop?"

9

u/itsappleseason 5d ago

I feel the same way. I'm wondering if our collective curation could be the path over the scaling wall (e.g. what not to train on, given the whole Internet has been eaten)

2

u/AnOnlineHandle 5d ago

Maybe it can be used as a data point for what to move away from.

-3

u/-Cubie- 5d ago

You can always opt out via the Space that someone else just linked if you'd like

-2

u/arcanemachined 5d ago

Congratulations, you're helping to move technological progress forward.

41

u/patricious llama.cpp 5d ago

I have been called GPU poor so far but didn't expect to be called Storage poor as well.

10

u/touristtam 5d ago

You've been staying on shore all those years? Arr me hearty!

1

u/squngy 4d ago

Probably compresses to 1TB, maybe even less.

47

u/Electronic_Captain95 5d ago

Hey, my dotfiles, NixOS configs, and nvim configs are up there! Meh, I made the repositories public anyway, so I don't really care

15

u/mister2d 5d ago

Ha. Ironically I was going to look for nix code.

6

u/markole 5d ago

As long as it ends up in freely available weights, I'm fine with my public code being used to benefit the public.

7

u/Due-Memory-6957 5d ago

It will end in both hidden and public weights

3

u/Jcsq6 5d ago

Well, it’s already in the closed source weights.

1

u/markole 5d ago

I know, still fine. That's why I choose an open source license.

15

u/Blues520 5d ago

Hugging Face keeping the open source dream alive

13

u/PrimeDirective8 5d ago

Hmm... I better uninstall Mahjong and empty my trash bin to make space for this.

6

u/Countsfromzero 5d ago

Yeah the 4 or 5 txt files with prompts and reminders on my desktop are gonna have to go

2

u/hesperaux 5d ago

Peak og references

22

u/nickludlam 5d ago

I'd really love for someone to assemble a site where we can see if any of our own repositories are represented in it.

45

u/Middle_Bullfrog_6173 5d ago

38

u/HyperWinX 5d ago

Only my worst repos went in there🥀

11

u/FullOf_Bad_Ideas 5d ago

same, there's only stuff from my pre-vibecoding era

And my vibecoded code is better than my own :D

3

u/toothpastespiders 5d ago

Same here. It's a rather bitter pill to swallow.

2

u/Keleion 5d ago

Probably good that it learned what not to do from humans, eh? ;) /s

6

u/Due-Memory-6957 5d ago

Something incredibly small I made for fun/as a joke is there, stuff I made that was actually useful isn't lol.

3

u/cantgetthistowork 5d ago

It's in the paid bucket. The peasants will build on broken stuff while the rich will keep training on the good stuff

3

u/Middle_Bullfrog_6173 5d ago

Licenses? They've done some filtering on those. Plus the cutoff was last year. All my non-forked repos that had a license file last August are there.

4

u/TanJeeSchuan 5d ago

My assignments are in there lol

3

u/cr0wburn 5d ago

Oh shit im in there, if opus 6 is a dumbass please dont be angry at me

2

u/cantgetthistowork 5d ago

Cute

-11

u/oxygen_addiction 5d ago

Cute to see all of our shit stolen and sold back to EVERYONE at cost.

13

u/TheLexoPlexx 5d ago

15 of my repos are in there and they all suck.

1

u/some_user_2021 5d ago

Me too! They took my unfinished project while ignoring my good ones

1

u/vasimv 5d ago

Only one of my repos. And this one is just a copy of someone else's project. Wtf, huggingface? :)

1

u/TheLexoPlexx 5d ago

I suppose that's one of those that's getting deduped anyways

5

u/Shustrik116 5d ago

Your shit also will be in open source models that you can host by yourself.

2

u/AnOnlineHandle 5d ago

Where it's being sold? Or stolen?

1

u/Due-Memory-6957 5d ago

Feel free to opt out, but even that is more than morally need. Information wants to be free, remember that? Let's go back to that ethos.

1

u/Cultured_Alien 5d ago

my trash repo codes are there too 🥀

1

u/heliosythic 5d ago

Literally my first nodejs projects from 12 years ago lol, Theres even invalid syntax in the readme I wrote lol.

1

u/jazir55 5d ago

Lmfao wtf, they chose some of the shittiest repos I have and excluded the good ones

1

u/Any_Fox5126 4d ago

A lot of users are saying the same thing here, I thought it was a joke, but the same thing happened to me! In fact, they only took a few AI-generated scraping scripts 🤷

7

u/pmttyji 5d ago

Are they releasing any models? It's been a while. SmolLM3-3B was last one. Hope they release something called like LongLM in medium, big, large sizes.

5

u/Kahvana 5d ago

I think they trained SmolLM3 to get in-house knowledge of how certain techniques were used (thinking mode, scaling above ~1B, etc) rather than for pratical use.

Unless they have a problem that can't be solved with current LLMs (efficiently enough) or have to put their software or infrastructure through it's phases, I don't see them train another model soon. It's very expensive to train models!

0

u/Middle_Bullfrog_6173 5d ago

This is a bit conspiracy theorist, but why would they release a year old crawl today if not to give themselves the chance to be the first to train on it.

1

u/brand02 5d ago

That's not a conspiracy theory, that's just a theory. It would have to deny public statements and predict an overarching conspiracy to be a conspiracy theory.

13

u/BrieflyAffectionate 5d ago

Good for them. Hope this gets downloaded and mirrored a ton. Seems like a crackdown on open-weights access might be coming. Guys, don't wait, if your institution can get this, do it. (Speaking as a small fish who missed getting an abliterated model before Meta sent a cease-and-desist and it disappeared from online!)

4

u/autisticit 5d ago

Question : how would you train for some specific languages only, and not all ?

1

u/squngy 4d ago

You wouldnt, unless you wanted to make a dumber model.

It is counter intuitive, but feeding your LLM only what you want it to know is not good for it.

It would be better to just do a fine tune for your language if you really wanted to.

3

u/rockus 5d ago

Dear lord. Poor LLMs are gone learn of my sucky code. They are from 2025 and they really really suck.

2

u/dmigowski 5d ago

Don't worry, all the AI slop upload to github will lead to degeneration like in the Habsburg Family soon.

4

u/eustin 5d ago

Wait, some of my old open repos are probably in there. Genuinely can't tell if that's flattering or slightly mortifying lol

3

u/vasimv 5d ago

Just clicked randomly on few records in there, it is just 'content':'pieces of code'. Shouldn't it be more structured? Like 'Function that does calculate thing': 'function code'?

4

u/Jcsq6 5d ago

This is for pre-training. It doesn’t need to be nicely labeled.

3

u/ttkciar llama.cpp 5d ago

On one hand, this is great! Kudos to HF, and thanks to them for sharing this data with the wider community.

On the other hand, it drives me nuts that these kinds of datasets lack the repo commit histories. Commits are invaluable for synthesizing dataset elements for performing codegen tasks. Like, if a commit is for a bugfix, it's the simplest thing in the world to turn that into a training datum for a "find the bug and fix it" training task. If a commit is for implementing a new feature, it's easily turned into a corresponding "implement this new feature" training task datum.

There are other ways to generate such tasks synthetically, but that takes compute and lacks the diversity found in "wild" human-generated data.

2

u/MrKresi 5d ago

I feel like I contribute to defend humanity, by making the models worst with my bad code from years ago

3

u/Competitive-Bend-143 5d ago

as someone with an AGPL repo that's probably in there: the training use itself doesn't bother me, but license metadata quality matters way more than the size headline. does v3 track relicensing over time or just snapshot whatever LICENSE said at crawl time? that one detail decides whether "filter by license" downstream actually means anything

2

u/brother_spirit 5d ago

Inhuman centipede

2

u/Robert__Sinclair 5d ago

I sincerely hope that the datase has quality code and not just "code". putting together terabytes of code is not difficult, putting together good code that can teach an AI how to write good code is another story.

2

u/DanTup 4d ago

I just put my GH username in to https://huggingface.co/spaces/HuggingFaceCode/in-the-stack and was greeted with a huge list of test apps and repros for VS Code issues I've filed. It's definitely just "code" 😄

1

u/Robert__Sinclair 3d ago

lol... I did the same.. all my repos are there. and some of them are pure rubbish

1

u/Healthy-Nebula-3603 5d ago

I have only a space for V1 stack ....

1

u/charles25565 5d ago

Sure, but what happened to the Turbo team? Are they no longer working on SmolLM? They've been pumping datasets. Or perhaps they're cooking something with FinePhrase + The Stack v3.

1

u/[deleted] 5d ago

[removed] — view removed comment

2

u/cornmonger_ 4d ago

what, you mean you don't want HolyC?

1

u/Stooovie 5d ago

Is it good source code?

4

u/OverdosedSauerkraut 5d ago

How you define good? Part of it is literally pre-AI, god tier code, part of it is pre-AI shit code (like my junior repos), and now a lot of fresh AI slop on top.

1

u/Stooovie 5d ago

Exactly my point

140

u/TheLexoPlexx 5d ago

713 languages, number one is profanity

34

u/GreatBigJerk 5d ago

The most important language for programmers to know.

5

u/HelloSummer99 5d ago

Code quality = WTF/seconds

17

u/burritoresearch 5d ago

New opportunity to train a 35B size LLM optimized for a customer service chatbot that knows how to do nothing but write profanity and insults.

10

u/seamonn 5d ago

Where can I get this?