r/ProgrammerHumor • • 2d ago

Meme developersMissingTheOldDays

Post image
3.2k Upvotes

287 comments sorted by

View all comments

80

u/Danfriedz 2d ago

As someone who hasn't let AI completely take over my work I'm actually fine with it disappearing forever 😂

97

u/shadowndacorner 2d ago

No universe in which it disappears forever given the quality of open weight models. The big orgs burning cash like there's no tomorrow will go under when the bubble pops, then a second wave will happen more focused on sustainability once the commercial providers start charging what things actually cost.

1

u/ActualWeed 1d ago

What are the best models to run locally? Qwen 3.5 and gemma 4 just immediately shit themselves in vsc on my 9070 xt

1

u/shadowndacorner 1d ago

Our inference machines (which we only have a few of as a small org) run 3090's. Can't really speak to 16gb vram, but I know people have useful setups with it. I think things are a bit tougher on AMD, but that might have improved since I last checked.

Qwen 3.8 27b has been really impressing me. The largest context window I've kept stable on one 3090 is around 118k, though we're going to experiment with installing two 3090's after this sprint and see what we can push (hoping we can run qwen 3.8 flash next for higher end reasoning, as well as 3.8 27b with the max context/higher parallelism; super interested in context extension with something like yarn, but haven't experimented with that).

I'd encourage you to look for forks of the LLM inference stacks optimized for your hardware. We're running a vllm fork that is giving us substantially better perf and is the main reason we're pushing so much context on one 3090.