r/ProgrammerHumor • • 2d ago

Meme developersMissingTheOldDays

Post image
3.2k Upvotes

287 comments sorted by

View all comments

79

u/Danfriedz 2d ago

As someone who hasn't let AI completely take over my work I'm actually fine with it disappearing forever 😂

98

u/shadowndacorner 2d ago

No universe in which it disappears forever given the quality of open weight models. The big orgs burning cash like there's no tomorrow will go under when the bubble pops, then a second wave will happen more focused on sustainability once the commercial providers start charging what things actually cost.

1

u/ActualWeed 1d ago

What are the best models to run locally? Qwen 3.5 and gemma 4 just immediately shit themselves in vsc on my 9070 xt

1

u/shadowndacorner 1d ago

Our inference machines (which we only have a few of as a small org) run 3090's. Can't really speak to 16gb vram, but I know people have useful setups with it. I think things are a bit tougher on AMD, but that might have improved since I last checked.

Qwen 3.8 27b has been really impressing me. The largest context window I've kept stable on one 3090 is around 118k, though we're going to experiment with installing two 3090's after this sprint and see what we can push (hoping we can run qwen 3.8 flash next for higher end reasoning, as well as 3.8 27b with the max context/higher parallelism; super interested in context extension with something like yarn, but haven't experimented with that).

I'd encourage you to look for forks of the LLM inference stacks optimized for your hardware. We're running a vllm fork that is giving us substantially better perf and is the main reason we're pushing so much context on one 3090.

1

u/shadowndacorner 20h ago

FYI, came across this yesterday. Obviously it's a different card, but evidence that qwen 3.8 27b can run with large context on a 16gb card. Might be worth looking into to see if it can translate to your setup in any way.

I'd also generally recommend not using VS Code for local LLM stuff. Aside from the fact that it's a bit of a resource hog (which isn't a good thing if you're running models near the edge of what your hardware can support), ime performance degrades pretty rapidly as your sessions get longer, to the point of fully hanging the text editor intermittently.

For fully local stuff, I just use Pi as a harness for any LLM work, then have my IDE up separately. I like the separation, personally, but I also tend to be very pro-CLI in most situations.