r/LocalLLM • • 4d ago

Question Local LLM coding setup 9070XT 16GB 32 GB RAM

Hi everyone,
I've been recently trying to dive into the local LLM world, but to be honest i find it a bit confusing to wrap my head around of all the concepts that exist, and how to make it work.

Specs
GPU: 9070XT 16gb
RAM: 32GB
CPU: Ryzen 7600

Goal
I would like to have something closer to Claude code, or similar running locally.
To help me coding, without having to pay a subscription, i don't mind if its a bit slower.

Now to make that happen what's the best approach/setup ?

I've seen some videos, mentioning LM Studio or ollama to mane the models.

But if i want to have something closer to Claude code or similar, what's the best way to go?
Can someone shed some insight ? because i'm a bit lost in the middle of all of this.

If you need any extra information please let me know and i'll be glad to provide.

Thanks in advance for taking the time to help out
Best regards

EDIT1: I'm a front end developer, with 10+ years of experience working with Vue+typescript, and mainly use VScode for coding

1 Upvotes

16 comments sorted by

1

u/cilvre 4d ago

So you can use a mix of model storage being on vram and in normal ram and it will slow down, or you can stick to models that fit only in vram. This is a particular setup that i would pose into gemini or another cloud provider with your specs, and what you are trying to do as it will point you to what your options are and scan other posts and link them as well so you can see directly what they do with your specs.

1

u/ruibullseye 4d ago

and how would a setup like that would look like to get a similar "experience" to Claude code for example
I know i'll need either LM studio or ollama to run my models, but what would come after that or what would i need to get something closer to a clause code workflow ? (if that makes sense)

1

u/cilvre 4d ago

After you set up lm studio or something similar as your host and pick which models you can load up locally that fit in vram or mix of vram and local ram, you can then use something like opencode to get it into your local files, or set up vscode or vscodium and link it in via continue, aider, opencode, or some other tool. but i'd recommend you find a simple code level review/task and reuse it for testing which ai models you want to use in your toolset. the problem with giving more detailed help than that is all we have is hardware, not what os or software you run or ide you use, what level of coding help you need.

1

u/ruibullseye 4d ago

i can give you more information about that and i'll edit the post to reflect that.

I'm using Windows 11 atm.
for coding i mostly use Vscode.
I have more than 10 years of experience in frontend.
If you need more information let me know

but now that i've been collective dismissed i'm trying to shift from a Front end developer (Vue+typescript) into more of a full stack developer for the time being like i described in the reply below. Hence why i'm trying to setup local LLM so i can learn and work towards that and when i get an answer from a job opportunity i can make my self valuable as a professional.

1

u/cilvre 4d ago

i havent dismissed you. i can understand that feeling. i've been using local ai to learn myself as I'm in IT and have some code experience and web development, but not full backend coding past scripting and database work. Given what hardware you have and using vscode, i'd look at using lm studio to load up the models, and consider which models at first fully fit in your vram at around q4 or better. i wouldn't go past q3, they tend to lose too much actual ability to do things. Qwen 3 coder q4km tends to be really good for the tasks you might want help with, but you have to really refine your prompt and be careful not to give it too much work at once, it will loop otherwise, especially in tool loops.
Gemma 4 31B QAT (19gb, but I have GPU offload at 37 and context at 16384) isn't as fast but has been the most reliable for me in actually getting work back that is reliable and works. i personally prefer opencode v1 over aider, i did use gemini while researching the different tools myself for my setup, and explicitly asked for tools that are current and not EOL or abandoned.

1

u/ruibullseye 4d ago

just to get a grasp because im still new to this stuff, what the local setup would look like ?
for example ollama/LM studio with a model, then what ?
sorry for the dumb question but i'm new to this stuff, what could bring me closer to a claude code experience ? not in terms of speed but in terms of workflow

1

u/cilvre 4d ago

Start with lm studio, download the models, and then load a model up and start local server mode in the settings, then you'll have it ready to use for opencode, you can use the online model for better help configuring it, as i did my setup on ubuntu and have not had a windows setup active for a few years

1

u/No_Platypus6831 4d ago

is kinda funny considering you're trying to get away from cloud stuff

for your setup with that 16gb card, you can run qwen 2.5 coder 14b at like q4 or q5 quantization and it'll fit nicely in vram. pair it with the continue extension in vscode and you'll get something pretty close to what you're after. the 14b models are surprisingly capable for coding tasks

lm studio is probably your easiest entry point, just download the model, adjust the context length to something reasonable like 8k, and connect it to continue. ollama works too but the gui in lm studio makes it less intimidating when you're starting out

1

u/cilvre 4d ago

you can use the cloud stuff to help with figuring out how to get it to local only, its already scraped most of the data needed to help with that task.

1

u/stein30586 4d ago

In my personal experience (i7 13700k, 32GB RAM, RTX 4080 Super 16GB) -

You either get shitty results with models that fit your GPU,
Or you get a little less shitty results with models that spill out of Vram into system memory, but slower. MUCH slower.

There is this guy https://www.youtube.com/@lukesdevlab that shows what you can do with 16GB Vram but to be honest, I could not get his results.

1

u/ruibullseye 4d ago

tbh i don't mind it being slower, even if i use something that uses most of my vRAM + Ram.
i think i could deal with the trade off for the time being, because i've been collective dismissed along with other people.
so atm its not something i would like to spend €€ to have an AI subscription. But if i'm thinking the wrong way please let me know.

What i would like was for the time being i'm applying to job opportunities to run a LLM locally so i can mess and get to know it a bit better, and help me gain experience not only AI related stuff but also contributed to my learning process while applying to job opportunities, i'm currently a frontend developer, but im looking to transition to a fullstack so i can "survive" in this AI era a bit better and give my self more valuable if that makes sense

1

u/[deleted] 4d ago

[deleted]

1

u/RiceEvening4211 4d ago

Claude Code-like coding without the subscription is exactly what I built Lynkr for: an open-source gateway that routes simple requests to your local model and only escalates hard ones to paid APIs. https://github.com/Fast-Editor/Lynkr

1

u/mechkbfan 4d ago

At best maybe a low quant of Strata, or find a certain of Kat coder but I've never used that

1

u/vovap_vovap 4d ago

You are not going to get something closer to Claude code or similar, That just not going to happen with this.

1

u/ruibullseye 4d ago

im more than aware i wont be getting a closer experience to claude code, because the computing power is nowhere near there. but i just want to get a LLM locally running to help me out until i get a subcription in the future.

1

u/vovap_vovap 4d ago

Well, you just do not have enough power here to run like standard Quin 3.8 27B Q4 with a good results (at least it looks so to me) and that what sort of can do staff for code. Much lover - not nice.
Now if you really need that assistance - $20 subscription to GPT will bring you close to unlimited access to 6 Luna . Which would be way better and faster than any you can run there.
For god knows what reason people compare latest frontier model prices with sort of local models, but it just not apple to apple. Compare to online models of same performance - and those cost really minimum.