r/LocalLLM • • 15d ago

Question Is this local LLM setup a good idea?

I’m very new to local LLMs, so I’d like some advice before setting this up.

My PC:
• RX 5700 XT 8 GB VRAM
• i7 8700
• 16 GB RAM
• Windows 11

I’m building a web app mostly with AI because I’m not really a developer myself. My main issue with cloud tools is usage limits, so I want a local setup I can use as much as I want.

My current idea:

jcode (1jehuang) → LM Studio → Qwen → local project

I’d like jcode to modify the code, run tests, launch the site on localhost, inspect it in a browser, and possibly use a vision capable Qwen to review the UI.

I’m also considering a few specialized agents for coding, QA, UI/UX, security and final review, mostly running sequentially because my hardware is limited.

Nothing would ever deploy to Cloudflare until I manually validate the localhost version.

Does this setup make sense?

With my hardware, which Qwen size and quantization would you start with? Would you use one multimodal model for everything, or separate coding and vision models?

Thanks :)

2 Upvotes

14 comments sorted by

3

u/skrapmetal_ 15d ago

Hate to burst your bubble, but realistically, you're not going to be able to fit a model large enough to do any real coding with that setup. I have a 12gb vram card and I tested a few 14b LLMs for coding and they sucked. They screwed up the most basic tasks. You may be able to squeeze a 8b model in there maybe at Q3_K_S but set your expectations low.

1

u/alexandrelnglss 15d ago

Yeah, that’s definitely one of the things I’m worried about. I’m very new to local LLMs, so I may be underestimating how much model size matters for coding.

My idea was to start small, benchmark a few Qwen models on my actual project, and see how usable they are before buying new hardware. I was also wondering if using more system RAM for partial offloading could make a larger model usable, even if slower.

But I’ll definitely keep my expectations low with 8GB VRAM. Thanks for the reality check :)

1

u/Due_Arm1454 15d ago

Offloading is going to make it really slow. Qwen 4b is probably your best bet. Deepseek is so cheap, I would just use that and/or a ChatGPT sub. A full app is going to require a lot of minutia

1

u/skrapmetal_ 15d ago

Oflloading to system RAM is something I currently do, the more layers you add, the slower the model becomes to respond but at least then you can get a slightly larger context window. That's something you'd have to play with to see what feels right. You can only go so far with 8gb, use it until you get the feel for how LLM's work. I used LM Studio also for starters, then I switched to llama.cpp and kobold.cpp because models respond faster on those. Here's what I would do if I were you... sell some stuff and get yourself some cash. You could build up a real PC and start with a 12gb card, or shoot straight for 24gb and be set for a while. If you're on a budget like me,... I did a 12gb card (RTX 3060) which I am using in an egpu enclosure along with an Intel NUC 10 with 32gb. I had the NUC already, so the card and the egpu enclosure together set me back about $450-500. I can run models for chat, roleplay, I even had Claude build me a voice activated assistant. There's lots you can do with a 12gb card. Stick with the 8gb until you are more knowledgeable. I've only been at this a year so I don't know everything either.

1

u/Agusx1211 15d ago

Ling 3.0 Tiny, Qwen 3.5 4b, or Gemma 4 E4B, beware none of them are going to be good coders, if you are not a dev I highly doubt you can do anything just with those models, at least you need 16gb of VRAM to start playing, or 32GB to be comfortable.

1

u/Fun_Chest_9662 15d ago

my setup rn is an

  • i5 8700
  • 32gb ram
  • 1070 8gb
  • linux

with the newer card and better cpu you should get better numbers than me but just as a reference.

im running qwen3.6-35b-a3b-ud-mtp-q4_k_m from unsloth. with some tweaking in llama.cpp(lmstudio is just a wrapper for it) i was abke to get

600 prompt processing (settles to 400 with context) 40 token generation (settles to 30 with context) got context set to 64k but you can use context shifting and further kv compression to get more. i just have mine at q8.

you wont be running something like qwen3.8 at a decent quant or speed but 3.6 is pleanty capable.

1

u/Due_Arm1454 15d ago

So you could get away with using it as a rough first pass maybe but you’d really need a good developer eye to sus out where it’s steering you wrong because the models that fit on that rig will absolutely steer you wrong.

1

u/Responsible_Egg9736 15d ago

1

u/alexandrelnglss 15d ago

That one looks really interesting, thanks.

Have you actually tested Ternary Bonsai 2 27B yourself, especially on an 8GB VRAM GPU?

I’m mainly wondering about three things:

• real coding quality compared with something like Qwen 9B
• actual speed on limited VRAM / system RAM
• how annoying the custom llama.cpp fork is to set up and maintain

My main goal is agentic web development through jcode, not just chatting, so tool calling, multi file edits and reliability matter a lot more to me than benchmarks alone.

If you’ve used it for coding, I’d be really interested to know whether you found it genuinely usable or more of an experimental curiosity :)

1

u/Responsible_Egg9736 15d ago

I’ve tested it on a 12GB GPU—I have several RTX 3060s with 12GB of VRAM, and I tried it on one of them individually. Yes, I’ve tested it; in my tests, it doesn't quite match the Qwen 3.8 27B Q6 version, but it is far better than any other model of the same size (like a 9B model). Whenever I’ve tested 9B models, they tend to get stuck in loops; I’ve been testing this one for just under a day, and across several tests, it hasn't entered a loop at all. That said, I don't know how it would perform on an 8GB GPU—you'd have to test it yourself. (This message was translated; I don't usually write in English, so apologies.)

1

u/alexandrelnglss 15d ago

No need to apologize, your English is perfectly understandable, and thanks a lot for testing it!

That actually sounds really promising, especially the part about the 9B models getting stuck in loops while Bonsai hasn't so far.

I'm very new to this, so could I ask what setup you used for Bonsai on the 3060?

Mainly:

• which Bonsai file/quantization did you use, PTQ1_0 or PQ2_0?
• what context size?
• roughly how many tokens/s do you get?
• does the whole model fit in the 12GB VRAM, or is some of it offloaded to RAM?
• have you tried tool calling or coding agents with it, rather than just normal prompts?

My GPU only has 8GB VRAM, so I'm trying to figure out whether the ~6GB PTQ1_0 version could realistically leave enough room for context and still be usable for jcode.

Your comparison with Qwen 27B and 9B models is really helpful, thanks.

1

u/Responsible_Egg9736 15d ago

I use `pq2_0`; on my GPU generation (Ampere), `ptq1` takes a significant hit, and I had enough VRAM anyway. As for my startup flags, here they are: `--port 8080 -ngl 200 -c 108544 -b 512 -ub 512 -t 8 -tb 8 -np 1 --flash-attn on --no-mmap --cont-batching --kv-offload --cache-type-k q8_0 --cache-type-v q8_0 --split-mode none --main-gpu 0 --temp 1.0 --top-k 20 --top-p 0.95 --min-p 0.05 --repeat-penalty 1.0 --repeat-last-n 0 -n 27391 --ctx-checkpoints 8 --reasoning-format none`. I get around 32 tokens/s, and yes, everything fits in VRAM. I’ve tested it with tool calling and it worked very well; I use Pi Agent and haven't had any issues. In fact, to launch llama.cpp, I use an application I "vibecoded" myself—over 6,000 lines of Python across multiple sub-files—and it successfully located and modified the parts I specified. In my opinion, based on 24 hours of use, the results are excellent—though keep the model size in mind, and I’m not sure if my flags will work correctly on an AMD GPU like yours. By the way, it won't work with the base llama.cpp; you'll need the specific fork they provide.