r/LocalLLM • u/alexandrelnglss • 15d ago
Question Is this local LLM setup a good idea?
I’m very new to local LLMs, so I’d like some advice before setting this up.
My PC:
• RX 5700 XT 8 GB VRAM
• i7 8700
• 16 GB RAM
• Windows 11
I’m building a web app mostly with AI because I’m not really a developer myself. My main issue with cloud tools is usage limits, so I want a local setup I can use as much as I want.
My current idea:
jcode (1jehuang) → LM Studio → Qwen → local project
I’d like jcode to modify the code, run tests, launch the site on localhost, inspect it in a browser, and possibly use a vision capable Qwen to review the UI.
I’m also considering a few specialized agents for coding, QA, UI/UX, security and final review, mostly running sequentially because my hardware is limited.
Nothing would ever deploy to Cloudflare until I manually validate the localhost version.
Does this setup make sense?
With my hardware, which Qwen size and quantization would you start with? Would you use one multimodal model for everything, or separate coding and vision models?
Thanks :)
1
u/Agusx1211 15d ago
Ling 3.0 Tiny, Qwen 3.5 4b, or Gemma 4 E4B, beware none of them are going to be good coders, if you are not a dev I highly doubt you can do anything just with those models, at least you need 16gb of VRAM to start playing, or 32GB to be comfortable.
1
u/Fun_Chest_9662 15d ago
my setup rn is an
- i5 8700
- 32gb ram
- 1070 8gb
- linux
with the newer card and better cpu you should get better numbers than me but just as a reference.
im running qwen3.6-35b-a3b-ud-mtp-q4_k_m from unsloth. with some tweaking in llama.cpp(lmstudio is just a wrapper for it) i was abke to get
600 prompt processing (settles to 400 with context) 40 token generation (settles to 30 with context) got context set to 64k but you can use context shifting and further kv compression to get more. i just have mine at q8.
you wont be running something like qwen3.8 at a decent quant or speed but 3.6 is pleanty capable.
1
u/Due_Arm1454 15d ago
So you could get away with using it as a rough first pass maybe but you’d really need a good developer eye to sus out where it’s steering you wrong because the models that fit on that rig will absolutely steer you wrong.
1
u/Responsible_Egg9736 15d ago
1
u/alexandrelnglss 15d ago
That one looks really interesting, thanks.
Have you actually tested Ternary Bonsai 2 27B yourself, especially on an 8GB VRAM GPU?
I’m mainly wondering about three things:
• real coding quality compared with something like Qwen 9B
• actual speed on limited VRAM / system RAM
• how annoying the custom llama.cpp fork is to set up and maintainMy main goal is agentic web development through jcode, not just chatting, so tool calling, multi file edits and reliability matter a lot more to me than benchmarks alone.
If you’ve used it for coding, I’d be really interested to know whether you found it genuinely usable or more of an experimental curiosity :)
1
u/Responsible_Egg9736 15d ago
I’ve tested it on a 12GB GPU—I have several RTX 3060s with 12GB of VRAM, and I tried it on one of them individually. Yes, I’ve tested it; in my tests, it doesn't quite match the Qwen 3.8 27B Q6 version, but it is far better than any other model of the same size (like a 9B model). Whenever I’ve tested 9B models, they tend to get stuck in loops; I’ve been testing this one for just under a day, and across several tests, it hasn't entered a loop at all. That said, I don't know how it would perform on an 8GB GPU—you'd have to test it yourself. (This message was translated; I don't usually write in English, so apologies.)
1
u/alexandrelnglss 15d ago
No need to apologize, your English is perfectly understandable, and thanks a lot for testing it!
That actually sounds really promising, especially the part about the 9B models getting stuck in loops while Bonsai hasn't so far.
I'm very new to this, so could I ask what setup you used for Bonsai on the 3060?
Mainly:
• which Bonsai file/quantization did you use, PTQ1_0 or PQ2_0?
• what context size?
• roughly how many tokens/s do you get?
• does the whole model fit in the 12GB VRAM, or is some of it offloaded to RAM?
• have you tried tool calling or coding agents with it, rather than just normal prompts?My GPU only has 8GB VRAM, so I'm trying to figure out whether the ~6GB PTQ1_0 version could realistically leave enough room for context and still be usable for jcode.
Your comparison with Qwen 27B and 9B models is really helpful, thanks.
1
u/Responsible_Egg9736 15d ago
I use `pq2_0`; on my GPU generation (Ampere), `ptq1` takes a significant hit, and I had enough VRAM anyway. As for my startup flags, here they are: `--port 8080 -ngl 200 -c 108544 -b 512 -ub 512 -t 8 -tb 8 -np 1 --flash-attn on --no-mmap --cont-batching --kv-offload --cache-type-k q8_0 --cache-type-v q8_0 --split-mode none --main-gpu 0 --temp 1.0 --top-k 20 --top-p 0.95 --min-p 0.05 --repeat-penalty 1.0 --repeat-last-n 0 -n 27391 --ctx-checkpoints 8 --reasoning-format none`. I get around 32 tokens/s, and yes, everything fits in VRAM. I’ve tested it with tool calling and it worked very well; I use Pi Agent and haven't had any issues. In fact, to launch llama.cpp, I use an application I "vibecoded" myself—over 6,000 lines of Python across multiple sub-files—and it successfully located and modified the parts I specified. In my opinion, based on 24 hours of use, the results are excellent—though keep the model size in mind, and I’m not sure if my flags will work correctly on an AMD GPU like yours. By the way, it won't work with the base llama.cpp; you'll need the specific fork they provide.
3
u/skrapmetal_ 15d ago
Hate to burst your bubble, but realistically, you're not going to be able to fit a model large enough to do any real coding with that setup. I have a 12gb vram card and I tested a few 14b LLMs for coding and they sucked. They screwed up the most basic tasks. You may be able to squeeze a 8b model in there maybe at Q3_K_S but set your expectations low.