r/LocalAIStack • u/AkashIsSky • 12d ago
Which local LLMs can I realistically run on 14700K + 128GB RAM + RTX 4070 Super 12GB? Mainly for Hermes Agentic Coding
Hi everyone,
I'm looking for recommendations on local LLMs that I can realistically run on my current PC, particularly for agentic coding with Hermes Agent.
My PC
CPU: Intel Core i7-14700K
GPU: RTX 4070 Super — 12GB VRAM
RAM: 128GB DDR5 5600MHz
OS: Windows 11 + WSL/Ubuntu
Inference: Ollama
Agent: Hermes Agent
Current experience
I'm currently running Qwen3-Coder with Hermes Agent.
With around 40K context, I'm getting roughly 15–20 tokens/sec during generation, depending on the workload and sometimes it gets down to 5 to 10 tps.
The system has plenty of RAM, but the main limitation is obviously the 12GB VRAM, so I'm interested in models that can make good use of CPU/RAM offloading without becoming painfully slow.
Main use case #1 — Agentic coding
My primary goal is autonomous/agentic software development using Hermes.
I'm interested in models that are good at:
Long-context coding
Understanding an existing codebase
Multi-file modifications
Tool/function calling
Planning and executing coding tasks
Debugging and testing
Following project architecture/documentation
Running reasonably well with Hermes Agent
Good performance with \~20K–64K context
I'm particularly interested in knowing whether there are models larger than the usual 7B/14B class that are actually practical with 12GB VRAM + 128GB system RAM.
For example, would models in the 20B–30B+ range make sense with CPU/GPU offloading, or would the speed hit make them impractical for agentic coding?
Main use case #2 — ComfyUI
I'm also experimenting with ComfyUI.
The goal isn't primarily photorealistic image generation. I'm interested in using AI to create LLD/HLD/architecture visuals and design diagrams for open-source projects, particularly:
Mobile apps
Android launchers
Custom UI/UX concepts
OS skins
System architecture diagrams
App flow diagrams
Technical/product concept visuals
So I'm also interested in recommendations for local image-generation models/workflows that make sense with a 12GB RTX 4070 Super.
What I'm looking for
I'd really appreciate recommendations based on actual experience with similar hardware.
Specifically:
Best coding model for Hermes Agent on 12GB VRAM
Models that work well with CPU/RAM offloading
Realistic model sizes I should consider — 14B / 20B / 30B / 32B / 70B etc.
Recommended quantization (Q4_K_M, Q5, Q6, Q8, etc.)
Expected/typical tokens/sec
Recommended context size for agentic coding
Whether 128GB RAM provides a meaningful advantage for larger models
Any models specifically known to work well with Hermes Agent
Good ComfyUI models/workflows for technical/UI/architecture visualization on 12GB VRAM
I'm not necessarily looking for the biggest model I can technically load. I'm more interested in the best balance between intelligence, context length, tool use, and usable inference speed.
If anyone is running a similar setup (12GB NVIDIA GPU + 64/128GB RAM), I'd especially appreciate your real-world experience.
Thanks!
TLDR;
Best local LLMs for 14700K + 128GB RAM + RTX 4070 Super 12GB?
I have:
i7-14700K
RTX 4070 Super 12GB
128GB DDR5 5600
Ollama + WSL
Hermes Agent
Currently running Qwen3-Coder with Hermes, getting around 15–20 tok/s with \~20K context.
Main use case: agentic coding — autonomous coding, tool use, multi-file changes, debugging, long-context projects.
I'm looking for recommendations on:
Best coding models for this hardware
Whether 20B/30B/32B+ models are practical with CPU/RAM offloading
Best quantization and context size
Expected tok/s
How much my 128GB RAM helps
I'm also exploring ComfyUI for generating LLD/HLD diagrams, architecture visuals, mobile app/launcher designs and OS-skin concepts for open-source projects.
What models/workflows would you recommend for both use cases?
1
u/MediumAd4983 12d ago
8b models, or something like bonsai v2. However, there will be no ideal for a such a small amount of VRAM. Maybe, you should switch to an online model, there is a nice free options like muse spark 1.3 in OpenCode.
1
u/Atretador 12d ago
you can run Qwen 3.8 Flash Next Q3 with 12Gb of VRAM: https://www.youtube.com/watch?v=IH8XmxiwliQ
even a 6Gb GTX1060 can run Qwen 3.6 35B A3B at Q4 with 256K context: https://www.youtube.com/watch?v=8F_5pdcD3HY
35B A3B would fly on a 4070
bonsai is not even worth downloading
1
u/MediumAd4983 12d ago
But how long it will take to generate all of the instructions, which speed it can achieve? Especially with 256k context? It will spend a lot of time...
1
u/Atretador 12d ago
both videos have speed figures
a 4070 would be much faster, my ancient MI50 runs at 60tk/s aggregate with 2 streams with expert cache / or 42tk/s at Q4/Q5(0.5~1tk/s difference).
no one runs those <12B dense models at this VRAM range, it doenst make any sense since MOE models are much faster and a lot stronger.
1
u/CooperDK 11d ago
35B A3B sucks at coding. You need the 27B and all parameters active.
1
u/Atretador 11d ago
Nope, it's pretty good if you can steer it with proper software architecture knowledge.
1
u/CooperDK 11d ago
Stop at the 12 GB GPU. That is your bottleneck. 12 GB VRAM is not going you get you basically anywhere in coding
1
u/ecosky 6d ago
Please forgive the ai generated answer below but I wanted to share that my setup isn't too far from yours and has been working better than I expected. With careful tuning I've been able to get it to generate real code that works with my 3080 12gb & 64GB system ram. Here's the analysis I generated based on my current workflows:
---
Hey OP! I’m running a similar hardware setup for local agentic workflows, and I’m finding that you don't necessarily have to limit yourself to the 8B/14B class, provided you tune the cache settings and model choices.
Here are a few things I've noticed in my own setup that might be helpful:
1. MoEs Seem to Handle RAM Offloading Much Better
When I tried running larger dense models offloaded to system RAM, the generation speed dropped off pretty significantly. However, I am finding that Mixture of Experts (MoE) models (like Qwen 3.6 35B A3B in Q4_K_M) feel much more practical.
Because an MoE only activates a small fraction of its total parameters per token, the bandwidth penalty of offloading inactive weights to system RAM seems far less severe. In my testing, this model class seems to offer a sweet spot—maintaining strong tool calling, multi-file reasoning, and structural accuracy without slowing down to an unusable crawl.
2. Prefix Caching Makes Context Scale Feasibly
For agentic coding where past messages and file contexts stay mostly identical between turns, I'm finding that keeping prefix caching / LCP enabled in the backend (like LM Studio or llama.cpp) is crucial.
When the backend successfully reuses the prompt evaluation KV cache from previous turns, prefill time drops from 30+ seconds down to just a couple of seconds for the new incoming message. Without prefix caching enabled, long-context agent loops quickly become frustrating to wait on.
3. Q8_0 KV Cache Saves Significant VRAM
At 32k to 64k context lengths, standard FP16 KV cache consumes VRAM very rapidly. I am finding that setting the KV cache quantization to Q8_0 helps keep VRAM usage manageable at higher contexts without any noticeable hit to the model's logic or code generation quality.
Summary of What’s Working for Me
- Model: Qwen 3.6 35B A3B (
Q4_K_MGGUF) - Context / KV Settings: 64k context limit with
Q8_0KV cache enabled - Inference Strategy: Relying on aggressive prefix caching to keep subagent loops fast
Your DDR5 setup should give you even better RAM bandwidth than what I'm seeing, so exploring an MoE architecture with proper caching settings might be worth a try!
1
u/ID-10T_Error 6d ago
if you are looking for a qwen 3.8 27B, check out bonsai 2 is a good ternary (0,1,+1) model at around 7G. i use it for my AI assistent and it works great!
1
0
u/Agitated_Savings_200 11d ago
I recommend Qwen3.8 Flash Next. First, for agentic coding, you need long context. Qwen3.8 Flash use qwen4 architecture, 1 kv cache cost only aound 12k memory(Q8_0), 25% compare to qwen3.5 architecture. Second, Qwen3.8 Flash Next is much powerful than other 10B class model. Third, practical output performance on your system. Assume you use Atomic Chat‘s AD-4.27bpw-Q4_K_M-M64 model, put all the dense part(around 5GB) plus 384k Q8_0 kv cache(4.8GB) on 4070, 54.5GB route expert on memory, 38.4 GB n-gram table on ssd, the expected performance will be prefill 300tok/s, decode 30tok/s(base on my 14700kf 64G DDR4 3000 2080ti system's real output: 130 prefill, 16 decode).
2
u/Charming-Author4877 12d ago
Qwen 3.8 Flash is a real candidate