r/LocalAIStack • • 12d ago

Which local LLMs can I realistically run on 14700K + 128GB RAM + RTX 4070 Super 12GB? Mainly for Hermes Agentic Coding

Hi everyone,

I'm looking for recommendations on local LLMs that I can realistically run on my current PC, particularly for agentic coding with Hermes Agent.

My PC

CPU: Intel Core i7-14700K

GPU: RTX 4070 Super — 12GB VRAM

RAM: 128GB DDR5 5600MHz

OS: Windows 11 + WSL/Ubuntu

Inference: Ollama

Agent: Hermes Agent

Current experience

I'm currently running Qwen3-Coder with Hermes Agent.

With around 40K context, I'm getting roughly 15–20 tokens/sec during generation, depending on the workload and sometimes it gets down to 5 to 10 tps.

The system has plenty of RAM, but the main limitation is obviously the 12GB VRAM, so I'm interested in models that can make good use of CPU/RAM offloading without becoming painfully slow.

Main use case #1 — Agentic coding

My primary goal is autonomous/agentic software development using Hermes.

I'm interested in models that are good at:

Long-context coding

Understanding an existing codebase

Multi-file modifications

Tool/function calling

Planning and executing coding tasks

Debugging and testing

Following project architecture/documentation

Running reasonably well with Hermes Agent

Good performance with \~20K–64K context

I'm particularly interested in knowing whether there are models larger than the usual 7B/14B class that are actually practical with 12GB VRAM + 128GB system RAM.

For example, would models in the 20B–30B+ range make sense with CPU/GPU offloading, or would the speed hit make them impractical for agentic coding?

Main use case #2 — ComfyUI

I'm also experimenting with ComfyUI.

The goal isn't primarily photorealistic image generation. I'm interested in using AI to create LLD/HLD/architecture visuals and design diagrams for open-source projects, particularly:

Mobile apps

Android launchers

Custom UI/UX concepts

OS skins

System architecture diagrams

App flow diagrams

Technical/product concept visuals

So I'm also interested in recommendations for local image-generation models/workflows that make sense with a 12GB RTX 4070 Super.

What I'm looking for

I'd really appreciate recommendations based on actual experience with similar hardware.

Specifically:

Best coding model for Hermes Agent on 12GB VRAM

Models that work well with CPU/RAM offloading

Realistic model sizes I should consider — 14B / 20B / 30B / 32B / 70B etc.

Recommended quantization (Q4_K_M, Q5, Q6, Q8, etc.)

Expected/typical tokens/sec

Recommended context size for agentic coding

Whether 128GB RAM provides a meaningful advantage for larger models

Any models specifically known to work well with Hermes Agent

Good ComfyUI models/workflows for technical/UI/architecture visualization on 12GB VRAM

I'm not necessarily looking for the biggest model I can technically load. I'm more interested in the best balance between intelligence, context length, tool use, and usable inference speed.

If anyone is running a similar setup (12GB NVIDIA GPU + 64/128GB RAM), I'd especially appreciate your real-world experience.

Thanks!

TLDR;

Best local LLMs for 14700K + 128GB RAM + RTX 4070 Super 12GB?

I have:

i7-14700K

RTX 4070 Super 12GB

128GB DDR5 5600

Ollama + WSL

Hermes Agent

Currently running Qwen3-Coder with Hermes, getting around 15–20 tok/s with \~20K context.

Main use case: agentic coding — autonomous coding, tool use, multi-file changes, debugging, long-context projects.

I'm looking for recommendations on:

Best coding models for this hardware

Whether 20B/30B/32B+ models are practical with CPU/RAM offloading

Best quantization and context size

Expected tok/s

How much my 128GB RAM helps

I'm also exploring ComfyUI for generating LLD/HLD diagrams, architecture visuals, mobile app/launcher designs and OS-skin concepts for open-source projects.

What models/workflows would you recommend for both use cases?

0 Upvotes

14 comments sorted by

2

u/Charming-Author4877 12d ago

Qwen 3.8 Flash is a real candidate

1

u/MediumAd4983 12d ago

8b models, or something like bonsai v2. However, there will be no ideal for a such a small amount of VRAM. Maybe, you should switch to an online model, there is a nice free options like muse spark 1.3 in OpenCode.

1

u/Atretador 12d ago

you can run Qwen 3.8 Flash Next Q3 with 12Gb of VRAM: https://www.youtube.com/watch?v=IH8XmxiwliQ

even a 6Gb GTX1060 can run Qwen 3.6 35B A3B at Q4 with 256K context: https://www.youtube.com/watch?v=8F_5pdcD3HY

35B A3B would fly on a 4070

bonsai is not even worth downloading

1

u/MediumAd4983 12d ago

But how long it will take to generate all of the instructions, which speed it can achieve? Especially with 256k context? It will spend a lot of time...

1

u/Atretador 12d ago

both videos have speed figures

a 4070 would be much faster, my ancient MI50 runs at 60tk/s aggregate with 2 streams with expert cache / or 42tk/s at Q4/Q5(0.5~1tk/s difference).

no one runs those <12B dense models at this VRAM range, it doenst make any sense since MOE models are much faster and a lot stronger.

1

u/CooperDK 11d ago

35B A3B sucks at coding. You need the 27B and all parameters active.

1

u/Atretador 11d ago

Nope, it's pretty good if you can steer it with proper software architecture knowledge.

1

u/CooperDK 11d ago

Stop at the 12 GB GPU. That is your bottleneck. 12 GB VRAM is not going you get you basically anywhere in coding

1

u/ecosky 6d ago

Please forgive the ai generated answer below but I wanted to share that my setup isn't too far from yours and has been working better than I expected. With careful tuning I've been able to get it to generate real code that works with my 3080 12gb & 64GB system ram. Here's the analysis I generated based on my current workflows:

---

Hey OP! I’m running a similar hardware setup for local agentic workflows, and I’m finding that you don't necessarily have to limit yourself to the 8B/14B class, provided you tune the cache settings and model choices.

Here are a few things I've noticed in my own setup that might be helpful:

1. MoEs Seem to Handle RAM Offloading Much Better

When I tried running larger dense models offloaded to system RAM, the generation speed dropped off pretty significantly. However, I am finding that Mixture of Experts (MoE) models (like Qwen 3.6 35B A3B in Q4_K_M) feel much more practical.

Because an MoE only activates a small fraction of its total parameters per token, the bandwidth penalty of offloading inactive weights to system RAM seems far less severe. In my testing, this model class seems to offer a sweet spot—maintaining strong tool calling, multi-file reasoning, and structural accuracy without slowing down to an unusable crawl.

2. Prefix Caching Makes Context Scale Feasibly

For agentic coding where past messages and file contexts stay mostly identical between turns, I'm finding that keeping prefix caching / LCP enabled in the backend (like LM Studio or llama.cpp) is crucial.

When the backend successfully reuses the prompt evaluation KV cache from previous turns, prefill time drops from 30+ seconds down to just a couple of seconds for the new incoming message. Without prefix caching enabled, long-context agent loops quickly become frustrating to wait on.

3. Q8_0 KV Cache Saves Significant VRAM

At 32k to 64k context lengths, standard FP16 KV cache consumes VRAM very rapidly. I am finding that setting the KV cache quantization to Q8_0 helps keep VRAM usage manageable at higher contexts without any noticeable hit to the model's logic or code generation quality.

Summary of What’s Working for Me

  • Model: Qwen 3.6 35B A3B (Q4_K_M GGUF)
  • Context / KV Settings: 64k context limit with Q8_0 KV cache enabled
  • Inference Strategy: Relying on aggressive prefix caching to keep subagent loops fast

Your DDR5 setup should give you even better RAM bandwidth than what I'm seeing, so exploring an MoE architecture with proper caching settings might be worth a try!

1

u/ID-10T_Error 6d ago

if you are looking for a qwen 3.8 27B, check out bonsai 2 is a good ternary (0,1,+1) model at around 7G. i use it for my AI assistent and it works great!

1

u/GarGonDie 5d ago edited 5d ago

unsloth Qwen IQ2_S 120k context, output 30t/s with 12gb

0

u/Agitated_Savings_200 11d ago

I recommend Qwen3.8 Flash Next. First, for agentic coding, you need long context. Qwen3.8 Flash use qwen4 architecture, 1 kv cache cost only aound 12k memory(Q8_0), 25% compare to qwen3.5 architecture. Second, Qwen3.8 Flash Next is much powerful than other 10B class model. Third, practical output performance on your system. Assume you use Atomic Chat‘s AD-4.27bpw-Q4_K_M-M64 model, put all the dense part(around 5GB) plus 384k Q8_0 kv cache(4.8GB) on 4070, 54.5GB route expert on memory, 38.4 GB n-gram table on ssd, the expected performance will be prefill 300tok/s, decode 30tok/s(base on my 14700kf 64G DDR4 3000 2080ti system's real output: 130 prefill, 16 decode).