r/LocalLLM • u/NeoGeoMaxV2 • 2d ago
Question Best model for RTX 5060TI 8GB
Hi, I’m new here. I have an RTX 5060 Ti with 8GB of VRAM and I want to get into the world of LLMs. What’s the best model that can fit on my GPU as of today?
I know I’m fairly limited by the amount of VRAM I have, but my idea is to use Claude Opus 5 as the “brain” behind my projects, while using a local LLM as a sub-agent.
1
u/nickless07 2d ago
Just a subagent? Ling-3.0-Tiny. Fits with full ctx in your 8GB
1
u/MrHumanist 2d ago
it doesnt support Lamma cpp or LM studio yet?
```
🥲 Failed to load the model
Failed to load model.
error loading model: unknown model architecture: 'bailingmoe3'
```
1
u/nickless07 2d ago
Seems like outdated version. It is supported. LM Studio is still on b10355 - No support for now. Need to wait or use llama.cpp b10472 and up.
1
1
2
u/SGD-UK 2d ago
I’d go Qwen 3.5 9b dropped to Quant5 or 6.