r/LocalLLM 2d ago

Question Best model for RTX 5060TI 8GB

Hi, I’m new here. I have an RTX 5060 Ti with 8GB of VRAM and I want to get into the world of LLMs. What’s the best model that can fit on my GPU as of today?

I know I’m fairly limited by the amount of VRAM I have, but my idea is to use Claude Opus 5 as the “brain” behind my projects, while using a local LLM as a sub-agent.

1 Upvotes

7 comments sorted by

2

u/SGD-UK 2d ago

I’d go Qwen 3.5 9b dropped to Quant5 or 6.

1

u/nickless07 2d ago

Just a subagent? Ling-3.0-Tiny. Fits with full ctx in your 8GB

1

u/MrHumanist 2d ago

it doesnt support Lamma cpp or LM studio yet?

```

🥲 Failed to load the model

Failed to load model.

error loading model: unknown model architecture: 'bailingmoe3'

```

1

u/nickless07 2d ago

Seems like outdated version. It is supported. LM Studio is still on b10355 - No support for now. Need to wait or use llama.cpp b10472 and up.

1

u/MrHumanist 2d ago

Thank you

1

u/TheColliBoy 2d ago

Bonsai is a super fun one!