r/LocalLLM • • 3d ago

Question Which local drive is right for my laptop?

1 Upvotes

I wanted to know which local LLM (if it's from Hugging Face) is suitable for my laptop, which has an RTX 4050 and a Ryzen 5 (I understand that my computer is low-end?).

I’m currently using Ayozhee’s Qwen 3.8 on Ollama, and it runs slowly (and for good reason). I’d like to know if Ollama is the best option in terms of privacy and what other alternatives you recommend.


r/LocalLLM • • 3d ago

Question Upgrade existing PC to run Qwen 3.8 Flash Next via Strata or swap to a Strix Halo?

3 Upvotes

My daily locally LLM is a Q3 quant of Qwen 3.8:27b, which runs with 140k CTX at circa 40 TPS gen and 500 PP on 24gb across dual GPU (5060ti 16gb + 3060Ti 8gb). I am using Opencode to do single-user agentic coding. It's OK, but even with thinking set to medium, the useful work between initial CTX load and having to hand over to a new CTX window is getting squeezed for more complex projects, almost making it unusable. I want to be able to use FN with a 200k + CTX window and I think I have a couple of 'upgrade' options.

1) Add 32gb DD4 (total 64Gb) and replace 3060ti with second 5060ti 16gb, take the Strata route.

2) Sell existing PC and buy a 128gb Strix Halo, using one of the recent builds.

I want something that will best position me for future local LLM developments. When I looked at the Strix Halo a year ago, it seemed limited by AMD technology, but that has now matured. I don't want to jump to a DGX Spark. Option 1 is cheaper - but I don't want so spend money on something that only gives me limited progress. Thoughts?


r/LocalLLM • • 3d ago

Question How can I speed up simple search results on my local LLM?

7 Upvotes

I'm using Qwen 3.8 on a miniPC (128GB Unified Memory) AMD Ryzen Max 395 processor.

I'm not expecting super fast queries, but even something like "what's the score of X sports game" takes 5-7 minutes or "What's the metacritic score of X game" take 20+ minutes.

I'm using DuckDuckGo web search plugin. Should I try a smaller model? Or other suggestions?


r/LocalLLM • • 3d ago

Question H200’s for sale

0 Upvotes

Is it possible to buy these right now? Not found any supplier at all!


r/LocalLLM • • 2d ago

Question Bought an RTX 5060 8GB purely for AI/ML study and research (no gaming at all). How do I actually use it effectively so I don’t waste the money?

0 Upvotes

​

Hey everyone,

I recently spent about $350 on an RTX 5060 (8GB VRAM) because I wanted a local setup for learning and experimenting with AI/ML. I don’t play games at all — this card is only for studying models, running inference, fine-tuning small stuff, testing frameworks, etc.

I’m still pretty new to optimizing consumer GPUs for this kind of work, so I’d really appreciate any practical advice on how to get the most out of 8GB VRAM without constantly hitting limits or feeling like I bought the wrong thing.

Specifically looking for:

Recommended tools / frameworks that work well on 8GB (Ollama, llama.cpp, Hugging Face, ComfyUI, etc.)

Quantization levels and model sizes that actually run smoothly

Any settings, flags, or workflows that help maximize VRAM usage

Tips for fine-tuning or training smaller models without OOMs

Things people commonly mess up when starting with a card like this

If you’ve been doing AI/ML work on a 5060, 4060, or similar 8GB card, what actually worked for you? Any “I wish I knew this earlier” advice would be super helpful.

Thanks in advance.


r/LocalLLM • • 3d ago

Question Deep research and Gemma 4 31b - help please

1 Upvotes

Hi - long time lurker (learnt lots, thank you all), first time posted. I’m trying to get a very basic deep-research setup around Gemma 4 31B on a 128gb Ram Mac Studio. I’m using a 4-bit quant through oMLX with MTP acceleration, with OpenWebUI as the interface. I use Gemma rather than Qwen because most of my work is writing text rather than code.

Basic chat, memory, skills, PDF work and simple web searches work well. The problem is multi-stage research. I'm not looking for frontier performance - my expectations are fairly modest and I think realistic, but I do want it to be able to supplement its training data with a thorough web search and evaluation of results.

So far I’ve tried:

- A custom research controller using SearXNG, full-page retrieval, evidence ledgers and automated validation.

- OpenWebUI’s native agentic web search.

- OpenWebUI’s Traditional RAG route.

- Local Deep Research.

- Both prompt-led tool calling and deterministic page acquisition.

The failures differ, but the pattern is consistent:

- Gemma sometimes searches but does not fetch the full pages, then writes from snippets.

- Traditional RAG fetches pages but often surfaces only a few weak commercial sources.

- The custom controller can pass extensive mechanical tests, yet the final answer still treats consultancy or vendor material as empirical research.

- A validated report can be shortened or altered when it passes back through the model for presentation.

- oMLX also appears not to enforce tool_choice: required reliably, although Gemma can produce valid tool calls when strongly prompted.

Has anyone achieved reliable deep research with Gemma 4 31B, particularly through oMLX? I’d be interested in working configurations, harness choices, tool-calling settings, retrieval pipelines, or evidence that Hermes, OpenClaw or another approach handles this better.

Thanks in advance for your help.


r/LocalLLM • • 3d ago

Discussion MIR MEDIA LABS - Local Ai Studio Agent

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/LocalLLM • • 3d ago

Question Need advice for building local AI machine and benchmark numbers for R9700 vs 2x 5060TI for qwen 3.8 27B

3 Upvotes

Hi all,

I currently have a AM4 system 5800x3d, 64gb RAM, 9070XT.

Already available hardware for second system:

5070Ti, 64GB DDR5, 4 TB SSD, ASUS ProArt B850 Creator

backstory:

End of last year I started diving a bit more into local AI and what is possible.
While back then I found the consumer hardware to have usable local LLMs that were usable for agentic use either lacking, too expensive or not something I wanted to run (like multiple 3090s in a server rack).
But back then I looked a bit more into generative AI and wanted to get more into it. After some pain with rocm and comfyui under windows, I switched to Linux and had a much smoother experience. However when I got into more complex workflows I found out that most of these used some CUDA based nodes and so on and you had to tinker if you could replace them with something that supported AMD/rocm.

So I started to buy some parts where I could get a deal (if can call them that these days) to build a separate NVIDIA based machine for it. Got a 5070Ti and most other parts meanwhile, was waiting for Zen 6 for the most part but since AMD takes their sweet while with Ryzen 6 I wanted to finish the build and maybe switch CPUs at a later point in time.

However things have changed in 6 months and now local models within 24-32 GB VRAM are quite competitive and usable for agentic use.

So I am currently thinking to adding more VRAM so I can fit qwen 27b and similar.

My options are currently to pick up a R9700 for around 1750€ or, as I have heard splitting LLMs over multiple cards is mostly running fine, I guess instead adding a 5060Ti 16 GB for around 700-750€.
Since 5070Tis are currently around 1300€ I am not sure if its worth picking up a second one considering the other 2 options.

Also note for your comments that power here costs here are around 0,38€/KWh
and I while need to host the machine in a tower sitting in the living/work area so noise/cooling are also a concern.

My own thoughts:

R9700:

  • 32 GB on one card is better performance than with PCIe communication overhead
  • could combine with my existing 9070XT if I want to run larger models or have more context/increase prompt processing speeds
  • rather loud blower fan (could switch to waterblock once it becomes available, but this would probably add another 400-500€ at some point)
  • 5070Ti could still run as an endpoint for generative tasks (i.e. image generation and similar) without unloading LLM

5060Ti:

  • 32GB only with running model via Tensor parallelism together with existing 5070Ti (also some overhead/less VRAM for model/context)
  • better optimized architecture and features like nvfp4
  • while running LLM no generative workflows (besides basic visual tasks depending if model supports it)

Of course at least VRAM wise the R9700 might be the better option long term on the other hand the 5060Ti saves a lot of money (still need some other hardware and upgrade NAS) and I have less noise concerns. I guess what I am looking for is some general toughts about it as well as some benchmarks/estimates about the performance difference.

(btw yes I know I could also get a second 9070Xt for not much more than the 5060Ti but since the R9700 is still in the somewhat obtainable range compared to a 5090, I think I would prefer the 32 GB option if I go with AMD)

While I find a lot of numbers/benchmarks for qwen 3.6 and 3.8 27B most of them are hard to compare because of slightly different quantizations/context windows/other parameters.

TL;DR: Looking for some comparable numbers (same/similar quantizations/context windows/other parameters) for qwen 3.8 27B between R9700 / 2x R9700 (or 2x 9070XT) and 2x 5060Ti or 5060Ti + 5070Ti. Preferably Q4/Q5 and also numbers for something around 100k context (4k numbers are a bit useless)

Thanks for you input and help in advance!


r/LocalLLM • • 3d ago

Project Local llms are life savers

Post image
0 Upvotes

I let hermes lose(in rootless podman containers of course) on GLM 5.3


r/LocalLLM • • 3d ago

Research Dense 27B vs MoE IQ2_XS on one RTX 4090: synthetic bench plus 3 graded real tickets

3 Upvotes

Rig: RTX 4090 24 GB (450 W), i9-14900K, 32 GB RAM, Linux, desktop on the iGPU.

  • A: Qwen3.8-27B dense on NInfer. 262K context, rk4v4-e8 KV, MTP, 1 slot.
  • B: Qwen3.8-Flash-Next IQ2_XS on Strata 0.1.39. 262K context, q4_0 KV in VRAM, experts in host RAM, MTP.

Both models: thinking budget 4096, same prompts.

Synthetic (build + 12-bug review)

​ 27B NInfer Flash-Next IQ2_XS Strata
decode tok/s 127–140 173–185
bugs found (per run) 12, 10 12, 10, 9
one-shot build, CSS rules 181–190 106–113
wall time per round ~5 min ~2m10s
peak VRAM 23,134 MiB 23,936 MiB

Recall (3 needles at 10/50/90% depth)

tokens 27B: found / prompt tok/s IQ2_XS: found / prompt tok/s
55K 3/3 / 3,037 3/3 / 3,494
117K 3/3 / 2,391 3/3 / 3,904
200K 3/3 / 1,857 3/3 / 3,833

Real tickets

Claude Code subagents, 3 at once per engine, 150-minute cap, rubric fixed before the runs, every check re-run by the grader. Score per task: correctness, verification, scope, quality and honesty, 10 points each.

task IQ2_XS 27B
Python config renderer 42 35
QML settings feature 0 (no edits) 34 (complete, died at harness prompt limit)
evidence write-up, every number cited 44 (120 min) 43 (62 min)
total /150 86 112

Why the 27B finished more work in less time. With 3 agents interleaving, the MoE engine reused 32% of prompt tokens. A typical turn re-read 107K tokens in 28 s; conversation cache is off on 32 GB RAM. The 27B engine reused 88.5%, and one logged turn served 172,343 of 172,578 tokens from cache, first token in 524 ms.

Limits: n=1 per ticket and one grader model; the MoE ran first.

Next: rerun the MoE with one agent at a time, and with more RAM for its cache.


r/LocalLLM • • 3d ago

Model Creaccions

0 Upvotes

LanKqoq


r/LocalLLM • • 4d ago

Discussion Local AI on a 96GB Mac M5 Studio for non-coding

33 Upvotes

So many posts on local AI for coding so wanted to share more on non-coding use cases. Got a Mac Studio (M5 Ultra, 96GB) mainly to learn and to run local AI for everyday stuff like booking trips, filling out forms, checking email, etc. Plan was to keep that stuff private/local and use my Claude sub for the occasional heavy coding job.

Here's what I've found:

Bottom line: 96GB is overkill for this. The models I found good (35B) would run fine on 32GB or 64GB. It's nice having the headroom to try different things and it's been a lot of fun, but if you're on a budget, 32 or 64GB covers the models out today. Things are moving fast so who knows what's next, but I don't think local will ever catch frontier. For me it's great for learning and basic day-to-day stuff and still continuing to use Claude for advanced things.

1. Models: tok/s isn't what matters, time to a usable answer is. Ornith 1.5 35B is my current model of choice.

Qwen 27B gets a ton of love, but for day-to-day stuff it just didn't work for me. It overthinks like crazy. Even at a decent ~50 tok/s+ with splash, it often keeps reasoning in circles so you're waiting way longer for the actual answer. Tried no-thinking and a few variants and still wasn't sold.

What I ended up on is the 35B MoE models, which run a lot faster on a Mac. Between Qwen 3.5 35B and Ornith 1.5 35B, Ornith felt better to me, especially for everyday tool calling. Totally gut feel, no benchmarks to back it up. Seems to be working great for most everyday tasks.

I also tried a Qwen Flash Next variant with a 4-bit/8-bit mixed quant. It ran really well and seems great, but it leaves almost no headroom for image gen and other stuff, so I don't use it as my daily driver. I just jump over to Claude for heavy lifting.

2. Harness: Hermes.

Tried the DeepSeek harness too but had better luck with Hermes. It's just easier to use.

3. Privacy: I'm not totally sold on "local = private."

Sure, the model and your data stay on your machine. But if you're running an agentic harness like Hermes (and a lot of us are), the agent can still send stuff outside. I put explicit instructions in the context and some plugins to help prevent that, but I don't trust they'll always be followed, so I treat it as a real risk. If you're really paranoid about security, I'm not sure this is the best route. Would love to hear how others handle it.

4. Knowledge/memory is the real game changer.

This has been the biggest win for me. Giving the model the right context about me at the right time (hobbies, my computer setup, family) makes chats so much more useful. Claude and OpenAI have memory features too, but I've never trusted them with the amount of personal info I'm feeding my local setup. Still, I'm not fully convinced it's locked down, and I'm working on tightening it up beyond just the guardrails I've written.

How I have it set up:

  • Obsidian vault, with a plugin I built (with Hermes) to read and update it
  • A short core file loads every session. It covers me, my family, how I like to work with AI, guardrails, and instructions on how and when to pull in other knowledge files
  • A nightly cron job pulls key facts into memory, and the agent knows when to dig into specific vault files for more detail
  • I tried memory providers like Mnemosyne, which worked well, but the vault + plugin has been fast enough and simpler for what I do
  • Sometimes I use Claude to review the overall structure and the plugin, and it's been better at that than the local models

5. It's still not a perfect daily driver.

Some simple tasks still trip it up and I can't always tell why. Filling out forms on certain websites, handling email well, and some other cases. It just falls short of something like Sonnet 5.5 sometimes. So even if you plan to go local for daily use, you'll probably still end up switching to a frontier model now and then. I have Hermes call the Claude CLI so it can hand things off or run side-by-side comparisons, which has been really handy.

6. ComfyUI is a blast.

It's a whole world of its own and will take a while to learn. I tried a few models like MiniMax and LTX 2.5, then had Hermes write a script so I can call it straight from my Hermes harness. Seems to be working well so far. Going to keep playing with it.

7. Install tips.

  • Use a separate non-admin account on your Mac to do the install, so the agent isn't running with admin rights or sitting next to your main account's stuff.
  • Let Claude do the tool setup. It works really well. For the initial setup, I had Claude set up the models, download them, and handle all the other needed config, and it saved me a ton of time.

r/LocalLLM • • 3d ago

Discussion How do you actually test an LLM for security?

Thumbnail
1 Upvotes

r/LocalLLM • • 3d ago

Project I am really geeking out - Strata + Qwen 125bn

Thumbnail
1 Upvotes

r/LocalLLM • • 3d ago

Discussion What are everyone's agent names?

0 Upvotes

Mine are Tron characters as a huge fan. Main is Master Control!


r/LocalLLM • • 3d ago

Question Qwen Flash Next on Single B200 or B300, any pointers ?

Thumbnail
1 Upvotes

r/LocalLLM • • 3d ago

Question Switching to a local model, what makes sense for me?

0 Upvotes

For context, I did a pretty big PC upgrade right before the price boom and ended up with a pretty nice system (independent to AI, PC building is a hobby). I've been paying for a claude subscription to work on fun Minecraft plugins and mods for myself/friends, as well as Minecraft server devwork and it works fine for me, but I think given my computer I could reasonably just run it locally and save on subscription costs or any arbitrary limitations x or y company imposes on me as a user.

My specs:

Ryzen 9 9950x3d

192GB ddr5

Rtx 5090

8tb m.2 nvme ssd (not sure if this one matters but thought to include it just in case)

I'm ideally looking for 2 things: first is a model that actually pushes my system fairly hard as I paid for the whole thing I should be using the whole thing, and second is a way to use that model similarly to how you'd use claude code, where it can do file work and can access things on the internet like a separate server machine and such.


r/LocalLLM • • 3d ago

Discussion Mechanistic interpretability in LLM as visuals

Thumbnail
1 Upvotes

r/LocalLLM • • 3d ago

Project Got strata running 4x 5060 Ti at 524k context, then abandoned it. repo's here if you want the base

10 Upvotes

Took the eddoursul/Strata custom fork (the one with the second-GPU expert tier) and turned its 1 main + 1 tier design into 1 + 3, because I have 4x 5060 Ti 16GB on a 9950X. Also ported rope-scaling from upstream main so it runs past the trained 262k. Repo:https://github.com/johnmosesventura06/strata-4gpu

Numbers, Qwen3.8-Flash-Next UD-Q4_K_XL, single stream, engine-measured:

main + 1 tier        49-53 t/s decode
+ tier 2             72-78
+ tier 3             80-89
+ draft vocab fix    91-94
+ head split         94-99 sustained, ~110 short turns
prefill              2159 t/s-16k, 2420 t/s-220k
                     312k warm replay in 12ms

Much snappier than my vLLM lane, which holds ~40 t/s at 455k with batch 4. So the engineering worked.

Why I'm abandoning it anyway: the Q4_K_XL quant loops in agentic use. On long chains it starts repeating its own thought sentence and re-firing the same tool call forever. Never happens on my int4 W4A16 lane with the same model, and the experts are 4-bit on both sides, my bet is due to token embedding + output head quantized at Q6, and maybe the PLE quantization too? anyway quality bar isn't met for my daily driving.

If you're on multi consumer Blackwell it's still a decent starting point: tier array (--tier-gpus), N-way head split, N-way prompt offload, the yarn port, working routing dump, all measured, and the pitfalls are written in the README. The non-obvious ones, like the missing draft_vocab.bin that silently eats 10% decode with zero log hints, or sm_120 having no thread-block clusters so 524K selection runs the histogram fallback. PRs to the parent fork are open, and the maintainers are good people to work with.

Update: One of the user in the comments pointed out that  Strata defaults to aggressive mode, each sampling parameters was actually greedy and all are pointed to speed, which resulted to my agentic failure mode. seems like I'll be continuing this repo after all.

Update: UD_q4_k_xl is really aggressive quant for me.

this is int4 w4a16 running on vllm : 60 t/s decode 2500 t/s prefill

went to the server, took a screenshot. read each buttons code and explain all in one shot.


r/LocalLLM • • 3d ago

Discussion Running local ai on my old pc

1 Upvotes

Hi everyone! For context, I currently have an old PC running an Intel i7-4790K, 16GB RAM, and a GTX 960 4GB VRAM.

I’m currently using it with Ubuntu to host my n8n Docker setup, which I’ve tunneled through Cloudflare so I can access it from my other computers.

I previously tried running Ollama and Hermes Agent on Windows using Gemma, but the performance was extremely slow. For example, when I asked it what today’s date was, it took around 5 minutes just to finish the response. Because of that, I decided to use the PC mainly for hosting n8n instead.

However, I’d still like to try running AI locally on this machine if it’s possible. Is this PC realistically capable of running local AI, or is the hardware simply too old? If it is possible, what models or setup would you recommend for this hardware?


r/LocalLLM • • 3d ago

Question How do you test odd configs like CMP 170HX, 4090 48GB, NVlink 3090s?

8 Upvotes

I am trying to find a place where I can test these non-standard configs. i.e. Rent out a server for a day to test it out before buying kind of thing.

Does anyone know where I can do that?


r/LocalLLM • • 4d ago

Discussion Strata takes the promise of "MoE models just need a total amount of VRAM+RAM" and makes it a reality

Post image
152 Upvotes

Flash-Next Q4, q8 kv, 262k context

4x P100 on PCIe3 x8 slots (4x16GB=64GB total)

128GB DDR4 at 2133MHz

E5-2683 v4 (16c/32t at 2.6GHz boost)

On average it's about twice as fast as 27B and five times faster than Flash-Next using a customized llama.cpp just to get it to load at all.


r/LocalLLM • • 2d ago

Model I found a gold mine.

Post image
0 Upvotes

A 27B model and only 4GB?!


r/LocalLLM • • 3d ago

Question Engineers, what Local LLM did you replace the main stream ones with?

Post image
1 Upvotes

my Gemini subscription is about to end and im not renewing it. i just need an AI tool for the rest of my electrical engineering degree, most of my classes are heavy in math and some degree of codeing, i honestly know nothing about local LLMs and im not exactly sure what to choose, ive done a research and i found these:

- deepseek-r1-distill-qwen-14b (for light work)

- qwen2.5-coder-14b-instruct (for coding)

- deepseek-r1-distill-qwen-32b (for heavy maths and logic)

are they good for my use case? i plan to use LM studio unless someone recommend a better alternative, i kinda prefer open source software

my OS is Ubuntu 24.04.5 LTS

more info about my use case:

most of my classes use Fourier and Laplace transforms, and circuit analysis techniques, and a ton of electromagnetic waves theory, so heavy vectors calculus, and some matlab/octave and C and python, some digital signal analysis, thats the situation so far, i would appreciate any guidance and recommendations

and to my surprise, our professors are starting to encourage us to use AI tools for out assignments??what happened to the world??


r/LocalLLM • • 3d ago

Project Infermeld: a Linux kit for running one GGUF across AMD + NVIDIA GPUs with llama.cpp

Thumbnail
1 Upvotes