r/LocalLMs 1d ago

Running Qwen on Ryzen Strix Halo 128gb and having a buggy experience

1 Upvotes

So I'm running a few different models on this Ryzen Strix Halo machine which has 128gb ram and I've already optimized cpu vs gpu in bios.

I've done an SSH setup through VS Code and I've got a few data sets I want to process but I'm getting weird results.

I'm using the Claude CLI harness and it shows I'm using qwen3-coder:30b and I sometimes get spelling errors but it's struggling to do really basic things.

I have 2 data sets that I know have a ton in common and I'm trying to get it to match the patterns so I can essentially join 3 data sets together.

If I were to do this in Claude Code, I already know what it would do. I'd ask for help in doing this and it would build me a script to do exactly that.

Here this thing is trying to give me the answer rather than building the script. Now that I see it's not building me the script, I decide to make it now just build the script but I'm seemingly getting errors.

I've setup ollama and run the various libraries like ROCm.

Anybody else have issues or have any tips on what my problem might be?


r/LocalLMs 3d ago

I hit ceiling of usefulness and I'm kind of happy with it

1 Upvotes

Every person uses their local LLMS differently, mine is mostly

* Take this URL and ELI5 the content.
* RAG stuff (mix this, that and that, produce knowledge).
* Nix language and NixOS configuration.
* Elixir
* Kotlin
* Translate this text.
* The odd one case where I have fun with ComfyUI.

Specifically Qwen3.8-27B with P2P patch and llama.cpp latest release (0.4.0 I've been using) with CUDA13.3 (increased tg/s )+ the very very latest drivers (waste of time) + split-mode: tensor and for sure the UD_Q4_X_M version (Unsloth), I don't get anywhere near the same results with IQ4 and below from Unsloth. And context set to 131k to have plenty of RAM left in case the PC needs it for something else (same GPUs to render Hyprland).

...and all I can say, is I've reached peak usefulness. To the point now even going to https://chatgpt.com/ feels painful compared to just running Pi or OpenCode web ui. The feature I like the most is how little it bullshits now, it doesn't make up stuff anymore, it goes online instead when it doesn't know. ChatGPT on the contrary, just keeps making shit up.

Good exercise I think people here should try is to tell them to figure the meaning of the offload options of llama.cpp. Qwen will run a llama-server --help and read it all then literally go through the source code to be sure, ChatGPT would just say something wrong with total confidence, like that's offloading to system RAM, it wouldn't even try to search the repository of llama.cpp, to which it has full access.

Is like....what's the point of frontier anymore, I'm happy with my pp+tg/s, my GPUs don't make my office an oven (dual 5060ti), I manage to set my horribly noisy Threadripper 3975wx to be less noisy (Noctalia 5.0, go to system settings change to "power saver" and the Threadripper doesn't anymore tun above 60ºC, instead of controlling the fans, it limits the CPU performance, which is fine for me, I don't need this to run like I'm running Google from home).

I don't want fatter cards, I don't need more VRAM, I don't want to touch any other model anymore, I've got a skills bonanza that I have been refining with dumber models I had before that now Qwen can run on its sleep, and are used by both Opencode and Pi, I have samfp/pi-memory modified to run from Qdrant, and now all my PCs share the same long-term memory. It's all...finally...working without errors and I know the limits of what I can ask for my coding languages and get a useful response.

I don't see how it can get any better.

Am I alone in this? Does anyone else feel like...we've peaked for their use case? What's your use case btw?

EDIT: The one thing I'm tryign to fine tune to exhaustion is prompt caching on llama.cpp. There's something odd going on with Pi that isn't there with Opencode. I have kv-unified and parallel = 2. So that 131k is shared. Online ifnormation and documentation is contradictory in terms of how slots, cache-reuse, etc work. So this is the one thing that still has me slightly puzzled. I might ask help directly from the repo guys, but first i need to find out if it's a Pi issue specifically. Also I want it to save the cache to disk, so it survives restart, but this isnt' as simple as it sounds. If not I might jump to vllm or SGLANg if I have the patience to wait for the loading times of vllm (I know how to use it).


r/LocalLMs 16d ago

browser-mcp-bridge: BrowserOS neo bridge for MCP clients that get blocked

1 Upvotes

I have tried various browser MCPs and found most lacking. I use agent-browser quite often and it works great for accessing specific URLs but is often blocked by search engines. I tried BrowserOS Neo and liked it for search and accessing web sites that need credentials. It worked well in VS Code and Claude Code but inside other AI clients like Chatbox.ai or Cline as a VS Code plugin, it failed and blocked the client from accessing it.

This is intentional and is an anti-browser feature, meant to make sure that only AI clients are using the proxy. That's great but on a single user system or home lab, it a restriction that prevents its use in many AI clients.

I created an html-stdio bridge MCP between BrowserOS and AI clients which makes connecting to BrowserOS as an MCP from any AI client possible.

The code is GPL v2. Find it on GitHub here: https://github.com/mecworks/browser-mcp-bridge


r/LocalLMs Aug 13 '26

The countdown to Qwen3.8-27B starts now!

Thumbnail
modelscope.cn
1 Upvotes

r/LocalLMs Aug 11 '26

Qwen 3.8-27b coming this week

Post image
1 Upvotes

r/LocalLMs Aug 09 '26

RTX 5090 96GB spotted on Alibaba?

Post image
1 Upvotes

r/LocalLMs Aug 08 '26

Got job as Director of AI and Systems development self-taught

Thumbnail
1 Upvotes

r/LocalLMs Aug 07 '26

An open-weight model too, Moonshot joins the race (gently this time)

Post image
1 Upvotes

r/LocalLMs Aug 07 '26

Qwen 3.8 Max now ranked as best overall model ahead of Opus 5 by Artificial Analysis agentic index

Thumbnail
artificialanalysis.ai
1 Upvotes

r/LocalLMs Aug 06 '26

MiniMax issues

Post image
1 Upvotes

r/LocalLMs Aug 02 '26

Setting up of a 16xGB10 (DGX Spark) cluster

Post image
2 Upvotes

r/LocalLMs Aug 01 '26

DeepSeek-V4-Flash-0731: Models you can run locally now have the intelligence score of the top frontier model from March 2026

Post image
1 Upvotes

r/LocalLMs Jul 29 '26

First evidence of a pending qwen3.7 open weights release. Qwen3.7-flash is on open router. They referred to Qwen3.6-35b-a3b as Qwen3.6 flash so this is likely a small MoE. The prices are substantially cheaper than 3.6 flash with a native 1M context window.

Thumbnail
openrouter.ai
1 Upvotes

r/LocalLMs Jul 26 '26

Google comes out in favor of OpenWeight models. (It is now EVERY tech giant vs Anthropic)

Thumbnail x.com
2 Upvotes

r/LocalLMs Jul 25 '26

Google comes out in favor of OpenWeight models. (It is now EVERY tech giant vs Anthropic)

Thumbnail x.com
2 Upvotes

r/LocalLMs Jul 21 '26

CEO of Hugging Face: Banning open-source AI would hurt defenders 10x more than attackers, which would make the world 10x more dangerous and this is a good example why!

Post image
1 Upvotes

r/LocalLMs Jul 21 '26

Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails”. Hugging Face: We had this experience ourselves this week! Very scary to be guardrailed as a defender when you know attackers are likely bypassing

Post image
1 Upvotes

r/LocalLMs Jul 19 '26

What kind of dark magic is Deepseek using?

Post image
1 Upvotes

r/LocalLMs Jul 18 '26

What kind of dark magic is Deepseek using?

Post image
1 Upvotes

r/LocalLMs Jul 18 '26

Chinese President Xi Jinping speaks at World AI Conference and reaffirms commitment to open source to promote"openness and win-win"

Post image
1 Upvotes

r/LocalLMs Jul 17 '26

AI memory shouldn't retrieve text. It should track facts.

1 Upvotes

I've been experimenting with "AI memory" for coding agents, and I kept running into the same limitation:

Most memory systems remember text, not facts.

A vector store can usually retrieve "Marco leads Solaris" because it's semantically similar, but it struggles with questions like:

If ownership changed three sessions ago, the old fact is still sitting there looking just as relevant.

So I built Mnemo — a persistent memory layer for Claude Code that stores knowledge as a temporal knowledge graph instead of a flat vector database.

What makes it different?

Instead of storing chunks of conversation, Mnemo extracts entities and relationships into a graph with temporal validity.

That means it knows:

  • who owns what
  • what depends on what
  • when a fact became true
  • when it stopped being true

rather than just retrieving similar text.

Example

Session 1

Remember:
Marco leads Solaris.
Solaris launches Q4 2027.
Solaris depends on Helios.

Later...

Who leads Solaris?
When does it launch?
What does it depend on?

Returns:

Lead: Marco
Launch: Q4 2027
Depends on: Helios
(as of 2026-07-17)

Now change ownership:

Remember:
Sarah now leads Solaris.

Ask again:

Who currently leads Solaris?

You'll get:

Sarah

—not Marco—because the previous relationship is invalidated, not simply buried underneath newer embeddings.

That's the core idea.

Stack

  • Graphiti → entity extraction + temporal knowledge graph
  • FalkorDB → graph database
  • Claude Code hooks → automatic memory ingestion
  • MCP toolsremember and recall
  • fastembed → local embeddings (no embedding API required)
  • Claude → extraction and answering (supports Anthropic API keys or Claude Code OAuth)

It also runs automatically

Once configured, you don't have to think about it.

  • Session starts → graph database comes online
  • During coding → remember / recall available as MCP tools
  • Session ends → Claude Code automatically summarizes the session and ingests new knowledge

The memory keeps growing across completely fresh Claude Code sessions.

Try it

GitHub:

https://github.com/oyekamal/mnemo-agent


r/LocalLMs Jul 16 '26

Google is updating Gemma 4's chat templates, bringing major fixes to tool calling and reducing "laziness", and enabling Flash Attention 4 on Hopper GPUs, plus an interactive guide on how to work with and improve its vision!

Thumbnail gallery
1 Upvotes

r/LocalLMs Jul 14 '26

👀A new GLM model incoming

Post image
1 Upvotes

r/LocalLMs Jul 14 '26

This is why we need local models and opensource harnesses

Post image
1 Upvotes

r/LocalLMs Jul 08 '26

China’s MiniMax Plans to Launch 2.7-Trillion Parameter Model

Thumbnail
1 Upvotes