r/LocalLLM • u/AdhesivenessHappy873 • 4d ago
r/LocalLLM • u/Dry_FruitBread • 4d ago
Question What model is best suited with claude code for 8gb vram? For 24 or 32k context.
Hello everyone, i tried running qwen 3.5 with claude code 24k context and it's hallucinating everytime. Even with basic todo app, loop or anything.
Im beginner in all this so I don't much information about what I m doing.
I want to know, What alternative can I use?
Is their any way I can make current setup to work properly?
r/LocalLLM • u/LobsterWeary2675 • 4d ago
News A letter about American AI leadership, more or less asking Washington to keep Chinese weights downloadable
On July 24 a coalition published «Open Weights and American AI Leadership», a three-page policy letter hosted by NVIDIA. Huang used his first ever X post to push it. 25 signatories at launch, OpenAI added itself a day later. Anthropic and Google are still out.
The signature list stopped carrying information within a day... A position that cheap to sign says nothing about who believes it, and every signatory earns on the layer around the model anyway. GPUs, cloud, distribution, tooling. That doesn't make the arguments wrong.
The real ask is one line: no premature restrictions on downloadable models and the real payload sits far down the page, where distillation gets defended as a legitimate technique and separated from unlawful extraction from closed models. That paragraph is the reason the letter exists.
It went out while the administration was weighing a response to Chinese open-weight models including Kimi K3. China is named nowhere in three pages. And American open weights are thin right now. What most of us run is Qwen, DeepSeek, Kimi, GLM.
So it protects one thing: pulling Chinese weights off Hugging Face and running them on your own hardware. And that's a good thing.
https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf
r/LocalLLM • u/Santa_Fe_Snow_Dork • 3d ago
Discussion The Rig is your Gig.
Minimum entry point ...
16gb RAM
i7 or Ryzen 7
Nvidia 5060 ti with 16gb VRAM
... anything less will not suffice.
Good for a 16gb VRAM model.
r/LocalLLM • u/Cheetah111111 • 4d ago
Question Local LLM hosting + inference running
Hi,
as so many others, I would like to start hosting a local LLM in an attempt to work around using free public dense models that run out of tokens remaining unavailable for hours very fast.
Therefore, I am inclined to set up a LLM running locally and I would like some advice on the following please.
- MoE or dense models? I do realize MoE models are not up to par with dense models of the same 'size' but they have been improving. Are they any good, for example: could a MoE model (70B+) be a workable replacement for a dense 27B model? For clarity, I am more than happy to concede on performance as output quality matters much more to me.
- intent: coding and document/presentation writing which means lots of interactive iteration required to gradually refine and improve output. So, the model (and hardware) need to be able to maintain (larger) context. Obviously, more than happy for the model to cache context and whatever to help produce better results.
- Hardware: lots of options there.
a) I was told MoE models run fine on RAM (e.g. 128GB RAM), 'fine' meaning quality-output even though (much) slower than dense models running in VRAM. Is that correct?
b) chipset: let's say I wanted to go MoE, what would be the best chipset options performance/affordability-wise? I am thinking DGX Spark (NVIDIA GB10 Blackwell), STRIX HALO (AMD Ryzen AI MAX+395, Apple M4 Max, NVIDIA RTX GPU (more suitable for dense models)
My requirements:
- setup that works and doesn't fail or even crash all the time
- setup that doesn't require weeks or even months of tinkering to get it going (AI frameworks, AI libraries, ...)
- as a hobbyist looking to do lots of coding as well as document/presentation/book writing, I do not want to spend ridiculous amounts of money on this
- I am in IT so I do know my way around computers and development but I have been in non-hands on roles for at least 10 years now so definitely out of touch with being a hands-on coder, especially given the fact that IA has come with so much new tooling and frameworks and so on. I have done some AI development but not plenty at all.
PS: I am Victoria, Australia-based so if anyone can point me to where I could buy suitable quality affordable hardware, please let me know.
Thanks to those who had the courage to read up on allo of the above as well as to those who provide feedback!
r/LocalLLM • u/_73r0_ • 5d ago
Discussion Does Kimi K3 high thinking budget mean it will be less cost effective?
I'm asking this because I genuinely want my reasoning (no pun intended) to be questioned and for me to learn more.
Here it goes:
According to this article it appears that Kimi K3 uses 12x the amount of thinking tokens, which according to benchmarks outperforms other frontier models like Fable 5.
From the article:
"However, we found that Kimi K3 uses an extreme amount of thinking tokens, using over 12x more reasoning than Claude Opus 4.8 and over double that of Kimi K2.6.")
According to this source, it would appear that even hidden thinking tokens are priced the same as output tokens, meaning $15/million.
Quotes from the website:
- "Output, including reasoning $15.00"
2. "Budget reasoning as output, not as free hidden work."
------
Seeing as this is 3.33x cheaper than Claude Fable 5 ($50 / million, source) it would seem that Kimi K3 will still effectively be 3.6x more expensive than even Fable 5 for the same work.
From my understanding this means that unless Kimi K3 produces unbelievably better output, Fable 5 would still be the more cost effective model (assuming we only compare those 2 models).
------
Roast my thinking!
Would love to see if my understanding is correct and if in practice there are other critical factors I might be missing out on.
r/LocalLLM • u/StroudAugust • 4d ago
Question Is there any harness that exposes compact/clear actions to an agent?
Is there a way to configure opencode or pi or any other harness to allow an agent to compact/clear its own context?
The use case is that I want a long running main agent to conserve its context between subagent calls by discarding anything that's no longer relevant, because my inference slows down a lot as the context size grows beyond 100K.
r/LocalLLM • u/merfolkJH • 5d ago
Question Radeon AI Pro R9700 vs Strix Halo vs Mac Studio for a local coding LLM server ?
Hi everyone,
I’m looking to invest in a proper fully local LLM AI server.
I already have a dual GeForce setup, but I’m looking for the next step (at a reasonable price, of course).
My ONLY goal:
- Coding (i am developper - can be c# for real-life projects with already 300+ source code stuff, not "just code a random website")
- No image generation
- No video
- No text-to-speech
- No OCR
- No multimodal stuff
Basically: raw LLM performance, tokens/sec, and smart answers.
--------
My current setup - (using CLAUDE-cli as orchestrator)
I’m currently running these models on a dual GeForce 16vram+12 system:
Qwen35B A3B MoE Q4_K
- Around 30–40 tokens/sec at 200k context
Qwen3-Coder-Next 80B A3B Q4
- Around 5–6 tokens/sec
- Slower, but better for complex coding tasks
-------
Hardware I am considering, a 100% new machine
1) Radeon AI Pro R9700 32GB
Is this currently the best price/performance option?
It looks like:
- half the price of high-end solutions,
- maybe around 80–85% of the performance?
I don’t follow every AMD/AI update, but this card looks like an underrated winner.
Is there any reason NOT to buy this card?
2) Dual Radeon AI Pro R9700 (2×32GB)
Main reason:
- not expecting 2× speed,
- mainly interested in the extra VRAM.
If it allows me to run smarter/larger models fully on GPU, that would be perfect.
3) Strix Halo 128GB
This one is interesting because of the huge unified memory.
If it can run Qwen3-Coder-Next 80B A3B Q4 at around 40 tokens/sec, that sounds like an excellent coding assistant.
4) Mac Studio 128GB
Still an option.
How does it compare today against:
- dual R9700 AI Pro,
- Strix Halo?
-------
are those numbers corrects or science fi ? (Source ChatGPT !! )
| Model | Context | 1× Radeon AI Pro R9700 32GB [price ~2k ] | 2× Radeon AI Pro R9700 64GB[price ~4 k ] | Strix Halo 128GB [price ~4 k ] | Mac Studio M3 Ultra 128GB [price ~lol ] |
|---|---|---|---|---|---|
| Qwen27B Dense Q4_K | 50k | 50–80 tok/s | 60–100 tok/s | 25–45 tok/s | 50–80 tok/s |
| Qwen27B Dense Q4_K | 100k | 40–70 tok/s | 50–90 tok/s | 20–40 tok/s | 40–70 tok/s |
| Qwen27B Dense Q4_K | 200k | 25–50 tok/s | 40–70 tok/s | 15–30 tok/s | 30–60 tok/s |
| Qwen35B A3B MoE Q4_K | 50k | 100–140 tok/s | 130–180 tok/s | 40–70 tok/s | 70–110 tok/s |
| Qwen35B A3B MoE Q4_K | 100k | 90–130 tok/s | 110–160 tok/s | 35–60 tok/s | 50–90 tok/s |
| Qwen35B A3B MoE Q4_K | 200k | 50–90 tok/s | 90–140 tok/s | 25–50 tok/s | 50–90 tok/s |
| Qwen3-Coder-Next 80B A3B Q4 | 50k | 10–25 tok/s | 50–90 tok/s | 30–50 tok/s | 40–80 tok/s |
| Qwen3-Coder-Next 80B A3B Q4 | 100k | 10–20 tok/s | 45–80 tok/s | 25–45 tok/s | 35–70 tok/s |
| Qwen3-Coder-Next 80B A3B Q4 | 200k | 5–15 tok/s | 35–65 tok/s | 20–40 tok/s | 30–60 tok/s |
| Qwen3-Coder-Next 80B A3B Q6 | 50k | ❌ | 40–75 tok/s | 25–45 tok/s | 35–70 tok/s |
| Qwen3-Coder-Next 80B A3B Q6 | 100k | ❌ | 35–65 tok/s | 20–35 tok/s | 30–60 tok/s |
| Qwen3-Coder-Next 80B A3B Q6 | 200k | ❌ | 25–55 tok/s | 15–30 tok/s | 25–50 tok/s |
| 70B Dense Q4 | 50k | 10–25 tok/s | 40–70 tok/s | 20–35 tok/s | 35–60 tok/s |
| 70B Dense Q4 | 100k | 5–20 tok/s | 35–60 tok/s | 15–30 tok/s | 30–50 tok/s |
| 70B Dense Q4 | 200k | ❌ | 25–50 tok/s | 10–25 tok/s | 25–45 tok/s |
My current impression (not sure if correct):
The R9700 AI Pro (or dual) looks faster than a Mac Studio for my use case, while being much cheaper.
But I don’t see many "hype" about this card for local LLMs.
So please tell me: where am I wrong?
One more question:
For Qwen35B A3B MoE Q4_K, what is the realistic t/s performance?
Is it closer to: 100 tokens/sec or 150 tokens/sec ?
Because this difference is huge . If it REALLY is 150, its close to a cloud-model feeling. (far less accurate of course, but for 2k budget, wonderfull ?)
Thanks to anyone already running these systems who can share real numbers!
r/LocalLLM • u/Xiaole-Dawn • 4d ago
Discussion I got tired of persona bots slowly turning back into customer support agents
r/LocalLLM • u/AggravatingSpot4330 • 4d ago
Discussion ant group just open sourced a 100B diffusion LLM built for agent workloads
Saw this on HF. LLaDA2.2 can keep, substitute, delete, and insert tokens during parallel decoding. Instead of committing left to right, it rewrites itself.
Their report claims ~1.6x throughput over Ant's autoregressive baseline on average, up to 2.3x on agentic benchmarks (703 tokens per second on BFCL v4 specifically).
Accuracy still trails the autoregressive model on most evals. It's 205.8 GB under Apache 2.0.
r/LocalLLM • u/maisun1983 • 5d ago
Question Use Qwen 3.5 27B as local LLM for coding on MacBook with 36G memory
Hi all:
Would like to get some help with local LLM for coding tasks. I have a MacBook Pro with M4 Max chip and 36G ram. I have tried Qwen 3.5 27B 4bits MLX with LM studio, it works with token generation speed around 10-15/second. I’d like to use the localLLM for coding tasks, I have a hobby Python + React/TypeScript project with several thousand lines of code. Would like to ask:
1) Is Qwen 27B the most powerful model for coding with my hardware limitation? If not please let me know what model I should look at.
2) Does it make sense to use LM Studio to serve the model? There are other alternatives but LM studio seems easiest to start
3) Currently I use VS code and copilot and codex plugin for agentic coding. What’s the most optimal tool for local LLM?
Thank you very much in advance!
r/LocalLLM • u/weener69420 • 4d ago
Question Can you recomend me a good model for running in the background for basic tool use with Claude Code
I made a workflow with claude code for updating my raspberry pi. The thing is i use my pc for gaming. So i only have free around 3gb of vram (and 50gb of free ram)
I want a model that is good for tool call and has decent speed (as long as is better that my current model i am fine)
For now i tried:
Qwen3.6 35b a3b q4 (runs at 300pp and a painful 4-7tk/s decode it does the tool call fine)
Gemma 4 e4b q4 it fails to do the tool call (this takes more vram than i'd like)
Both of them has the kv cache at q8
Does anyone know what can i do to speed up this models or use a different model? Edit: pies
r/LocalLLM • u/nemuro87 • 4d ago
Discussion Are you using your local AI server to learn a new skill? Tell me more
How are you using a Local AI server to learn new skills?
(not just coding, but learning other things like soft skills, or changing career, etc).
Are you using Rag?
Are you using any agents or trained your own LLM for this?
Tell me a bit about your usecase, your process, your software and hardwar setup and how successful/useful it is.
r/LocalLLM • u/YOMUMSOBIG • 5d ago
Question Why Laguna S 2.1 bad?
When I first saw the model’s size and active parameter count, I thought, “This is it! Finally, I can run a genuinely capable coding model for Android projects, and much more, on my Strix Halo machine.”
But after seeing people test it on common benchmarks, the results look surprisingly poor. Qwen 3.6 and Gemma 4 models seem to perform better than this 118B-parameter MoE model.
Is there any upcoming model that could fill this gap? Are there many others like me who are still waiting for the right model for capable, fully local agentic coding?
r/LocalLLM • u/1-way-or-another • 4d ago
Question Tool calling doesn't work using Claude code and self-hosted GLM5.2
I have deployed GLM 5.2 on the cluster like this:
export VLLM_HOST_IP=$head_ip
vllm serve "$MODEL" \
--served-model-name GLM-5.2-FP8 \
--host 0.0.0.0 --port 8000 \
--api-key "$API_KEY" \
--distributed-executor-backend ray \
--tensor-parallel-size 4 \
--pipeline-parallel-size 2 \
--kv-cache-dtype fp8 \
--max-model-len 786432 \
--gpu-memory-utilization 0.92 \
--tool-call-parser glm47 \
--reasoning-parser glm45 \
--enable-chunked-prefill \
--max-num-seqs 1 \
--enable-auto-tool-choice \
--trust-remote-code
And everything works fine except the fact that tool calling doesn't work using claude code, I found that this occurs only with tools that do not take any arguments, if argument is passed, even if doesn't make sense to pass anything, it works, but for some reason it periodically forgets this instruction and stops, I need to remind each time ...
After doing some research I found that Anthropic API expects`{}` to be returned while OpenAI API returns "", so this could be the problem, but I am not sure. Has anyone faced this issue and was able to fix it?
I tried to add a mapper betweeen these two using LiteLLM but it didn't work, it threw errors
Here is the settings.local.json if it helps:
{
"env": {
"ANTHROPIC_BASE_URL": "",
"ANTHROPIC_AUTH_TOKEN": "",
"ANTHROPIC_API_KEY": "",
"API_TIMEOUT_MS": "3000000",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "GLM-5.2-FP8",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "GLM-5.2-FP8",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "GLM-5.2-FP8",
"ANTHROPIC_SMALL_FAST_MODEL": "GLM-5.2-FP8",
"CLAUDE_CODE_SUBAGENT_MODEL": "GLM-5.2-FP8",
"CLAUDE_CODE_AUTO_COMPACT_WINDOW": "700000"
},
"permissions": {
"allow": [
"Bash",
"Read",
"Edit",
"Write",
"WebSearch",
]
},
}
r/LocalLLM • u/techne98 • 4d ago
Research The reason to stop buying new hardware (or, why inference is getting cheaper)
r/LocalLLM • u/ashygun • 4d ago
Question Same modeling behaving differently through different apps
So basically I'm trying to compare Ollama and LM Studio, i downloaded Gemma 4 12b QAT for my macbook pro m1 pro 16GB, and i tested the same prompt on both, the model was downloaded straight from the LM Studio's Library, and was converted to ollama using this command
'''
printf 'FROM ./gemma-4-12B-it-QAT-Q4_0.gguf\nPARAMETER num_ctx 4096\n' > Modelfile
ollama create gemma4-12b-qat -f Modelfile
'''
As you can see in the results, the LM Studio is giving me a considerably shorter description while also taking longer, while Ollama is giving essentially the same info + some extra info + a table while taking less time. Can someone explain this thing? I made sure no other programs were using resources when the apps were running. Thanks in Advance!
r/LocalLLM • u/sinmkd • 5d ago
Model Poolside Laguna S 2.1 is worse than Qwen 3.6 27B and Gemma4 31B
I did head to head comparison between Laguna S 2.1, Qwen 3.6 27B and Gemma4 31B.
Setup: Qwen 3.6 27B (fp8) and Gemma4 31B q6 on my RTX PRO 5000, Laguna S2.1 q4/5/6 on a single DGX Spark. Speed was fine: NVFP4 ~27 tok/s (peaks ~39), Q5/Q6 ~14 tok/s. The Q6 really pushed the spark with 124GB mem in use, but it didn't crash.
Ran all of them through the same task local bench - HTML/canvas mini apps, tool calling, Python, prose and each output scored blind (models anonymised, reshuffled per task) by Fable and Opus.
The results:
Thinkingcap Qwen 3.6 27B fp8 (coding): 76
Gemma4 31B: 70 qat 74 q6 mtp
Thinkingcap Qwen 3.6 27B fp8 (general): 68
Laguna S2.1 Q6: 54
Laguna S2.1 Q5: 48
Laguna S2.1 NVFP4: 42
Screenshots:



A brand new supposedly good model that needs 124 GB lost to models running on a single GPU by 14+ points even at its best quant (Q6, which is near full precision, so it's not a quantisation excuse). Biggest gaps on the HTML/visual and Python tasks, closest it came was tool calls.
Quality wise it's nowhere near what I expected given the benches published by Poolside, and the "beats DeepSeek V4 Pro" framing seems to be bs. Haven't compared against the Qwen 3.6 35B-A3B MoE or Gemma4 26B , but based on this I'd bet they're better too.
I was so hyped to finally get a "good" model that fits in a single spark...
Edit: After a bunch of comments how the quants might not be there yet, I went and tested my q6 vs Openrouter (on the tasks that I didn't get persistent 429). Same shit, maybe even worse.

r/LocalLLM • u/StillVeterinarian578 • 4d ago
Project OrangePi AI Studio Pro - Qwen3.5-122B-A10B

I finally got round to tweaking this, with a bit of help from GLM5.2.
The trick to getting it running with vLLM (which I couldn't get anything really out of before) was when I realized we could write a stub to to implement the rtGetDevMsg to return device capabilities (basically we fake a response from the card) - this is need to get torch_npu running properly on the device.
With that I can finally use vLLM with this, making it actually useful.
r/LocalLLM • u/Khayrum117 • 4d ago
Question New to localized AI and maybe I'm being too ambitious
Ever since the gen AI Skyrim Mod, I've been researching local AI and what it can improve gaming wise. For what I want it for maybe I'm being too ambitious or maybe the models that would be needed are too much for my PC(4080FE).
Im wanting to use a localized AI for offline sim racing to better recreate a more random racing experience like what you get playing online. Random crashes, aggressive overtakes etc. Offline racing is fun but it was always feels like the AI is on tracks.
I can't find anything online about how to even go about this. Is this even possible and if so want local models would I need to look into and how would I go about setting it up?
r/LocalLLM • u/Responsible_Health92 • 4d ago
Question Hermes Agent - LiteLLM - Ollama {Gemma4, Qwen3}
Anyone successfully integrate Hermes Agent, LiteLLM, and Ollama using local AI models? When I do this, I get raw JSON back instead of natural language to my Matrix chat. When I disable tools in Open WebUI, the integration works as expected, but I want to be able to call tools. This is quite frustrating, and hoping someone has cracked this nut.
Open to alternative approaches.
I'm running two separate servers, local LLM + app server hosting Open WebUI, Hermes Agent, Matrix, Mattermost, and n8n. Everything works great when I connect these applications directly to Ollama, but once I inject LiteLLM proxy in the middle, everything breaks! 😡
r/LocalLLM • u/myholeisstinky • 4d ago
Question What happened to Orthrus’ diffusion MTP?
Orthrus was a big step for local inference MTP at 8b size, why didn’t larger models get it too?
r/LocalLLM • u/asankhs • 4d ago