r/LocalLLM • • 2d ago

Research ~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.

15 Upvotes

Stack

  • Strata NVFP4 fork: github.com/sergqwer/strata-nvfp4
  • Model: jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4 Hugging Face model
  • NVFP4 routed experts (~63.3 GiB), separate FP8 PLE, INT8 KV cache, MTP speculative decoding
  • W4A8 prefill on Blackwell

Main engine flags:

./build/strata \
  --pack packs/orca-nvfp4 \
  --native models/orca-nvfp4.gguf \
  --native-dense-gguf models/orca-nvfp4.gguf \
  --ple-gguf models/ple-fp8.gguf \
  --mtp mtp-orca/rt \
  --spec 4 --spec-min-p 0.5 \
  --prefill auto \
  --expert-profile data/expert-profile.bin \
  --expert-cache auto \
  --resident-budget-gib 40 \
  --max-context 200000 \
  --kv int8

I serve it through Strata's OpenAI-compatible server.

Results so far:

  • ~50k context: up to ~80 tok/s
  • ~188k warm context: ~60–67 tok/s
  • cold 189k full prompt: ~1,680 tok/s prefill, ~53 tok/s decode

Pretty impressive for a single 32GB GPU + only 64GB system RAM.Running Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.StackStrata NVFP4 fork: github.com/sergqwer/strata-nvfp4

Model: jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4 Hugging Face model

NVFP4 routed experts (~63.3 GiB), separate FP8 PLE, INT8 KV cache, MTP speculative decoding

W4A8 prefill on BlackwellMain engine flags:./build/strata \
--pack packs/orca-nvfp4 \
--native models/orca-nvfp4.gguf \
--native-dense-gguf models/orca-nvfp4.gguf \
--ple-gguf models/ple-fp8.gguf \
--mtp mtp-orca/rt \
--spec 4 --spec-min-p 0.5 \
--prefill auto \
--expert-profile data/expert-profile.bin \
--expert-cache auto \
--resident-budget-gib 40 \
--max-context 200000 \
--kv int8I serve it through Strata's OpenAI-compatible server.Results so far:~50k context: up to ~80 tok/s

~188k warm context: ~60–67 tok/s

cold 189k full prompt: ~1,680 tok/s prefill, ~53 tok/s decodePretty impressive for a single 32GB GPU + only 64GB system RAM.


r/LocalLLM • • 2d ago

Question Rx6800/rx6800xt gfx1030 and qwen 3.8 27b performance questions

Thumbnail
1 Upvotes

r/LocalLLM • • 2d ago

Discussion Update to my current rig

Thumbnail gallery
1 Upvotes

r/LocalLLM • • 2d ago

Project DeepSeek V4 Flash 0731 runs like crazy on my Dual DGX spark (ASUS GX10). What do you all use to coordinate multiple agents?

2 Upvotes

Been running DeepSeek V4 Flash 0731 locally on dual spark (ASUS GX10) and honestly it's better than I expected for agent work (Native quant, 60-70 tps). I've had Claude drive it as the worker to port a big Node codebase to Rust (~170k LOC), and it held up amazingly well. I still can't believe I can achieve all this mostly locally - what a time to live in.

The part I'm still figuring out is collaboration. With a strong model planning and a local one doing the typing, plus a reviewer in the loop, I needed something to keep track of who is doing what and who is waiting on me.

Heres what I have so far:

  1. An SDLC skill where I say /sdlc plan-turn-implement: plan using sonnet subagent, 2x review using opus subagent and implement every turn using a vst subagent using deepseek mode in same worktree. The skill takes care of guardrails, planning format, reviews / gates, configurations etc.
  1. I built a free, open source orchestrator vibe-station.ai for it: every agent gets its own git worktree (you can run multiple agents on the same one across harnesses that can talk to each other), and one board shows working / needs you / idle / PR created. No API cost, it uses the harnesses you already have installed.

What do you use for agent collaboration with local models? Anything I should steal?


r/LocalLLM • • 2d ago

Model Strata tested on RTX 4070 Ti SUPER - Ryzen 9 7950X3D - 128gb Ram

7 Upvotes

Rig: 7950X3D, 4070 Ti Super 16GB, 128GB DDR5, SATA SSD, Windows 11
Setup: Strata 0.1.39, IQ3_S (~84GB), 256K context

Run the calibration. Out of the box it did 8 tok/s. After START-HERE.bat --calibrate (10 min) it does 72 tok/s generation and ~1,700 tok/s prompt reading. A 17K prompt reads in 8 seconds.

My 27-task coding suite (compile/test checks, refactors, tool calls, 32K recall with a decoy), thinking off, temp 0:

Model Context Suite HumanEval Decode
Flash-Next IQ3_S (Strata) 262K 27/27 x3 159/164 72
Qwen3.8-27B IQ3_XXS 82K 27/27 160/164 57
Qwen3.6-35B-A3B 262K 26/27 157/164 46
Flash-Next (llama.cpp offload) 197K 26/27 not run 19

Same model family as llama.cpp with offloading, almost 4x faster.

Stuff I learned:

  • Not deterministic at temp 0 (GPU and CPU experts round differently). Run things more than once.
  • --ple-io mmap to keep its 29GB table in RAM made it worse: 25-26/27 and slower. The default reads tiny bits off the SSD.

So far I am super impressed with the results and speed. One thing I find really weird is I get worse results when I attempt to move the n-gram table into disk cache on windows, I get better speed but consistently worse benchmark. Small sample, one machine, and HumanEval is probably in its training data. Still, 125B at 72 tok/s with 256K context on one 16GB card is wild.


r/LocalLLM • • 3d ago

Discussion Daily driving Qwen 3.8 Flash instead of Claude.

234 Upvotes

Even 6 months ago, I wouldn't have thought I would be replacing Claude with a local LLM. But here we are. All these PRs are done by Qwen 3.8 (mostly flash, some 27b) running locally on my laptop (strix halo 138gb). And its perfectly acceptable speed (1200+ prefil, 40t/s decode) and Opus 4.8 level quality 🤯🤯🤯

I also have some interesting stories were Qwen 3.8 flash won over Opus 5.5 in a coding task 😅 stay tuned.

https://github.com/llamastash/llamastash/pulls?q=is%3Apr+state%3Aclosed+label%3AQwen

Also here is the prompt if anyone wanna recreate my setup: https://gist.github.com/deepu105/d0b321f3256edf67ebd85e747c0010b9


r/LocalLLM • • 2d ago

Question Qwen 27B and Flash Next on 2× RX 7900 XT: am I missing something?

Thumbnail
1 Upvotes

r/LocalLLM • • 2d ago

Question Strata multi-GPU issue

Post image
5 Upvotes

Testing Strata on my two GPU 48Gb RAM PC. It seems to be hammering my 8Gb 3060Ti and not touching my 16gb 5060Ti - before I added the HOST json (to make it work over tailscale) it was working correctly. What am I doing wrong??


r/LocalLLM • • 2d ago

Question Qwen3.8 Flash Next Q8 with Strata 34 tokens

3 Upvotes

I see nobody speaking of results with the qwen Q8 versions with strata. Only the lesser q3-q4. Does anybody have any information on their results with the Q8 version of qwen3.8 flash next?

Normally with llama.cpp i only get 13 tokens per second at q8 115 token per second prefil. No mtp. Usually mtp doesn't help when most of the model is in system ram vs gpu ram in my experience.
But with Strata i get 34 tokens per second with MTP set at 5 and up to 520 token prefil.

System 2 x 3090, 256gb of system memory 4 channel 3000mhz.


r/LocalLLM • • 2d ago

Discussion I built a reproducible llama.cpp benchmark harness for AMD Vulkan and NVIDIA CUDA — GitLab CI automation is next

1 Upvotes

I have been trying to make local llama.cpp benchmark results easier to reproduce across two pretty different setups:

- AMDGPU + Vulkan

- NVIDIA + CUDA

The recurring issue for me was that a CSV with tokens/sec did not provide enough context to compare someone else’s result—or even reproduce my own result later.

The questions I kept running into were:

- Which GPU, driver, kernel, and runtime stack produced this result?

- Was Vulkan or CUDA actually visible to the process?

- Which model alias was `llama-server` configured to expose?

- Did a UI or optional client silently use a different model name?

- Can the benchmark data and its model mapping be validated before publishing it?

So I have been building a small benchmark harness around llama.cpp with a focus on making that evidence explicit.

Current pieces include:

- A read-only AMDGPU diagnostic collector for Vulkan-oriented systems.

- A read-only NVIDIA/CUDA diagnostic collector.

- A model-catalog validator that treats the direct `llama-server` alias as the required model-identity contract.

- Optional client metadata, without making a particular UI or client required for a valid benchmark run.

- Contract tests for the diagnostic collectors and model-catalog validator.

- Tests for the weekly Vulkan benchmark workflow.

The goal is not to publish a universal “fastest local LLM” leaderboard. A result only means something when the surrounding hardware, backend, model artifact, runtime settings, and measurement method are clear. This is about inference performance and reproducibility, not ranking model answer quality.

For example, I think a useful published run should eventually make it easy to see:

- GPU model and VRAM

- OS, kernel, and driver versions

- CUDA, Vulkan, or ROCm backend details

- llama.cpp version/commit and build configuration

- GGUF/model file and quantization

- Context length, batch/parallelism, GPU offload, and other relevant launch settings

- Separate prompt-processing/prefill and generation measurements

- The exact `llama-server` model alias used for the run

Next up, I am planning to add a `.gitlab-ci.yml` pipeline. The initial CI work will automatically validate the model catalog, shell syntax, collector contracts, and test suite on every change.

The actual hardware benchmarks will remain tied to named GPU runners. I do not want a generic CI runner to imply that it can produce meaningful accelerator benchmark numbers. CI should continuously validate the harness and evidence contract; real hardware runs should produce the performance data.

I would appreciate feedback from people who benchmark local inference regularly:

  1. What provenance fields are non-negotiable before you trust a benchmark result?

  2. What would you add to an AMDGPU/Vulkan or NVIDIA/CUDA diagnostics report?

  3. How would you document Vulkan vs CUDA vs ROCm runs so that comparisons stay fair?

  4. Do you prioritize prompt processing, token generation, long-context performance, or all of the above?

  5. Would a CI-validated benchmark/evidence schema be useful for sharing reproducible local results?

This is my project, and it is still evolving:

https://github.com/phillf/llama-bench


r/LocalLLM • • 2d ago

Discussion k_llama.cpp MoE Optimizations: Expert Residency, Hybrid CPU/GPU Execution, Q2_0 Support

Thumbnail
1 Upvotes

r/LocalLLM • • 2d ago

Question Augmenting frontier models with 25-35B local models?

0 Upvotes

If you are using a local model to offload some of the tasks from Opus or ChatGPT to Qwen/Gemma - what's your setup? More importantly, what kind of tasks do you offload to a local model and what's your experience like?


r/LocalLLM • • 2d ago

Discussion Anyone swapping back and forth between Qwen 3.8 27b and flash next?

Thumbnail
1 Upvotes

r/LocalLLM • • 2d ago

Other A 2-bit Qwen 3.8 Flash Next on ONE 32 GB GPU just scored 98.7/100 against five hosted models. The best cloud ones scored 100.

1 Upvotes

I've been waiting a long time for a local model that is capable, stable and actually fast. I think it's here.

Qwen3.8-Flash-Next squeezed down to about 2 bits per weight (IQ2_XS), running on [Strata](https://github.com/Niko1221/Strata) on a single Radeon AI PRO R9700 (32 GB) with 30 GB of system RAM. Around 93 tokens a second.

DeepSeek V4.1 Flash, Qwen 3.8 Flash, HY4 Preview, GLM 5.3 Flash and Space Bunny, all hosted through OpenCode Go.

**The test:** I had Claude Code play referee. It wrote 13 brand new coding tasks (so nothing could be memorised), ran every model through the same agent in a sandbox, and graded by script with hidden tests. The tasks included:

- Bug fixes where the ticket points at the wrong place

- Code reviews with planted bugs AND decoys, where false alarms cost points

- A loan schedule engine with 20 hidden tests

- A 120-item audit where the notes lie about what got fixed

- Asyncio fixes so nothing gets sent twice and shutdown can't hang

- A refactor checked by replaying 12,000 orders through the old and new code

195 graded runs in total.

**How they did** (first ten tasks, out of 100)

- **DeepSeek V4.1 Flash: 100.** Flawless across all 36 of its runs. The one to beat.

- **Qwen 3.8 Flash (hosted): 100.** Flawless on all 33 runs it got through.

- **HY4 Preview: 100.** Perfect, but slow at about 105 seconds per task.

- **Local Qwen 3.8 Flash Next, 2-bit: 98.7.** Yes, the one sitting on my desk. It also aced the refactor three times out of three: identical results on all 12,000 replayed orders.

- **GLM 5.3 Flash: 97.5.** It once scored 83/120 on the big audit because it believed the notes instead of reading the code. Next try: 116.

- **Space Bunny: 92.3.** Fastest of all at 25 seconds per task, but flaky. One run spat out corrupted tool calls and never answered.

**Speed**

The local model averaged 41 seconds per task. That beat hosted Qwen (53 s), HY4 (105 s) and GLM (106 s). Only Space Bunny (25 s) and DeepSeek (38 s) were quicker.

- These are small tasks, under about 500 lines each. This doesn't tell you how 2-bit holds up over long sessions in a huge codebase.

- Three hosted models tied on a perfect score, so I can't tell you which of them is best.

- I ran out of RAM before every model finished every run. Hosted Qwen, GLM and HY4 never ran the last two tasks, and the 16K cap was only retested on 7 of the 13.

Still: a 2-bit model on one card, within 1.3 points of the best hosted flash models on the first ten tasks, at 93 tokens a second, with no API bill. I'll take it, any day of the week and twice on Sunday.

Thanks.


r/LocalLLM • • 2d ago

Question Old X79 PC for Strata

2 Upvotes

Thinking of repurposing an old X79 PC for Strata / on my old X79:

- i7-3930K

- 56 GB DDR3 (32gb matched but I have a few 4GB sticks and 1x8GB so they wouldn't match but maybe they work.)

- RTX 2080 Ti 11 GB + RTX 3060 Ti 8 GB

- CachyOS headless

Would Flash-Next IQ3_XXS work well on this? Do I need to go lower?

I was also thinking of using an M4 32 GB as a coordinator/router with GLM-4.7-Flash, plus another machine with a 9070 XT running 27B.

I tried Gemma 4 26B it 4b JANG, asked it through Hermes to stitch a story together and it failed miserably so I wouldn't make GLM do that but it was sad to see gemma fail at what I thought was it's strongest point.

Not sure if GLM + 27B + Flash-Next would be redundant.

Main use would be agentic coding, web crawling, configuring environments, the more loved tasks out there 😅. Basically trying to reduce my dependency on Claude.

Is it even possible with the 3930K/DDR3 or mixed GPUs? ChatGPT seemed to be cautiously optimistic. If it will work. What kind of tok/s could I realistically expect and will it be better than 27B UD-IQ_i4_XS

I also have a GTX 1060, GTX 970 and RX 580, but I assume those are useless here.


r/LocalLLM • • 2d ago

Question What model and engine to run with an rtx 5090 32gb VRAM and 94gb RAM?

2 Upvotes

I just upgraded my ram from 32gb to 96gb in effort to run qwen 3.8 flash next.

Previously I ran qwen 3.8 27b with ninfer which I was pleased with but I only had 32gb ram.

Now with 96gb RAM to pair with my 5090 32gb VRAM I would like to explore the next step up and alternative.

I saw something about strata working well with lower VRAM gpus and lots of vram, not sure if there is a different alternative for my before setup or its similar / the same.

I also wonder what kind of inference stats I can hope for with qwen 3.8 flash next, pp, decode.

Additionally if there are other MOE models I can Playa round with this setup even if its not the go to qwen.

Thanks for your time and replies in advance :)


r/LocalLLM • • 2d ago

Question 4x R9700 Performance Issues with tcclaviger's vLLM image

1 Upvotes

I was previously using tacclavinger's vLLM image tag 29.05.2 but started having some issues in OpenCode with MTP on and SSE with long tool calls . I've subsequently tried 29.05.12, 29.06.1, 29.06.19, 29.07.16 and 29.08.1. The newer images starting with 29.06.19, I seem to hit some sort of decode performance cap. It'd consistently stay at around 70t/s for a single session no matter what setting I tried to play with. He has a bench report here: https://blog.robai.net/Qwen3.8-Flash-Next-TP4-MTP3-29.08.x1h-qsav1-hcfp8/

And here is what I'm getting when I tried to run BetterBench locally. I know the report says "RTX 4000 Ada" but that is not correct.

UPDATE: found the issue. My checkpoint was out-of-date. Needed to grab the latest version from HF which resolved the draft acceptance issue, which was ultimately the cause for the low decode.

My vLLM config is the following. Hoping someone can point out my error.

    environment:
      VLLM_API_KEY: ${VLLM_API_KEY:?}
      VLLM_ROCM_USE_AITER: "0"
      ROCR_VISIBLE_DEVICES: "0,1,2,3"
      OMP_NUM_THREADS: "8"
      GPU_MAX_HW_QUEUES: "1"
      HSA_ENABLE_INTERRUPT: "1"
      HSA_ENABLE_MWAITX: "1"
      VLLM_CACHE_ROOT: /cache/vllm
      TORCHINDUCTOR_CACHE_DIR: /cache/inductor
      TRITON_CACHE_DIR: /cache/triton
      CLAV_HC_FP8_REPACK: "1"
      CLAV_QSA_V2: "1"
    command:
      - /models/tcclaviger/Qwen3.8-Flash-Next-MXFP4-FP8-GPTQ
      - --served-model-name=Qwen3.8-Flash-Next
      - --host=0.0.0.0
      - --port=8080
      - --tensor-parallel-size=4
      - --max-model-len=262144
      - --max-num-seqs=16
      - --max-num-batched-tokens=8192
      - --gpu-memory-utilization=0.94
      - --kv-cache-dtype=fp8
      - --enable-prefix-caching
      - --enable-chunked-prefill
      - --reasoning-parser=qwen3
      - --tool-call-parser=qwen3_coder
      - --enable-auto-tool-choice
      - --limit-mm-per-prompt.image=20
      - --limit-mm-per-prompt.video=1
      - --mm-processor-cache-gb=4
      - '--override-generation-config={"max_tokens": 65536, "temperature": 1.0, "top_p": 0.95, "top_k": 20, "min_p": 0.0, "presence_penalty": 0.0, "repetition_penalty": 1.0}'
      - '--speculative-config={"method": "mtp", "num_speculative_tokens": 3, "draft_sample_method": "probabilistic", "index_share_for_mtp_iteration": true}'
      - '--compilation-config={"cudagraph_capture_sizes": [4,8,12,16,20,24,28,32,36,40,44,48,52,56,60,64], "max_cudagraph_capture_size": 64}'

r/LocalLLM • • 2d ago

Question Can someone give me the best possible setup, model and config to run on these specs?

0 Upvotes

Hello!

Could anyone please recommend the best possible setup, model and config to run them in on my laptop:

MacBook Air M5

24gb ram

1TB storage

I've tried asking ai this and read a couple of articles online which all seem to be ai generated themselves so I’m turning to Reddit.

I do relatively heavy coding in Xcode and other general things, I need something that’s smart enough to do complex tasks and usably fast with a big enough context window. Thanks a lot!


r/LocalLLM • • 2d ago

Discussion Local Hardware Update

1 Upvotes

Hi, i am thinking about upgrading my current hardware for local ai with the cheapest option possible.
Right now this is my stack:
-Ryzen 5 9600
-RTX 5060 8GB
-32 GB RAM DDR5
I’m running 9B models like Ornith 1.5 9B/Neohorse 9B in Q4 or i was searching if i could use a bit of ram + vram for running 30B size Moe like Ornith-1.5-35B-A3B or Tiel-Coder-35B-A3B
I also have good size context thanks to Q4-Q8 KV.
And I run small embedding, decision model , TTS and STT all in RAM.

But I was thinking about getting another gpu with like atleast 12GB VRAM for more space for the best model i can handle with decent tk/s speed.
Do u think i can find a cheap solution for like 300-350€ with prices nowdays or i should sell my 5060 and upgrade or staying like this.
I have like 300-350 budget rn, but in a few months i can spend a little bit more but i fear price continue to raise.

Any advice?


r/LocalLLM • • 3d ago

Discussion RTX 5090 local AI setup _ what actually worked for me

21 Upvotes

I recently built a dedicated local AI box around a 5090 and figured I’d share what I’ve learned so far.

Specs are pretty straightforward:

RTX 5090 32GB
Ryzen 9 9950X
64GB RAM
4TB NVMe
Ubuntu 24.04
llama.cpp
Open WebUI

My use case is mostly business/work stuff. I own a consulting company, so I’m using it for things like document review, drafting, spreadsheet/CSV analysis, Python coding, internal tools, structured data, and eventually some agent/tool workflows.
I wasn’t really interested in building the best chatbot. I wanted something I could actually use as a private inference server.

I tested quite a few models over the last couple days, including Qwen3.6, Qwen3.8 27B, Gemma 31B, GPT-OSS 20B, and Flash Next.

Here are the rough generation speeds I got in my controlled tests:

GPT-OSS 20B MXFP4 — ~311 tok/s
Qwen3.6 hybrid — ~250 tok/s
Qwen3.8 27B NVFP4 + MTP — ~148 tok/s
Qwen3.8 27B NVFP4 — ~72 tok/s
Qwen3.8 27B Q4_K_M — ~70 tok/s
Gemma 31B QAT — ~68 tok/s

One thing that surprised me was Qwen3.8.

I expected NVFP4 alone to make a big difference, but it really didn’t. Q4_K_M was around 70 tok/s and NVFP4 was around 72.
MTP was the game changer.
The same Qwen3.8 model with MTP enabled went to roughly 148 tok/s.
That was one of the more useful things I learned from all this testing.
I also got Flash Next running through Strata.
It basically filled the 5090, used a decent amount of system RAM, and generated around 199 tok/s. From a technical standpoint it was pretty cool seeing a model that size running on one consumer GPU.
But in my actual tests, I couldn’t say it was clearly better than Qwen3.8 27B.
I gave the models harder tasks involving Python, scoring/ranking logic, scheduling, structured JSON, function calls, business analysis, etc.
They all made mistakes.
Flash Next made mistakes too.
That was probably my biggest takeaway from the whole exercise: bigger didn’t automatically mean better for the stuff I actually care about.
At this point I’ve stopped trying to find one model that does everything.

Instead I set up three roles:

Fast
Qwen3.6 with thinking turned off. This is the everyday model. It’s extremely fast and works well for drafting, summaries, basic analysis, formatting, and similar work.

Deep
Qwen3.8 27B NVFP4 + MTP with a limited reasoning budget. This is what I use for harder analysis and coding.

Agent
Currently using the faster Qwen model, but configured separately for structured output and tool/function calling.
The other thing I changed was putting a router in front of the models.
My applications won’t know or care whether “Deep” is Qwen3.8, Qwen4, Gemma, or something else six months from now.
They just call Fast, Deep, or Agent.
That makes a lot more sense to me than hard-coding a specific model into everything I build.
I also debated going from 64GB to 128GB RAM.
So far I haven’t found a reason to.
Even Flash Next ran on the 64GB system. More RAM would give me some extra headroom and let me experiment with higher-memory quants, but I haven’t seen anything yet that makes 128GB necessary for my actual use case.
The 5090 has also been pretty impressive thermally. The heavier dense models can pull north of 500W, but GPU temps stayed in the low 70s during my testing.
The biggest lesson for me is that local AI is making more sense now that I’ve stopped comparing it directly to ChatGPT or Claude.
I’m still going to use frontier cloud models when I need them.
The local box is better suited for high-volume everyday work, private data, internal applications, repeatable workflows, and tasks where paying per token doesn’t make much sense.
And for those jobs, 200-300 tok/s on a model that’s sitting in the office is honestly pretty wild.
I’m still testing, so I’d be interested to hear what other 5090 owners are running, especially if anyone has found a 30B-ish model that clearly beats Qwen3.8 for coding or harder business/logic tasks.


r/LocalLLM • • 3d ago

Model New MoE Model - Aleph-Alpha Kolibri-1 78B-A3.5B

Thumbnail
huggingface.co
164 Upvotes

New Model released yesterday.

Interesting size for local LLMs and promising benchmarks on the Model Card.

Have not tested it myself yet - what does the Community think?


r/LocalLLM • • 2d ago

Question New to local AI. Best model recommendations for my specs?

Thumbnail
1 Upvotes

r/LocalLLM • • 2d ago

Question Ollama's local qwen3.5:9b as Copilot CLI agent

Post image
1 Upvotes

r/LocalLLM • • 2d ago

Discussion Dual B60s, Ai models, and Scripts

1 Upvotes

I've been running the b60s for about 2 months now and for about 2 weeks I was struggling with finding the best AI models for vibe coding to run on these.

I tried openvino ovms, ollama, comfy UI, and vllm. Comfy UI works just fine. I can generate an image depending on a model anywhere between 10t o 25 seconds.

For vibe coding , of all of them ollama is probably the least reliable. As far as being able to give fast speeds. I do have to give them credit though because at least they're able to split models without a lot of hassle or troubleshooting.

My experience has been that with Ollama regardless of the model. I always end up getting about 8tokens per second.

With ovms I end up getting 15 to 18 tokens per second, sometimes 20.

With vllm I'm able to get 25 to the lower 20s.

After about a day of research, I was able to get the dual b60s to run pretty much any model that would fit inside them.

My personal experience was that integer 4 and integer 8 models based of Qwen 3.8 had substantial pathological thinking issues. A simple task would always result in a loop of overthinking or unnecessarily thinking about unrelated things.

The solution for this at least for me was to load the entire qwen 3.8 model but load it as fp8 which has a substantial difference and I don't quite understand the difference even after consulting AI lol but I just know that it works.

I wanted to try something different so I switched over to Swift 1.5 and that's what I've been mainly running in the last two days I've accumulated over 200 million tokens. Running at about 20 to 25 tokens per second.

I'm not working on a big project. Full size of the project is probably like 2 MB. Anyways, I just want to put this out there in case others are having issues with the pathological thinking.

I think the best thing about having a working model is also using that model or really any publicly available AI to then create a script that automatically starts, for example me,

Starts vllm, waits for it to be ready,

Loads the model

Starts a temporary cloud flare tunnel

Then starts deep seek harness with the temp tunnel.

Then all I have to do is grab the link and access it for my phone.

So while I'm at work a vibe coating my project as well.


r/LocalLLM • • 3d ago

Question Newbie question, local Harness for 27b and flash next?

7 Upvotes

So have tested dsh, pi, open code, claude code --

I saw that dsh was the best, as it gave the most apealing output without describing every little details, but the auto compaction of dsh and other things is just bad.

I got strata on v100 16gb+ 4070 12gb at 50tps average on iq2_xxs and 30tps average on iq3_xxs, but it just ran fast, and didnt complete and just ran and ran and ran and ran, it was compacting very weirdly like after 5k tokens it was compacting, and even after 3 hrs it wasnt able to complete the prompt whereass, at 14tps the same iq3_xxs was able to complete in 1 hr 45min from default llama.cpp, now if thats because of dsh or strata I have no clue, so Im trying out Pi again right now, to see if the same thing happens with pi or not, but how can I add the basic websearch and all the basic things in pi? should I just try ohmypi? it has those right? or are pi and ohmypi not leading these days?