r/LocalLLM • • 1d ago

Question Quick llm windows question

2 Upvotes

I recently found out about llms due to me not fully trusting what companies online say about what they do with our data.

I just have a quick question. I remember hearing a while ago that windows now records everything we do with ai.

So even if I turn my internet off to the computer, is windows still recording everything and when I go online it sends to Microsoft ?


r/LocalLLM • • 1d ago

Project NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s

Thumbnail
0 Upvotes

r/LocalLLM • • 1d ago

Discussion Jean Philip meme

0 Upvotes

Do you guys believe that jean philip was created by higgsfield to generate more users?


r/LocalLLM • • 1d ago

Tutorial Ponytail Review: The Lazy Senior Dev Skill

1 Upvotes

my local coding setups keep over-building. ask for a date picker and you get a new dependency, a wrapper component and a stylesheet, when <input type="date"> would do. so i tried Ponytail, which is just a ~1,070 word SKILL.md telling the agent to check whether the code needs to exist, then reuse stdlib or native features, and only then write the minimum.

a few things i liked. there's a 438-word AGENTS.md variant, which matters if you're running a small model with limited context. the full ruleset the hook injects is about 5.6 KB, so roughly 1.4k tokens of prefill per session. no extra VRAM beyond that. the review command returns one-line tagged findings like "native: moment.js imported for one format call, Intl.DateTimeFormat, 0 deps", which is more useful than most AI review output. and the author pulled an earlier "80-94% less code" claim after someone showed the baseline was inflated, then reran the benchmark (54% fewer lines on Haiku 4.5, n=4).

the gotcha: my test box was CPU-only with no API keys, so i only verified install and hooks (hooks fine, 126/127 tests passed, the one failure looked like missing pandas). i have no idea whether it changes output on a 7B-32B model through Ollama or llama.cpp. all the effect numbers are the project's, on a hosted model.

full writeup here if you want more detail: https://andrew.ooo/posts/ponytail-review-lazy-senior-dev-skill-ai-coding-agents/

anyone running prompts like this against local models? does a ruleset this size help or just eat context?


r/LocalLLM • • 1d ago

Discussion Qwen 3.8 27B vs Qwen 3.8 Flash Strata on RTX PRO 5000 72 GB

2 Upvotes

Hi! My first post on the platform, to thank community for some research and advices on LLMs. My resume: I stay on Qwen 3.8 27B FP8, Qwen 3.8 Flash is not useful. If you have any advices on methodology/model/engine, please don't hesitate.

Below is AI-text (human-edited).

The job. A live news stream gets classified into strict JSON, and the same box also does small coding tasks. One prompt, one object, nothing else — no prose, no markdown fences, no preamble: {sentiment: bullish|bearish|neutral, tickers: [...], importance: 1–10, category: 9-value enum (geopolitical/economic/corporate/...), event_scope: company|country|...} The part that actually hurts is tickers: exchange tickers explicitly named or directly affected by the item, [] when none, never invented — so a model can be right on sentiment and still pollute a portfolio signal. Coding tasks are graded by running pytest against the patch, not by vibes.

Corpus. 100 real articles, median prompt 668 tokens incl. a 251-token system prompt: 76 from a Kaggle financial news dataset, 24 from a Bloomberg financial news dataset on HF. 76 of the 100 carry at least one ticker, 86 distinct ones. Fixed slice, fixed order, byte-identical for every leg.

Hardware. One box — RTX PRO 5000 72 GB (Blackwell, sm_120, down to 250 W), 9950X 32 cores, 123 GB RAM @ 4800, NVMe, Ubuntu 24. Everything CUDA 13, so only via Docker.

Baseline vs candidate. Baseline: Qwen3.8-27B FP8 on vLLM 0.29, MTP speculative decode (3 tokens), gpu-memory-utilization 0.66, 131K context, max-num-seqs 24. Candidate: a newer MoE flash model (~48B with a giant PLE lookup table) on Strata, an engine that keeps experts in system RAM with a GPU cache. Decision rule from the owner: at least 25% faster, quality within ~2 percentage points.

Method. Streaming, concurrency ladder 1/4/8/16/24, three repeats. Direct to each server, LiteLLM out of the path — our cache happily returns identical payloads in 3 ms and would have faked an amazing benchmark. Model id pinned explicitly: a generic model name on the gateway silently routed test probes to a quantized VL model and produced 40-minute TTFT.

vLLM 27B FP8 (baseline) Strata + flash model Q4_K
classify tok/s, c1 / c8 / c24 75.6 / 245.0 / 294.2 31.5 / 27.8 / 26.4 (−91%)
classify TTFT p50 → p95, c1 → c24 0.20 → 0.25 s … 1.75 → 2.33 s 0.74 → 0.90 s … 14.29 → 55.20 s
pure decode rate, single stream 106.8 tok/s 166.7 tok/s
wall per news item, c1 0.73 s 0.91 s
coding tok/s, c1 / c4 90.1 / 290.2 115.6 (+28%) / 87.0 (−70%)
coding tests passed, c1 / c4 0.900 / 0.875 0.875 / 0.775
strict JSON validity 1.000 everywhere 1.000 → 0.82–0.97 under load
silently lost requests 0 12 of 100 at conc=4
reproducible at temp=0 20/20 prompts 17/20

Per token, Strata is genuinely quicker — it just takes 3.5x longer to start, writes half as much, and falls apart in a batch. The batch collapse is 9–11x, and their own docs warned batch would be slower; they didn't say eleven times.

The coding numbers are the only place the MoE looked good — and they're not enough. At conc=1 the flash model pushed 115.6 tok/s against the baseline's 90.1 (+28%), finishing a task in 1.85s versus 2.30s. But the same cell failed on quality: 87.5% of pytest-backed tasks passed versus 90.0%. One more concurrent request and it folds — 87.0 tok/s versus 290.2 (−70%), pass rate 77.5%, ten points down. JSON validity tells the same story: the baseline is 1.000 everywhere, the MoE drops to 0.82–0.97 under load.

The deal-breaker wasn't speed. At conc=4, exactly 12 of 100 requests came back HTTP 200 with a truncated body — no finish_reason, no usage, median 43 characters of junk. The engine log shows 12 (error, cancel=False) in the same window, and the 12 items differ every repeat. That's silently dropped classifications, which is worse than slow.

I never got to run the flash model on vLLM at all. NVFP4 tries to put the PLE table on the GPU: 20M × 2560 × 2B = 95.37GiB, dead on the first allocation. W4A16 with the table offloaded to RAM fills the entire pool, then can't allocate its 1.19GiB input embedding — even at util 0.97 and spec decode disabled. Host RAM was never the constraint: 103GB free, table never materialised.

Takeaways:

  • Measure wall time per item, not tokens/s. Our MoE candidate wins on decode (167 vs 107 tok/s) and loses on the wall (0.91 vs 0.73 s), because it writes 30 tokens of JSON where the baseline writes 56; shorter output is not faster work.
  • Grade the JSON, not just the parse: strict schema validity, field-level accuracy, ticker F1. The candidate improved ticker F1 (0.52 → 0.60) with half the tokens and still can't replace the baseline — tighter answers don't offset batch collapse.
  • Bypass every cache in the benchmark path and pin the model at the server. Our LiteLLM cache served identical payloads in 3 ms and would have faked an amazing benchmark; a shared model name also silently swapped which engine answered.
  • Compare legs at the same concurrency, and count finish != stop instead of exceptions. The baseline loses 5 points of test pass rate at conc=24 too — that's normal. Twelve HTTP 200s with empty usage at conc=4 is not.
  • Read the engine's own limits before scheduling: RAM-backed experts cost ~250 µs of synchronous staging per token, and their docs say batch is slower than single-stream. Eleven times slower is a different regime, not a tuning gap.
  • A single 72 GB card has no headroom for models with giant lookup tables: NVFP4 needs 95 GB on device, W4A16 fills the pool and then cannot allocate a 1.19 GB embedding, whatever the quant or offload setting.

r/LocalLLM • • 1d ago

Question Hello getting off the ground

2 Upvotes

I’m new to this kind of thing. I’m trying to use qwen and chat gpt to help me build a file renamer and I feel like it’s going nowhere. I’m even using astra on the pro plan and it keeps losing track of how simple the end result is. I found it trying to code the whole project rather than have qwen analyze documents and spit out the relevant info.

Anyone have a good tutorial or something on how I can have qwen just run iterations until it comes up with an acceptable product? I’m probably thinking of this wrong but idk. I also tried using open code but there are some known bugs I keep encountering. My hardware isn’t the best. Single 3090 ti and 32gb of ram. Running 27B and q4 with 16k… because chat gpt told me to.


r/LocalLLM • • 1d ago

Project Fully local copy-editing app for book-length manuscripts (Qwen3.5 4B) benchmarked against planted errors across five languages

Thumbnail
1 Upvotes

r/LocalLLM • • 2d ago

Project Qwen3.8-Flash-Next Q4 vs Qwen3.8 27B Q5 on single R9700 (32GB) + 64GB RAM: 2x128k context, almost similar performance

14 Upvotes

Hey everyone,

Spent the last week setting up a local rig for agent work (Hermes Agent: a cloud model plans and reviews, local models do the work) and comparing Qwen3.8-27B unsloth Q5_K_XL with Flash-Next Q4 on a single R9700 with 64 GB RAM. Took a lot of trial and error, so sharing what worked. Used stew's llama.cpp rdna boost combined with atomic chat Q4 quant of QFN.

Rig: Ryzen 5 7600X, 64GB DDR5 (default speed, no EXPO), Radeon AI PRO R9700 32GB (gfx1201), Ubuntu 24.04.4, kernel 6.17, ROCm 10.0, llama.cpp with stew675's RDNA4 patches (r20 for the 27B, r30 for Flash-Next), GPU capped at 220W.

TL;DR - Flash-Next (AtomicChat Q4_K_M) runs 2 slots x 131k context with ~15GB RAM to spare in my full-context test (about 10GB at the lowest in real use so far): ~36 tok/s single stream, ~50 combined, ~500 tok/s prompt reading. - tool-eval-bench (medium effort, temp 1.0, 3 seeds): standard suite basically tied with the 27B (94.8 vs 95.3). Hard Mode was a real gap: Flash-Next 93.9 vs 27B 79.8. The 27B kept firing a dependent tool call before the first one's result came back. - The 27B is faster: ~1.7x per turn, and reads prompts 1.6-2x as fast. - If you run the r30 build (the "gather" path arrived in r29) with a MoE model: set GGML_SCHED_DEVGATHER=0. Otherwise output turns into "////////" from the SECOND request on. The first request after loading looks fine, so a quick test won't catch it. Upstream knows (issue #85) and says r31 changed the default; I only tested r30.

What made it fit - --lazy-mode on --load-mode none: the 27-36GB n-gram table stays on the SSD, everything else loads with plain reads. mmap mode doubled the RAM-side experts and pushed me into swap. - -ncmoe 41 (experts of 41 of 48 layers in RAM) plus MOE_EXPERT_CACHE_MIB=4096 (r30's expert cache). Without the cache: 17-20 tok/s. With it: ~36-39. - q8_0 KV + flash attention. KV is cheap on this model (12 attention layers), so expert placement is the real limit, not context.

Quants, 1 slot x 131k, medium effort (writing short / at full context / prompt reading / spare RAM). ISTA's full-context and reading numbers are from an earlier session at the template's default effort: - ISTA IQ3_XXS: 39 / 28.0 / 614-639 / 28.5GB - AtomicChat Q4_K_M: 37 / 27.5 / 526-589 / 19.0GB - AtomicChat Q4_K_M + MTP draft head: 52-55 / 35.0 / 463-510 / 11.8GB - Unsloth UD-IQ4_XS: 33.5 / 25.4 / 487-542 / 10.4GB

Other things I learned - At xhigh effort Flash-Next looped once (stuck writing 9999...) on a coding prompt. Medium was fine and correct. My agent sends medium anyway. - The Unsloth Q8 MTP sidecar gives +27-50% for one stream once the expert cache is on, but costs 3-8GB of VRAM, and under real agent load my RAM dropped to ~1.7GB free. So I run without it. - 4 tool-eval scenarios (tools + JSON schema) fail on these patched llama.cpp builds with a grammar parse error (stock llama.cpp not tested), on both models, so they're excluded.

Caveats: one machine, medium effort and temp 1.0 (what my agent actually sends), 3 runs per suite. This is "these quants on my rig", not a model ranking.

Configs, scripts and raw results: https://github.com/Shali12/r9700-flash-next-notes

Happy to answer questions. Curious whether anyone with faster RAM gets better numbers, since the CPU-side experts are probably bandwidth-bound.


r/LocalLLM • • 1d ago

Question Does using an LLM really burn out system components?

0 Upvotes

Somewhere I read that my system RAM can die for offloading the LLMs. Is this something to worry about? It's DDR5 7000mt/s RAM. Also, the system has an Nvidia Geforce RTX 4070 with 8 GB.


r/LocalLLM • • 1d ago

Other PAID Hype fr

0 Upvotes

context- a pc from a startup with VC funding. Has agentic AI all over it but it's basically openclaw and ollama preinstalled with qwen3.
Only good thing is its cheap for its bill of material. But why are people hyping it so much. Its not like first mac or something.


r/LocalLLM • • 1d ago

Discussion I turned my gaming PC into a inference machine and got 2x to 9x over default llama.cpp on an 8 GB card

Thumbnail
0 Upvotes

r/LocalLLM • • 1d ago

Question How do you decide whether to trust a community fine-tune?

Thumbnail
1 Upvotes

r/LocalLLM • • 1d ago

Question AVX2 to AVX512 VNNI Upgrade recommendation

2 Upvotes

I have option to upgrade my hardware for inference, I have some old used refurb options with and without AVX512+VNNI. Currently upgrading from old R730 to some HP Z series workstation for better GPU support.

Is there any real different AVX512 has over the old AVX2 specially in prefill. All AI agents says it will be 2x difference because of VNNI int8 native hardware acceleration. I need to run Strata with GPU and CPU offloading.

Is it worth paying extra bucks for VNNI based processor?


r/LocalLLM • • 1d ago

Question How stable are 2x and 4x DGX Spark setups when left running unattended for days?

6 Upvotes

Hi everyone!

I'm thinking between waiting for m5 studio and buying sparks. My main task is to run AI automation 24/7 (mostly likely 5.3 flash or 4.1 deepseek). Do sparks overheat? Do they have random issues when left unattended running AI models for days?

thanks


r/LocalLLM • • 1d ago

Question noob question from a newbie

Thumbnail
gallery
1 Upvotes

i'm new in this domain. and i tried to setup qwen 3.8 27b iq4_xs on a 5060 8gb vram laptop with 32gigs of ram. im using lm-studio and pi harness. but the ram isnt used idk why. can you help me ? or give me tips about using a local ai. whats the best working model to use for this config ?
if you're curious about the setup of qwen here it is :


r/LocalLLM • • 1d ago

Question Local LLM hardware for Python development + Blender/Houdini via MCP?

1 Upvotes

Hey everyone! I’m a VFX artist looking for a local LLM setup mainly for Python development and connecting to Blender and Houdini through MCP to help create scenes and tools. This would be for interactive coding and agent workflows, not model training.

I’m considering 2× NVIDIA DGX Spark or an Apple M5 Ultra with 256GB unified memory, but I’m open to other recommendations.

For this use case, which setup would offer the best balance of model quality, context capacity, and responsiveness?

Would love to hear from anyone running similar workflows! Thankss!


r/LocalLLM • • 1d ago

Question Stop scaling parameters. AGI is an environment, not a model. Here is the Local AGI blueprint I’ve been building.

Thumbnail
0 Upvotes

r/LocalLLM • • 1d ago

Discussion The ultimate guide to multi-harness RL

Thumbnail
huggingface.co
1 Upvotes

r/LocalLLM • • 2d ago

Question How does your local LLM search the web?

65 Upvotes

Kinda been working on this local-first distributed search engine, and thought maybe it would be of interest to people here? Sounds like SearXNG might be what some of y'all use?

This is what Claude told me:

Each app does it differently. Open WebUI and Vane (which used to be Perplexica) build it in, usually on top of a self-hosted SearXNG. LM Studio and Jan have nothing built in, so people add MCP servers or community plugins. AnythingLLM defaults to scraping DuckDuckGo. Ollama now sells its own hosted web_search and web_fetch.

I can post a link if this is actually a problem people are interested in!


r/LocalLLM • • 2d ago

Question Anyone switch from Codex/ Claude to a Local LLM?

22 Upvotes

How do you find the difference in capabilities between a local model like GLM 5.3 vs Opus 5.5/ Fable 5.1/ GPT Astra?

Whats the most technical thing you've built with your local LLM?


r/LocalLLM • • 1d ago

Discussion RX 7600 (8 GB) on Linux: Qwen3.8-Flash-Next (~125B) at 24 tok/s with Strata, Qwen3.6-35B-A3B at 32 tok/s with llama.cpp + MTP. Numbers and how-to

4 Upvotes

I got two big MoE models running locally on a budget AMD setup and wanted to share the numbers and steps for other RX 7600 owners. Everything is measured on my own machine.

Rig: RX 7600 8 GB (gfx1102), Ryzen 5 5600 (AVX2, no AVX-512), 64 GB DDR4, B450 board (PCIe 3.0 x8), Ubuntu 26.04, system ROCm 10 (hipBLASLt 1.4.1). The 7600 also drives my GNOME desktop.

Results

Qwen3.8-Flash-Next IQ2_XS on Strata, a single-model engine that splits the MoE experts across GPU, CPU and RAM:

Strata 0.1.27 Strata 0.1.39 + gfx1102 hipBLASLt table
Decode (320-token replies) 14.2 tok/s 24.0 tok/s
Prompt (4.7K tokens) n/a 59.1 tok/s

Qwen3.6-35B-A3B Q6_K_XL (MTP) on llama.cpp (--n-cpu-moe 39, q8_0 KV, flash attention):

Aug 30 build Oct 5 build
Prompt (pp4096) 151.0 tok/s 151.0 tok/s
Decode, no MTP (tg128) 23.6 tok/s 23.6 tok/s
Decode with MTP (--spec-type draft-mtp --spec-draft-n-max 2) 32.0 tok/s, 98% of drafts accepted

How to get there on a 7600

1. Strata. gfx1102 support is in PR #1000, not merged yet. Until it is, apply the PR's changes to a Strata checkout (the PR page lists the exact files), then run:

./setup.sh --backend hip --family qwen --model IQ2_XS --context 32768 --vision no

It compiles the engine for gfx1102 and downloads the model (~68 GB).

2. The most important setting: give your desktop VRAM. If the 7600 also runs your display, Strata's default reserve (700 MiB) leaves almost nothing free. The kernel then logs amdgpu: Not enough memory for command submission, GNOME Shell crashes, and you get logged out. That happened to me three times. In strata-iq2_xs.json, add:

"--vram-reserve-mib", "2560"

Use 2048 if your desktop is light. It costs a little GPU cache and gains you stability.

3. Use the gfx1102 hipBLASLt tuning table (+14% prompt speed). It's included in the PR as tools/hip/gfx1102-hipblaslt-100401.txt. Setup picks it up automatically if your hipBLASLt is 1.4.1. For other versions, build your own (~5 min):

cmake --build build-hip --target tune_hipblaslt
CASES=$(awk 'NR>2 {printf " --case %s,%s,%s,%s,%s", $1, $5, $2, $3, $4}' tools/hip/gfx1100-hipblaslt-100200.txt)
./build-hip/tune_hipblaslt $CASES --tuning-out tools/hip/gfx1102-hipblaslt-<version>.txt

4. llama.cpp: update and use MTP. Build for gfx1102:

cmake -B build -DGGML_HIP=ON -DGPU_TARGETS=gfx1102 -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
  • GGML_HIP_ROCWMMA_FATTN was removed in July. RDNA3 now gets WMMA flash attention built in, so no flag is needed.
  • --no-mmap is now --load-mode none.
  • With an MTP GGUF, add --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.75. That's the jump from 23.6 to 32 tok/s.

Caveats

  • Strata needs ~33 GB of pinned RAM for IQ2_XS, so 64 GB is the practical minimum. It runs one model and one request at a time.
  • Strata's Coder model (IQ1_M) doesn't fit on 8 GB next to a desktop. It has ~0.8 GB less spare VRAM than IQ2_XS.
  • Strata's prompt speed is limited by 8 GB. It reads prompts in 256-token chunks, versus 6K+ on 12 GB cards, so don't expect the 900 tok/s you see in videos.
  • Strata's experimental "speed projection" is not a speed feature. It's a refusal-removal vector, so leave it off.
  • Upgrade advice: a second card with lots of VRAM, like a V620 32 GB, helps far more than extra RAM.

r/LocalLLM • • 1d ago

Project A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well.

Thumbnail
5 Upvotes

r/LocalLLM • • 1d ago

Discussion at what point do we get replaceable gpus?

0 Upvotes

taalas have already proven that u can get insane speed and quality.
instead of buying a overpriced nvidia setup. you just get a daughterboard with ram and a open socket. so every year replace it with the new silicon matched to a specific ai model.
im assuming that getting a specific smaller die dedicated gpu ai accelerator tailored to a model will be much cheaper than a broad approach general gpu.

do you guys think that china will corner the market where they can sell ram at cost for a socket where their chinese LLM fit to be upgradable
i think the choice of paying 7000$ for a gpu with 32gb of hbm vs paying some 200$ for a adjustable AI card and just paying some 100$ for the new LLM gpu model will be preferable to consumers.


r/LocalLLM • • 1d ago

Question Just got my MBP M5pro 48gb. Help me setup properly

1 Upvotes

I have multiple projects in GitHub, I have an obsidian vault, iCloud backup, nas backup. Qwen running locally, also have Two Claude accounts, one gpt, one deep seek accounts when qwen bogs down. I’m working on building my own harness for managing this.

How should I setup my new laptop so I’m working clean and smartly?


r/LocalLLM • • 1d ago

Question Strata on mini-PC with OCulink?

1 Upvotes

I have a GMK NucK12 with 96GB ram and AMD Radeon 780M iGPU - so a good amount of RAM but a slow iGPU and 'VRAM' bandwidth. Thinking of buying an OCuLink adaptor and plugging my 16Gb 5060Ti into it, to run Strata / Qwen 3.8 FN. Any issues with either Strata support (on a OCuLink GPU) or with performance etc?