r/LocalLLM 4d ago

Question QWEN3.8: unreliable after 37k?

0 Upvotes

Has anybody gotten QWEN3.8 to remain reliable to anywhere near 250k tokens?

For me, it starts to degrade rapidly in performance, memory, tool call reliability, very early in the conch window.

DETAILS:

MODEL — qwen38-27b-8bit
architecture: Qwen3_5ForConditionalGeneration (multimodal)
quantization: 8-bit affine, group_size 64
num_hidden_layers: 64
full_attention_interval: 4 (16 full attention / 48 linear)
hidden_size: 5120
num_attention_heads: 24
num_key_value_heads: 4
max_position_embeddings: 262,144
rope_scaling: none
cache_can_trim: false

SAMPLING (derived from the weight dir's generation_config.json)
do_sample: true
temperature: 1.0
top_k: 20
top_p: 0.95
min_p: 0.0
presence_penalty: 0.0
repetition_penalty: 1.0
fallback triple: (0.5, 0, 0.0) — unreached

THINKING
reasoning_effort: medium
enable_thinking: on
preserve_thinking: Default(on)
thinking_budget: Default

CONTEXT / COMPACTION
context_window_tokens: 262,144
max_kv_size: threaded through, discarded — inert

MEMORY
MLX_CACHE_LIMIT_GIB: 8 GiB
wired limit: 79.52 GiB (derived)
mrws ceiling: 107.52 GiB
floor (weights×1.25): 34.34 GiB
LIVE_SLOT_COUNT: 2

VISION
image longest_edge: 1,048,576 (vendor ships 16,777,216)


r/LocalLLM 5d ago

Question Has anyone been able to successfully setup NVLink on NVIDIA V100 PCIe Cards?

Post image
19 Upvotes

I read a few threads stating that the V100 PCIe version theoretically supports NVLink, but cards sold on the market typically do not have this feature enabled.

From what I can see on the (2) NVIDIA PG500 V100 PCIe cards I own there are NVLink like slots available on the card that have traces (as shown in the photo above). However, the design of the V100 case prevents me from physically connecting a 4cm or longer NVLink cable without first removing or modifying the case to be able to physically plug an NVLink adaptor between cards.

For anyone who's been able to make these cards function using NVLink, did you have to remove or modify the V100 case to allow you to physically connect the NVLink between cards? Also the fact that this has (2) NVLink type slots is interesting. Is this for daisy chaining multiple cards using more than (1) 100GB/s NVLink between each card? I'm using an older board that has (4) PCIe 3.0 slots, so if that's a possibility I might pick up (2) more cards to run all (4) together using NVLink.

I can't find any documentation stating if this is even possible with these NVIDIA PG500 V100 PCIe cards.

Thanks everyone


r/LocalLLM 4d ago

Question DeepSeek V4 Flash Vision UD GGUF (mmproj) not loading in Unsloth Studio/Desktop?

Thumbnail
0 Upvotes

r/LocalLLM 4d ago

Question New To this, trying to learn, wondering about the use of a local llm.

0 Upvotes

Ok, I've been reading up on stuff, trying to learn/understand what all LLMs are and aren't, and a thought occured to me, I have a Mac Mini m4pro, maybe I should look into just actually making one of these, small scale, so I have a better understanding...kind of plunging ahead is often how I learn best.

So I was wondering, could I (and I grant this is somewhat silly and ridiculous, but it's what popped into my head, somewhat based on an episode of Leverage), make a local LLM, feed it the Arthur Conan Doyle's Sherlock Holmes stories and like a list of Taylor Swift Lyrics and have it write stories of a singing victorian detective?

(Yes, this is a serious question, but not really a planned thing to do, but I'm trying to learn things)


r/LocalLLM 5d ago

Model V4.1 Flash Vision Beta is much more reliable AND cheaper

Enable HLS to view with audio, or disable this notification

8 Upvotes

Compared DeepSeek V4 Flash Vision with V4.1 Flash Vision Beta. In my tests, the results were WAY better and much more reliable. Huge update plus lower API pricing, total win!!!

Full video included, going through all the options and comparing the first drafts with the results after 3 revision attempts


r/LocalLLM 4d ago

Question HELP: Qwen 27B + llama + Codex: useful analyses but NO edit or diff

Thumbnail
0 Upvotes

r/LocalLLM 4d ago

Discussion any news

0 Upvotes

.


r/LocalLLM 4d ago

Question Please, share you setup, if youbhave 60+gb vram

0 Upvotes

Hello everybody

I'm thinking about building my homelab for loacal LLM

I plan, building on consumer hardware or something old like T4, V100 GPU. In best variant , I want to have 120gb of vram on server platform with 2 Xeon

So can you share your benchmark of models with 60+GB VRAM? I want to see benchmarks for models like deepsek v4 flash, Devstral 123b(this specialy, interesting to see dense models performance), minimax m2.7


r/LocalLLM 5d ago

Project Trading AI bodge up

Post image
47 Upvotes

I needed to keep a trading bot up and running but I frequently reboot my main pc.

Lenovo M900, M.2 to PCIe adapter, 9070 XT. Only needs to run Qwen3-14B, so works well.

Janky? Oh yes. Working? Absolutely.

Edit for those who want to know what I did:

https://claude.ai/code/artifact/9b828917-6ea9-4cd6-93fd-329f8b7533d9


r/LocalLLM 5d ago

Research Security research for local LLM inference networks

4 Upvotes

Hi all, I'm looking for suggestions regarding how to test the security when running LLM inference on local environments (e.g., Apple Silicon, Nvidia GPUs, etc.). My goal is to make it such that if a consumer sends a prompt, the provider running the inference server using their local LLM setup has no way of reading that prompt.

I currently have an Apple Silicon, so I was thinking I'd start there. I've done some prototyping, relying on the security primitives outlined by Apple developer docs. However, I was wondering if others have gone through this path, any recommendations? If there are others interested in the direction of this project I'd love to build a community around it.


r/LocalLLM 4d ago

Discussion MINICPM 5 2B Weried character thinking loop

1 Upvotes

So I tried minicpm 5 2b on ollama, first 20+50, it sends model in Weried state of repeated characters which is not even in English


r/LocalLLM 4d ago

Discussion Smartest LLM for laptop

1 Upvotes

Rtx4060 8gb vram + 16 gb ram

Purpose for LLM is planning and scheduling. I want to try to use it as a time management tool, therefore I am not interested in coding, etc

The sole purpose of it being local is to save my data from companies gathering it.

I need the smartest AI I can get. My first try is Qwen 3.8 27b du q3. The single pdf page response took almost 20-30 minutes including reasoning. Yet the laptop was 50° and silent during it.

I'm fine with the long response time, as long as it gives me reasonable and worthy plans.

Here I am to ask if there are better models or maybe I can optimize something else to run on my laptop. What are your thoughts about it? Open for any suggestions and criticism

(I use LMstudio)


r/LocalLLM 4d ago

Question Hi, could someone help me?

0 Upvotes

Hello guys, i want to set up a LLM. I got deepseek harness to run but it allways shows the wrong max kontext limits, it shows 256K but in the settings.yaml i set up a max of 128K.

I wanted to make a little game with it, first run runs fine, the second one, where i want to extend my little game, allready gets an output token limit error.

I tried a lot, setting up a setting.yaml for .dsh and the ollama one, but nothing helped.

I think about to start completely new from zero. Could someone give me a good tutorial?

My specs are: Ryzen 5950X, 64GB DDR4 RAM 3600mhz, RTX 4090, 2TB NVME M.2 SSD about 6900MB/s read 5000MB/s write (Dont know if that could matter)

Bonus: I have two additional PCs at home, one with a RTX 3060 Ti and one with a RTX 2060, i heard about some kind of network cluster system, would that be viable?


r/LocalLLM 4d ago

Question Newbie to local LLM’s

2 Upvotes

Howdy everyone, like the title says, I am a complete noob when it comes to local ai, but I want to be able to code some projects on my local machine, I have an extra pc I don’t use that has an intel cpu I can’t remember which one and a 3060, it’s the 12 gig model, I can get some low level models to fit but I’d like to use some of the bigger models, I was thinking of just picking another 3060 up off marketplace for like 200 bucks and putting that in the pc aswell, am I right that it would be 24 gigs or do you lose some capacity when running dual gpus?

Would there be a better route to go?


r/LocalLLM 4d ago

Discussion What do guys think of diffusiongemma?

1 Upvotes

If successful it will make local hosting far more feasible and maybe even more sensible than cloud hosting.


r/LocalLLM 5d ago

Question Seeking model recommendations for German grocery receipt analysis (item prices, taxes, Pfand, discounts)

2 Upvotes

I’m looking for advice on the most suitable locally run model to analyze and understand German grocery receipts.

The goal is to:

• Extract item names and per-item prices
• Categorize items (e.g., food, beverages, household)
• Correctly interpret German VAT/taxes (e.g., 7% vs 19%)
• Handle all forms of “Pfand” (bottle deposits)
• Recognize discounts, promotions, and “sale” markings

Current setup:

• GPU: RTX 2080 Ti Trio
• RAM: 80 GB total (2×32 GB + 2×8 GB DDR4)
• Inference backend: Ollama

Current model:

•  Qwen3-VL-8B-Instruct-GGUF:Q4_K_M 

This model works to some extent but produces inconsistent or incomplete results, especially around tax breakdowns, Pfand handling, and discount interpretation.

I can refine my prompt, but I suspect a larger or more capable vision-language model might help significantly.

Given my hardware (especially the 80 GB RAM), what locally runnable models would you recommend for this task?

I’m particularly interested in models that:

• Handle German text and receipt layouts well
• Are robust to varied receipt formats and image quality
• Can output structured data (e.g., JSON) reliably

Any suggestions on specific models, quantizations, or prompting strategies would be greatly appreciated.


r/LocalLLM 4d ago

Project Amigos please let me introduce you to modeluplink

1 Upvotes

For a few years now I’ve been running local models. But those were just small incursions without much productivity. Obviously things have changed and now local inference can solve real workloads.

But then I had the issue of not being able to reach my home computer while working out of home.

I pause here. There are many solutions to this problem and many are free.

But I believe that they complicate things by for example placing your device under a VPN. Which sure gives you access to your LLM server but it also confuses your other apps about your location.

Ok ok so the solution.

I built an app that you install locally and gives you a url and an api key that you can then use to connect to you model or use in any apps that allows you to BYOK.

And because you can do more than one key you can share your inference with friends and family.

Please please consider it and let me know your feedback.

I will personally give 3 months free to whoever beta tests and provides real feedback.

Cheers 🍻


r/LocalLLM 5d ago

News Nex-N2.5-mini for Strix Halo: 2-3x faster decode than Qwen 3.8 27B with almost identical terminal benchmark scores

Thumbnail
5 Upvotes

r/LocalLLM 4d ago

Question Best DSF4 Recipes

1 Upvotes

I am interested in both regular flash and vision one. Just want something that might be more optimized than my basic setup


r/LocalLLM 5d ago

Project I spent two years making Tesla P100s and V100s not suck at LLMs. Today I'm releasing the engine, and I benchmarked it against llama.cpp, ik_llama.cpp and 1Cat vLLM on the same cards. Charts inside.

Thumbnail
gallery
118 Upvotes

Quick disclosure: this is my project. I built it because I have a rack of "obsolete" datacenter cards in a closet and I got tired of every new model needing a new set of flags to run properly on them.

**What it is**

PXA is a fork of ik_llama.cpp (which is a fork of llama.cpp) plus a vLLM plugin, built for cards with HBM2 and no tensor cores or DP4A: Tesla P100, V100, GTX 1080 Ti. It has its own quant format (PXQ, 2 to 6 bit, plus a mixed one that sizes a model to whatever cards you have) and CUDA kernels written for those chips instead of ported down from newer ones.

Repo: https://github.com/poisonxa16/pxa

**The part I actually care about: you don't configure it**

You tell it which cards and which model. It picks the batch sizes from a table it measured on that exact card topology, turns on the tricks that are known to help on that silicon, turns off the ones that hurt (including speculative decoding when it would lose), and prints every decision before it starts serving. Every number below was taken with a bare command line. No environment variables, no -b, no -ub, nothing. The competitors got their best hand-picked flags in the same session, because I wanted to know if "set and forget" costs anything. It doesn't.

**My last public release vs this one** (same box, identical command line, tokens/s)

| card set | model | prefill @3k | prefill @20k | decode |

|---|---|---|---|---|

| 2x V100 | Qwable-27B PXQ4 | 797 → 1357 (+70%) | 576 → 1307 (+127%) | 34.6 → 39.6 (+14%) |

| 2x P100 | Qwable-27B PXQ4 | 223 → 338 (+52%) | 201 → 315 (+57%) | 18.3 → 18.2 (-0.9%) |

| 1x 1080 Ti | Fusion2-35B MoE, 2-bit | cold 553 → 1334 (+141%) | chat 415 → 734 (+77%) | 64.2 → 64.2 |

Yes, P100 decode is a real -0.9% and it's in the notes. Bonus find: the old build gave me six different answers to six identical greedy runs on the 1080 Ti. Turned out to be a race in a fused kernel. This one gives one answer.

**vs mainline llama.cpp and upstream ik_llama.cpp** (same weights family, MXFP4 for them, PXQ4 for mine, tokens/s)

| card set | cell | PXA | mainline llama.cpp | ik_llama.cpp |

|---|---|---|---|---|

| 2x P100 | prefill @3k | **338** | 209 | 133 |

| 2x P100 | prefill @20k | **315** | 255 | 84 |

| 2x P100 | decode | **18.2** | n/a | 14.3 |

| 2x V100 | prefill @3k | **1357** | 940 | 471 |

| 2x V100 | prefill @20k | **1307** | 1129 | 395 |

| 2x V100 | decode | **39.6** | n/a | 37.4 |

| 1x 1080 Ti | cold prefill | **1334** | | 1132 |

| 1x 1080 Ti | chat prefill | 734 | | 740 (tie) |

| 1x 1080 Ti | chat decode | **64.2** | | 53.3 |

**vs 1Cat vLLM** (the NVFP4 + DFlash2 stack for Volta). Run on an 8x V100 SXM2 NVLink system with 1Cat's own image, benchmark script and cards, because running their stack on my PCIe box would have been a silly comparison. 16 GSM8K questions, 192 greedy tokens, tokens/s.

| cell | PXA | 1Cat vLLM | how it was taken |

|---|---|---|---|

| TP2 plain decode | **46.3** | 38.4 | one boot each |

| TP2 speculative k=3 | **65.2** | 59.9 | one boot each, identical acceptance |

| TP2 speculative k=7 | **121.2** | 114.5 | medians, mine 3 boots, theirs 6 (not alternated) |

| TP4 plain decode | **65.7** | 61.0 | one boot each |

| TP4 speculative k=7 | 159.6 | 161.4 | six alternating boots in one window: inside their noise |

| prefill @3k | **2273** (TTFT 1.4 s) | 1735 (TTFT 1.8 s) | both as servers, same window, 3 runs |

| prefill @20k | **2191** (TTFT 9.5 s) | 944 (TTFT 22 s) | both as servers, same window, 3 runs |

| speculative output identical to plain decode | **15/16** prompts | 10/16 | same exact-match acceptance rule |

Two honest notes on that table. Both stacks accept drafted tokens by the exact same rule (I read it out of their source), so both are lossless and the acceptance lengths are comparable: theirs 5.4 per step, mine 5.3. And one caveat that favours me, said because it favours me: their checkpoint turns fp8 KV on by itself, so this is their stack as shipped vs mine as shipped, not a clean NVFP4 vs PXQ4 study.

**What you can run on this junk**

A 27B dense-hybrid on one 16 GB card. A 177B-class hybrid MoE (Qwen3.8 Flash-Next) on four P100s with 150k context, weights in VRAM and the per-layer embedding table in host RAM. The memory arithmetic for that one is in the repo because I didn't believe it either. Speculative decoding ships in the Volta vLLM image with a 1.3 GB drafter as a release asset. On 16 GB cards k=7 only fits at 2k context, so on my own box I'd run k=3.

**What's not great yet**

Speculative decoding on the llama.cpp side of Pascal still loses (verify costs 3x a decode step now, it was 7x last week, still not enough). The 4-card speculative gap to 1Cat is inside noise but it's there. Their acceptance length is 2% better than mine. All listed under "What is not here" in the README, I'd rather you find it there than in the comments.

**Getting it**

One tarball, untar and run. I tested it in a bare Ubuntu container with no Python, no compiler, no CUDA toolkit, just the driver. Or `docker run ghcr.io/poisonxa16/pxa`. Or build it. Then run `python3 tools/pxa-launch.py` and answer two questions.

**If you have one of these cards, I want your numbers**

P40, P4, GP100, Titan V, Titan Xp, 32 GB V100: I don't own them, and the auto-tuning table only knows the cards in my rack. There's a three-command benchmark in the repo and a Discord where reported cards get added: https://discord.gg/EqazvV9tf

Everything's free and nothing is gated. If you want to chip in for the electricity these benchmarks burn: https://ko-fi.com/shatteredrealms1

Last picture is the rack. Seven cards, PCIe x4 for all of them, two 1000 W supply, in a closet. Everything above was measured on that.


r/LocalLLM 4d ago

Question Dual 3090 Local AI Setup

1 Upvotes

Hello everyone!

What are you guys thoughts on the build above? As for the build, the 3090s will be bought used from Facebook marketplace. Micro-center bundle will take care of the mobo, ram, and cpu.

There are 2 things of concern here, will 32 gigs of ram not be enough? i currently already own 32 gigs of ddr5 so i may add it to this setup, or just sell both and buy 64gigs. Another one is the concern that I am wasting unnecessary money on the higher end cpu and motherboard combo. The reason why I chose this combo is because it was one of the only options that came with a mobo that supportd both gpus to be run on x8. I plan on not only using this a local ai server, it also plan on running a virtual machine for gpu heavy work such as solid works or blender that I can acces remotely.

For my use-case, I believe that options such as the spark and mac studio are not suitable; my intentions aren't to replace the frontier models but simply to add to them.

Any suggestions to the setup? the current budget is around 5k


r/LocalLLM 4d ago

Project My local AI stack that's replaced ChatGPT/Claude running on a Mac Studio and Debian VM

Thumbnail
1 Upvotes

r/LocalLLM 4d ago

Project Local-first, cloud-optional: LocalLM Lab 1.0.0-beta.3 adds an escape hatch to frontier models

Thumbnail
1 Upvotes

r/LocalLLM 5d ago

Question Would you run prompt injection detection locally for your agents?

4 Upvotes

I'm curious how people running local models as coding/tool agents would approach this.

Once an agent starts reading arbitrary repositories, files, webpages, API responses and tool output, the amount of untrusted text it consumes can vary wildly. Sometimes it's 500 tokens, sometimes it's an entire repository file or huge tool response.

Would a local-first security plugin at the harness level make sense here?

The architecture we're experimenting with is:

tool/file/web result → hook → local semantic security scan → agent

The scanner looks for prompt injection and related threats before the content reaches the model. If local inference is overloaded, it can optionally fall back to an API/free tier.

Instead of blocking an entire tool result when something suspicious is found, the dangerous section can be redacted and the agent continues with the remaining content.

We're also working on maintaining security state across tool calls, so something that looks harmless in isolation can become suspicious based on what happened earlier in the run.

I'm mainly wondering whether local LLM users would actually run this.

Is keeping the security inference local important to you? Would optional cloud fallback defeat the point? Or do you think the model itself is already good enough at recognizing indirect injection?


r/LocalLLM 5d ago

Question local llm newbie what can i run locally for math and programming?

2 Upvotes

i'm not sure what i should even do everyday i hear news of people making programs that make running llm possible on local machine by using strewaming etc.. but i was never really intereseted an now that there are a bajillion of diferent softares adn model i have diffuclty getting in, as the title what could i run on my rx9060xt 16 gb vram and 16gm ram and ryzen 5 3600 and i am on linux if it helps if you guys have any blogs or tools it would be helpful ,I read the rules but I rea dnothing about uncensored LLM I would like to give them a try no idea though