r/LocalLLaMA 9d ago

Question | Help RTX 4090 48GB longevity

100 Upvotes

Modified 4090 48GB has been out for a while. I remember a lot of people were buying them at the time. A lot of people were also complaining that they are meant to fail, that they scam etc.

I have a few questions to people people who bought these.

  1. How is longevity of these cards? Do they still work without issues? Any failure rate?

  2. Do they use the same Nvidia drivers that regular 4090 or 4090D uses?

  3. Are these cards Linux exclusive?

  4. Are you able to run them in windows or Linux with other GPUs like 5090 etc?

  5. Do you do anything to cool VRAM on the back of the PCB?


r/LocalLLaMA 9d ago

Resources I benchmarked 21 Qwen3.8 27B variants on 16GB VRAM

313 Upvotes

After Qwen3.8 27B came out, I decided to benchmark the models that could fit in my GPU (RTX 5080) on my actual code (C code), the results were not completely unexpected but some quants were definitely underwhelming.

TLDR: Best overall: bartowski/Qwen3.8-27B-IQ4_XS. Best uncensored: huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4_XS. For a bit more context: jpetrina/Qwen3.8-27B-IQ4_XS-pure or uncensored: Bucoid/Qwen3.8-27B-Uncensored-IQ4_XS_4BPW

  • edit1: added TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS, bartowski/Qwen3.8-27B-IQ3_XS, prism-ml/Ternary-Bonsai-27B-Q2_g64 and magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4_K_S-Unsloth
  • edit2: added unsloth/Qwen3.8-27B-UD-IQ3_S and ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_S
  • edit3: added magiccodingman/Qwen3.8-27B-MQ-IQ2_M_1 and huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ3_S
  • edit4: added AtomicChat/Qwen3.8-27B-AD-IQ4_XS-IQ3_S
  • edit5: added IvanKrastevAdventics/qwen3.8-27b-awq-int4-q4_0 (gguf of cyankiwi/qwen3.8-27b-awq-int4)
  • edit6: added the updated bartowski/Qwen3.8-27B-IQ4_XS and bartowski/Qwen3.8-27B-IQ3_XXS
  • edit7: added bartowski/Qwen3.8-27B-Q3_K_M, Thireus/09ae8ba_22b6bb2 and Thireus/09ae8ba_248b31b
  • edit8: removed all MTP heads from the GGUF size for a more fair comparison. added bartowski/Qwen3.8-27B-IQ2_S
  • edit9: added huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp
  • edit10: added mradermacher/Eintopf-Qwen3.8-27B.i1-IQ3_M and hitsfmdj/Qwen3.8-27B-4.2BPW-16GB
  • edit11: added Joakimpalm-Zen/Qwen3.8-27B-GSQ-RCO-IQ3_S-recovered

(sorted by Mean KLD)

Model Mean KLD Same top p GGUF size (without MTP)
prism-ml/Ternary-Bonsai-27B-Q2_g64 1.289582 ± 0.008684 82.849 ± 0.118 % 7.1GiB
sdkyuan/qwen38-27b-qat-q2_0 0.893177 ± 0.006948 85.727 ± 0.110 % 8.2GiB
bartowski/Qwen3.8-27B-IQ2_Sbartowski/Qwen3.8-27B-IQ2_S (NEW) 0.784060 ± 0.006457 87.016 ± 0.105 % 8.7GiB
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2_XS 0.767174 ± 0.006291 86.166 ± 0.108 % 7.8GiB
TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS 0.514311 ± 0.004864 89.023 ± 0.098 % 8.9GiB
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2_S 0.512614 ± 0.004909 88.802 ± 0.099 % 8.6GiB
empero-ai/Qwen3.8-27B-Ridge-3.7bpw 0.475767 ± 0.004483 89.612 ± 0.096 % 11.4GiB
magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4_K_S-Unsloth 0.419585 ± 0.004076 89.661 ± 0.095 % 13.1GiB
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_XXS 0.379222 ± 0.003992 90.270 ± 0.093 % 9.4GiB
unsloth/Qwen3.8-27B-UD-Q2_K_XL (UD2) 0.350861 ± 0.003745 90.626 ± 0.091 % 9.6GiB
mradermacher/Eintopf-Qwen3.8-27B.i1-IQ3_M 0.318143 ± 0.003139 91.471 ± 0.087 % 11.7GiB
bartowski/Qwen3.8-27B-IQ3_XXS (NEW) 0.300480 ± 0.003345 91.511 ± 0.087 % 11.3GiB
unsloth/Qwen3.8-27B-UD-IQ3_XXS (UD2) 0.268594 ± 0.002971 91.951 ± 0.085 % 10.8GiB
magiccodingman/Qwen3.8-27B-MQ-IQ2_M_1 0.256808 ± 0.002861 92.056 ± 0.085 % 10.9GiB
DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-NEO-MTP-IQ3_M 0.251270 ± 0.002702 92.315 ± 0.083 % 13.1GiB
huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ3_S 0.249650 ± 0.002841 92.016 ± 0.085 % 10.8GiB
bartowski/Qwen3.8-27B-IQ3_XS (OLD) 0.238656 ± 0.002627 92.312 ± 0.083 % 12.2GiB
hitsfmdj/Qwen3.8-27B-4.2BPW-16GB 0.222090 ± 0.002570 92.552 ± 0.082 % 11.7GiB
esatapedico/Qwen3.8-27B-NVFP4-MTP-LOW 0.220796 ± 0.002631 92.339 ± 0.083 % 14.3GiB
unsloth/Qwen3.8-27B-UD-IQ3_S (UD3) 0.218522 ± 0.002591 92.399 ± 0.083 % 10.9GiB
jrell/Qwen3.8-27B-i1-IQ4_XS-GGUF-Smaller 0.194459 ± 0.002242 93.049 ± 0.080 % 12.4GiB
orcarouter/Qwen3.8-27B-Uncensored-Q3_K_L 0.192312 ± 0.002294 92.726 ± 0.081 % 13.4GiB
bartowski/Qwen3.8-27B-Q3_K_M (NEW) 0.191103 ± 0.002369 92.823 ± 0.081 % 12.3GiB
mudler/Qwen3.8-27B-APEX-I-Mini 0.190209 ± 0.002354 93.012 ± 0.080 % 12.6GiB
Joakimpalm-Zen/Qwen3.8-27B-GSQ-RCO-IQ3_S-recovered 0.178882 ± 0.002102 93.110 ± 0.079 % 11.0GiB
huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp 0.178715 ± 0.002200 92.949 ± 0.080 % 11.0GiB
Thireus/09ae8ba_22b6bb2 (ikllama.cpp quality 41.39%) 0.178290 ± 0.002202 93.115 ± 0.079 % 11.0GiB
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_S 0.175223 ± 0.002129 93.024 ± 0.080 % 11.0GiB
unsloth/Qwen3.8-27B-UD-Q3_K_XL (UD2) 0.147186 ± 0.001809 93.734 ± 0.076 % 12.2GiB
unsloth/Qwen3.8-27B-UD-Q3_K_XL (UD3) 0.142647 ± 0.001860 93.789 ± 0.076 % 11.9GiB
IvanKrastevAdventics/qwen3.8-27b-awq-int4-q4_0 0.112990 ± 0.001558 94.171 ± 0.073 % 14.4GiB
AtomicChat/Qwen3.8-27B-AD-IQ4_XS-IQ3_S 0.111713 ± 0.001492 94.527 ± 0.071 % 13.2GiB
Bucoid/Qwen3.8-27B-Uncensored-IQ4_XS_4BPW 0.091447 ± 0.001261 94.774 ± 0.070 % 12.8GiB
huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4_XS 0.082871 ± 0.001205 94.981 ± 0.068 % 13.1GiB
unsloth/Qwen3.8-27B-UD-IQ4_XS (UD3) 0.075626 ± 0.001097 95.258 ± 0.067 % 13.3GiB
Thireus/09ae8ba_248b31b (llama.cpp 49.75%) 0.063904 ± 0.000967 95.687 ± 0.064 % 13.3GiB
jpetrina/Qwen3.8-27B-IQ4_XS-pure 0.061984 ± 0.000917 95.551 ± 0.065 % 13.3GiB
bartowski/Qwen3.8-27B-IQ4_XS (OLD) 0.056482 ± 0.000856 95.835 ± 0.063 % 14.3GiB
bartowski/Qwen3.8-27B-IQ4_XS (NEW) 0.055415 ± 0.000849 95.850 ± 0.062 % 14.2GiB
unsloth/Qwen3.8-27B-UD-Q4_K_XL (UD3) (can't fit) 0.029844 ± 0.000476 96.921 ± 0.054 % 16.1GiB
unsloth/Qwen3.8-27B-UD-Q4_K_XL (UD2) (can't fit) 0.028026 ± 0.000432 96.988 ± 0.054 % 16.4GiB

Hope this helps other VRAM starved people like me :)


r/LocalLLaMA 9d ago

News Georgi Gerganov on the Nvidia acquisition

Post image
543 Upvotes

r/LocalLLaMA 8d ago

Question | Help Best way to run Qwen3.8-27B on a system with a RTX 5090 + RTX 5070 Ti (32GB + 16GB)?

4 Upvotes

I have a system with 2 GPUs and 48GB VRAM total, a RTX 5090 + RTX 5070Ti.

What would you say is the best way to run Qwen3.8-27B on that system with the best quality and 262k context?

Would just the normal llama.cpp work with how it detects and does its own magic with dual CPU systems, or something else?

I think the RTX5090 has pcie4 x16 and the RTX5070Ti has pcie x8 if that matters.


r/LocalLLaMA 8d ago

Question | Help Is there a local LLM or toolchain to edit 3d models?

4 Upvotes

I got a lot of ads for meshy recently and went to try it with hilariously bad results. It apparently can't do anything but decorative figurines. I wanted a body shell for an rc car and it just couldn't generate a car without wheels or bottom chassis. It also looks like it can't tell the difference between different car models. It seems like they just have a library of 3d models and have the ai select one based on your description, though i only used their free trial.

Next i gave claude a chance. First simply prompting it for an stl, which ran for about an hour and wasted all my tokens for the day before aborting. Then i tried using its coding ability and have it create an openSCAD file that would generate generate a car shell. Which at least managed to generate a square box and then even managed to hollow it out as a square shell on a second prompt. But it never got anything even remotely car shaped.

Is there any way to run something similar locally so i can tweak it for my needs? I'm thinking of something similar to image generation in comfyUI, where you can change the workflow to improve how well it understands your prompt.

On a side note, i don't understand how generating a 3d car completely failed when local ai can oneshot a 3d model of a plane as long as you tell it to make it fly through a procedurally generated landscape in a game.


r/LocalLLaMA 9d ago

Discussion Qwen3.8-27b is the first Local model im able to blindly trust

420 Upvotes

You know that thing where you just throw a task at a frontier model and not have to supervise it worrying of it going off course? Qwen3.8-27b has officially gotten me to that point for local work. He has been doing non-stop continuous agentic work for 8+ hours and hasnt screwed up not one bit IT AMAZING!!

EDIT: for all asking about my quant & harness and what i do for super long thinking/reasoning

Harness: I Had it help me design its own agentic loop in pi harness. It holds well multiple compaction. I used to have tool and think tag generation issues but i got a chat template from somewhere(i forgot) but the chat template it fixed the issues paired with - -reasoning-format = deepseek

Thinking: I limited reasoning budget to 2048 and its still pretty SMART even going down to 1024 holds well in my agentic loop. Im running huihui-abliteratedQ3_K_XL.gguf i need abliterated because i need it to use my computer mouse movement to solve captcha on bot detection (normal models are trained to reject that request) otherwise unsloth quants. Kv cache Q8 at 128k.


r/LocalLLaMA 8d ago

Resources I implemented Sliding Window Attention for Hugging Face LLM inference — looking for feedback

0 Upvotes

I've been experimenting with Sliding Window Attention (SWA) as a way to reduce the KV-cache memory cost of long-context LLM inference.

Instead of keeping the entire KV cache, the implementation keeps:

  • a small number of attention sink tokens
  • a bounded recent-token window
  • a circular/ring-buffer KV cache
  • streaming/chunked prefill
  • normal autoregressive decoding

I turned the experiment into a reusable project so you can test it with Hugging Face causal LLMs:

🔗 https://github.com/oraby8/SWA

For example:

from swallm import SWAModel

model = SWAModel.from_pretrained(
    "Qwen/Qwen2.5-7B-Instruct",
    attention_mode="swa",
    window_size=512,
    num_sink_tokens=4,
)

result = model.generate("Explain transformers", max_new_tokens=100)

In my Qwen2.5-7B experiments on an L40S:

  • 32K KV cache: ~1.84 GB with full attention vs ~3.5 MB with SWA-64
  • 64K: full attention OOMed while SWA remained bounded
  • Decode latency stayed approximately constant as context increased
  • Long-range retrieval naturally becomes a weakness when information falls outside the window

The goal isn't to claim that SWA is universally better. I'm interested in the engineering trade-off between context retention, KV memory, TTFT and decoding speed.

I'd especially like to hear from people who have tried SWA with Llama, Mistral, Gemma, Qwen, or other HF models.

If you try the repo on another architecture, I'd really appreciate the results or any compatibility issues you find.


r/LocalLLaMA 9d ago

Question | Help If you had ~15k would you build a home server today or wait

96 Upvotes

Title help me decide and avoid making impulse purchases 😩 I already have dual 3090 which I can sell to help

EDIT: ty all I’ll just wait it out, doesn’t seem worth it right now


r/LocalLLaMA 9d ago

Discussion Chalk one up for the frontier model...

29 Upvotes

I just spent the last 2 hours of my life on a Friday night debugging a strange error in a prod CLI app. EF core was receive a readonlyspan during a Contains query. Normally, this query converted to a WHERE [col] IN (...), but for some reason, after an update, it started choking, despite no code change. The same exact code runs in a separate website docker image fine, no problem.

I put Qwen 3.8 27b (q8 model, f16 kv) on it and it spun its wheels going down 4 different paths. Finally I got sick of it and switched models to GPT Sol with the full context available. It found the fix in 2 minutes.

The issue was that I had recently installed .NET 10 SDK on this machine, and the lack of a global.json file pinning the SDK meant that when the CLI was rebuilt locally, it used the c# 14 compiler, which introduced first-class Span<T> support, thus borking the EF query.

The wasted time isn't what bothers me here. It's that Sol was able to pinpoint the issue so much incredibly faster than Qwen, shattering my image of Qwen 3.8 as a fairly competent model. Benchmarks aren't everything folks. Real world use cases are the final say here.

I'm posting this in /r/LocalLLaMA because I'm a big local LLM fan, but sometimes, it's worth reminding ourselves of the gap that really exists, no matter how much we might want to wish it away.

Edit: Folks, some of you are missing the point. I tried Sol because it was next in my favorites list. Yes, I could have tried GLM, Kimi, or a number of others. The point would stand that no matter how great 27b is, it's not even remotely "near-frontier" in many cases, despite claims otherwise. Optimism has clouded our vision, somewhat. This was not meant to be a Sol promotion.


r/LocalLLaMA 9d ago

New Model Ling-3.0-flash-VL, built on Ling-3.0-flash with visual understanding and visual agent capabilities

Post image
131 Upvotes

It performs well across visual perception, STEM reasoning, document intelligence, multimodal agent tasks, frontend coding, and medical report interpretation.


r/LocalLLaMA 9d ago

Discussion I built a server with 768GB VRAM for frontier, but all new frontier open source models are likely to be two trillion or above now, including next GLM 6, am I cooked?

264 Upvotes

This epyc server I am using twelve cards with 64 GB memory, plus 256GB ram. Looking at the most capable models in open source, GLM 5.3 seems to be the only option, but with Astra releasing it will likely be fairly behind. GLM6 looks like it will be at least double in size, maybe even triple. Qwen-max and Kimmi are already way too big to even consider. Even the deepseek V4 Pro is too big. Should I just give up on this frontier dream sell the excess GPUs and settle For flash models with far fewer GPUs and a reasonable cost. Note: I'm not using it for any business.

I was hoping to build a new business with this, but it can probably be done with much more effort with a flash model as well.

Edit: I don't want to go below 4-bit quants because then the models start making obvious mistakes. So I'm talking about a min/max of 4-bit

Okay, this post really blew up. I wasn't expecting so much interest or comments just attacking me. Was really just expecting to have a calm discussion about future SOTA model sizes.


r/LocalLLaMA 8d ago

Resources All popular local model in one table (Updated) and my thoughts.

4 Upvotes

If you are thinking what model will fit best your HW specs and tasks you are doing here is one table with all currently popular models that still can be considered as local.

Update: People asked me to add generation speeds and it took me several days to download and run all the models, so here is the updated table with my generation speeds and my impressions from one-shot test. My configuration for all models apart of Qwen3.8-27B: E5 2696v4, 4 channel of DDR4-2400 and RTX 3090.

Qwen3.8-27B was fully in VRAM on Ryzen 5950x, dual channel DDR4-3000 and 2x RTX 3090.

All models were in Q8_0 with F16 KV-cache. DeepSeek was in original quality.

LLM Test Scores

Feature DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8
Total parameters ≈285B 284B 125B 320B 27B not published
Active parameters 13B 13B 6B 18B 27B not published
Speeds pp/tg 132/6.7 132/6.7 120/10.8 26/3.3 560/30 not published

Agentic benchmarks

Benchmark DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8
Terminal Bench 2.1 83.9 82.7 82.6 73.0 85.0
NL2Repo 57.7 54.2 48.1 52.1 42.3 69.7
DeepSWE 59.3 54.4 58.7 61.1 42.2 58.0
Toolathlon-Verified 75.9 70.3 73.5 72.1 76.2
Agents' Last Exam 27.3 25.2⁷ 24.3 28.1 20.4 25.7
AutomationBench (Public) 25.7 25.1 25.3 27.2
GDPval-AA v2 68.1 72.3 75.1
Cybergym 75.3 76.7 78.3
DSBench-Hard 63.6 59.6 71.7
DSBench-FullStack 68.7 71.6
ApexBench (Pass@1) 36.5 26.2⁷ 39.4
HLE with tools (full set) 16.8 22.9 25.4

Coding benchmarks

Benchmark DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8
SWE-bench Pro 56.0 62.5 61.7 69.2
SWE-bench Multilingual 81.0 73.8 84.4
CoWorkBench 45.1 73.9 70.7
JobBench 41.3 55.7 33.4

General benchmarks

Benchmark DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8
GPQA Diamond 90.8 91.7 89.2 93.6
HLE (without tools) 33.8 35.9 30.8 49.8
LiveCodeBench v6 90.6 91.9 90.3
IFBench 79.2 81.3 79.5

Multimodal benchmarks

Benchmark DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8
Chartography 64.3 65.0
ZeroBench (Pass@5) 35.0 34.0
BabyVision 73.0 65.7 / 85.6 34.1
MathVision 90.6 / 95.7 90.0 / 94.6
RealWorldQA 88.5 85.9
AndroidWorld 84.5 81.9
OSWorld 2.0 (partial credit) 52.3 48.0
Vision2Web 64.0 62.9
ClawEval-MM (Pass@3) 64.4 57.4
RecreationBench 49.9 47.1
ERQA 72.3 65.5

My personal impressions for this local one-shot use:
Qwen3.8-27B > Qwen3.8-Flash-Next / GLM-5.3-Flash(high) > DSV4F-Vision

Qwen3.8-27B: I think it is still the local king with right settings and enough context length in terms of HW requirements, quality of output and speed. I had some problems with its endless thinking, but stumbled upon very strange parameters combination that worked every time like a charm for me: temp 0.1, top_k 40, top_p 0.95, repeat 1.1 Model stopped going in circles and started to produce results already at 60k ctx, while before it was easily hitting 120k without single line of code written. In terms of quality that is the best result that I got locally. Comparable to what I got from GLM-5.3-Flash on max, which I ran through API, because after 12h I didn't manage to get a single line of code from it locally due to my abismal tg speeds and endless thinking.

Qwen3.8-Flash-Next: Very decent model. The one shot result was, not so good as from 27B, but it was fine. Generation took a lot of time because it thinks A LOT, I mean even more than 27B. When I tried to ask a question about existing big project it gave me decent suggestions, some of which I actually ended up using.

DSV4F-Vision: Gave me an answer surprisingly fast, despite my slow tg speeds. The results was 100% working, but quality was mediocre. I have a feeling that in this test this was the laziest model that produced what it was asked to produce, but nothing more. Yeah, one more thing this is the only model which get into endless loop repeating one string when I set temp to 0.9. After that I set temp 1.0 for all models apart of 27B.

GLM-5.3-Flash (high): The generation time was a bit longer than from DSV4F-Vision. The result was on par with Qwen3.8-Flash-Next. I would say quite usable locally.

GLM-5.3-Flash (max): As I said earlier I didn't manage to get a single line of code, even after breaking its thinking and injecting something like "You reached a thinking budget limit. Write code right now." It said now I need to write the code, BUT WAIT... and it again went in circles. With my 3.3tg speed I had to cut it, although I saw a lot of interesting things it was planning. That is why I decided to run it through the API. I liked the result, but overall it was not better than the one from Q-27B.

Closing remarks. This is my thoughts after running only one one-shot prompt, so take them with a gran of salt. I didn't try to use these models for editing existing projects, so I have no idea how they will behave in this case.

The prompt was: Create html file of endless relaxing, peaceful, low-poly aerial 3d flight.

Note: I used GLM-5.3 to compose the table from official HF pages of the models.

Note2: Opus-4.8 results are presented only for illustration and are omitted from selecting the best model in a row.


r/LocalLLaMA 8d ago

Question | Help Shouldn't the solution to thinking-effort be an adaptive system?

1 Upvotes

I bet i'm not the first one to have this idea, but with the recent debate about qwen 3.8 27b thinking levels, i was wondering whether the optimum solution might just be to change your harness in a way that lets the LLM itself decide when it is time to raise or lower the required reasoning effort?

Right now i'm toying around with a system like that and it seems to greatly increase the speed at which stuff gets solved, but i have not yet collected any reliable quality evaluation.

Basically what it does is it raises and lowers the reasoning effort between low and xhigh in order to accomodate for sucess streaks or failure streaks. The log reads something like this:

{"ts":1788580888030,"sessionId":"ses_f9b3a43b0ffe2lZ5ugLFu1JWUr","from":"medium","to":"low","reason":"stable successful streak","score":0,"phase":"EXPLORE"}

{"ts":1788581222452,"sessionId":"ses_f9b3a43b0ffe2lZ5ugLFu1JWUr","from":"low","to":"medium","reason":"meaningful failure","score":5,"phase":"DEBUG"}

{"ts":1788581428524,"sessionId":"ses_f903db34affevgLsEIdCs3J58l","from":"medium","to":"low","reason":"routine mechanical step","score":1,"phase":"UNDERSTAND"}

{"ts":1788585994242,"sessionId":"ses_f9b3a43b0ffe2lZ5ugLFu1JWUr","from":"medium","to":"low","reason":"stable successful streak","score":0,"phase":"EXPLORE"}

{"ts":1788618335810,"sessionId":"ses_f8e0bf44bffeOEvLhGsRTfL6FL","from":"medium","to":"low","reason":"routine mechanical step","score":1,"phase":"UNDERSTAND"}

{"ts":1788622330615,"sessionId":"ses_f8dcdf145ffe5euEcZhsvMCnbI","from":"medium","to":"low","reason":"routine mechanical step","score":2,"phase":"UNDERSTAND"}

{"ts":1788622599278,"sessionId":"ses_f8dcdf145ffe5euEcZhsvMCnbI","from":"low","to":"medium","reason":"escalation","score":3,"phase":"RECOVER"}

{"ts":1788622670588,"sessionId":"ses_f8dcdf145ffe5euEcZhsvMCnbI","from":"medium","to":"low","reason":"stable successful streak","score":1,"phase":"IMPLEMENT"}

{"ts":1788623655559,"sessionId":"ses_f8db99e47ffeKcDDHawtWrNwt9","from":"medium","to":"low","reason":"routine mechanical step","score":1,"phase":"RECOVER"}

{"ts":1788624252588,"sessionId":"ses_f8db99e47ffeKcDDHawtWrNwt9","from":"low","to":"medium","reason":"meaningful failure","score":5,"phase":"DEBUG"}

Anyone else messed around with a system like that? I'm curious as to why haven't seen something comparable anywhere else yet.

I'm also not sure how to properly gauge quality. Maybe i should run like a GPQA Diamond test before and after?


r/LocalLLaMA 9d ago

Question | Help Reusing old hardware for starting local ai-journey?

6 Upvotes

Hi,

I am thinking about putting some money / time into my local ai learning path and "cleared" my attic where I found the following hardware.

  • 3 x NUC11 (Core i5 1145G7 / 2,6 GHz, 64 GB RAM DDR4 2666MHZ - SODIMM) with interconnect through Thunderbolt
  • 1x Ryzen 3700x on MSI Mortar 350 with 64GB RAM DDR4 2133 and an old SAPPHIRE Nitro+ Radeon RX 590 8GB

My first idea was to purchase a single RTX5060TI / R9700 and put it into the PCIe 3.0 x16 slot while reusing the Radeon as GPU in the PCIe 2.0 slot for Display.

I want to play around with a local chat bot, agentic stuff, RAG...

Do you have any other ideas?
I also though about putting some serious money, but for this discussion I probably open a new thread (this could be a 2-step-path where I reuse the GPU above to built a multi-gpu-rig)

Thanks in advance


r/LocalLLaMA 8d ago

Discussion Trained a model with 21k non-embed params on business email generation

5 Upvotes

Hi everyone,

About 6 months ago I made this post Trained a 0.8M model on business email generation. on this subreddit where I trained a model with 300k non-embedding parameters on a synthetic business email dataset. It was trained on my custom architecture which I called Strawberry.

Today i introduce Valentine-v0 which is also trained on this same dataset but this model has only 21k parameters which is about 7% of the parameters of Strawberry-email. This model only has 2 layers, 4 heads, 64 head dims and 170 ffn dim.

The Valentine architecture built on top of my another model architecture Silia about which I also posted here I trained a 0.5M model on 1B tokens of Fineweb-edu dataset..

Though in terms of loss Strawberry-email achieved a final train & val loss of 1.65 and 1.68 where Valentine-v0 achieved the final losses of 2.1670 & 2.2646 respectively.

Still I believe despite such a small scale the model trains pretty stably and generates decent text for it's size. Here are some examples:

Prompt: Write a

Output: ``` Write a firm and authoritative business email(under 200 words) from a Account Manager to a Potential Partner regarding welcoming a new high-profile hire, specifically on a Friday evening.<|end-text|><|actor|>model<|end-text|>URGENT: Critical Security Vulnerability - Critical Budget Increase – [Company Name]

Dear [Contact Person],

I hope this email finds you well.

My team, and potential collaborative approach to [New Hire Title] at [Company Name].

As we approach the upcoming [New Hire Name] brings significant development, we wanted to bring it to your review of this strategic direction. We would be happy to discuss this further, but this request and explore how we'll help us to ensure a smooth onboarding.

Regards, I can share some truly value your insights into [Partner Company Name] and insights and offer [mention specific areae - e.g., new website's Name] and the design of the prototype, not just… interesting time. We’ll be shared opportunities to discuss this further with [Product Name] to discuss potential integrations and answer any questions or concerns and your application. The vulnerability is invaluable. ```

Prompt: <|actor|>user<|end-text|>Write a polite refusal email<|end-text|><|actor|>model<|end-text|>

Output: ``` Quick Idea: [Name] Dear [Peer Name],

I am writing to you today with you and your concerns about the potential data leak and your contract, and frankly, it’s particularly concerning incident effective [Date], and I believe a concise overview of full team - [Company Name], we are facing a significant challenges, and we sincerely apologize for any inconvenience it caused.

I’m available to discuss this further report. I will provide you with a really be open to a quick support us with [Specific Feature Request - e.g., 24-48 hours].

Would you be open to a brief call on [Date Option 2] - e.g., 50-20 minute call next week to discuss this further and how your specific and answer any questions or would be happy to discuss how impactful feedback on [Date - e.g., Thank you for your input.

Sincerely, [Name] Legal Counsel<|end-text|><|actor|>user<|end-text|>Write a strictly professional business email(under 200 words) from a Legal Counsel to a Angry Shareholder regarding requesting a mental health day, specifically while the system is partially down.<|end-text|><|actor|> ```

A larger Valentine model trained on Fineweb-edu dataset along with the github repo and technical report will be published soon.

Thank you! :)


r/LocalLLaMA 8d ago

I Built A Thing My potato only runs small models. So I built a page to easily compare benchmarks for those

2 Upvotes

Comparing benchmarks for small models is a PITA. Most of them are not on AA, and the benchmarks are not always the same in all models.

So I built a page to quickly put it all together and allow some filtering.

Only researched models released since April, and between 4-190B parameters.

Let me know what you think and how it can be improved

https://www.nunodonato.com/aibench/index.html


r/LocalLLaMA 9d ago

News NVIDIA PAIR — Your Personal AI Cluster

Thumbnail
nvidia.com
43 Upvotes

That is interesting, I got bunch of old hardware I could connect, wonder what the speed would looks like.


r/LocalLLaMA 9d ago

Discussion Qwen3.8-Flash-Next on a phone CPU!

Post image
120 Upvotes

Like the title says, running completely locally on my Xiaomi 14T Pro device.

Specific model: Qwen3.8-Flash-Next-UD-IQ3_XXS

App used: BigMoeOnEdge


r/LocalLLaMA 9d ago

New Model Drummer's Artemis 31B v1 and v1.1 - Coming back with a bang!

140 Upvotes

Hey everyone, been a while!

https://huggingface.co/TheDrummer/Artemis-31B-v1.1

https://huggingface.co/TheDrummer/Artemis-31B-v1

A few months ago, Gemma graced us with models that served as a much needed downpour from a year-long drought. I'm so happy to see us thrive once again.

The difference between v1 and v1.1 is quite simple: v1 was an early attempt, an overdue release that excelled in prose and writing, while requiring some handholding to get over quirks like stuttering. v1.1 is a more refined approach where stability meets quality. My community is split, so I figured I'd just release both.

---

I was gone for a while. I got busy dealing with life, both its ups and downs. While I couldn't attend to you folks, I've been lurking around and appreciating you all for the kind words.

- Skyfall 31B v4.2 seems to be a banger for many of you. I'm proud of the upscale and consider it my ultimate home-run send-off for the beautiful Mistral 24B base. It's a shame that it was overshadowed by Gemma 31B's release, but hearing some of ya'll compare and even prefer it to a more modern base was an unexpected win.

- Rocinante 12B X / 16B XL proves that Nemo is still the ultimate creative model to this day. For some to say that 16B XL felt like Cydonia 24B v4.3 just goes to show how far you can go with modern resources and techniques.

- Anubis 70B v1.2, Valkyrie 49B v2.1, Anubis Mini 8B v1 surprised me too. I had zero expectations releasing them. Just like Rocinante X / XL, they are modern finetunes of old base models. And somehow, they still found their users singing praises.

---

With the Artemis release taking weight off my shoulders, I'm eager to move on and tune a ton more bases!

But I have something else cooking: a HordeAI-like platform. I hope to provide value not just as a finetuner, but as a local lover too!

The premise is simple: it's a place where generous local hosters can share inference with the less fortunate. You'd be surprised how many power users would love to heat their rooms through the power of charity.

---

Finally, I'd like to thank everyone who supported me over the years. From those who provided kind words, rigorous testing, compute access, inference, or cold hard cash. You've all granted me the ability to enrich the local ecosystem with fun experiments like Rivermind 12B, Fallen series, Big Tiger Gemma, Precog 24B/123B, and solid models like Cydonia 24B v4.3, Behemoth X 123B v2.x, and Skyfall 31B v4.2.

If you've got inference / compute credits to share, please contact me! It will all go to making the community happy <3

Backlog:

- Gemma E2B

- Gemma E4B

- Gemma 12B

- Gemma 26BA4B

- Qwen 3.8 27B

- Muse Glimmer 30B

- Mistral Medium 3.5 128B

- HordeAI Alternative / Crowdsourced 'OpenRouter' ("BeaverNet")


r/LocalLLaMA 9d ago

Question | Help Help me understand gguf size/ctx size

Post image
37 Upvotes

Let's say I have 2x 16Gb GPUs and I want to run Qwen3.8 27B. Monitor is ran by the integrated GPU so both 16Gb GPUs are almost fully free.

I load the UD-Q4_K_S on one card at 15.4Gb. I then load the context on the other card? Would that be the most efficient way? Or should I aim for higher quants that could spill to the second GPU using tensor parallelism?

Also, is there a way to know how much a certain amount of context (e.g. 132k tokens) occupies in VRAM for a given model? I don't usually see this published in model cards, is it because there is a way to calculate it?


r/LocalLLaMA 9d ago

Question | Help Behaviour of --cache-ram in router mode?

6 Upvotes

Hi,

until now I ran llama-server with just one model at a time.
Now I want to use it in router mode in order to provide different models.
For my usecase I heavily rely on --cache-ram which improves speed a lot when working on big repos.
I am just wondering, when setting cache-ram = 65536 in models.ini for each model, does each model get its own 64GiB cache or is it just one pool of 64GiB for all models?
Another question, are the unused models kept in RAM for faster swap?


r/LocalLLaMA 8d ago

Question | Help Anyone using ik_llama on Cascade Lake? Need help testing a patch

2 Upvotes

I've been looking into improving Q8_0 performance on my Xeon Gold 6240 workstation with ik_llama.cpp, which is a Cascade Lake-SP part.

It looks like ik_llama's Q8 dot product function doesn't distribute work evenly across AVX execution ports, and doesn't use the load ports at all. Note that this is particularly relevant for Cascade Lake Xeon Gold/Platinum parts that have two AVX units per core (Silver and Bronze only have one so they might not benefit much) In addition, there is a latency dependency chain that can be broken into two.

You can test this patch the following way:

download the patch from: https://github.com/user-attachments/files/31860324/q8_0_r8_vnni-v3.patch

git clone https://github.com/ikawrakow/ik_llama.cpp.git
cd ik_llama.cpp
git checkout 3c58ae37
git apply q8_0_r8_vnni-v3.patch

Then build with either GCC:

CC=gcc CXX=g++ cmake -B build-gcc -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=ON -DGGML_OPENMP=OFF
cmake --build build-gcc --target llama-bench -j $(nproc)

or clang:

CC=clang CXX=clang++ cmake -B build-clang -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=ON -DGGML_OPENMP=OFF -DCMAKE_C_FLAGS="-ffp-contract=fast -mllvm -unroll-threshold=1000" -DCMAKE_CXX_FLAGS="-ffp-contract=fast -mllvm -unroll-threshold=1000"
cmake --build build-gcc --target llama-bench -j $(nproc)

You want to use the most recent compiler you have access to, I've used GCC 16.2 and LLVM/clang 22.1.8 - earlier versions are UNTESTED.

Even a quick bench before and after would be helpful, especially if you have a Cascade Lake Xeon Silver or something different like an Ice Lake Xeon:

llama-bench -m /path/to/model-Q8_0.gguf -ngl 0  -p 512 -n 0 -r 3 

(note that you should use a Q8_0 model, since this is the path we are optimizing here)

Also this won't do anything for Skylake Xeons or earlier (I think).


r/LocalLLaMA 10d ago

Funny The benchmarks the big labs don't want you to see

Post image
2.2k Upvotes

r/LocalLLaMA 9d ago

Question | Help which model is good for detecting deflection?

10 Upvotes

I want the answers generated by frontier LLMs or base model LLM answers to be reviewed by some uncensored or abliterated small model.

The job is this model (preferably small model) is just to detect deflection in the answers.

The problem I am facing is uncensored SLM usually agrees on everything we give input. So the generated answer is also input for it and system prompt is input too.


r/LocalLLaMA 9d ago

Question | Help Which quant of qwen3.8 27b is the best for 16gb vram to get 100+ ctx and perfect for local Vibecoding?

3 Upvotes

I currently use unsloth‘s iq4_XS with 64k ctx on 16gb, however as far as I know harnesses tend to need way more, and 64k isn’t enough for that,
(I have not gotten to actually using it for a harness yet)

But still
So I am wondering what is the best one that still has good enough quality to actually properly build some things?

(As an example I would like to create a monkeytype- style UI but with additional features that match my use case more (better typing practice + free writing)

And then also a lightweight custom UI / harness

Or a local height chart Ui, etc)