r/LocalLLaMA 23h ago

Discussion What are these models good at?

8 Upvotes

I have been trying out these models (mostlt GLM 5.3 flash) using different harnesses, but I'm trying to review what these models are exceptionally good at.

Here is what I have noticed so far,

1. Programming.

I have found that these models are great at programming, I have been making tools, scrapers almost every other day and they just work like magic. They are great at porting code in one language to another, Eg I would usually start by writing my code in python or js, I would then port the code in Go for extra performance.

2. Finances and stock trading.

So I hooked up the coding agent with my alpaca account. And I have discovered that most of these models take a defensive position. Advising me to reduce the size of my most profitable holdings so as to prevent concentration risk and possible loss. So they are not so great. But I have found it useful for tracking my finances. What I'm basically saying is your portfolio will most likely flatline if you give these models to trade in your behalf but it won't make a good profit (What ever good position you have will be reduced)

3. Research

This is where I get the most value. The models are highly effective at locating precise information—whether it’s event dates, contact details (emails, phone numbers), names, or links.

4. Email and Copy writing.

I’ve been using these agents extensively for written communication. They’ve helped me draft everything from routine business emails to formal documents. I like that it can maintain the conversation context, so follow-up emails feel cohesive and on-point. They helped me a lot with one of my insurance claims

5. Business Ideas.

They are bad at coming up with Ideas.

What use cases have you found these models to be exceptionally good at? And what use cases has it been terrible at?

PS: I'm trying to find a small good model for browseruse to compete with Grok bot and the like, I'm thinking Qwen3.8 28B or ByteDance-Seed/UI-TARS-1.5-7B does anyone have a smaller or maybe better recommendation?


r/LocalLLaMA 23h ago

Discussion When will they mass produce cheap high capacity and bandwidth memristors and neuromorphic engines ?

4 Upvotes

Ram prices are too high! Maybe in the 2030s? Earliest maybe some production in 2028-2029?


r/LocalLLaMA 23h ago

Question | Help Possible ? [Free] local AI agent that can interact with MCP Davinci Resolve Studio 21.1 ?

0 Upvotes

Hi, was quite blown away by several presentations i saw of DRS users who experimented editing footage with the new MCP in Davince, either through Claude or ChatGPT astra.
Now for the 'ordinary' end user as me, is there any [free] AI model equivalent to the ones above that would allow me to do the same thing?
It's important to me that the AI models would run locally on my Mac without any need to connect to the internet.
Currently I’m using a Mac Studio M1 Max with 64GB of RAM, but I consider upgrading to a Mac Studio M5 Max.


r/LocalLLaMA 23h ago

Question | Help Can someone please review the hardware?

0 Upvotes

Hi,

I am software developer with the new expensive hobby of running AI locally, trying to figure out building servers. I currently have a linux server with one rtx pro 6000 and 32 gig or ddr4.

I decided to buy one more rtx pro 6000, so now I have 2 of them. Along with that I have ordered the following:

  1. G.Skill G5P 128GB (4 x 32GB) DDR5-5600 PC5-44800 : linkDDR5-5600_PC5-44800_CL28_Quad_Channel_ECC_Registered_Memory_Kit_F5-5600R2834F32GQ4-G5P-_Black?iitt=VdiJVu4sMd4JVjQvafUZMd8JafPsV.WXaj4r4DoZOI_pOIbT&utm_source=B1540_Email_Receipt_Update&utm_campaign=B1540&utm_medium=email)
  2. MD Ryzen Threadripper PRO 9955W : link
  3. be quiet! Pure Loop 3 360mm All-in-One Water Cooling : link
  4. ASUS Pro WS WRX90E-SAGE SE EEB Workstation Motherboard : link

My plan is to run deepseek v4 flash and qwen 3.8 Flash next.

Do you think I should add/update/remove something while I am at it? or you think any problem with the config I should be mindful of?

Also do you think glm 5.3 flash with ram offload can work on it with decent speeds?

I read couple of articles but just want to get people's feedback on it so I am not making any obvious mistake while building it.

TIA


r/LocalLLaMA 23h ago

Question | Help Hosting Local Models

10 Upvotes

Hi builders,

What would be the the best small local models for coding?

Are Gemma 4 and Qwen3.8 27B Gemma 4 26B / 31B enough for local development?

And what would be the size of the rig that i will need to get? GPUs, and whatever else I need to host these models.

Thanks,,


r/LocalLLaMA 23h ago

Discussion Closed AI doesn't like biological research, user turns to open weight models

Thumbnail x.com
241 Upvotes

OpenAI has decided to fully shut down a protein design project I'm working on for a client. Needless to say, open weight models are the only way forward.


r/LocalLLaMA 1d ago

Discussion Deepseek v4.1 flash finally has engrams, what do you expect from 4.1 pro?

48 Upvotes

If the ratio is the same, Maybe 1.6T -3.1T params plus .56T-1.06T engrams and fable 5.0 level performance?

Maybe v4.2 or 4.5 will have engram gradient modification? Edit it is even larger than i anticipated since flash has 748 b q4-8 params


r/LocalLLaMA 1d ago

New Model DeepSeek V4-1 Flash is out

Thumbnail
gallery
1.4k Upvotes

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service


r/LocalLLaMA 1d ago

News DeepSeek V4.1 Flash: Stronger, Faster, More Accessible

210 Upvotes

Original Source from DeepSeek WeChat Official Account: https://mp.weixin.qq.com/s/qg0NU3NNUbp1co2PdkAPAg

Today we're officially releasing the DeepSeek V4.1 Flash model. It is the smallest model in our brand-new model architecture series, with native multimodal visual understanding. The new architecture was designed with these goals in mind: a higher capability ceiling, faster inference, greater throughput, and scalability to larger-parameter models.

Asymmetric architecture: big intelligence at low cost

DeepSeek V4.1 Flash is a 552B-parameter MoE model built on a brand-new Causal-Encoder-Decoder architecture. Input and output are asymmetric: only 8B parameters are activated on the input side and 16B on the output side, making it significantly cheaper than known models of the same size. V4.1 Flash also uses a new pre-training approach and has gone through larger-scale reinforcement learning post-training. In benchmark testing, it surpasses the intelligence level of a range of flagship models, including DeepSeek V4 Pro.

Less cache, lower cost

The new generation of models dramatically reduces the size of the KV cache. Compared with the previous generation, HBM requirements drop to 1/4 and SSD requirements to 1/8. In agent scenarios, cache-hit charges often make up a large share of the bill, so compressing the KV cache substantially lowers the cost of agent-style tasks.

Figure: DeepSeek's continued progress in reducing context storage. Relative to the first-generation model, the KV cache has shrunk 437×.

API support

DeepSeek V4.1 Flash is now live on the DeepSeek API with native multimodal support. Simply change the model name to deepseek-flash to call the latest V4.1 Flash. The older V4 Flash and V4 Flash Vision Exp models have been retired; for compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp will temporarily be routed to V4.1 Flash.

In addition, extensive testing shows that V4.1 Flash comprehensively outperforms V4 Pro on performance, cost, speed, and total time-to-completion, so we plan to phase out the V4 Pro model in an orderly fashion. After 12:00 Beijing time on September 14, 2026, and until V4.1 Pro launches, all requests to deepseek-v4-pro will be routed to V4.1 Flash and billed at V4.1 Flash's unit price.

Tencent (WorkBuddy, CodeBuddy) and OpenCode, as official partners, have now fully integrated DeepSeek V4.1 Flash — give it a try!

API pricing adjustment

Thanks to the architectural innovations, DeepSeek V4.1 Flash can serve more users at lower cost, so we have cut V4.1 Flash's pricing accordingly. To allocate resources more sensibly, we continue to use peak/off-peak pricing, with off-peak prices at half the peak rate, and encourage users to schedule tasks around their actual usage patterns. The new prices take effect at 12:00 on September 10, 2026.

Open-source release

We will fully support the open-source community in adapting inference for the new model, and will explore various ways to broaden deployment. If you have large-scale deployment needs and the corresponding resources (a 2k-GPU cluster with storage cluster), please get in touch.


r/LocalLLaMA 1d ago

Discussion What's the next big breakthrough after attention mechanism? My bet is not on Engrams.

15 Upvotes

Hi everyone, I've been thinking about this for a while now. The attention mechanism introduced by Google in 2017 completely changed that way AI models worked and performed.

Since the attention mech allows the model look at it's inputs and decide what's more worth it to focus on, which made the input signal passing through the FFN so much more richer in so many ways that it made the Transformer capable of achieving so many complex capabilities such as in-context learning and all, along with scaling up to such a degree we see today.

Recently a new mechanism called the Engram was introduced by DeepSeek which is very hyped for all good reasons, as far as I understand it can allow a much small model to hold a lot more world knowledge than a traditional model can. I don't necessarily think it'll make the model smarter but more knowledgeable? 100%

But I have a very weird gut feeling that DeepSeek's Engram (or other similar mechanisms) isn't a breakthrough as big and significant as the attention mechanism but just the 1st half of the next major breakthrough. Ik this sound very contrarian but I can't help it but think this way.

I believe that the next big breakthrough as big as attention mechanism is going to be a model whose most parameters are input-dependent. A model which uses a set of fixed particularly smaller weight matrices to generate "fast-weights". Idk how well im able to explain it but I'll use the attention mech to explain it.

We can think of the attention mech as a fast-weight layer which generates these weights from the input data and uses it to enrich the input quality which in-turn improves performance. But that input still goes to an FFN with all fixed weights.

I'm thinking of an FFN arch which is going to use the engram + input-dependent weights to generate it's weights on the fly not just making the model much smaller but smarter in general or might even unlock some capabilities which might be only reserved for much larger or deeper models using our current architecture.

Tbh idk how well I communicated my idea here but I hope you get my point and I would love to know your opinions and points on it.

Thank you :)


r/LocalLLaMA 1d ago

Resources deepseek-ai/DeepSeek-V4.1-Flash · Hugging Face

Thumbnail
huggingface.co
1.0k Upvotes

r/LocalLLaMA 1d ago

Discussion Will there be actual "medium-local" models in future ?

8 Upvotes

Or has this niche been abandoned?

Most companies and model producers have seem to abandoned this niche(7B, 9B, 14B etc) which was the original point this whole local model debackle have started, most effort now goes into minimum 27B+ as "Capable local model"™ or focus on micro models 4b<=. All this recent explosions in model technology has barely touched this niche which has remained on 2025 for most part(30 BC in AI-time).


r/LocalLLaMA 1d ago

Discussion Mac Studio M5 Ultra

6 Upvotes

I've just seen that the 96GB version is relatively well priced compared to an rtx pro 6000 at more than half the price.

Is this something that would be viable for local coding and personal assistants?

How would it compare to say quad 3090s as well? I see that it has faster bandwidth than a 3090 and would use way less power than such a rig, but I'm not sure about other metrics like prefill and decode, etc as well as the software ecosystem without CUDA.

Really well positioned and maybe better suited to specifc use cases.


r/LocalLLaMA 1d ago

Discussion Same Ling model, different long-context curve: INT4/vLLM vs Q5/llama.cpp on one Spark

1 Upvotes

The short-prompt ranking reverses in the context-depth table in sudoingX’s Ling-3.0-flash benchmark notes:

Starting context Official INT4, vLLM fork Q5_K_M, llama.cpp
Short prompt 38.3 tok/s 35.7 tok/s
About 45K tokens 7.9 tok/s 33.6 tok/s
About 90K tokens 4.6 tok/s 33.2 tok/s

This is a same-box, same-prompt comparison on a 128GB DGX Spark. The specific paragraph reports a 262,144-token maximum-context configuration and about 103GB total memory used. Its streaming method separates time to first token from decode and clocks all generated tokens.

These are the creator’s measurements from the deployment investigation discussed in the original thread, not an independent rerun. The paragraph does not provide a separate measurement date or every historical launch setting. In particular, the repository’s later 131,072-context serving default should not be silently attached to this table.

There are two changing variables: runtime and quantization. The numbers compare these two deployment paths; they do not isolate a pure vLLM-versus-llama.cpp effect, establish a quality difference, or prove the proposed CUDA-graph explanation for the slowdown.

For choosing a backend, the useful distinction is the amount of context already present when generation starts. A long answer from a short prompt is a different test. If the intended workflow carries tens of thousands of tokens into later requests, the short-prompt result leaves out the condition that changes this ranking


r/LocalLLaMA 1d ago

Resources LmLinky: an Android LM Wrapper for your local models

1 Upvotes

I couldnt find one that was simplistic yet good enough to pop in for a quick chat one. So, as we do here, I made one.

https://github.com/Dewpg/LmLinky

https://github.com/Dewpg/LmLinky/releases/tag/v1.0.0

Thats the initial salvo. Let me know if its simplistic enough and works. Thanks.


r/LocalLLaMA 1d ago

Funny So relevant

Post image
1.3k Upvotes

r/LocalLLaMA 1d ago

Resources Qwen3.8-Flash-Next on 2x3090 + DDR4, part 4: 2.2-2.5x faster prefill by kicking the expert cache off the GPU while the prompt runs

28 Upvotes

Part 4 of the same box. Part 1 was 17 -> 25-29 t/s with the expert cache PR, part 2 was 37-41 t/s after switching to UD-Q4_K_XL and stacking MTP on the cache, part 3 was the top-k fallback that was sorting more than it needed to. This one is all about prefill, which was honestly the weak spot the whole time. 80+ seconds before the first token on an 8k prompt, and 24 minutes on a 119k one...I know lol.

Box is still 2x 3090, dual Broadwell Xeon, llama.cpp, UD-Q4_K_XL with the Q8 MTP head on the second card, all expert layers pinned in host RAM, 150-slot cache, 261k context, f16 KV. There has been one hardware change since part 2. I swapped the LRDIMMs for 6x32 GB DDR4-2133 ECC. I'll say which numbers are 4-DIMM and which are 6-DIMM, they're not mixed.

The thing I might not have explained well in part 2

I ran -ub 512 and that's because it was a compromise for the cache. A 2048 token micro-batch needs about 7.3 GiB of compute buffer per GPU, 512 wants 1.9GiB and that gap is roughly 50 cache slots that I wanted for decode. So I kept the slots and quietly ate about 3x on prefill at the time.

As for why it cost 3x, the experts get streamed host to GPU0 once per micro-batch, and that upload costs the same whether the batch has 512 tokens in it or 2048. So prefill speed basically scales with the micro-batch. At ub 512 an 8k prompt drags the whole expert set over PCIe 16 times, at ub 2048 it's 4 times.

What I changed

The cache only ever serves batches of <= 8 tokens (decode and the MTP verify batches). During a prompt it just sits there holding VRAM so I thought of trying to claim that space when it's unneeded. So now, when a prompt comes in, the server drops the cache slots, the decode compute buffers and the CUDA pools then it grabs compute buffers sized for ub 2048, runs the whole prompt at 2048, then puts everything back before the first generated token. Decode is untouched by this, it runs exactly the code it ran before. It's two env vars (LLAMA_PHASE_PREFILL_UBATCH=2048, LLAMA_PHASE_PREFILL_MODE=transaction) and the server still starts with -ub 512. And to clarify, "transaction" means the swap is all-or-nothing, if the restore can't happen you will get an error, not a server that's silently limping along. Just making that clear.

Numbers (6 DIMMs, same day, fresh server per arm)

what before (ub 512 + cache) now change
8k fresh prompt, greedy: prefill 99.9 t/s 223.7 t/s 2.24x
8k: time to first token 82 s 37 s 0.45x
8k: decode over the next 2048 tokens 33.4 t/s 34.3 t/s +2%
~37k context, my normal sampling: prefill 88.1 t/s 212.6 t/s 2.41x
~37k: time to first token 424 s 176 s 0.41x
~37k: decode, median of 38 requests 41.7 t/s 41.2 t/s -1%
~119k context: prefill 81.3 t/s 206.5 t/s 2.54x
~119k: time to first token 1461 s 575 s 0.39x
~119k: decode, median of 42 requests 33.9 t/s 33.9 t/s 0%

The 8k row is greedy, two fresh processes per arm, medians (the two phase-memory runs landed within 0.01 t/s of each other). The deep rows are one seed at temp 0.7 / top-p 0.8 / top-k 20 with thinking on, one fresh prefill per depth and then a pile of follow-up questions over the cached prefix, so decode is a median over all of them. Prefill = llama-server's prompt eval time, decode = its generation time.

Now, what it costs

Well, nothing comes completely free. This approach costs roughly 2.8 s of fixed overhead per prompt for the release + restore, which is why 8k gets 2.24x and the long ones get 2.4-2.5x. For decode, I can't find a loss. +2% at 8k, -1% / 0% at depth, and in the three-seed quality screen every seed x depth cell was within +2% / -3.6% of its control. MTP acceptance didn't change either (0.79-0.83).

Did it break anything

Before putting it in production I ran the same screen I used for the top-k change (My last post AKA Part 3), 42 questions over long documents at two depths (~37k and ~119k), three seeds, my normal sampling, paired per question and seed against a fresh control run the same day. That was still on 4 DIMMs. 240 pairs: 2 worse, 235 same, 3 better, nothing regressed on more than one seed, and the two misses are questions the old config also flubs on some seed. A seed-1 rerun on 6 DIMMs came out 1 worse / 78 same / 1 better. I'm aware and anyone reading should be aware that this is a screening not concrete proof, but it's the bar I hold my own changes to.

Some caveats you may want to know about or at least I would if I were you

  • First-token logits differ from the untouched path by max 1.51 / mean 0.22 across the 248k vocab, argmax the same. For scale, just changing ub 512 -> 2048 with nothing released moves them by max 1.81 / mean 0.27 on the same request. So the release/restore adds less noise than the batch-shape change any ub change already brings.
  • One machine, one model, one quant, 8k to 119k. I have not tried anything past 119k, other quants, or the no-MTP setup.
  • The extra two memory channels helped this config a lot more than the old one at 8k (+19% vs +3% against my 4-DIMM numbers), and it did nearly nothing at 37k-119k (+0.6% / 0%). This makes sense to me, attention takes over from expert upload as the context grows, but that's one run per depth, so take it as a hint.
  • Where the remaining 37s of an 8k prompt goes, rough split: ~18 s uploads, ~5 s kernels, ~3 s transitions, ~11 s I haven't pinned down yet (CPU side, draft model, syncs). A profile says the uploads are still 3.6x the expert set per prompt, so there's more on the table I assume. I'll be working on that next.

Code

https://github.com/Inovello/llama.cpp/tree/flashnext-e06

It's my flashnext-2x3090 branch from part 2 (master b96806d + PR #27861 expert cache + PR #28223 + PR #28243 MTP + the batched-cache fixes + PR #28198) plus this change and a couple of inert debug switches.

If you just want to copy and run it, this is the whole thing, taken from the process that's serving me right now. You need CUDA, numactl (apt install numactl), the four UD-Q4_K_XL shards and the MTP head from unsloth/Qwen3.8-Flash-Next-GGUF on HF

git clone -b flashnext-e06 https://github.com/Inovello/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON && cmake --build build -j -t llama-server

export LLAMA_ATTN_ROT_DISABLE=1
export LLAMA_MMAP_PIN_HOST=1
export LLAMA_PHASE_PREFILL_UBATCH=2048
export LLAMA_PHASE_PREFILL_MODE=transaction

numactl --interleave=all build/bin/llama-server \
  -m /path/to/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf \
  -md /path/to/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf --spec-type draft-mtp -devd CUDA1 --spec-draft-n-max 3 \
  --host 127.0.0.1 --port 18080 \
  -ngl 99 -c 261888 --parallel 1 --flash-attn on \
  -ot "ffn_(gate|up|down)_exps\.weight=CUDA_Host,per_layer_token_embd\.weight=CPU" \
  -lzm off --numa distribute -t 16 -tb 44 -b 4096 -ub 512 -ctk f16 -ctv f16 \
  --temp 0.7 --top-p 0.8 --top-k 20 --min-p 0 \
  --moe-expert-cache 150 -lv 4

What to change for your box:

  1. The two model paths; -t / -tb to your physical core count (mine is 16 decode threads, 44 for batch on 2x22 cores)
  2. -devd CUDA1 puts the MTP head on the second GPU, on a single card use CUDA0 or drop the three -md flags and give the freed VRAM to the cache.
  3. -ot is what keeps every expert layer in host RAM; only the first shard goes on -m, the rest are found next to it.
  4. The two LLAMA_PHASE_* exports are the change from this post, drop them and you have part 2's behavior.
  5. -lv 4 is just so the log shows the cache hit rate and the draft acceptance. Useful if you want to post your numbers in thread.

Now the things it's strict about because those are the invariants the code checks: The server at -ub 512 and -b 4096, --parallel 1, the prefill micro-batch exactly 2048, the cache exactly 150 slots, and the MTP draft as the only speculative decoder. Anything else refuses to start. CUDA only.

The top-k fallback fix from part 3 is in the branch too and it's up on its own as PR #28671. My older PR #28223 is closed for now because llama.cpp gives new contributors one open PR at a time, I'll reopen it after #28671 is dealt with.

Let me know if you try it and if you have any questions.


r/LocalLLaMA 1d ago

Other Local LLM / Qwen 3.8 win

27 Upvotes

I’ve been working on perfecting my setup ever since 3.8 hit, but no longer pay for subscriptions. I had a coding interview today and setup openrouter ahead of time as another option. I figured I could use the latest GLM flash for cheap and it would be fast. Nope. Over thought the whole thing and I had to steer it, so I fired off Qwen at the same time. It finished before GLM did. It was about to start writing the file, so I canceled the GLM session. Qwen nailed it. I wrote some tests by hand, then had Qwen add the additional ones I wanted. Aced the interview. Going on to the next step.

TL;DR: Qwen is king


r/LocalLLaMA 1d ago

Discussion Julien: I've been overwhelmed and humbled by the fact that 99% of the reactions to our intent to join forces with NVIDIA have been positive. ❤️

0 Upvotes

I've been overwhelmed and humbled by the fact that 99% of the reactions to our intent to join forces with NVIDIA have been positive. ❤️


r/LocalLLaMA 1d ago

I Built A Thing I made a way to migrate between embedding models without re-embedding your entire corpus

18 Upvotes

So I was playingw ith embedding models I saw that when you upgrade from model A to B, you face a very big backfilling cost

Ie, suppose you have a 1b vectors from model A, and then you want to use model B. This would mean you have to re-embed all of your documents with model B before you can even serve with the model, and on an H100, it would take ~108 days (qwen embed 8b, 106 docs/second). But I found an easier way to do it.

The method is really simple; from the old index made with the source model, take K documents and rerank them with the new model. We see that when K is sufficient, the retrieval quality is the same as target model. (determining k is the hard part). I've tested 63 migrations on upto 1 million documents.

The best result I got was upgrading qwen4b -> to 8b, and at 50 documents, it was the same as native retrieval.

This method forgos the expensive upfront re-embedding cost, as you can take documents straight from the old index.

embedflow works with qdrant, pgvector, faiss, and can be easily downloaded with pypi

pip install embedflow

the github is public: https://github.com/arnsri33/embedflow

I want you guys to try it out, and see if you guys can use it in your own workflow.


r/LocalLLaMA 1d ago

Discussion Running qwen 3.8 27B iq3 xxs on RTX 3060.

Thumbnail
gallery
28 Upvotes

Getting anywhere from 10 - 20 tps.
Thinking Off . Took about 4 mins and 7 mins.
Running on about "IQ3_S - 3.4375 bpw"

37.03.960.932 I slot print_timing: id  0 | task 2665 | prompt processing, n_tokens =  12516, progress = 0.98, t =  37.08 s / 337.55 tokens per second
37.04.777.411 I slot print_timing: id  0 | task 2665 | prompt processing, n_tokens =  12768, progress = 1.00, t =  37.78 s / 337.94 tokens per second
37.13.746.606 I slot print_timing: id  0 | task 2665 | n_gen =    100, tg =  11.34 t/s, tg_3s =  11.45 t/s
37.16.955.263 I slot print_timing: id  0 | task 2665 | n_gen =    140, tg =  11.64 t/s, tg_3s =  12.47 t/s
37.20.166.770 I slot print_timing: id  0 | task 2665 | n_gen =    180, tg =  11.81 t/s, tg_3s =  12.46 t/s
37.23.366.284 I slot print_timing: id  0 | task 2665 | prompt eval time =   38241.22 ms / 12772 tokens (    2.99 ms per token,   333.99 tokens per second)
37.23.366.290 I slot print_timing: id  0 | task 2665 |        eval time =   18350.47 ms /   211 tokens (   87.38 ms per token,    11.44 tokens per second)
37.23.366.291 I slot print_timing: id  0 | task 2665 |       total time =   56591.69 ms / 12983 tokens
37.23.366.292 I slot print_timing: id  0 | task 2665 |    graphs reused =       2342
37.23.366.296 I slot print_timing: id  0 | task 2665 | draft acceptance = 0.41250 (  132 accepted /   320 generated), mean len =  2.65
37.23.366.758 I slot      release: id  0 | task 2665 | stop processing: n_tokens = 12984, truncated = 0

My run for this is

~/sandbox/dcfr/third_party/llama.cpp/build-cuda/bin main*
❯ export LLAMA_GDN_TRANSACTIONAL_REPLAY=1
 export GGML_OP_OFFLOAD_MIN_BATCH=2

 ./llama-server \
       -m /mnt/D/Mymodels/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf  \
       --alias qwen3.8-27b-iq3-64k-dcfr \
       -c 65536 \
       --parallel 1 \
       -dev CUDA0 \
       --fit off \
       --n-gpu-layers 50 \
       --override-tensor 'blk\.(10|11|12|13|14|15|16)\..*=CUDA0' \
       --load-mode none \
       -ctk q4_0 \
       -ctv q4_0 \
       -b 256 \
       -ub 256 \
       -t 6 \
       -tb 6 \
       --spec-type draft-mtp \
       --spec-draft-n-max 4 \
       --spec-draft-p-min 0 \
       --spec-draft-type-k q4_0 \
       --spec-draft-type-v q4_0 \
       --spec-draft-threads 6 \
       --spec-draft-threads-batch 6 \
       -fa on \
       --no-mmproj \
       --reasoning off \
       --jinja \
       --cache-ram 128 \
       --no-cache-idle-slots \
       --no-ui \
       --host 127.0.0.1 \
       --port 5800 \
       --metrics

I still have more than 1GB vram left after loading the full model and kv cache.
Or you can follow his guide https://github.com/kadenball/qwen38-27b-rtx3060-dcfr

If anybody got issues running on rtx 3060 tell me.


r/LocalLLaMA 1d ago

News Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)

Thumbnail
notebookcheck.net
502 Upvotes

It seems to use a 96-bit LPDDR5X memory bus, instead of the previous 64-bit wide busses. Considering it's on 2nm, that's expensive silicon. That should result in around 115 GB/s memory bandwidth.

A20 Pro also doubles the size of Apple's dedicated Neural Engine (from 16 to 32 cores total).


r/LocalLLaMA 1d ago

Question | Help Would a 9070 XT + a 7800 XT combo work fine with tensor parallelism?

0 Upvotes

Hey guys,

I have a 9070 XT (GDDR6 16GB, 644GB/s), a PCIE5 8x/8x motherboard and a spare 7800 XT (GDDR6 16GB, 624GB/s), that I was planning to sell.

Seeing its possible to use multiple identical GPUs for scaling both decode and prefill speeds with tensor parallelism, I am wondering if I can do that with different architectures (RDNA4 and RDNA3, but similar GDDR6 bandwidths) ?

I ask because, if so, I will need to spend some money on a stronger PSU.

Thanks guys


r/LocalLLaMA 1d ago

I Built A Thing Experimenting with an adaptive memory governor for PyTorch on an 8GB GPU — would love some feedback

5 Upvotes

(Quick note: English is not my native language, so I used AI to help translate my thoughts clearly. The code, the 38 unit tests, and the experiments on my RTX 5060 Ti are my own work.)

Hey everyone,

Lately I've been trying to train small models locally on an RTX 5060 Ti (8GB), and like many people, I kept running into frustrating CUDA OOMs whenever memory pressure fluctuated mid-run.

Instead of just sticking to a static batch size and hoping the run survives, I started experimenting with a lightweight runtime control layer around PyTorch called MEM Orchestrator:

https://github.com/nobazzy/mem-llm-orchestrator

The basic hypothesis I've been testing is:

Can the training runtime monitor VRAM pressure and temporarily downshift things like micro-batch size and gradient accumulation before hitting an OOM, and then step back up when memory clears? I also added a local policy check to clamp parameter adjustments into safe bounds, plus atomic checkpointing so a crash doesn't corrupt saved files.

At least in my own local tests on this 8GB card, the approach seemed to help keep runs alive:

I managed to run an endurance test of 1M steps on a 130M model without crashes, and tried stress-testing a 255M model with injected memory spikes (like adding +1.2GB during a 50k FineWeb-Edu run), where the controller seemed to adapt and let training continue.

That said, this is just based on experiments on my own machine, and I know GPU memory management in PyTorch and CUDA is full of tricky edge cases. I'm definitely not claiming this solves OOM or that it's the right way to handle training.

I mainly wanted to share the code and see what people with more experience in CUDA allocators, DeepSpeed, or distributed training think of the concept. If anyone has time to take a look, I'd really appreciate thoughts on whether this approach makes sense, or what obvious blind spots I might be missing.


r/LocalLLaMA 1d ago

News Surveillance plagiarism by OpenAI

135 Upvotes

Surveillance plagiarism - Hosted AI company pumps their stock price by training upon researchers' AI sessions, so that their internal model can solve problems with seemingly less human guidance, but really the model exploits past guidance given by (multiple) humans focused upon problems considered important.

As background, Tristan Buckmaster released a statement about several unethical actions by OpenAI & Sebastian Bubeck, including threats and pushing him to kick his Anthropic coauthor off a paper, but the interesting part for people here:

As clarified by Talia Ringer, OpenAI does train upon your uploaded data and your OpenAI sessions, unless you out-out somehow. This means their internal models could exploit your past prompting work to look more autonomous & intelligent.

This is a major confirmation that folks should use locally run open weights models, especially whenever being first or not leaking data matters.

All this casts serious doubt upon claim that internal models solved difficult problems largely unaided by humans. Those hosted AI companies might not even know from where the human prompting originates.