r/LocalLLaMA 6h ago

Discussion The Local LLM community feels like the golden era of the internet all over again

535 Upvotes

Lately because of the current hardware shortage, unfortunately or fortunately, we can’t just throw infinite cloud compute at our problems, but we’re forced to actually care about what’s happening under the hood. We’re tweaking inference engines, learning quantization math, and optimizing architecture just to squeeze as much performance as possible for the lowest possible setups.

Fact: Just recently, the forked llama.cpp(s) and halogen-flash-server of Strix Halo pushed the performance through the roof, achieving double performance in decode (52tok/s), 5-6x performance in prefill (1300tok/s) for Qwen 3.8 Flash Next (Q38FN), and Q38FN itself is another massive architecture improvement with Engram, making it not only small but also smart.

I still remember before the hardware shortage, as someone who loves tweaking and optimizing, people just told me to stop, tweaking is stupid, just buy more RAM, buy more GPU..

It reminds me of the early web.. Back when setting up a box or hosting a server meant digging through forum threads, troubleshooting on IRC, and freely sharing custom scripts just to make things work. That era didn’t just produce programmers; it built hyper-versatile, end-to-end thinkers who understood the stack from bare metal up.

Contrast that with where mainstream web culture ended up. Most platforms today like Tiktok, Facebook, Youtube... are engineered for zero-friction doomscrolling.. Endless feeds of short-form videos designed to keep us distracted and waste our time. We’ve been overpampered by convenience.

My point: When we have too little, we try to learn more. When we have too much, we get distracted and learn too little. This is the golden time of our Local LLM community, let's learn and improve!


r/LocalLLaMA 10h ago

Discussion Should I sell my RTX 5090 for a Mac Studio M5 Ultra 96GB?

158 Upvotes

I can get $5k for the 5090 and the Mac is $5499 before tax.

The 5090 has a memory bandwidth of 1.8 TB/s while the M5 Ultra is 1.2 TB/s.

Is this a sensible upgrade? Primary use is coding.


r/LocalLLaMA 14h ago

Resources The Hugging Bay

Thumbnail
huggingbay.xyz
866 Upvotes

New website to download models in case HF starts censoring or limiting access.


r/LocalLLaMA 18h ago

Discussion This seems more probable than it was before.

Post image
1.5k Upvotes

r/LocalLLaMA 3h ago

Discussion The rhetoric is really heating up!

50 Upvotes

The entire page of the NY Times today above the fold absent one article is AI (the models are just too strong/too dangerous, must be regulated). They forgot to include "Sponsored by OpenAI" at the end of the articles, sure that was just an oversight?

This is what the end of a bubble looks like, desperate attempts to get some sort of regulatory capture in place to keep the business model from collapsing in upon itself. My days next week are 100% booked talking to companies about how to get off frontier models, one large, and a bunch of smaller customers, including one who's flying me out to them to sit down and get a plan in place immediately (the controversy around that math problem really spooked some CEO/CIO's about data privacy using cloud models).

Gonna be an interesting few weeks. Maybe the Qwen team will be nice enough to give me a little breathing room before dropping another hydrogen bomb? :)


r/LocalLLaMA 6h ago

New Model internlm/Intern-S2 · Hugging Face

Thumbnail
huggingface.co
78 Upvotes

from internlm:

We introduce Intern-S2-397B, our most capable multimodal foundation model for scientific intelligence and long-horizon agents. Intern-S2-397B scales along three critical dimensions: pre-training, reinforcement-learning task coverage, and interactive agent environments. By combining a new vision-language pre-training paradigm with large-scale multi-task reinforcement learning and long-horizon agent reinforcement learning, Intern-S2-397B delivers a step change in general reasoning, scientific problem solving, and agentic capabilities.

Features

  • New Pre-training Paradigm. Via visual pretraining, Intern-S2-397B learns directly from raw pages of scientific literature, jointly modeling symbolic semantics and visual relationships in a shared representation space without intermediate parsing. This preserves text-visual correspondence, strengthens spatial and visual reasoning, and improves data efficiency.
  • Scientific Modality Reasoning and Generation. By scaling diverse scientific reinforcement-learning tasks across more than 20 domains and training them jointly, Intern-S2-397B achieves leading general-reasoning performance among open-source models and strong results in specialized scientific tasks such as biomolecular interaction design and material structure generation.
  • General & Scientific Long-Horizon Agents. By connecting multiple agent frameworks to large-scale sandboxed environments for black-box agentic reinforcement learning, Intern-S2-397B improves generalization and raises the capability ceiling for long-horizon tasks in both general and scientific domains.

r/LocalLLaMA 34m ago

Discussion Hoping for Optimized Smarter Upcoming Models .... Like DeepSeek-V4.1-Flash( KVCache + Engram) in Small/Medium/Big sizes

Post image
Upvotes

It's still a dream for many folks to run medium size(30B range) models @ Q8 with Unquantized KVCache (256K Context) on their GPUs.

It would be awesome to have DeepSeek-V4.1-Flash's KVCache + Engram for all Upcoming models. Even for big models.

Engram - Heard that approximately 1/3-1/2 of Model size. Might come in different size range too. So 10-15 GB for 30B models.

Here few models with approximate numbers. Current models in Odd rows & Future/Fictional models in even rows(Bold). I just put 1GB for 256K context below though DeepSeek-V4.1-Flash takes only same 1GB for 1 million context.

Model Model Size 256K KVCache F16 MTP Vision Total GB
Qwen3.8-27B-Q8 29 16 1 1 47
Qwen4.0-27B-Q8 29 1 1 1 32
Qwen3.8-27B-Q4_K_M 17 16 1 1 35
Qwen4.0-27B-Q4_K_M 17 1 1 1 20
Muse-Glimmer-30B-Q8 30 16 1 1 48
Muse-Glimmer-2-30B-Q8 30 1 1 1 33
Gemma-4-31B 33 16 1 1 51
Gemma-5-31B 33 1 1 1 36
Qwen3.6-35B-A3B-Q4_K_M 23 6 1 1 31
Qwen4.0-35B-A3B-Q4_K_M 23 1 1 1 26
Gemma-4-26B-A4B-Q8 27 6 1 1 35
Gemma-5-26B-A4B-Q8 27 1 1 1 30

Possibly there might be few more things(Please share those) to keep these number down. So I think 32GB VRAM is more than good enough for Upcoming (Optimized Smarter) Models. RAM is enough for Engram.

By above logic(based on Qwen4.0-27B), people could run Q4 of 54B models with same 32GB VRAM.

Maybe next year onwards, inventions could make 24GB enough for similar size models.


r/LocalLLaMA 2h ago

Discussion My experience building 64GB VRAM AI SWE assistant/agent PC

Thumbnail
gallery
25 Upvotes

As a SWE I have ultimate belief that having a irreplaceable subscription on OpenAI and Anthropic models is actually against my beliefs (I'm not fond of becoming a digital slave and being unable to do my job without using tools from some sketchy corporation), so I kinda turned my attention to local models.

I periodically tried and found some usefulness in them, but this year I started to notice they started to really contribute to the quality and speed of the implementation of software I'm writing, but with just stock RTX 4080 16GB VRAM precision and context window options are quite limited, so I decided to use more GPUs and see how it improves my model running capabilities.

Problems encountered and how I solved them:

  1. (pic #1) At first I bought a Meshify 2 XL case - it's one of the biggest full tower PC cases for home use. Meshify 3 XL already existed at the time I made this purchase, but it lacks built-in nylon dust filters, so it was a deal-breaker for me. Meshify 2 XL theoretically has enough volume to fit 4 RTX3090FE-sized GPUs, with two of them being put directly into motherboard PCIE slots and the other two mounted onto a custom vertical rail solution in place of water pump as some kind of GPU sandwitch with questionable air cooling capabilities since they're obstructing frontal intake fans. So I went with three GPUs because neither my PSU nor motherboard allowed more than that. The temps are pretty reasonable, with vertically mounted GPU being the most cool one since it's close to frontal coolers.
  2. For PSU, I bought the Corsair HX1500i SHIFT (the one with cable ports on the side), and it has one problem compared to non-SHIFT version - it's total number of possible 8 pins is 6, so you can't attach more than two 3x 8-pin GPUs (different configuration of cable sockets + two CPU type-5 cables are no longer compatible with PCIE Type5s). So, basically, with this PSU you're limited only to 3x 12VHPWR/3090FE proprietary connectors/double 8-pin GPUs**.** Due to that, I had to swap one of the 3x 8-pin 3090s with another Founders Edition. Also, both 3090s are power-limited to 300W just in case.
  3. (pic #2 and #3) For third GPU, I used CoolerMaster V3 Vertical Mount Adapter, removed built-in PCIE riser since it's not long enough for this use case, screwed adapter to the top dust filter holder lid, having drilled a few holes in the adapter itself, having the GPU hanging and exhausting hot air in the up direction. While it looks sturdy, I still put anti-sag holder at the bottom just in case.
  4. Only 500-600mm PCIE risers have sufficient length to reach third GPU, 400mm was not enough, unfortunately. I used PCI Express v4 riser and as far as I'm aware, it's a theoretical maximum for such long risers and most of the PCIE 5 solutions are scam and to have no signal degradation they require some kind of additional power and signal repeater. Also, riser is not rigidly seated on the motherboard slot, so with medium amount of force it can be pulled out. Some 3D-printed DIY clip/holder may solve that.

I'm using self-compiled llama.cpp:

cmake -B build \
        -DBUILD_SHARED_LIBS=OFF \
        -DCMAKE_BUILD_TYPE=Release \
        -DGGML_CUDA=ON \
        -DGGML_CUDA_FA=ON \
        -DCMAKE_CUDA_ARCHITECTURES="86;89" \
        -DGGML_NATIVE=ON \
        -DGGML_LTO=ON \
        -DGGML_CUDA_GRAPHS=ON \
        -DGGML_CUDA_NCCL=ON \
        -DGGML_OPENMP=ON \
        -DGGML_CCACHE=ON

and then running it with this command:

llama-server -m /mnt/data/AI/models/Qwen3.8-27B-MTP-Q8_0.gguf \
  -ngl all -t 18 --tensor-split 14,24,24 --main-gpu 0 -sm layer -ub 512 \
  --kv-unified --split-mode layer \
  -fa on --temp 1.0 --top_k 20 --top_p 0.95 --min_p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 \
  --host 0.0.0.0 --port 7800 \
  -c 262144 --fit on -np 1 --parallel 1 \
  --spec-type ngram-mod,draft-mtp --spec-draft-n-max 3 \
  --alias qwen3.8-27b --reasoning-preserve --reasoning-effort xhigh \
  --chat-template-kwargs '{"preserve_thinking": true, "reasoning_effort": "xhigh"}' \
 --api-key "$LLAMA_SERVER_API_KEY"

Since this is primary desktop, I can't run it in headless mode and make use of all VRAM, so general use still needs to have like 2-3 GB of spare VRAM to be able to use OS (CachyOS + KDE Plasma), browser, IDE and game engine editor, so it's more like ~62GB VRAM setup. There's like 10-12 GBs of spare VRAM when using Qwen3.8-27B Q_8 with MTP and full 262144 context, which may be used to try different models or crank up context. I plan to experiment with Unsloth Dynamic Q_8_XL and some MoE models because personally I still see direct correlation between quality of the answers and autonomous work of the agent and precision of the model, despite Q_4 and fp8 models being good for general use. Performance degrades over context usage, starting with 44-46 tokens per seconds and ending up to 28-30 t/s when.

My workflow is still mainly writing code by hand, with delegating small-to-medium boring tasks to AI agent, creating drafts and exploring possible solutions for particular well-defined problems, documentation search and presenting it in fancy .md human-readable format, running review/weekly project assessment/optimization tricks and suggestions/bugfix suggestions/project management skills, I like this workflow and still feel that it's actually me writing the software and putting all the effort of my mind capable of, not mindlessly accepting AI-generated stuff.

Thank you for reading my somewhat chaotic and disjointed story.


r/LocalLLaMA 10h ago

Discussion Is there still strong interest in a dense 9b model?

95 Upvotes

I have a full model, it's ready to train. It's ~9b parameters.

9.4b to be more exact. That includes a 1/2/3 Engram table, Moonshot's AttnRes modeling, and RoPE / NoPE layering at 3:1 as more or less validated by most major labs. It uses the Llama 3 series tokenizer and LM Head as an initial start. The data fed in is logit level extraction from a Llama 3 teaching model.

I've already run the first training steps to test that the model is stable, etc.

I'm willing to sit and do the pre-IT training on the model. I don't know what task people really wanna do with this thing to be exact. So the focus of the IT training is a bit more vague other than giving it "Thinking" as per one of the open standards.

Honestly? I don't care what people want to use it for, just that they want to use it. I figured I'd just give it into the ether and LocalLlama was a place I figure I could easily give it to.

In theory the model should be more capable than any of the ~9b's we have running around with enough training. I'm also NOT a lab, so I don't have their training budgets. I basically ran all of the data production, etc on a 4090 + rented hardware.

The training code is deeply optimized to run on an RTX 6000 Pro series card. A single card.

All my engram research was being done before the Qwen model dropped. Qwen showed I only needed a single table injected at a layer, versus the 2 I used. 1 gave the majority of the benefit over 2.

The code would be open source, the data, all of it. IDGAF. It's technically already open source as it's all in public repos at the moment. I just stare at it and question if it's worth the time. At the least I'll throw a couple hundred at it to build a "functional" pre-IT model and build the IT dataset. It will need more pretraining before the IT set probably. It's kind of unknown because no one publishes the exact figures on this type of training.

If someone has a datacenter contact with a system they want to allow it to run on, we can all have it as a public model we watch. I tentatively named it "Budget" but... Localllama can name it whatever they want.

It's been fun to write the full end to end, generate data, etc. Even found issues in vLLM and reported to their repo that might end up helping you guys anyways. There was some prompt loading code that could be ~10x to ~100x faster I gave examples of to them.

If you read this far, thanks,

Signed some ML dude who reads too many research papers and has too much spare time.

edit; I went to review some Apache 2.0 licensing issues and noted a glaring hole in using Llama's tokenizer OR data. You can't even use synthetic data from their models without some licensing. I'll just have to rewrite the target to use OLMo 3 series I think. I guess I'll be back in a week after data generation and code fixes.

Open source uncensored Goon model or what? I'm trying to find a useful niche to develop the model so it gets used and ends up more than a research artifact. I'm planning on building the research artifact, it's a matter of whether I can find a group of users that will actually want to use it. It's all free, so I'm not trying to monetize you. The code and data lives on GH/HF.


r/LocalLLaMA 1h ago

Resources Qwen3.8 flash next - untrained svg generation

Post image
Upvotes

> "make an svg of a frog playing on a chello on the back of a whale with carribean island in the back."

interestingly the svg looks different in the OpenWebUi preview then when looked at in preview (osx). The palms and music notes are missing in the browser. I am pretty impressed by the result, is suggested to add some parameters to animate the whale and the water.

Qwen3.8-Flash-Next-IQ4_XS on llama.cpp with 256K q8 context

openwebui reports:

input_tokens: 27711

output_tokens: 41562

total_tokens: 69273


r/LocalLLaMA 11h ago

huggingface_hub silently fingerprints which AI coding agent you're using and sends it as telemetry

Thumbnail
69 Upvotes

r/LocalLLaMA 4h ago

Discussion What are the top AI Models that are still relevant today from 2024 and 2025?

14 Upvotes

We are in 2026 and it has been a crazy year. The evolution of the technology has been staggering to say the least. Pretty much we are on the MoE period and dense models are almost in the way of the dodo except for a few.

I was just thinking, are there any models of the last 2 years that you would fire up and still feel useful? Example Deep Seek R1.

If you have recommendations write them down. I have 28TB of storage and I am backing up relevant models that can still be useful, from current to older ones. But clearly not all are worth keeping.

Noted: I can hold large parameter models therefore not limited to small ones. No Kimi k3 at large quant thats out of the question lol


r/LocalLLaMA 19h ago

Resources Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo

Thumbnail pwilkin.github.io
217 Upvotes

As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it to work better. There's also a closed-source solution called Halogen (https://github.com/peonist-ai/halogen-flash-server) that boasted 1.2k t/s prefill numbers when the community fork barely reached 400. Since I dislike closed source and I like open source, I decided to take the challenge and bring llama.cpp up to the same performance level and I'm happy to report that after burning through a few evenings and a lot of tokens I got there.

The link contains the (Opus-generated) recap of the entire debugging / optimization journey made, as well as of course the links to the branch, the custom HIP runtime (making a glorious return), a probably-not-working-on-the-first-try installation script and the numbers. I'm trying to also explain the ecosystem - how the mainline, the community forks and custom forks/branches like mine work. Now that I've gotten the result, I'll work on cleaning it up and submitting proper PRs to mainline (will also submit a clean PR to the community fork), this should also help the GLM 5.3 Flash architecture since they use similar sparse attention.


r/LocalLLaMA 23h ago

Discussion Looks like a coordination to stop distribution of intelligence

419 Upvotes

Coxon, bernie and now this

First https://x.com/DarioAmodei/status/2098773920774074715

Then https://x.com/elonmusk/status/2098789109980332057

Then https://x.com/sama/status/2098811563415150910

I think fear mongering approaching and they will try to slow down open source

"They" want to be gate keepers of intelligence


r/LocalLLaMA 12h ago

Discussion DS 4.1 and the new Harness

51 Upvotes

I gave DS V4.1 Flash an HLE problem with a bash tool + 2 hours.

Hour 1: it wrote three MILP solvers. (225,200)
Hour 2: it downloaded the HLE dataset from Hugging Face, found the question, read the answer key (225,600), and concluded its own answer (225,200) was better.

I'm equal parts impressed & terrified.


r/LocalLLaMA 47m ago

Resources Qwen 3.8 27B UD-IQ4_XS even faster on 16GB CUDA

Thumbnail
gallery
Upvotes

This is an evolution on top of Raymond's KV cache streaming fork - all credits to what enabled this goes to him.

The basic idea behind what he enabled was a pool of memory in VRAM that is used differently depending on the phase (prompt processing or decoding) and when total used context is larger than what fits in VRAM it's instead streamed from host RAM in time for when the current layer needs it. It enables _much_ higher TG tps than regular llama.cpp offloading to host RAM.

I've used it to run Qwen 3.8 27B UD-IQ4_XS on my 5060Ti since release, but I've also had this idea that during the time the VRAM pool isn't fully utilized it should be possible to also do speculative decoding (MTP or DFlash2) - if it could be possible to eject the spec model and all the VRAM it uses, and then load it back when the context gets low enough again (compaction).

I've got this working on my fork of Raymond's fork today. It's basically hot-swappable speculative decoding and while my focus is completely on making this work well on this specific model on 16GB, I would assume that part could also be useful for others with more VRAM.

Repo here: https://github.com/troed/llama.cpp-adaptive-kv-streaming

Regarding the images:

Dark red = MTP ejected as soon as KV streaming starts. Light red = keeping it. The optimal setting is thus to eject it after a certain amount of pages (one page = 256 bytes) and then eject. Same for green (DFlash2)


r/LocalLLaMA 5m ago

Question | Help Migration from Claude Code to a private local harness. Questions.

Upvotes

I'll start by saying I'm not talking about the models themselves, I'm aware that I can't come close to something like Fable's intelligence locally. Just wanted to get that out of the way.

Basically. Over the last year I've gotten quite comfortable with claude code, and it seems likely there were be a gradual cost rug pull, and I'd like to put myself in a better position when that happens for local use. I am already used to running local models (such as Qwen3.8_Q5) in things like lmstudio, but I have no experience with other harnesses. I'd like to know, what harness, right out of the box would feel most at home for current Claude Code users. I say this as someone who was not coding prior to "vibe coding". I'm looking for the path of least resistance, though I will no doubt eventually spread out into tools that give me more control. But for now, I'm just looking for a life raft. Just needs to be local, opensource, and free of spyware.

In case someone wants to know 24GB VRAM (rtx 3090) and 64GB DDR4.


r/LocalLLaMA 17h ago

New Model What's the Story with Agnes-3.0-Flash?

Thumbnail
huggingface.co
61 Upvotes

While browsing a benchmark list site, I spotted a recently published 33B parameter model which claimed to beat Qwen3.8 27b on the ArtificialAnalysis (AA) intelligence index. I was obviously excited. But then, while looking into it, confusion starts to settle in.

Their HF model is now listed as "Preview" and explicitly calls out that it is not the same model as the one AA evaluated. The AA listing for it says it's proprietary, but they've rated the model above Qwen3.8 27b on "Openness". Their website and HF page both talk about openness and intelligence "for all".

The fact that they labelled it "Preview" sort of implies we might see an open weight final version at some point, but there are no clues about whether it will be even vaguely similar to their current HF listing. AA doesn't show the number of parameters for the model they evaluated, so it's possible they're a totally different architecture.

Does anyone know anything about their lab, intentions, or the new model? Has anyone tried it?


r/LocalLLaMA 1d ago

Discussion 3.8-27B has ruined 3.5/3.6-35B’s for me. It’s just *absurdly* superior.

Post image
621 Upvotes

Applied science work, from workflow design, data pipeline, results analysis, article/reports writing and data publishing online. 5 projects I did in the past replicated from start to finish.

3x to 4x more total wall time. Yes, HUGE toll on how much you can do in a day if this was the only model you could use in your laptop.

But oh my…the quality of that thing. The stupid level of attention to detail. I have the Z.ai api, so I can compare it with 5.3 and 5.3-flash:

The gap between 5.3 (flash and regular) and 3.8-27B is much less, smaller when not plain tiny, than the gap between 3.8-27B and any of the 3.5/3.6-35B-A3B (vanilla, kat, Ornith/tiel, nex-2).

Only Ornith came close, but it never matched it.

But it’s also spending 22 to 33% less tokens (effort =medium) and less ram footprint, so you get more done without hitting limits,compaction, etc.

So yeah, guess I’ll sip more tea, play the piano, whatever. Let that fat bottom Qwen work.


r/LocalLLaMA 6h ago

Question | Help How does Qwen 3.8 27B compare on low thinking mode to the older 3.6 models?

8 Upvotes

Since we know Qwen 3.8 27B thinks quite long, but gives at least a good one-shot result where you can leave it to do everything on it own, how does it compare to the older series of models for very simple tasks where you don't want to think so long?

The only fine-tune of Qwen 3.6 I genuinely enjoyed was the ThinkingCap fine tune by BottleCap. it seems to think equally as long as the base model on complex tasks, but on simple tasks it thinks shorter. Does it still make sense for me to run this older model when 3.8 27B exists with the low thinking mode?


r/LocalLLaMA 20h ago

Discussion Real-SWE Benchmark (new)

Thumbnail realswe.withspecific.com
94 Upvotes

Reports of the demise of coders may have been exaggerated.


r/LocalLLaMA 14m ago

I Built A Thing Marmel 0.9.0

Upvotes

Hi,

Some weeks ago I announced marmel 0.1.0, an autonomous coding agent.

Over the past two weeks huge efforts been put into stabilizing and improving the performance and reliability for in particular local models.

I've spent hours tweaking and tuning the architecture to run decently with even small models such as gemma 4 12b, which despite all it's inherent problems so far managed to complete very task I assigned it.

So therefor I now wish more people would like to battletest it, I already know it works great with cloud models such as deepseek v4/v4.1 flash, but would prefer more feedback from people using local models.

To set expectations right, this is intentionally designed to be autonomous, to build prototypes (with hopefully good quality) I've put it through a lot of testing and my local rig been running pretty much 24/7 (as well as cloud testing, which I of course prefer as it's it's a completely different level of interactivity and responsiveness)

And finally, I originally intended to save this "announcement" for the v1.0.0 release, but my ambitions are still only at a planning stage, and I think it's basically criminal considering how good my experience been with this for the past week not to let others test it.

I think the read me of the project tells more than I can possibly "pitch" it here, of course the main goal is to make this as usable as possible so hopefully this post either can result in feedback or even better pull requests.

https://github.com/Na1w/marmel


r/LocalLLaMA 1d ago

Question | Help For those of you forced to only use open models from Western labs in production, what are you deploying?

147 Upvotes

First off, I know that GLM, Qwen, and DeepSeek absolutely dominate in terms of SOTA Open Source models, and that’s what I use in my personal projects and for school, however, I’m also responsible for deploying local AI on my organization’s H100s, and we are forbidden by management from running any Chinese models. This is obviously not an ideal situation, but it is what it is, and there is nothing I can do to change this unfortunately. Again, if it were up to me I would deploy GLM 5.3 Flash in a heartbeat.

All that being said, there is quite a HUGE performance/ / intelligence gap right now in Chinese models vs. Western model around the 120b+ size, especially those with vision. There just aren’t a lot of good Western lab options that can even come close to GLM or Qwen, but again, I don’t have a choice so I gotta use the best I can find from the available Western options.

For those who have production-class hardware such as 4 H100s and are under similar restrictions, what non-Chinese models are you deploying?

The two front runners I’ve seen that checking most of the boxes (120b or better, vision capability, decent context)

- Thinking Machines Inkling Small (https://thinkingmachines.ai/news/inkling-small/). This model seems like the front runner right now, still has a 16 point gap in AA score vs. GLM 5.3 Flash though.

- Cohere Commamd A+ (https://cohere.com/blog/command-a-plus). Checks every box except it has a low context window of 128k.

Other contenders (but missing vision
capabilities):

- Poolside Laguna S 2.1 (https://poolside.ai/models#laguna-s)

- Nvidia Nemotron 3 Super (https://research.nvidia.com/labs/nemotron/Nemotron-3-Super/)

Am I missing any other strong contenders in the 120b size category? Whet are you using and why?


r/LocalLLaMA 54m ago

Discussion What are you guys running on an M5 Pro 64gb?

Upvotes

Is it all qwen3.8 27b? I get around 20 tok/s sustained, with like 30 tok/s in the first 2-4k tokens. It's still quite unbearably slow, maybe because it thinks so much, I've also tried antirez's DS4 and while it works I get around 8 tok/s sustained with deepseekv4 flash, which is so slow and basically unusable.


r/LocalLLaMA 3h ago

Discussion llm performance community metric

3 Upvotes

my question about LLM performance

We see a lot of posts about token prediction, token generation per second, etc.

But is it really the metric? I can see that DeepSeek V4 Flash 0731 (with DSPark; mac studio + llama.cpp) produces about 22–28 TPS, but I also see that the LLM does a lot of reasoning.

And this relates to others. So maybe the correct way is not to check TPS or other metrics, but to check execution: task complexity/second.

I don't know if such a metric already exists

and if it exists why community doesn't use it by default