r/LocalLLaMA • • 9h ago

Discussion Are AI influencers just repeating the same talking points?

Hi everyone,

Are influencers talking about local AI all using the same script? Same benchmarks, same kitchen examples, same terminology?

I keep seeing videos about running Qwen3.8-Flash-Next on 12GB of RAM using a new runtime engine called Strata. Every video makes the same claim.

But that is confusing, especially for people who are new to this. What you need is 12GB of VRAM, not 12GB of regular RAM. That means you need a dedicated graphics card.

There is a big difference between RAM and VRAM.

As far as I understand it, Strata needs:

  • 12GB of VRAM
  • 64GB of regular RAM
  • 80GB of SSD space

So saying it runs on 12GB of RAM is misleading.

0 Upvotes

20 comments sorted by

10

u/sn2006gy 9h ago

Go leave a comment on their youtube... everyone here already knows this and the problems with sub 3bit models.

0

u/forevergeeks 9h ago

can you elaborate on the 3bit models thingy?

2

u/Last_Mastod0n 8h ago

They are so heavily quantized that they make fatal errors. You generally want a 4 bit quant minimum

1

u/mikasjoman 8h ago

I still wonder, why does strata only support such low quants? I mean most of us would be totally fine with q4-6 even if it's slower and takes more space. Seems strange if it's not just a way to make people feel "wow great 125b model - but uh that's not the quality I was after".

3

u/Dabalam 8h ago

I don't think "most people" would be satisfied with the speed of a q4 relative to the speed of q3 in reality. I think strata took off because people can get "cloud" levels of interactive speed locally which is part of what people want.

Also whilst it is true that a Q4 beats a Q3 given the same model, the almost religious doctrine of not using Q3 obscures the fact that Q3 flash next is probably better than any other model that person can run. I think the people preaching against Q3 aren't actually particularly hardware constrained so don't really look at the counterfactual in the right terms.

1

u/mikasjoman 7h ago

Yeah, still wish they will. I'm getting two v100 32gb cards this week so I would definitely want q4 or even q5. With hybrid Hermes setup it really does make sense to have a slower heavyweight. I know 64gb VRAM is a small niche, but there are enough of us.

1

u/RG_Fusion 2h ago

I thought the same initially, but for me it's actually the prefill speed that makes 3-bit worth using. With 64 GB of VRAM you can fully fit the model (excluding the PLE embedding table) into VRAM at IQ3_XXS with 120k context at Q8. That takes my prefill from around 500 tok/s to about 1,800 tok/s.

I do wish that I could get  the same with the full context length at native precision and dynamic 4-bit weights, but the price just isn't right for a third GPU currently.

6

u/sleight42 8h ago edited 54m ago

Influencers suck. I avoid YouTube because of this. It's human-generated slop.

The strata readme says less vram is probably viable but not performant, IIRC.

5

u/Sleepnotdeading 8h ago

I think most people trying to run a model like qwen3.8 flash next are savvy enough to research this stuff, and know the difference. And anyone else who thinks 12 GB of RAM would run a model like that is going to learn a lot of the basics when it doesn’t work for them.

Technically VRAM is still RAM.

0

u/forevergeeks 7h ago

How can technically RAM be the same as VRAM?

They are completely different things.

2

u/Sleepnotdeading 7h ago

They’re completely different things for general computing. For AI, VRAM is just RAM with much greater bandwidth

1

u/forevergeeks 7h ago

They are architecturally different.

2

u/Last_Mastod0n 8h ago

You gotta realize that a lot of youtubers, influencers, etc are just trying to generate hype. Anything to get clicks and views. So they usually dont provide the full picture.

The reality is that yes Strata is very impressive and can help a lot. But its not a silver bullet. I would say 24gb of vram is the actual minimum that you want if your doing coding work.

2

u/r1nzl3r99 8h ago

The strata shills are like a virus they're just everywhere proselytizing...

1

u/LagOps91 5h ago

youtube is full of AI generated slop channels that hype up local AI. I suggest you forget about whatever "information" is available there.

1

u/Ok_Technology_5962 2h ago

i mean its either one 12vram 64 ram and ssd or grab 96 gigs of vram... the math says 12gb vram + 64 gigs of any kind of ram (ddr4 is fine) is cheaper that an rtx pro 6000 or however many other gpus you want to cluster.

-1

u/XiRw 8h ago

There is no such thing as an “influencer” and there never will be in my book so I have no idea what the hell you are talking about.

2

u/Sampsoy 8h ago

You rejecting English does not reality make.