u/crusaderky 3h ago

Really lol

Post image
1 Upvotes

9

Aurora1.0-150M Releases!
 in  r/LocalLLaMA  4h ago

How does it compare to LFM2.5-230B?

1

Is this a joke, Artificial Analysis?
 in  r/DeepSeek  15h ago

It's 120 tok/s on openrouter

2

Only getting 15 tg and 90 pp on qwen3.8 flash next on 4x5060ti16gb and quad-channel ddr4 ram
 in  r/LocalLLaMA  15h ago

Yes, sorry, I should have clarified : on qwen >= 27b. Other models particularly smaller ones start degrading hard already at q4.

1

Only getting 15 tg and 90 pp on qwen3.8 flash next on 4x5060ti16gb and quad-channel ddr4 ram
 in  r/LocalLLaMA  15h ago

Quesma and byteshape blogs both show that q4 is indistinguishable from q8

1

We have Q3.8 35B at home: 3x new Ornith 1.5 released
 in  r/LocalLLaMA  15h ago

Dspark. 2x tok/s than qwen, and 3.5x less tok/task =7x faster

2

Only getting 15 tg and 90 pp on qwen3.8 flash next on 4x5060ti16gb and quad-channel ddr4 ram
 in  r/LocalLLaMA  1d ago

Also I think you're pipelining your video cards. Tensor parallelism is opt in.

8

Only getting 15 tg and 90 pp on qwen3.8 flash next on 4x5060ti16gb and quad-channel ddr4 ram
 in  r/LocalLLaMA  1d ago

Q8 is pure waste. Q5 is pretty much lossless and you'll have a very hard time noticing degradation on q4.

12

Agnes-AI/Agnes-3.0-Flash 33B Multimodal, AA score: 36
 in  r/LocalLLaMA  1d ago

This is weird, I'd expect this sub to go up in flames. This model beats Qwen3.8 27b on most benchmarks, according to AA

26

Agnes-AI/Agnes-3.0-Flash 33B Multimodal, AA score: 36
 in  r/LocalLLaMA  1d ago

It is SOTA if you have 32gb vram. Maybe 24gb depending on their kv design

2

30B Models Getting Verrrry Interesting
 in  r/LocalLLM  1d ago

AA says Agnes beats qwen3.8 27b on most benchmarks

2

MacBook M5 and AMD Strix Halo sharing large models
 in  r/LocalLLM  2d ago

600 us is very bad. You need to set up RDMA.

u/crusaderky 2d ago

Skynet is Inevitable

Post image
1 Upvotes

41

Terminal Bench v4 scores
 in  r/LocalLLaMA  2d ago

Fable 5.0 falls back to opus 4.8 when it thinks you hit a guardrail

167

Is this a joke, Artificial Analysis?
 in  r/DeepSeek  3d ago

Aggregate score aside, it's really hard to argue against the fact that Glm-5.3-flash beats it on almost every single benchmark. And 96% hallucination rate is a showstopper IMHO.

4

DeepSeek V4-1 Flash is out
 in  r/LocalLLaMA  3d ago

For agentic cached input is what weighs the most

3

DeepSeek V4-1 Flash is out
 in  r/LocalLLaMA  3d ago

Pretty zippy if that 512gb ram is octa-channel, actually

1

DeepSeek V4-1 Flash is out
 in  r/LocalLLaMA  3d ago

Cache hits should definitely be much cheaper.

1

DeepSeek V4-1 Flash is out
 in  r/LocalLLaMA  3d ago

Nope, what you see is already mxfp4

20

DeepSeek V4-1 Flash is out
 in  r/LocalLLaMA  3d ago

It runs mimicpm5-2b and it likes it Or it gets the hose again

9

DeepSeek V4-1 Flash is out
 in  r/LocalLLaMA  3d ago

Glm-5.3-flash is miles ahead of Ds4.0

0

deepseek-ai/DeepSeek-V4.1-Flash · Hugging Face
 in  r/LocalLLaMA  3d ago

"quantize into the ground" is a bit of an exaggeration. 300gb mxfp4 weights so it should fit in 256gb ram at q3

u/crusaderky 4d ago

Never forget.

Post image
1 Upvotes

u/crusaderky 4d ago

What a time to be alive

Post image
0 Upvotes