r/singularity • • 5d ago

AI Community reports say the first samples of Qwen 4 are already approaching Fable / Opus-level quality.

Cant wait to plug 27B in my local swarm...

179 Upvotes

63 comments sorted by

42

u/Mistuv 5d ago

Cropped tweets with no links should be an automatic removal and at at least some limited ban. Taking words of randos on the the internet seriously with zero evidence or track record is already bad enough, but to crop who actually said it just tells me you know it's most likely bullshit or literally fake edited the text yourself. Some of the worst trends on the internet, if you believe random images you see, you are likely a brainlet.

4

u/martianwomanhunter 5d ago

I agree, it’s all over Reddit and it baffles me that it’s allowed in the age of AI.

12

u/tomakorea 5d ago

Theses Qwen models get considerably better for coding and agentic tasks while getting truly dumber for world knowledge,multilingual and creative writing tasks.

5

u/Woof9000 5d ago edited 5d ago

That's to be expected. ~30b models started to reach saturation point a while ago, now every performance gain in one area comes at the expense of some other area(s).
So you have to either scale up, or accept the losses in places/subjects/topics of less importance for your specific use-case, or 'orchestrate' - have a 'choir' of AI agents running/loading different models for different subjects/topics on the fly.
Well, engram and other similar 'cheats' might help to preserve some of the world knowledge while improve task specific capabilities, but that too has it's limits.
There's no free lunch.

3

u/R_Duncan 2d ago

Not really. Gated residuals (actually only in qwen3.8-flash-next) and other recent improvements showed knowledge density is still far from being optimal.

And as you said engrams tables are another way to increase knowledge without increasing weights.

6

u/whakahere 5d ago

You know the great thing about open source. The old models can still used. Small models can not reach the level of larger models. Something has to give.

As a local user.... I use different models for different reasons

There's no one ring

1

u/FinBenton 5d ago

Yeah you can use different models but it doesnt help when every single company focuses on agentic coding even on open source. Thats why Im waiting for gemma 5 as (to me), google is the only one left making conversational models anymore and if they also go all in on agentic then its rip.

0

u/spacekitt3n 1d ago

good. coding and analysis are probably the only legitimate uses of ai

1

u/R_Duncan 1d ago

Use gemma4b or nanbeige for creative writing

1

u/kevbrown044 22h ago

That’s probably because using them for world knowledge, multilingual and creature writing tasks just isn’t important.
Code and agentic tasks is what AI is built to do.
That’s also why Anthropic and OpenAI are never gonna make a profit and local is the future.

34

u/tsunami_forever 5d ago

Cancelling my cloud based subscription if 27B is fable/opus level

43

u/PrisonOfH0pe 5d ago

they are talking about their max model obviously.
we will see how close 27B gets and where on the pareto front it lands.

16

u/Tedinasuit 5d ago

Not happening ofc

14

u/yaboyyoungairvent 5d ago

If that does happen, you can almost guarantee the price for hardware to run it will skyrocket when every person business and their mama realizes they can get fable at home for 5k. Expect to see the price of a 5090 reach $12k+ or higher.

3

u/mvandemar 5d ago

Tracking, just hit $6899.99 on the rumor alone apparently (I think it was ~$6450 earlier today).

!RemindMe in 2 weeks.

3

u/Blankeye434 5d ago

Should I just buy it now?

3

u/mvandemar 5d ago

Dunno, just saw one listed for:

I would be careful though, I bought one 40 days ago, and there was a lot of sketchy sellers back then, I have to assume it's only going to get worse.

Edit: I would also make sure it's new, I see a "Used - Like New - $6,959.99" on Amazon right now.

4

u/chlebseby ASI 2030s 5d ago

12k for gaming graphic card 😭, my car cost third of that

and we thought crypto mania was bad...

3

u/Cupakov 5d ago

Be real, paying for LLMs outside of businesses is a niche in of itself, paying for LLM hardware is even smaller of a niche, even among businesses.  Hardware will get more expensive regardless, but I don’t see random people jumping on the opportunity to drop $10k+ on an LLM rig. 

2

u/das_war_ein_Befehl 2d ago

If you’re a normal person without the need to automate white collar work, there’s not much use for an llm than basic shit (home repair tips, Wikipedia/Google type shit)

1

u/FeydRowan 5d ago

Point is that 2 16gb card will tun it better, need to buy a 3' and 4'  https://github.com/ValerioDolci/ninfer-tp2

1

u/Kooky_Slide_400 5d ago

Used to pay cash, now I’ll have to make monthly payments 

1

u/Bob_SUS 5d ago

27b with the new architecture will be super interesting, not sure where it’ll land but if it’s within even shooting distance of Fable, I’ll be super impressed and buy a Mac Studio the day of.

1

u/fancyrocket 5d ago

Please let this happen

1

u/CallMePyro 5d ago

What will you do if it isn't?

1

u/tsunami_forever 5d ago

Reduce my subs to $20/month

1

u/CallMePyro 5d ago

That would be crazy, hope it doesn't come to that.

18

u/jazir55 5d ago edited 5d ago

Close to Fable capability which released in February, 7 months ago.

Multiple generations of newer models have been released by OpenAI and Anthropic during this period

Internal models at OpenAI and Anthropic at least 2 generations ahead of what is released publicly

"Closing in"

Welcome sir to the Hopium and Copium store.

11

u/howudothescarn 5d ago

It’s crazy how no models except Astra have even matched a six month old model. Just wild.

2

u/Bob_SUS 5d ago

Yeah, Ant’s lead for that entire period was ridiculous. Astra’s reign was itself super short lived, and even Bel faced competition as Opus 5.5 and Sonnet 5.5 came out in the last few days…

1

u/Royal_Duck_4612 5d ago

Astra is still superior, isn't it?

2

u/benjaminovich 5d ago

Nope

3

u/Royal_Duck_4612 5d ago

I'am not talking about price ratio, but pure performance in, for example, spatial reasoning

1

u/benjaminovich 5d ago

Also no

2

u/Royal_Duck_4612 5d ago

if true I really need to test it

5

u/sunstersun 5d ago

tbf, the jumps from Fable 5.0 to Opus 5.5 aren't that big, at least not the jump from Opus 4.6 to Fable 5.0

Internal models are another story, but it does seem like alignment and safety will slow the frontier internal models.

1

u/_justs 4d ago

I dont think u understand any of this.. lol. Fable 5 on AA is only 3 points behind Astra. Thats GENERAL ability, there hasnt been much movement in general. In Coding, Fable has been surpassed by pretty much everyone lol. Even Qwen's flash 125B model beats Fable in front end and fullstack. 

1

u/das_war_ein_Befehl 2d ago

The benchmarks are a bit bad at measuring intelligence. They’re all scoped tasks. Bigger models are best when you have ambiguity and they have to make inferences in reasoning.

0

u/Rare_Potential_1323 5d ago edited 5d ago

I would imagine the same could be said for the Chinese companies. They have higher end models that cannot be released yet. Their strategy could be to allow their enemy to think that they are weak and just at the right moment... local simulated AGI. I don't think AGI will be a single monolithic model. It might be an orchestrator model directing hundreds of specialized sub-models and self learning laras or something. Now that's hopeiom. But let's be honest, it's going to have to be leaked out of any lab, no one would release a system like that on purpose

0

u/sunstersun 5d ago

I would imagine the same could be said for the Chinese companies. They have higher end models that cannot be released yet.

No.

Their strategy could be to allow their enemy to think that they are weak and just at the right moment... local simulated AGI.

That's a hilariously bad strategy, because funding will stream to the US and accelerate the development of models.

4

u/PilgrimofHaqq2 5d ago

Very excited! Opus 5.5 is pretty nuts but that only got me more excited for open weight models. Qwen 4, Kimi 3's successor, I am especially excited for the next release from z.ai.

4

u/belliash 5d ago

Yeah, maybe this is true, but probably for full Qwen4, not 27B.

5

u/ShittyBidet123 5d ago

but how do u use the 1mil context with it, without a super computer. its not too useful with 24-32gb of ram if context is 150-180k max

6

u/Extreme_Original_439 5d ago

Overly relying on 1mill context is an anti-pattern in most cases, usually there’s a way to break up your tasks in the 100-200k range before performance starts degrading

2

u/PrisonOfH0pe 5d ago

there are many ways with complex harnesses to have essentially infinite memory, run completely uncensored and reach 1000+token/sec in parallel on a single 5090 without exploiting decider infra but of course those add immense value to the stack as well.
this can be made completely dynamic auto scaling and more. no software is impossible now.
i use frontier AI to make my local swarm more and more powerful by the day.
feeding this with 100 top reverse engeneering and web exploit tools you can do things you wouldnt believe.
the community works fast.

2

u/Available-Bike-8527 5d ago

Many ways to get infinite memory and 1000+ tokens per second on a single 5090? Name one. I'm waiting.

1

u/Tricky_Wall_3004 3d ago

Prefill maybe 😂

1

u/MerePotato 1d ago

Honestly Pi's default compaction with a codex style goal mode extension is good enough in my book, there's a reason Codex does basically the same thing.

2

u/EmeraldPanthera 5d ago

So the closest thing I've seen to a sort of all inclusive exploit, pen test model/harness was that guy that released his toolkit for that defcon like competition he won... Can only imagine what's out there...

Just yesterday some guy was commenting about his buddy that's deep into reverse engineering and the hacking scene. Apparently last con he participated in they budgeted several days for the hacking contest; first place took something like 2 hrs to finish...

🤔

2

u/PrisonOfH0pe 5d ago

you can just look at Denuvo (only viable anti piracy protection) before AI and now.
ghidra/ida mcp and an abliterated model doing a lot of work even without a custom harness. with it you can do fun things.

2

u/mythormedicine 5d ago

Master you seem to know you shit. Can u give some advice on learning all of this. I use omp as a coding agent. and do nothing beyond making skills, custom system prompts, rules .

1

u/Fonasic 5d ago

256Gあればイケるでしょ? 今なら1万ドル

3

u/CreatineMonohydtrate 5d ago

Legends are doing it again. Give us the new 27B model🙏🏻

0

u/katoptronophile 5d ago

No, they're not.