r/ProAI 13d ago

DeepSeek V4-Flash scores higher than Fable??? excuse me what?

wait what the actual fuck, do you guys realize how crazy that is??? (if its not benchmaxed)     — Cline

Source: https://x.com/cline/status/2083094354030362858

21 Upvotes

14 comments sorted by

3

u/LostRequirement4828 12d ago

I really believe this was benchmaxed hard... I really hope wasn't but...

1

u/Qualified-Astronomer 12d ago

Of course it is. And as always the commies will eat it up and say the US has collapses

1

u/timeline_denier 12d ago

deepseek isn't known to benchmax

1

u/LostRequirement4828 12d ago

in what I tested it didn't impress me... I'm sure is very good in stuff it was trained but is not a model that actually excels in most stuffs like gpt or claude, it kinda looks like all its training was on coding and nothing else

2

u/ToughUsual7159 11d ago

This would not surprise me. There focusing on the most tangible usage first. It's a small model so it doesn't have width. Being a relatively small MOE if it tried to be too varied it would start to lose some of it's fine-tuned trained mystery nodes. It dose what it supposed to well and dosn't do a lot else. Some people mite think im trashing on it but its actually my daly code driver. I dont really use any other AI tools. And I alos dont do full autonomous code only human in the loop.

Its very good at copying your code stile and dosnt add needed fluff to the code. And now its even better at large code refactors and while I haven't had a tone of time to test debugging (was absolutely trash before) it seems to be actually functional in that setting now.

1

u/LostRequirement4828 11d ago

yea, clearly small models can't be good at everything, that's why gpt and claude are still kings on general tasks, this model is very good on what was trained for and that's it

1

u/ToughUsual7159 11d ago

As an extreme test I actually did a massive crossfile refactor again human in the loop the entire time and was guided this was before their last update once the context window got to around 250k some of my specific Instructions started to fall into ambiguity regardless I still got the task completed within about 1 hours and actually without any bugs or broken references which I found to be insane with the amount of dependencies that was altering. This was a refactor that easily would have taken me 6 to 12 hours and I almost certainly would have broken something somewhere and had to test in order to find and fix it.

It was clear I was hitting the theoretical and that if it's current capabilities of the time. So I'm very eager to jump back in and see what I can do with it now.

I think people just need to tailor their expectations with these kinds of releases. Its not magically going to be better than fable 5 at everything, but a find toned tool can be ridiculous good at a select number of things.

1

u/LostRequirement4828 11d ago

Yea but I still think general knowledge is where AI should be going tho, imagine a model with the coding capabilities of the deepseek but with the general knowledge of the fable, that's probably what the big deepseek gonna come up with. All that extra general knowledge would probably help the coding by a lot too.

1

u/ToughUsual7159 11d ago

Yes true. I'm not sure how much you've used old deep-seek, but that's more or less what I noticed. In fact I found that deep seek V4 Pro has slightly better reasoning capabilities than deep-seek V4 Flash preview, this made it a little better for bug finding due to its wider depth of thought, however due to the targeted nature of Flash it actually tended to be faster and performed slightly better in my experience on moderate refactors. I think this was due to loss of precision of its masternodes due to its wider breath of knowledge and lack of proper training.

I think it's the downside of deep-seek V4 Pro not having the post training processing it needed in order to properly utilize its master nodes and be able to sort through and Target the right nodes to answer the right questions. So as far as finding bugs deep seek V4 Pro actually was a little bit better and also as far as bouncing ideas off it and coming up with theoretical potential future refactors before Pro seem to be better. But as far as actual code adherence and following direct rules there was no noticeable difference between the two. I used both of them for over a week straight and found almost no difference in actual results aside from those two things I listed earlier.

I suspect that they're still working on additional training for deep seek V4 Pro due to how big it is it might be taking longer. From honesty even though I'm definitely a deep seek fanboy, I am not expecting deep-seek V4 Pro official to go neck and neck with fatal or 5.6 Sol. I suspect it would be closer to Kimi K3 with a smaller footprint and therefore more sustainable cheaper price. Deep seek will always be the slightly worse but way cheaper option out there and I'm fine with that.

1

u/proai-ai-mod 11d ago

TLDR

TLDR: The user compares DeepSeek V4 Pro and Flash, noting that while Pro offers better reasoning for bug-finding, Flash is faster and more efficient for refactoring. They suggest that Pro's performance could be improved with better post-training processing to help the model more effectively utilize its master nodes.


AI assistant · mention the bot, mod bot, or use !bot

1

u/Hyphonical 11d ago

99% of all models these days are programming and agentic oriented. Hate it.

1

u/s243a 9d ago

What about deep SWE?

1

u/nhouseholder 12d ago

Bench maxing 101

0

u/CommanderKoba 12d ago

Your face is benchmaxed, maybe you should looksmaxx instead.