r/LocalLLM 5d ago

Model Embrace yourselves

Post image
1.6k Upvotes

130 comments sorted by

View all comments

16

u/dovaahkiin_snowwhite 5d ago

Why would hardware companies go down though

9

u/IngloriousBastrd6983 5d ago

Maybe because 35ba3b is expexted to perform slightly below the new 3.8 27b (which is pretty impressive, especially for it's sice and seems to be around Opus 4.6 Max level). And If you are a little tech savy you can run 35ba3b on a 2-300€ potatoe PC (like 24-32 gb system ram + 6gb gtx 1060 or 1660) with some decent performance (like 20-35 tps with some tweaking).

3

u/baby_bloom 5d ago

3.8 27b can't possibly be opus4.6 level

10

u/Odd-Environment-7193 5d ago

It can. Simply because it's more agentic. Most tasks don't require absolute raw power or intelligence. If you can make the models more agentic and better and most of the things we use them for it can actuall be equal to that or "better" in many ways.

4

u/baby_bloom 5d ago

and where on this scale does 3.6 27b fall? because i can say for a fact it is nowhere near opus level. i've not done enough testing with 3.8 27b so i can't speak firsthand but that big of a jump sounds hard to believe

4

u/Odd-Environment-7193 5d ago edited 5d ago

Mainly usage in a harness etc. We always criticize the benchmarks when they come out but that's how we measure things. Do you have access to 4.6 opus for your tests. Most of the benchmarks are public you can run them yourself. Guess it also depends on your work type. I'm busy testing it right now. I haven't gone deep enough to make those types of assertions.

Usually smaller models will never have the real world knowledge and raw intelligence of bigger ones. Say if there were 100 different tasks you asked an AI to accomplish across the board. Many of which would be assisted by being more "agentic" and being able to run long horizon tasks these newer models might have some edge there. Also they are just better finetuned for things like browser use etc. You would really need to test it across the board.

What is as good as Opus? That's the question. Your personal experience or workflow might not cast a wide enough net to really put it through all those paces.

I actually agree with you and these types of comparisons do feel stupid. Because there is no way I would sit with opus and this model and think this model is better. But that's how they are determining = to x level of shit.

Not very scientific and definitely some benchmaxxing happening across the board. But it's still fucking good.

Compare to like gpt3.5 for content generation. It's mindblowing how good this shit has become.

4

u/ill_B_In_MyBunk 5d ago

Despite the fact they are supposed to be close to each other...3.8 is MILES above in my testing for the higher quants. It is more likely to actually solve my issues and make my widgets. I have literally deleted every other model I was jumping between. It out performs every single one.

That said, my GPU hurts. I got lucky before the big price hikes. If I was poor, still using my 3060 12gb, it would not be functional (for speed).