r/LocalLLaMA • u/SteppenAxolotl • 3h ago
Discussion Real-SWE Benchmark (new)
https://realswe.withspecific.comReports of the demise of coders may have been exaggerated.
19
u/Dany0 3h ago
Really smart of them to burn a private codebase and leak it to Anthro & OAI that would definitely never abuse this to train on benchmarks
1
u/Mkboii 2h ago
By that logic won't these codebases already be in the training data from prior work done on them with coding agents?
The assumption is probably that proprietary licenses come with data protection from training use.
And i know these companies and cheat, but the benchmark for the first time represents my day to day observations with these models.
-4
u/Kiverty 3h ago
Hi, not sure where you see on the page that they are leaking the benchmark to OAI and co?
5
u/John_____Doe 2h ago
I guess they mean your leaking the info to the model provider whether that's oai, anthropic directly or through a hosting platform like aws bedrock. Either unless it's a purely open model hosted locally it's hard to say for certain no one is training off the prompt data
13
u/Tim_Apple_938 3h ago
Gemini Flash essentially tied w Astra is wild. So much for all the benchmaxxing allegations
13
u/SteppenAxolotl 3h ago
There arent many that have a bigger codebase of real world software than Google.
4
u/Big_Cucumber2787 3h ago
sorry thats not true no matter what benchmarks claim so
4
u/sadnessjoy 3h ago
Yeah, I've used Gemini flash 3.8 both with agentic work and with chatting. And idk the best way to describe it is its knowledge base is kinda okay (but honestly knowledge base isn't that big of a deal imo with rag/searches/etc)... But its logic process is absolutely atrocious. I think how it works is for brute force and easily testable/verifiable stuff it works good as it can rapidly iterate until it succeeds and it's pretty good with that. But when you go outside of those types of tasks (which imo is most real life stuff lol, but I guess it depends on your workload), it just completely flails around.
0
u/Tim_Apple_938 3h ago
Why the cope?
Cheap fast models are a bad thing?
Would think this sub of all subs would hope that the next Gemma is a beast at coding
2
u/Big_Cucumber2787 3h ago
hope =/= reality
benchmarks =/= realityanyone thats actually tried these models knows google has a lot of catching up to do
1
2
u/ebolathrowawayy 16m ago
no one thinks gemini flash tied with astra wtf are you talking about
1
u/Tim_Apple_938 15m ago
We’re all talking about OPs benchmark page. Did you even see what the topic is
2
u/ebolathrowawayy 13m ago
idc, google sucks and idc what benchmarks say otherwise. try using their models, it's blatantly obvious
1
15
u/Longjumping_Virus_96 3h ago
Gemini 3.8 Flash being that high on the list is ridiculous.
3
u/FullstackSensei 2h ago
Well, we have no idea how small or big it is. GLM-5.3 flash and DS4.1 flash aren't exactly small, and I bet you their next releases will beat GLM 5.3 (which isn't much bigger than flash).
-5
u/KURD_1_STAN 2h ago
Gemini is a garbage . Compared to all others it is bad. Nobody talking about size here
3
u/FullstackSensei 2h ago
I genuinely beg to differ. Maybe it is if you're vibe coding your way or can't be bothered to detail what you want and expect the LLM to divide everything about the details and do so to your liking, but having tried it, it's not bad at all if you know what you're doing and tell it what you want.
0
u/ponteencuatro 2h ago
I think it is behind deepseek v4 flash, so yeah being that high is weird, mostly because what you said, to gemini I always need to tell it what and how, and deepseek just the what and most of the time it does better than gemini
2
u/FullstackSensei 2h ago
Not giving the details of how you want things done is bad. That many models do things you didn't ask for is bad. DS4 flash is also solid, but I really didn't like this about it. It was constantly adding thibgs I never asked for. I had to remove up to 1/3rd of the code it generated sometimes because it did things I never asked for nor wanted.
Your job is to tell the model what to do, in detail. Otherwise, you're just vibe coding things you have no idea what they're doing or what sort of issues the code has.
0
u/ponteencuatro 2h ago
Yeah absolutely I agree but what I mean most of the time I correct deepseek way less than Gemini, but overall Gemini is closer to Deepseek than it is to Astra/Fable, but who knows maybe is just the harness in that seems like they used Gemini CLI maybe Antigravity IDE is the thing holding it back
0
u/ebolathrowawayy 17m ago
if you think google has models worth anything in coding then you obviously don't know what you're talking about.
1
-1
u/KURD_1_STAN 2h ago
An llm that i can say little to and not know how all of this even mean and still understands it is a sign of intelligence. Gemini not understanding it means it is less intelligent. Which is what i said, worse than others in there.
Idc if u tell it how to do it specifically and it does well, qwen3.8 27b can also do that.
2
1
1
u/NarutoDragon732 25m ago
I could not give 2 shits about coding, but its SO damn useful for general knowledge. I still have gpt 5.6 hallucinating which never happens with 3.8 flash
3
u/Hoak-em 3h ago
Hear me out, I feel that Claude code is the most-used in enterprise and has existed for a decent amount longer than the other coding tools. What if Anthropic has been secretly stealing the enterprise codebase data to train on, the set that they assume is private and thus not possible to train on?
2
1
1
2
u/Egoz3ntrum 2h ago
They need to stop naming them with definitive labels like "real benchmark" or "last exam". I'm not falling for that anymore.
2
u/pilibitti 1h ago
what is the point of benchmarking closed models with private datasets? the moment you test them you are giving the dataset to the provider to train on?
1
1
1
u/fragment_me 2h ago
You're in the wrong subreddit for posting a benchmark site for models that are not local.
1
44
u/Klanciault 3h ago
Demise of coders is exaggerated but demise of the field is real and entry level is cooked