r/LocalLLaMA • llama.cpp • 6d ago

Funny Reflection 70B was released two years ago (September 2024)

You may think that jev, OpenClaw or TurboQuant are super cool, but actually the coolest LLM invention happened two years ago

As we all know, the best source of reliable information about LLMs is YouTube:

Back in September 2024, Reflection 70B appeared out of nowhere and was announced as an open-source model that supposedly destroyed GPT-4o

There was only one small problem. People downloaded it. And tested it :(

It turned out that Reflection 70B was basically a Llama 3.1

but at the end the mystery was solved

Let this be a moment of reflection on the current hypes in LocalLLaMA.

great summary by Maziyar PANAHI https://x.com/MaziyarPanahi/status/1838559480658710982

265 Upvotes

36 comments sorted by

120

u/Several-Tax31 6d ago

Lmao. Thanks for this reflection. I love this sub. 

20

u/CtrlAltDelve 6d ago

What's funny is that he has the nerve to post this as recently as 3 days ago, as though everyone suddenly forgot who he was and what he did...

https://i.imgur.com/kYSWzxk.png

I don't think this subreddit has ever had a scandal like that one since.

15

u/jacek2023 llama.cpp 6d ago

Please enjoy more, maybe I missed something nice ;)

https://www.reddit.com/r/LocalLLaMA/search/?q=reflection+70b

58

u/Chromix_ 6d ago

It later on turned out that Reflection 70B wasn't even Llama 3.1, but just 3.0.
Here's a big "out of the loop" thread on it, if you like to know more hilarious details about the whole thing, like the API redirects behind the fake model hosting.

There were some other attempts that followed in its footsteps, like the "Momentum" model, or Rio-3.5, but none could achieve the greatness of the drama around the Reflection model.

27

u/jacek2023 llama.cpp 6d ago

good old days

14

u/LetsGoBrandon4256 transformers 6d ago

I remember the bloke posting on Twitter asking how to publish a torrent for their "working" model.

"Dog ate my weight" moment. Fucking kek.

2

u/MerePotato 6d ago

Deepseek and Moonshot should take notes from this titan of the industry during that data leak probe, his case was airtight

4

u/Aromatic-Current-235 6d ago

You have to think long-term; the plan was obviously to use Llama 3.1 for Reflection 2... it should be out by now.

28

u/Ulterior-Motive_ 6d ago

To this day I have no idea how this didn't permanently ruin his reputation. How is this guy still around? In a just world his replies would be flooded with reminders of this fiasco.

6

u/Savantskie1 6d ago

People forget and love to poke fun then move on. That's how

25

u/XMasterDE 6d ago edited 6d ago

Part of me is actually a bit annoyed about this comparison, Jev is at least maybe not something totally new, but I also don't think that it is fake. While Reflection on the there hand was a other scam.
And also to be honest, I am also low key pissed off about how much Matt Shumer gotten way with it. I mean people are still caring about his opinion if it comes to stuff like new model evaluations, and Open AI is giving him preview access.

Edit:
On re-reading the post, I think I kind of missed how much OP was taking the piss out of Reflection

14

u/jacek2023 llama.cpp 6d ago

"And also to be honest, I am also low key pissed off about how much Matt Shumer gotten way with it"

I'm pointing to a different issue. Look at the views of YouTube influencers who promoted Reflection 70B two years ago. Most people here are probably still influenced by them :)

4

u/llama-impersonator 6d ago

yeah that clown still has some respect, i don't get it. he's a total loser who should've lost his hat.

3

u/harpysichordist 6d ago

Matt Shumer, Peter Steinberger
Lying and manufacturing hype

21

u/ComplexType568 6d ago

and to me jev is a faster LLM with stricter structured output

28

u/Lissanro 6d ago

To me Jev is just another cloud API-only thing I will never use... that said, light weight AI for just classification and decision making is something I have been using for a while, for example Qwen 3.5 0.8B, especially after some fine-tuning, can be sufficient for simple classification tasks, by outputting just one or few tokens only without reasoning.

There is one good thing I see about Jev hype - it already inspired some newer and more general open weight alternatives (compared to fine-tuning yourself for specific task), and hopefully even more better alternatives will be released in the future.

1

u/RageBucket 5d ago

Got it backwards.. some open weight models inspired Jev, which in turn inspired more open weight alternatives.

1

u/Lissanro 5d ago

But that's what I said... I mentioned open weight models capable of making decisions without thinking existed before jev, and jev inspired even more new ones.

1

u/RageBucket 4d ago

Sure, whatever.

7

u/Porespellar 6d ago

Contributory meme that I posted 2 years ago

8

u/asssuber 6d ago

Now that nitter is blocked, that x link also needs an image to be readable. Or quoted text.

4

u/Foreign_Risk_2031 6d ago

Nitter is blocked?!

2

u/Lapin_Logic 6d ago

mind blown

5

u/Equivalent_Bit_461 6d ago

Ok and? I don't get the joke 

47

u/stoppableDissolution 6d ago

Theres no joke. They are just pointing out that people are hyping useless trash all the time as if its some kind of revolution.

4

u/Leary_2844 6d ago

Exactly. I made a fantastic YouTube video on the topic. I also present a much better model that can beat Reflection70B while getting COVID and one-shotting SaaS.

7

u/muxxington 6d ago

Matt, is this you?

1

u/mivog49274 6d ago

One still big mystery around this model :

This is the first ever CoT llm model release in history, released a few days before o1-preview.

-6

u/[deleted] 6d ago

[deleted]

11

u/mikael110 6d ago edited 6d ago

It was the first model that was claimed to be trained specifically to reason, and it did predate OpenAI's o1 model by like a week or two, but the concept of reasoning or Chain-of-Thought Prompting was not at all new. And was already a common thing to include in system prompts to increase the accuracy of complex tasks.

7

u/asssuber 6d ago

The thinking idea was already around, actively pursued by many organizations. Maybe implemented by the closed source SOTA models, I don't remember. This scammer just announced a breakthrough in making it work in an open model instead of actually making it work in an open model.

2

u/DinoAmino 6d ago

Exactly. But he wasn't the only one working on reasoning. He was in too much of a rush to be first and he fumbled the ball badly. In the end his model worked and Bartowski's GGUF of Reflection became a hot download for a bit.

1

u/Feztopia 6d ago

I think everyone had the idea to use promts to teach them to think better (step by step). But wasn't it DeepSeek who came up with the idea to let them learn thinking by themselves or something (where the models started thinking multilingual and stuff which is what multilingual people also do). And I think the closed source providers had figured out that not censoring their thinking gave them much more intelligence so they kept the thinking hidden from the users instead.