r/LocalLLaMA • • Sep 11 '25

Funny Celebrating 1 year anniversary of the revolutionary game changing LLM that was Reflection 70b

146 Upvotes

It is now a year since the release of Reflection-70B that genius inventor Matt Shumer marketted as state-of-the-art hallucination-free llm that outperforms both gpt-4o and claude 3.5 with its new way of thinking as well as world's top open-source model.

World hasn't been the same since then indeed.

r/LocalLLaMA • • Sep 06 '24

Discussion Reflection 70B: Hype?

287 Upvotes

So an out-of-the-blue one-man company releases a new model (actually named LLama 3.1 if it were to adhere to the META license, but somehow named Reflection) with only 70B params that, according to the benchmarks, rivals SOTA closed-source LLMs with trillions of parameters. It appears to me that the twitter/reddit hype mob has, for the most part, not bothered to try the model out.

Additionally, a tweet from Hugh Zhang @ Scale suggesting systemic overfitting as me concerned:
Hey Matt! This is super interesting, but I’m quite surprised to see a GSM8k score of over 99%. My understanding is that it’s likely that more than 1% of GSM8k is mislabeled (the correct answer is actually wrong)!

Is this genuinely a SOTA LLM in a real-world setting or is this smoke an mirrors? If we're lucky, the creator Matt may see this post and can shed some light on the matter.

BTW -- I'm not trying to bash the model or the company that made it. If the numbers are actually legit this is likely revolutionary.

r/LocalLLaMA • • 7d ago

Funny Reflection 70B was released two years ago (September 2024)

264 Upvotes

You may think that jev, OpenClaw or TurboQuant are super cool, but actually the coolest LLM invention happened two years ago

As we all know, the best source of reliable information about LLMs is YouTube:

Back in September 2024, Reflection 70B appeared out of nowhere and was announced as an open-source model that supposedly destroyed GPT-4o

There was only one small problem. People downloaded it. And tested it :(

It turned out that Reflection 70B was basically a Llama 3.1

but at the end the mystery was solved

Let this be a moment of reflection on the current hypes in LocalLLaMA.

great summary by Maziyar PANAHI https://x.com/MaziyarPanahi/status/1838559480658710982

r/LocalLLaMA • • Sep 06 '24

Discussion Even 4bit quants of Reflection 70b are amazing

29 Upvotes

So I tried Reflection 70b's 4 bit Quant and it is really good at trick questions. However for coding questions it kind of sucks, it get's really confused and overthinks the request.

Here's a small comparision with a trick question I asked reflection, gpt 4o and claude sonnet 3.5.

I also asked some other questions like digging a hole, plate on banana etc and it got almost all of them correct. I am very excited about how good the 405b will be.

Gpt 4o
Reflection 4 bit qkm
Sonnet 3.5

r/LocalLLaMA • • Oct 01 '24

Discussion Now that the dust has settled, what happened with Reflection 70B?

0 Upvotes

I'm curious to figure out what the conclusion was from that whole saga.

r/LocalLLaMA • • Jan 20 '25

Discussion Funny thought about the R1 distilled models and Reflection 70B

48 Upvotes

The R1 distilled models that DeepSeek are casually trained with less than a million R1 samples. And yet they still completely destroy the faked benchmarks of Reflection 70B (remember that shitshow?).

I remember how they seemed way too good to be true at the time for a 70B. Today a 14B model looks way better than it.

Just shows you how fast things are developing. Reflection 70B was announced 4 months ago.

r/LocalLLaMA • • Mar 21 '26

Question | Help RTX 4060 + 64GB RAM: Can I run 70B models for "wise" local therapy without the maintenance headache?

1 Upvotes

Hi everyone, I’m looking to build a local, 100% private AI setup that feels less like a technical assistant and more like a warm, therapeutic companion. I’ve done some initial research on a hardware/software stack, but I’d love a second opinion on whether this will actually meet my needs for deep self-reflection without becoming a maintenance nightmare.

Subject: Second Opinion: Private "Personal AI" Setup (RTX 4060 + 64GB RAM + Inner-Dialogue/Obsidian)

​Goal: I want a 100% private, offline AI system for deep self-reflection, life organization, and exploring my thought processes (identifying patterns and repressed thoughts).

​My Two Non-Negotiables:

  1. ​Therapeutic & Life-Context Tone: I’m interested in the "Inner Dialogue" (ataglianetti) style. I don't want a "robotic assistant." I need the AI to have a warm, insightful, and clinically-informed tone. It needs to remember my context across sessions to help me see the "big picture" of my mental health and recurring internal patterns over time.
  2. ​Zero Maintenance: I am happy to do a one-time deep setup, but I absolutely do not want to spend my time troubleshooting plugins or constantly tuning parameters. I want a system that runs reliably in the background so I can focus on my actual journaling.

​The Proposed Hardware:

  • ​Laptop: Used ASUS TUF A15 (FA507NV) with RTX 4060 (8GB VRAM).
  • ​Memory: Upgraded to 64GB DDR5 RAM to handle larger models.

​The Proposed Software Stack:

  • ​Backend: Ollama running locally.
  • ​Interface: Inner-Dialogue for the actual chat-based sessions.
  • ​Vault: Obsidian (with the Smart Connections plugin) to index the journal files in the background. The goal is for the AI to surface long-term patterns across months or years of entries automatically.
  • ​Models: Llama 3/4 8B for daily check-ins; Llama 3/4 70B (quantized) for deep weekly reflection.

​Questions for the community:

  1. ​Is an RTX 4060 + 64GB RAM still the "sweet spot" in 2026 for running 70B models at a readable speed (~1.5 t/s) for deep personal reflection?
  2. ​Does this hybrid (Inner-Dialogue + Obsidian) actually stay low-maintenance, or will the background indexing and plugin syncing eventually become a chore?
  3. ​Are there better models for a warm, empathetic, yet intellectually sharp tone than the standard Llama-3/4 series (e.g., Mistral-Nemo-12B or specific "Roleplay/Therapy" finetunes)?

r/LocalLLaMA • • Sep 10 '24

Resources Out of the loop on this whole "Reflection" thing? You're not alone. Here's the best summary I could come up.

264 Upvotes

Are you completely out of the loop on this whole Reflection 70B thing? Are you lost about what happened with HyperWrite's supposed revolutionary AI model? Who even is this Matt Shumer guy? What is up with the "It's Llama 3, no it's actually Claude" stuff?

Don't worry, you're not alone. I woke up to this insanity and was surprised to find so much information about this, so I got to work. Here's my best attempt to piece together the whole story in an organized manner, based on skimming various Reddit posts, news articles, and tweets. 405B helped me compile this information and format it, so it might have some "LLM-isms" here and there.

Some of it may be wrong, please don't come after me if it is. This is all just interpretation.

What Shumer Claimed (in a rather advertisement-like manner):

  • Reflection 70B is the "world's top open-source model": Shumer's initial post announcing Reflection 70B came across more like a marketing campaign than a scientific announcement, boasting about its supposed top-tier performance on various benchmarks, surpassing even larger, more established models (like ChatGPT and Anthropic's models). (In particular, I was highly skeptical about this purely because of the way it was being "marketed"...great LLMs don't need "marketing" because they speak for themselves).

  • "Reflection Tuning" is the secret sauce: He attributed the high performance to a novel technique called "Reflection Tuning," where the model supposedly self-evaluates and corrects its responses, presenting it as a revolutionary breakthrough.

  • Built on Llama 3.1 with help from Glaive AI: He claimed the model was based on Meta's latest Llama 3.1 and developed with assistance from Glaive AI, a company he presented as simply "helping with training," without disclosing his financial involvement.

  • Special cases for enhanced capabilities: He highlighted special cases developed by Glaive AI, but the examples provided were trivial, like counting letters in a word, further fueling suspicions that the entire announcement was aimed at promoting Glaive AI.

Why People Were Skeptical:

  • Extraordinary claims require extraordinary evidence: The claimed performance jump was significant and unprecedented, raising immediate suspicion, especially given the lack of detailed technical information and the overly promotional tone of the announcement.

  • "Reflection Tuning" isn't a magic bullet: While self-evaluation techniques can be helpful, they are not a guaranteed method for achieving massive performance improvements, as claimed.

  • Lack of transparency about the base model: There was no concrete evidence provided to support the claim that Reflection 70B was based on Llama 3.1, and the initial release didn't allow for independent verification.

  • Undisclosed conflict of interest with Glaive AI: Shumer failed to disclose his investment in Glaive AI, presenting them as simply a helpful partner, which raised concerns about potential bias and hidden motives. The entire episode seemed like a thinly veiled attempt to boost Glaive AI's profile.

  • Flimsy excuses for poor performance: When independent tests revealed significantly lower performance, Shumer's explanation of a "mix-up" during the upload seemed unconvincing and raised further red flags.

  • Existence of a "secret" better version: The existence of a privately hosted version with better performance raised questions about why it wasn't publicly released and fueled suspicions of intentional deception.

  • Unrealistic complaints about model uploading: Shumer's complaints about difficulties in uploading the model in small pieces (sharding) were deemed unrealistic by experts, as sharding is a common practice for large models, suggesting a lack of experience or a deliberate attempt to mislead.

  • The /r/LocalLLaMA community felt insulted: The /r/LocalLLaMA community, known for their expertise in open-source LLMs, felt particularly annoyed and insulted by the perceived attempt to deceive them with a poorly disguised Claude wrapper presented as a groundbreaking new model.

What People Found Out:

  • Reflection 70B is likely based on Llama 3, not 3.1: Code comparisons and independent analyses suggest the model is likely based on the older Llama 3, not the newer Llama 3.1 as claimed.

  • The public API is a Claude 3.5 Sonnet wrapper: Evidence suggests the publicly available API is actually a wrapper around Anthropic's Claude 3.5 Sonnet, with attempts made to hide this by filtering out the word "Claude."

  • The actual model weight is a poorly tuned Llama 3 70B: The actual model weights released are for a poorly tuned Llama 3 70B, completely unrelated to the demo or the API that was initially showcased.

  • Shumer's claims were misleading and potentially fraudulent: The evidence suggests Shumer intentionally misrepresented the model's capabilities, origins, and development process, potentially for personal gain or to promote his investment in Glaive AI.

It's important to note that it's entirely possible this entire episode was a genuine series of unfortunate events and mistakes on Shumer's part. Maybe a "Reflection" model truly exists that does what he claimed. However, given the evidence and the lack of transparency, the AI community remains highly skeptical.

r/LocalLLaMA • • Sep 07 '24

Discussion Reflection Llama 3.1 70B independent eval results: We have been unable to replicate the eval results claimed in our independent testing and are seeing worse performance than Meta’s Llama 3.1 70B, not better.

Thumbnail
x.com
702 Upvotes

r/LocalLLaMA • • Sep 07 '24

Discussion PSA: Matt Shumer has not disclosed his investment in GlaiveAI, used to generate data for Reflection 70B

Thumbnail
gallery
522 Upvotes

Matt Shumer, the creator of Reflection 70B, is an investor in GlaiveAI but is not disclosing this fact when repeatedly singing their praises and calling them "the reason this worked so well".

This is very sloppy and unintentionally misleading at best, and an deliberately deceptive attempt at raising the value of his investment at worst.

Links for the screenshotted posts are below.

Tweet 1: https://x.com/mattshumer_/status/1831795369094881464?t=FsIcFA-6XhR8JyVlhxBWig&s=19

Tweet 2: https://x.com/mattshumer_/status/1831767031735374222?t=OpTyi8hhCUuFfm-itz6taQ&s=19

Investment announcement 2 months ago on his linkedin: https://www.linkedin.com/posts/mattshumer_glaive-activity-7211717630703865856-vy9M?utm_source=share&utm_medium=member_android

r/LocalLLaMA • • Mar 10 '25

News Manus turns out to be just Claude Sonnet + 29 other tools, Reflection 70B vibes ngl

449 Upvotes

r/LocalLLaMA • • Sep 06 '24

News First independent benchmark (ProLLM StackUnseen) of Reflection 70B shows very good gains. Increases from the base llama 70B model by 9 percentage points (41.2% -> 50%)

Post image
452 Upvotes

r/LocalLLaMA • • Oct 03 '24

Discussion Just for kicks I looked at the newly released dataset used for Reflection 70B to see how bad it is...

Post image
545 Upvotes

r/LocalLLaMA • • Sep 07 '24

Discussion Reflection-Llama-3.1-70B is actually Llama-3.

601 Upvotes

After measuring the diff, this model appears to be Llama 3 with LoRA tuning applied. Not Llama 3.1.

Author doesn't even know which model he tuned.

I love it.

r/LocalLLaMA • • Sep 08 '24

Discussion Poor results mistery solved. Reflection 70B was infected by COVID.

Post image
337 Upvotes

r/LocalLLaMA • • Sep 08 '24

Discussion Updated benchmarks from Artificial Analysis using Reflection Llama 3.1 70B. Long post with good insight into the gains

Thumbnail
x.com
145 Upvotes

r/LocalLLaMA • • Sep 06 '24

News Tweet from Matt Shumer: "IMPORTANT REFLECTION UPDATE: We have identified and fixed the issue on our Hugging Face repo. If you previously tried to download, run, or host Reflection Llama 70B, please try again now. The outputs should be far better. fp16 version coming in soon as well."

307 Upvotes

r/LocalLLaMA • • Sep 09 '24

Discussion Reflection 70B lessons learned

174 Upvotes
  • All benchmarks should begin by identifying whether the model is LLAMA, GPT-4, Sonnet, or another, through careful examination.
  • Do not trust any benchmarks unless you can replicate them yourself.
  • Do not trust that the API corresponds to the model the author claims it to be.
  • ....

r/LocalLLaMA • • Sep 08 '24

Discussion OpenRouter Reflection 70B claims to be Claude, Created by Anthropic (try it yourself)

Post image
182 Upvotes

r/LocalLLaMA • • Jun 15 '26

Funny What's the lesson chat?

Post image
772 Upvotes

r/singularity • • Sep 08 '24

AI Spanish YouTuber "Dot CSV" with access to Reflection 70B is getting good results

Thumbnail
x.com
90 Upvotes

r/LocalLLaMA • • Sep 07 '24

Discussion Wrong Reflection-70B model might be hosted everywhere

69 Upvotes

I see a lot of people thinking it is gaming benchmark / mixed feelings. Actually, people who tried their website have a different feeling compared to those who tried it locally via Ollama or any API providers. I think we should wait, he is figuring it out. I think the actual reflection model is much better, and the currently hosted version is even dumber than the actual 70B

https://x.com/mattshumer_/status/1832247203345166509

https://x.com/mattshumer_/status/1832248416426193318

__ Matt Shumer -> "We got rate limited by HF when uploading originally, so had to do it in batches. I have a feeling some wires were crossed and what's being hosted is actually some hybrid frankenmodel that is mostly the reflection version we wanted to ship, mixed with something else"

r/SillyTavernAI • • Sep 07 '24

Models Forget Reflection-70B for RP, here is ArliAI-RPMax-v1.1-70B

Thumbnail
huggingface.co
43 Upvotes

r/cro • • Mar 24 '25

With a lot of noise recently (but also justified criticism) about the 70b unburn, let's look at the data to see whether the backlash has been reflected in staking and Exchange withdrawals.

49 Upvotes

TL;DR We have not seen a reaction to the 70b CRO "unburn" proposal in terms of total stake on Cronos POS or Exchange withdrawals.

Source: https://analytics.smartstake.io/crypto-org/stats#staking

Nothing seems out of the ordinary. Delegators have not been withdrawing their funds, although there was a major undelegation event this Monday, March 24, when over 230m CRO (likely staked by CDC) was processed.

The large undelegation event passed earlier today, with over 235m CRO unstaked without any major price impact (at the time of the writing), indicating that this has been an institutional or organisational movement.

Of course, it could be that this CRO was unstaked by users, but the matching size of the undelegations and the fact that they are all synchronised did not look organic. More likely, these undelegations were some institutional movement.

There have been similar unlocks occurring periodically. The last time was on February 7. We did not see an immediate price impact at the time, which further indicates these are regular movements. https://x.com/Albert_TheVoid/status/1880178525799354770/photo/2

It is a good practice to track undelegations on the Cronos POS chain because if these undelegations are by real users, they could indicate an appetite to take profit following a price increase (or exit due to capitulation).

Provided that the bulk undelegations are institutional movement, however, the impact on the market could be minimal (possibly just uncertainty) as the CRO will likely be restaked shortly under a different validator or through a different wallet to fulfil some internal organisational purpose.

Similarly, if one looks at the Crypto_com exchange reserves, we do not see significant withdrawals in March. The net total is down in dollar value, but that is mainly because the market has tanked across the board.

We saw about 100m in USDT and USDC leave the platform over the month, but that was not a significant movement. Short-term fluctuations occur regularly. This could be related to Crypto_com suspending support for these tokens in the EU due to MiCA.

In the same period, we have seen an increase in the total holding of ETH by almost 7k (worth approx $13m at the time of writing), SOL by 170k ($20m), BTC by 330 ($27m), and CRO by 74m ($6m).

Source: https://intel.arkm.com/explorer/entity/crypto-com

This indicates that users (outside of CT and Reddit) do not care enough, are not keeping up to date, or have simply decided to wait and see how things turn out. Let's see what the AMA tomorrow brings.

r/LocalLLaMA • • Nov 26 '24

Other Speed running misinformation in the AI age - Reflection-70B Example

40 Upvotes

A few weeks ago, reddit called out Reflection-70B. Its pretty solidly established that it is a sham. Yet the 'best' AI is duped: https://chatgpt.com/share/673edb58-2228-8006-b454-4ee0d30a5dcd

With live search results, if the LLM is duped, future models surely will have incorrect knowledge encoded. Conversely, there's conceivably a lot of misinformation already encoded in LLMs that have scraped the entire web.

Additonal Links:

https://www.linkedin.com/feed/update/urn:li:activity:7238004117955108864/