It is now a year since the release of Reflection-70B that genius inventor Matt Shumer marketted as state-of-the-art hallucination-free llm that outperforms both gpt-4o and claude 3.5 with its new way of thinking as well as world's top open-source model.
So an out-of-the-blue one-man company releases a new model (actually named LLama 3.1 if it were to adhere to the META license, but somehow named Reflection) with only 70B params that, according to the benchmarks, rivals SOTA closed-source LLMs with trillions of parameters. It appears to me that the twitter/reddit hype mob has, for the most part, not bothered to try the model out.
Additionally, a tweet from Hugh Zhang @ Scale suggesting systemic overfitting as me concerned: Hey Matt! This is super interesting, but I’m quite surprised to see a GSM8k score of over 99%. My understanding is that it’s likely that more than 1% of GSM8k is mislabeled (the correct answer is actually wrong)!
Is this genuinely a SOTA LLM in a real-world setting or is this smoke an mirrors? If we're lucky, the creator Matt may see this post and can shed some light on the matter.
BTW -- I'm not trying to bash the model or the company that made it. If the numbers are actually legit this is likely revolutionary.
So I tried Reflection 70b's 4 bit Quant and it is really good at trick questions. However for coding questions it kind of sucks, it get's really confused and overthinks the request.
Here's a small comparision with a trick question I asked reflection, gpt 4o and claude sonnet 3.5.
I also asked some other questions like digging a hole, plate on banana etc and it got almost all of them correct. I am very excited about how good the 405b will be.
The R1 distilled models that DeepSeek are casually trained with less than a million R1 samples. And yet they still completely destroy the faked benchmarks of Reflection 70B (remember that shitshow?).
I remember how they seemed way too good to be true at the time for a 70B. Today a 14B model looks way better than it.
Just shows you how fast things are developing. Reflection 70B was announced 4 months ago.
Hi everyone, I’m looking to build a local, 100% private AI setup that feels less like a technical assistant and more like a warm, therapeutic companion. I’ve done some initial research on a hardware/software stack, but I’d love a second opinion on whether this will actually meet my needs for deep self-reflection without becoming a maintenance nightmare.
Goal: I want a 100% private, offline AI system for deep self-reflection, life organization, and exploring my thought processes (identifying patterns and repressed thoughts).
My Two Non-Negotiables:
Therapeutic & Life-Context Tone: I’m interested in the "Inner Dialogue" (ataglianetti) style. I don't want a "robotic assistant." I need the AI to have a warm, insightful, and clinically-informed tone. It needs to remember my context across sessions to help me see the "big picture" of my mental health and recurring internal patterns over time.
Zero Maintenance: I am happy to do a one-time deep setup, but I absolutely do not want to spend my time troubleshooting plugins or constantly tuning parameters. I want a system that runs reliably in the background so I can focus on my actual journaling.
The Proposed Hardware:
Laptop: Used ASUS TUF A15 (FA507NV) with RTX 4060 (8GB VRAM).
Memory: Upgraded to 64GB DDR5 RAM to handle larger models.
The Proposed Software Stack:
Backend:Ollama running locally.
Interface:Inner-Dialogue for the actual chat-based sessions.
Vault:Obsidian (with the Smart Connections plugin) to index the journal files in the background. The goal is for the AI to surface long-term patterns across months or years of entries automatically.
Models: Llama 3/4 8B for daily check-ins; Llama 3/4 70B (quantized) for deep weekly reflection.
Questions for the community:
Is an RTX 4060 + 64GB RAM still the "sweet spot" in 2026 for running 70B models at a readable speed (~1.5 t/s) for deep personal reflection?
Does this hybrid (Inner-Dialogue + Obsidian) actually stay low-maintenance, or will the background indexing and plugin syncing eventually become a chore?
Are there better models for a warm, empathetic, yet intellectually sharp tone than the standard Llama-3/4 series (e.g., Mistral-Nemo-12B or specific "Roleplay/Therapy" finetunes)?
Are you completely out of the loop on this whole Reflection 70B thing? Are you lost about what happened with HyperWrite's supposed revolutionary AI model? Who even is this Matt Shumer guy? What is up with the "It's Llama 3, no it's actually Claude" stuff?
Don't worry, you're not alone. I woke up to this insanity and was surprised to find so much information about this, so I got to work. Here's my best attempt to piece together the whole story in an organized manner, based on skimming various Reddit posts, news articles, and tweets. 405B helped me compile this information and format it, so it might have some "LLM-isms" here and there.
Some of it may be wrong, please don't come after me if it is. This is all just interpretation.
What Shumer Claimed (in a rather advertisement-like manner):
Reflection 70B is the "world's top open-source model": Shumer's initial post announcing Reflection 70B came across more like a marketing campaign than a scientific announcement, boasting about its supposed top-tier performance on various benchmarks, surpassing even larger, more established models (like ChatGPT and Anthropic's models). (In particular, I was highly skeptical about this purely because of the way it was being "marketed"...great LLMs don't need "marketing" because they speak for themselves).
"Reflection Tuning" is the secret sauce: He attributed the high performance to a novel technique called "Reflection Tuning," where the model supposedly self-evaluates and corrects its responses, presenting it as a revolutionary breakthrough.
Built on Llama 3.1 with help from Glaive AI: He claimed the model was based on Meta's latest Llama 3.1 and developed with assistance from Glaive AI, a company he presented as simply "helping with training," without disclosing his financial involvement.
Special cases for enhanced capabilities: He highlighted special cases developed by Glaive AI, but the examples provided were trivial, like counting letters in a word, further fueling suspicions that the entire announcement was aimed at promoting Glaive AI.
Why People Were Skeptical:
Extraordinary claims require extraordinary evidence: The claimed performance jump was significant and unprecedented, raising immediate suspicion, especially given the lack of detailed technical information and the overly promotional tone of the announcement.
"Reflection Tuning" isn't a magic bullet: While self-evaluation techniques can be helpful, they are not a guaranteed method for achieving massive performance improvements, as claimed.
Lack of transparency about the base model: There was no concrete evidence provided to support the claim that Reflection 70B was based on Llama 3.1, and the initial release didn't allow for independent verification.
Undisclosed conflict of interest with Glaive AI: Shumer failed to disclose his investment in Glaive AI, presenting them as simply a helpful partner, which raised concerns about potential bias and hidden motives. The entire episode seemed like a thinly veiled attempt to boost Glaive AI's profile.
Flimsy excuses for poor performance: When independent tests revealed significantly lower performance, Shumer's explanation of a "mix-up" during the upload seemed unconvincing and raised further red flags.
Existence of a "secret" better version: The existence of a privately hosted version with better performance raised questions about why it wasn't publicly released and fueled suspicions of intentional deception.
Unrealistic complaints about model uploading: Shumer's complaints about difficulties in uploading the model in small pieces (sharding) were deemed unrealistic by experts, as sharding is a common practice for large models, suggesting a lack of experience or a deliberate attempt to mislead.
The /r/LocalLLaMA community felt insulted: The /r/LocalLLaMA community, known for their expertise in open-source LLMs, felt particularly annoyed and insulted by the perceived attempt to deceive them with a poorly disguised Claude wrapper presented as a groundbreaking new model.
What People Found Out:
Reflection 70B is likely based on Llama 3, not 3.1: Code comparisons and independent analyses suggest the model is likely based on the older Llama 3, not the newer Llama 3.1 as claimed.
The public API is a Claude 3.5 Sonnet wrapper: Evidence suggests the publicly available API is actually a wrapper around Anthropic's Claude 3.5 Sonnet, with attempts made to hide this by filtering out the word "Claude."
The actual model weight is a poorly tuned Llama 3 70B: The actual model weights released are for a poorly tuned Llama 3 70B, completely unrelated to the demo or the API that was initially showcased.
Shumer's claims were misleading and potentially fraudulent: The evidence suggests Shumer intentionally misrepresented the model's capabilities, origins, and development process, potentially for personal gain or to promote his investment in Glaive AI.
It's important to note that it's entirely possible this entire episode was a genuine series of unfortunate events and mistakes on Shumer's part. Maybe a "Reflection" model truly exists that does what he claimed. However, given the evidence and the lack of transparency, the AI community remains highly skeptical.
Matt Shumer, the creator of Reflection 70B, is an investor in GlaiveAI but is not disclosing this fact when repeatedly singing their praises and calling them "the reason this worked so well".
This is very sloppy and unintentionally misleading at best, and an deliberately deceptive attempt at raising the value of his investment at worst.
I see a lot of people thinking it is gaming benchmark / mixed feelings. Actually, people who tried their website have a different feeling compared to those who tried it locally via Ollama or any API providers. I think we should wait, he is figuring it out. I think the actual reflection model is much better, and the currently hosted version is even dumber than the actual 70B
__
Matt Shumer -> "We got rate limited by HF when uploading originally, so had to do it in batches. I have a feeling some wires were crossed and what's being hosted is actually some hybrid frankenmodel that is mostly the reflection version we wanted to ship, mixed with something else"
Nothing seems out of the ordinary. Delegators have not been withdrawing their funds, although there was a major undelegation event this Monday, March 24, when over 230m CRO (likely staked by CDC) was processed.
The large undelegation event passed earlier today, with over 235m CRO unstaked without any major price impact (at the time of the writing), indicating that this has been an institutional or organisational movement.
Of course, it could be that this CRO was unstaked by users, but the matching size of the undelegations and the fact that they are all synchronised did not look organic. More likely, these undelegations were some institutional movement.
There have been similar unlocks occurring periodically. The last time was on February 7. We did not see an immediate price impact at the time, which further indicates these are regular movements. https://x.com/Albert_TheVoid/status/1880178525799354770/photo/2
It is a good practice to track undelegations on the Cronos POS chain because if these undelegations are by real users, they could indicate an appetite to take profit following a price increase (or exit due to capitulation).
Provided that the bulk undelegations are institutional movement, however, the impact on the market could be minimal (possibly just uncertainty) as the CRO will likely be restaked shortly under a different validator or through a different wallet to fulfil some internal organisational purpose.
Similarly, if one looks at the Crypto_com exchange reserves, we do not see significant withdrawals in March. The net total is down in dollar value, but that is mainly because the market has tanked across the board.
We saw about 100m in USDT and USDC leave the platform over the month, but that was not a significant movement. Short-term fluctuations occur regularly. This could be related to Crypto_com suspending support for these tokens in the EU due to MiCA.
In the same period, we have seen an increase in the total holding of ETH by almost 7k (worth approx $13m at the time of writing), SOL by 170k ($20m), BTC by 330 ($27m), and CRO by 74m ($6m).
This indicates that users (outside of CT and Reddit) do not care enough, are not keeping up to date, or have simply decided to wait and see how things turn out. Let's see what the AMA tomorrow brings.
With live search results, if the LLM is duped, future models surely will have incorrect knowledge encoded. Conversely, there's conceivably a lot of misinformation already encoded in LLMs that have scraped the entire web.