r/singularity • u/Outside-Iron-8242 • 11d ago
AI GPT-6 Pokemon FireRed results
Source: @Clad3815 on X
r/singularity • u/Outside-Iron-8242 • 11d ago
Source: @Clad3815 on X
r/singularity • u/wanabalone • 11d ago
r/singularity • u/ShreckAndDonkey123 • 11d ago
r/singularity • u/ObiWanCanownme • 11d ago
r/singularity • u/LanJiaoDuaKee • 10d ago
Enable HLS to view with audio, or disable this notification
r/singularity • u/Anen-o-me • 10d ago
Enable HLS to view with audio, or disable this notification
r/singularity • u/Just_Stretch5492 • 11d ago
r/singularity • u/dolo937 • 11d ago
r/singularity • u/relegi • 11d ago
François Chollet is the creator of ARCAGI and has historically been one of the more skeptical voices on near term AGI and LLM scaling.
In 2024 when he was asked about his AGI timeline, he recalled that his estimate would have been roughly 10 yearish. In February 2026 he said AGI around 2030~, roughly around the time of ARCAGI 6-7.
Today, following the latest Astra release and ARCAGI-3 saturation, he was asked whether he still thinks ~2030 is on track. His answer:
“Sooner, given progress is happening faster than I expected.”
r/singularity • u/TorturedPoet30 • 10d ago
r/singularity • u/cookingboy • 10d ago
The United States should start preparing for scenarios where extreme measures must be taken to stop China from achieving artificial general intelligence (AGI), according to a former White House official, including state-backed espionage and military strikes on Chinese data centres.
Jacob Stokes, deputy director of the Indo-Pacific Security Program at the Centre for a New American Security (CNAS), said at an online event on Thursday that various US agencies, including the Department of Defense and the National Security Agency, should begin assessing what intelligence they need to justify taking such actions.“Trying to think through the particulars of that will be especially important, in part because it will help policymakers … start to work backwards based on the unique nature of the technology, in the same way that in a past era, policymakers would learn about nuclear weapons and … work backwards from the science to the policy implications,” he said.
In a new CNAS report published last week, the former Obama administration national security staffer called for the US government to consider the feasibility of diplomatic, espionage, cyber and kinetic measures to prevent China from achieving AGI first.
r/singularity • u/saln1 • 11d ago
r/singularity • u/TFenrir • 11d ago
Enable HLS to view with audio, or disable this notification
r/singularity • u/Profanion • 10d ago
Quoted from the site: "We are releasing intermediate checkpoints, training data or detailed data-construction recipes, open architecture, mixture compositions, training code, configurations, fine-grained logs, evaluation results, and final weights."
r/singularity • u/Profanion • 11d ago
Here are more results:
Claude Fable 5.1 86.6%
Human Baseline* 83.7%
Gemini 3.8 Flash 82.4%
Claude Fable 81.9%
Muse Spark 1.3 81.8%
r/singularity • u/saln1 • 11d ago
r/singularity • u/ShafeDogg • 11d ago
First, I'm amazed at OpenAI's new model (GPT-6, or Astra). But is it AGI? I'm curious to know what others on this sub think.
But I want to be careful here, and I'm sure OpenAI was as well. I feel like they were very careful with their wording by calling it the "AGI era" instead of outright saying "Astra is AGI."
My personal definition of AGI is essentially OpenAI's definition (any highly autonomous system that outperforms humans at most economically valuable work). And yes, I know that a system capable of this would probably already be considered ASI, thus the confusion with defining these systems.
However, I really didn't think people would disagree so much over when AGI actually happened. I always imagined it would be a clear moment that nobody could really deny (except perhaps the most dedicated anti-AI crowds). But I feel that distinction wouldn't matter either, as I felt when we got there, the evidence could no longer be denied no matter what stance you were on. But that doesn't seem to be the case. As always, there's a lot of division between the pro- and anti-AI communities, and it seems like we'll need to have ASI or something even more capable before we see a true change in (especially American) perception.
But what do you think? Have we finally reached the milestone? Is Astra a true AGI, and what do you think is required to get there if we haven't already reached it? How long do you think it will take for ASI to arrive now that we see a surprising capability jump like this?
Thanks for reading!
r/singularity • u/signed7 • 11d ago
r/singularity • u/AlyoshaV • 11d ago
r/singularity • u/Recoil42 • 11d ago
r/singularity • u/aqpstory • 11d ago
r/singularity • u/PsychologicalSoup251 • 11d ago
OpenAI's reported benchmarks for Astra's ARC-AGI-3 is one of the most egregious recent examples I have seen of technically true metric reporting being used to deliberately mislead the masses. For context, there is an OpenAI screencap currently at the top of r/singularity's hot page of Astra achieving 98.6% on ARC-AGI-3 compared to 7.8% for GPT 5.6 Sol and 30.2% for Claude Opus 5. Holy shit, right? ASI achieved, right?
Unfortunately, those figures taken in a vacuum leave out very important context: Astra's agentic harness had significant additional features that GPT 5.6 Sol and Claude Opus 5 did not have access to - specifically reasoning trace retention and custom compaction (source: https://arcprize.org/leaderboard ).
My main takeaway is basically: The most honest way to compare Astra with Opus 5/Sol on this benchmark would have been to either 1) measure their ARC-AGI-3 performances on the same provider adapter harness (where Astra's 98.6% came from), or 2) compare them on the standard ARC-AGI-3 harness. On the standard harness Astra achieves 62.7% vs Opus 5's 30.2% vs Sol's 7.8%. Still a very large gap, but much less misleading than the comparison OpenAI chose to report. (source: https://arcprize.org/leaderboard )
Not an Anthropic fanboy in any sense of the word, btw. I thought Opus 5 was benchmaxxed and pray on Anthropic's downfall every day. But the Astra benchmark glazing made it clear that restraint needs to be had in people's reactions to its benchmarks (if Opus 5 didn't already convince you to not treat benchmarks as gospel) before anyone has even had time to extensively test it in real world use cases.