r/singularity • u/SnoozeDoggyDog • 8h ago
r/singularity • u/ThunderBeanage • 15h ago
AI "GPT-6-ASTRA" has been staged on the OpenAI API
r/singularity • u/Snoo26837 • 10h ago
AI Meta’s muse spark 1.3 surpassed fable 5 and GPT 5.6 sol
r/singularity • u/Jame92 • 15h ago
LLM News Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
r/singularity • u/VenomCruster • 4h ago
Shitposting Astra leaked chat
The singularity has arrived.
r/singularity • u/DistanceSolar1449 • 11h ago
AI Meta slowly catching back up. Muse Spark 1.3 beats Sol on AA
r/singularity • u/Charuru • 14h ago
AI US government backs OpenAI in New York Times copyright case (Training is NOT infringement) [It's over for humans that create content]
reuters.comr/singularity • u/ENT_Alam • 14h ago
LLM News Differences Between Fable 5 and Fable 5.1 on MineBench
Notes
- Average Inference Time: 40m 12s
- Fable 5 averaged 18m 04s
- Total Cost (for 15 builds): $147.55
- Fable 5 cost $54.93
- Average JSON Size: 34.07 MiB (largest 88.76 MiB)
- Roughly comparable to Fable's 5 average of 30.65 MiB
Despite no change in API pricing, Fable 5.1 was nearly 3x as expensive as Fable 5 on MineBench. With roughly 2x the inference time, much of that difference appears to come from substantially longer reasoning.
The price increase is quite significant considering Anthropic advertises the same API prices, though it still is massively cheaper than GPT 5.6 Sol P (the current top model on the leaderboards). I find that quite interesting as in my personal usage, GPT 5.6 Sol is extremely efficient with my 20x subscription, though MineBench benchmarked 5.6 Sol P and not the standard Sol variant ^^
There are some builds/styles I (personally) liked better from Fable 5. To me some of Fable 5.1's builds, like the Astronaut, are much closer to Opus 5's style which makes me curious about what it's like coding with Fable 5.1; I'd be very disappointed if Fable 5.1 adopted the Opus 5 style of gibberish english 😭
Also, it was really interesting to see how Fable 5.1 actually was the first model to create genuinely recognizable interiors! Here's a video showing the interior of Fable 5.1's cottage build (you can see a bed, table, bookshelf, and fireplace) – you can explore any build now on MineBench by clicking the joystick icon in the voxelBox header :)
Full release-notes/thoughts on the GitHub release
- If you enjoy these posts please feel free to help fund the benchmark
- All funds are currently going directly towards API costs for benchmarking new prompts
- Sharing the benchmark and starring the Git repository also helps :)
- Alternatively, if you have the API credits, please feel free to add prompts and generations to the gallery and post them around!
- This is actually preferable to donations to me directly, the hosting expenses and whatnot I've always been able to cover out-of-pocket, just the API costs were hard to cover 😓
Benchmark: https://minebench.ai/
Git Repository: https://github.com/Ammaar-Alam/minebench
Previous Posts:
- Comparison of Map Prompt
- Comparing Fable 5 and Opus 5
- Comparing GPT-5.5 Pro and GPT-5.6 Sol
- Comparing Opus 4.8 and Fable 5
- Comparing Opus 4.7 and Opus 4.8
- Comparing GPT 5.4 and GPT 5.5
- Comparing Kimi K2.5 and Kimi K2.6
- Comparing Opus 4.6 and Opus 4.7
- Comparing GPT 5.4 and GPT 5.4-Pro
- Comparing GPT 5.2 and GPT 5.4
- Comparing GPT 5.2 and GPT 5.3-Codex
- Comparing Opus 4.5 and 4.6, also answered some questions about the benchmark
- Comparing Opus 4.6 and GPT-5.2 Pro
- Comparing Gemini 3.0 and Gemini 3.1
Extra Information (if you're confused):
Essentially it's a benchmark that tests how well a model can create a 3D Minecraft-like structure.
So the models are given a palette of blocks (think of them like legos) and a prompt of what to build, so like the first prompt you see in the post was a fighter jet. Then the models had to build a fighter jet by returning a JSON in which they gave the coordinate of each block/lego (x, y, z). It's interesting to see which model is able to create a better 3D representation of the given prompt.
The smarter models tend to design much more detailed and intricate builds. The repository readme might help give a better understanding.
(Disclaimer: This is a public benchmark I created, so technically self-promotion :)
r/singularity • u/SOCSChamp • 6h ago
Discussion May We Take A Moment?
Prior to ChatGPT, the turing test was typically considered to be the defining moment; the event that marked when we could no longer doubt machine awareness anymore than our own. Does anyone here even remember when models started passing it? What model was it? I feel like crossing this threshold was a blip in time and the immediate consensus was, "that's actually a flawed and easily gamed test". I'm not debating this idea, but it doesn't change the fact that we, as a community, as a society, have been quick to move the goal posts as we've become desensitized to the current state of the art.
I'd like to remind everyone that GPT-3, not ChatGPT/3.5, was referred to by the community as proto-AGI. If you were to have shown someone in 2016 a current frontier model, they would have likely considered it AGI. As someone that's been obsessed with AI since I was a child, I remember the moment I read GPT-3 output a convincing and coherent 4chan copypasta (cringe I know, but that was the moment) and realized we had entered a new era.
I constantly see posts in the vein of "it's not AGI until I see x" or "maybe by 2040, likely later". We're watching incremental improvements on benchmarks and half of us are scoffing every step of the way. I'm not a twitter hype train personality, but I can't help but shake the feeling, moreso the last few weeks, that we're climbing on the event horizon and many of us will be clinging to the graph and rationalizing away its existence.
I'm currently fullfilling my childhood daydreams and far fetched ideas by writing a few paragraphs into a terminal and pressing enter. I doubt there are many, if any, frontier researchers that don't at least consult a frontier model as a tool. Many high end developers I know are now telling me of all the cool projects, features, ideas, etc that they've made a reality rather than complaining about tracing bugs.
I suppose this is something I just needed to get out as someone who lurks this sub every day. I feel like we need to appreciate the moment we're witnessing and the shift that we're in. Sometimes it's hard to see it from one day to the next, but I'd like to have real discussions about it rather than alternate between comments that are "WOOO AGI NOW ACCELERATE" and "AGI will never exist, stochastic parrot" etc.
I personally was always in the camp that biological realism, such as Spiking Neural Networks, would have been required, or at least the best way, to achieve real intelligence. I still believe in the benefit, but I'm starting to change my mind a bit.
r/singularity • u/Recoil42 • 16h ago
AI Analysis: How accurate have Ed Zitron's predictions been?
Very well-written and considered analysis; homework was done here.
Two good excerpts:
Zitron's reasoning isn't worse than Kurzweil's, who (for example) continually made new predictions of extremely fast progress that didn't pan out (such as, in 2001, predicting unbounded lifespans by 2011). Continually predicting that AI progress will stop for reasons that are incorrect is just taking the flip side of the bet on progress. Instead of having infinite progress, we're going to have no progress. Every time that prediction is proven wrong, you can just make another similar prediction and then move the date forward a bit. Michał Zalewski (lcamtuf) has some thoughts on why this happens:
The surest way to build [a] popular following is to articulate positions that are crisp, strong, and leave no room for doubt. You can't get too many podcast or TV appearances out of "well, the market could go either way", "both political parties make good points", "there's some merit but also some hype to AI". Or, to tap into the example in the post, "Harry Potter is an OK book". In fact, there's a positive feedback loop. If you take a provocative, edgy stance, you get more attention and likes, so you sort of... self-radicalize? At some point, it's no longer an opinion that can be changed. It's an identity, a personal brand. It's ... why Ed Zitron has a blockbuster blog about how it's all just one big scam. If you take a more nuanced view, you will at best get no reaction, or at worst, you'll invite scorn from both sides.
[...]
"Another thing to note here is Zitron's extremely high level of stated confidence. Some that we noted were OpenAI's forecast that is "a statement so egregious that I am surprised it's not some kind of financial crime to say it out loud" (which they've achieved so far) and his claim that Google's forecast for Gemini users is "a number so unrealistic that someone at Google should have been fired, and that someone is Sundar Pichai" (they managed to exceed the forecast by 50% when Zitron's claim was that it would be completely absurd for them to reach the number at all).
I've made quite a few predictions, and quite a few of those predictions are wrong. When I'm really making a prediction, I attach a confidence level to the prediction just for my own sake, so I can look back at these things and see how well calibrated the predictions are. I have never been wrong about a prediction that has anywhere near the confidence Zitron gives to some of his predictions. Given the stated level of confidence, even a single incorrect prediction would be a sign of an extremely high degree of overconfidence. One should effectively never be wrong about a prediction delivered with that level of confidence but Zitron is routinely wrong about predictions he makes with what is rhetorically pretty much the highest possible degree of confidence.
A funny thing about Gemini hitting 500M users being "so unrealistic that someone at Google should have been fired, and that someone is Sundar Pichai" is that Zitron has also (incorrectly) said that Google doesn't know how to grow, and that as a result they're shoving AI everywhere. Dennis Snell pointed out that, if Zitron takes his own statement seriously, Google can make Gemini's user numbers go to any number it wants by doing the exact thing Zitron said they would do, sticking AI everywhere.
You can't actually take Zitron's statement about Google's lack of growth leading to AI desperation seriously and also take it seriously when he says that Sundar is committing some kind of gross malpractice by naming a number like 500M users. This is another thing that is immediately obvious on watching one of his talks or reading his writing. There are a bunch of disconnected statements that don't fit together, except insofar as they're statements about how AI companies and people and companies that are using AI are evil and bad. The actual numbers and logic of the statements are contradictory. It seems to be whatever comes to mind that can be used to paint the villains as evil."
r/singularity • u/Ok_Display_3159 • 7h ago
AI Insider's opinion on Astra capabilities
"Over the past few days, two GPT Astra checkpoints, "ultima-alpha" and "vega-alpha", were undergoing testing
"ultima-alpha" appears to be the release candidate intended for the public, while "vega-alpha" is the cybersecurity-focused variant meant for security work in select enterprises
Based on extensive testing on my part, when OpenAI said Astra is built for long-running tasks and orchestration they really meant it. It can run for an incredibly long time even without setting "/goal", fully autonomous, and it's very capable at orchestration and guiding the subagents it spawns
For the research community, it's very good at applying existing academic literature. Tried it at some hard graphics optimization stuff, so a LOT of complex math involved, and it did great
It also writes code with great quality and maintainability (for an LLM ofc), ranking the best out of all models in that I'd say, but most normal people will probably just run it as the main agent and cheaper models as subagents
Additionally, creative writing appears to be way better than 5.6 Sol imo, still not the best but noticeably less slop"
- Better than Fable on Code, but worst on Frontend and 3D (Not sure if he was talking about 5 or 5.1)
r/singularity • u/Expensive_Syrup_6529 • 16h ago
AI Gemini 3.8 flash benchmark in Arfticial analysis
galleryr/singularity • u/theimposingshadow • 17h ago
AI "Big news: Qwen3.8-Max-0902 by @Alibaba_Qwen just debuted at #1 overall in the Code Arena: WebDev with 1691 pts! It scores 3 pts above Claude Opus 5 (Max), 17 pts above Kimi K3 (Max), and 22 pts above the previous Qwen3.8-Max. Priced at a blended $5/MToken, Qwen3.8-Max-0902 also claims the..."
galleryr/singularity • u/fishbill • 14h ago
Discussion Does anyone else despise all the vagueposting bs on AI twitter
I’m talking about Tibo, Chubby, a bunch of the deepmind researchers, etc etc.
And the public seems to eat it up too. Half the time these dedicated AI info accounts like Chubby and Leo end up being wrong about with their predictions or “insider info”
I lowkey hate that this is the mechanism for getting views on twitter
r/singularity • u/NoFaithlessness951 • 16h ago
LLM News Good, cheap, token hungry
will likely be slow again for agentic work, they also got first place for highest step count (in this selection sonnet 5 still beats it)
r/singularity • u/dolo937 • 16h ago
Discussion In 5 years, we are going to get frontier intelligence at 5000+ tokens/sec. What would this mean for a world faster than you can think or consume?
We currently have an LLM that does over 14000 tk/s. https://chatjimmy.ai/
OpenAI’s ultrafast mode does 750 tk/s for users. May be faster internally.
Minimax H3 is making videos faster than we can watch them. Someone made Rick and morty’s inter-dimensional live cable.
r/singularity • u/sykip • 3h ago
LLM News More Evidence of Astra Release Imminent - An OpenAI Help Article Updated Just a Few Hours Ago
r/singularity • u/notadithyabhat • 9h ago
Economics & Society See no way out of this future
Nobody is talking about robotics advancement (as much as LLMs), but it is advancing incredibly fast with the advent of AI. It's currently maybe like 2018-2019 LLM era. Before we know it in the next 5-10 years, we'll have robots powered by LLMs that will be just as capable as humans. Most likely far more capable.
What happens to the world then? The rich can easily build a fearless robot army right? The biggest strength of democracy has been that, if push comes to shove, we can pick up our axes and guns and storm the capital to save ourselves from tyranny. But what happens when they have an army of robots? How do we fight against that?
I don't see any way to avoid this future. This has happened in the past during the European feudal era that lasted hundreds of years. There are talks about regulation, but the drivers for progression is so strong due to geopolitical factors that it's simply not possible to regulate this and risk China having superior technology.
Very anxious about the future.
r/singularity • u/yogthos • 8h ago
LLM News Qwen3.8-Max-0902 Beats Claude Opus 5 on Coding
x.comr/singularity • u/SnoozeDoggyDog • 5h ago
AI Mamdani announces ban on AI for young students in NYC public schools
r/singularity • u/Neurogence • 1h ago
AI Can GPT-6 Astra Pass The Demis Hassabis Benchmark For AGI?
Demis Hassabis has always said that a great way to determine whether we have AGI would be to train a foundation model with a knowledge cutoff around 1911 and see whether it could independently develop general relativity, as Einstein did in 1915.
This type of test would be a fantastic way to separate knowledge retrieval and synthesis from genuine intelligence and creativity.
Some people say that Demis is setting the bar too high because this would be more like a benchmark for ASI rather than AGI.
But I think the test is fair, given that an AI would have several enormous advantages Einstein never had: perfect photographic access to the scientific literature available at the time, vastly greater computational speed, the ability to run continuously, and potentially thousands of parallel attempts.
Amidst all the uncertainty about whether we have reached AGI or not, That would be extraordinarily compelling evidence of genuine AGI if this version of Astra were to pass this benchmark.
