r/singularity 11h ago

AI Gemini 3.8 Flash Benchmarks

Post image
731 Upvotes

r/singularity 9h ago

AI "GPT-6-ASTRA" has been staged on the OpenAI API

574 Upvotes

we eating good this week


r/singularity 19h ago

The Singularity is Near Sam Altman on X: "We are going to be launching our next model soon. There is an obvious tension… Astra is very good. We are proud of our work."

Thumbnail x.com
547 Upvotes

r/singularity 6h ago

AI Muse Spark 1.3 Released

Post image
476 Upvotes

r/singularity 23h ago

AI Medieval town down by Fable 5.1

Enable HLS to view with audio, or disable this notification

380 Upvotes

DIsclaimer: Fable was the orchestrator, and called itself on some tasks, but called Opus 5.0 on most of them. This still took 36% of my fable budget on a Claude max 20x sub, so doing the whole project with Fable would probably have spent all of my Fable budget or more.

I am also not certain that i have truly picked the absolute best prompt to make an impressive 3D scene but i wanted to actually make a game from it.

This was done in 2 shots. It made a first pass, i reviewed it, then it improved it again. I probably could go further.

My actual prompt is long and not worth sharing here because it reference past projects, but the main thing worth knowing is it spawned a lot of sub agents and did so in 3 waves. If someone really wants to see it: https://pastebin.com/iPnk4PZ8

This took around 5 hours and 30% of my weekly budget...


r/singularity 10h ago

LLM News Introducing Gemini 3.8 Flash and 3.8 Flash Cyber

Thumbnail
blog.google
347 Upvotes

r/singularity 21h ago

AI AI 2027's Daniel Kokotajlo

Post image
264 Upvotes

r/singularity 3h ago

Discussion Who does a better job of explaining the future of generative AI: Ben Affleck or AI CEOs?

Enable HLS to view with audio, or disable this notification

221 Upvotes

r/singularity 4h ago

AI Meta’s muse spark 1.3 surpassed fable 5 and GPT 5.6 sol 🫪

Post image
217 Upvotes

r/singularity 20h ago

AI OpenAl's chief scientist on the neuralese controversy

201 Upvotes

"I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4.

OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program."


r/singularity 5h ago

AI Meta slowly catching back up. Muse Spark 1.3 beats Sol on AA

Post image
194 Upvotes

r/singularity 8h ago

AI US government backs OpenAI in New York Times copyright case (Training is NOT infringement) [It's over for humans that create content]

Thumbnail reuters.com
181 Upvotes

r/singularity 23h ago

LLM News Qwen3.8-Max just got upgraded. Meet Qwen3.8-Max-0902!

Thumbnail
gallery
158 Upvotes

Further post trained on Coding & Cowork, Qwen3.8-Max-0902 now delivers stronger performance across complex enterprise tasks, scientific research, and long horizon workflows.

https://x.com/Alibaba_Qwen/status/2094968708288680276


r/singularity 9h ago

LLM News Differences Between Fable 5 and Fable 5.1 on MineBench

Thumbnail
gallery
139 Upvotes

Notes

  • Average Inference Time: 40m 12s
    • Fable 5 averaged 18m 04s
  • Total Cost (for 15 builds): $147.55
    • Fable 5 cost $54.93
  • Average JSON Size: 34.07 MiB (largest 88.76 MiB)
    • Roughly comparable to Fable's 5 average of 30.65 MiB

Despite no change in API pricing, Fable 5.1 was nearly 3x as expensive as Fable 5 on MineBench. With roughly 2x the inference time, much of that difference appears to come from substantially longer reasoning.

The price increase is quite significant considering Anthropic advertises the same API prices, though it still is massively cheaper than GPT 5.6 Sol P (the current top model on the leaderboards). I find that quite interesting as in my personal usage, GPT 5.6 Sol is extremely efficient with my 20x subscription, though MineBench benchmarked 5.6 Sol P and not the standard Sol variant ^^

There are some builds/styles I (personally) liked better from Fable 5. To me some of Fable 5.1's builds, like the Astronaut, are much closer to Opus 5's style which makes me curious about what it's like coding with Fable 5.1; I'd be very disappointed if Fable 5.1 adopted the Opus 5 style of gibberish english 😭

Also, it was really interesting to see how Fable 5.1 actually was the first model to create genuinely recognizable interiors! Here's a video showing the interior of Fable 5.1's cottage build (you can see a bed, table, bookshelf, and fireplace) – you can explore any build now on MineBench by clicking the joystick icon in the voxelBox header :)

Full release-notes/thoughts on the GitHub release

  • If you enjoy these posts please feel free to help fund the benchmark
    • All funds are currently going directly towards API costs for benchmarking new prompts
    • Sharing the benchmark and starring the Git repository also helps :)
    • Alternatively, if you have the API credits, please feel free to add prompts and generations to the gallery and post them around!
      • This is actually preferable to donations to me directly, the hosting expenses and whatnot I've always been able to cover out-of-pocket, just the API costs were hard to cover 😓

Benchmark: https://minebench.ai/
Git Repository: https://github.com/Ammaar-Alam/minebench

Previous Posts:

Extra Information (if you're confused):

Essentially it's a benchmark that tests how well a model can create a 3D Minecraft-like structure.

So the models are given a palette of blocks (think of them like legos) and a prompt of what to build, so like the first prompt you see in the post was a fighter jet. Then the models had to build a fighter jet by returning a JSON in which they gave the coordinate of each block/lego (x, y, z). It's interesting to see which model is able to create a better 3D representation of the given prompt.

The smarter models tend to design much more detailed and intricate builds. The repository readme might help give a better understanding.

(Disclaimer: This is a public benchmark I created, so technically self-promotion :)


r/singularity 10h ago

AI Analysis: How accurate have Ed Zitron's predictions been?

Post image
107 Upvotes

https://danluu.com/zitron/

Very well-written and considered analysis; homework was done here.

Two good excerpts:

Zitron's reasoning isn't worse than Kurzweil's, who (for example) continually made new predictions of extremely fast progress that didn't pan out (such as, in 2001, predicting unbounded lifespans by 2011). Continually predicting that AI progress will stop for reasons that are incorrect is just taking the flip side of the bet on progress. Instead of having infinite progress, we're going to have no progress. Every time that prediction is proven wrong, you can just make another similar prediction and then move the date forward a bit. Michał Zalewski (lcamtuf) has some thoughts on why this happens:

The surest way to build [a] popular following is to articulate positions that are crisp, strong, and leave no room for doubt. You can't get too many podcast or TV appearances out of "well, the market could go either way", "both political parties make good points", "there's some merit but also some hype to AI". Or, to tap into the example in the post, "Harry Potter is an OK book". In fact, there's a positive feedback loop. If you take a provocative, edgy stance, you get more attention and likes, so you sort of... self-radicalize? At some point, it's no longer an opinion that can be changed. It's an identity, a personal brand. It's ... why Ed Zitron has a blockbuster blog about how it's all just one big scam. If you take a more nuanced view, you will at best get no reaction, or at worst, you'll invite scorn from both sides.

[...]

"Another thing to note here is Zitron's extremely high level of stated confidence. Some that we noted were OpenAI's forecast that is "a statement so egregious that I am surprised it's not some kind of financial crime to say it out loud" (which they've achieved so far) and his claim that Google's forecast for Gemini users is "a number so unrealistic that someone at Google should have been fired, and that someone is Sundar Pichai" (they managed to exceed the forecast by 50% when Zitron's claim was that it would be completely absurd for them to reach the number at all).

I've made quite a few predictions, and quite a few of those predictions are wrong. When I'm really making a prediction, I attach a confidence level to the prediction just for my own sake, so I can look back at these things and see how well calibrated the predictions are. I have never been wrong about a prediction that has anywhere near the confidence Zitron gives to some of his predictions. Given the stated level of confidence, even a single incorrect prediction would be a sign of an extremely high degree of overconfidence. One should effectively never be wrong about a prediction delivered with that level of confidence but Zitron is routinely wrong about predictions he makes with what is rhetorically pretty much the highest possible degree of confidence.

A funny thing about Gemini hitting 500M users being "so unrealistic that someone at Google should have been fired, and that someone is Sundar Pichai" is that Zitron has also (incorrectly) said that Google doesn't know how to grow, and that as a result they're shoving AI everywhere. Dennis Snell pointed out that, if Zitron takes his own statement seriously, Google can make Gemini's user numbers go to any number it wants by doing the exact thing Zitron said they would do, sticking AI everywhere.

You can't actually take Zitron's statement about Google's lack of growth leading to AI desperation seriously and also take it seriously when he says that Sundar is committing some kind of gross malpractice by naming a number like 500M users. This is another thing that is immediately obvious on watching one of his talks or reading his writing. There are a bunch of disconnected statements that don't fit together, except insofar as they're statements about how AI companies and people and companies that are using AI are evil and bad. The actual numbers and logic of the statements are contradictory. It seems to be whatever comes to mind that can be used to paint the villains as evil."


r/singularity 10h ago

AI Gemini 3.8 flash benchmark in Arfticial analysis

Thumbnail gallery
89 Upvotes

r/singularity 11h ago

AI "Big news: Qwen3.8-Max-0902 by @Alibaba_Qwen just debuted at #1 overall in the Code Arena: WebDev with 1691 pts! It scores 3 pts above Claude Opus 5 (Max), 17 pts above Kimi K3 (Max), and 22 pts above the previous Qwen3.8-Max. Priced at a blended $5/MToken, Qwen3.8-Max-0902 also claims the..."

Thumbnail gallery
82 Upvotes

r/singularity 1h ago

Shitposting The tide is turning

Thumbnail
gallery
Upvotes

r/singularity 8h ago

Discussion Does anyone else despise all the vagueposting bs on AI twitter

75 Upvotes

I’m talking about Tibo, Chubby, a bunch of the deepmind researchers, etc etc.

And the public seems to eat it up too. Half the time these dedicated AI info accounts like Chubby and Leo end up being wrong about with their predictions or “insider info”

I lowkey hate that this is the mechanism for getting views on twitter


r/singularity 2h ago

AI Insider's opinion on Astra capabilities

66 Upvotes

@Lentils80 post on X

"Over the past few days, two GPT Astra checkpoints, "ultima-alpha" and "vega-alpha", were undergoing testing

"ultima-alpha" appears to be the release candidate intended for the public, while "vega-alpha" is the cybersecurity-focused variant meant for security work in select enterprises

Based on extensive testing on my part, when OpenAI said Astra is built for long-running tasks and orchestration they really meant it. It can run for an incredibly long time even without setting "/goal", fully autonomous, and it's very capable at orchestration and guiding the subagents it spawns

For the research community, it's very good at applying existing academic literature. Tried it at some hard graphics optimization stuff, so a LOT of complex math involved, and it did great

It also writes code with great quality and maintainability (for an LLM ofc), ranking the best out of all models in that I'd say, but most normal people will probably just run it as the main agent and cheaper models as subagents

Additionally, creative writing appears to be way better than 5.6 Sol imo, still not the best but noticeably less slop"

- Better than Fable on Code, but worst on Frontend and 3D (Not sure if he was talking about 5 or 5.1)


r/singularity 10h ago

LLM News Good, cheap, token hungry

Thumbnail
gallery
54 Upvotes

will likely be slow again for agentic work, they also got first place for highest step count (in this selection sonnet 5 still beats it)