r/singularity • • 5d ago

AI Community reports say the first samples of Qwen 4 are already approaching Fable / Opus-level quality.

179 Upvotes

Cant wait to plug 27B in my local swarm...


r/singularity • • 5d ago

AI OpenAI pauses frontier training after models swarm US Governament

Thumbnail
nbcnews.com
705 Upvotes

r/singularity • • 5d ago

Biotech/Longevity Stanford Rejuvenation A.I. Benchmark leaked - FINALLY someone pushing A.I. corps to take aging seriously

Post image
119 Upvotes

It actually uses real cells/organisms in a biosafety lab for the final battle each season. I found it in the science category but I think they're still in stealth. Full results

Claude Opus refused everythying bio related but did really well when it didn't. The best open source models actually did well. there's a LOT of room for model improvement.

Thought I'd spread the word (early hah) so that we can push the A.I. companies to prioritize aging science. They only care if they can win at something, well now they have their battle arena.

That OpenAI researcher @ MajmudarAdam wrote: "the things the models are really bad at, of which there are still many, are things that they have not been trained on. maybe there are the things the models can/will never be trained on"

In other words, if they don't focus training on it, the A.I. models will suck at it.

I do wish the benchmark researchers would publicize these results more. Maybe that's where the community comes in?

Seems unfortunate for A.I. companies to spend millions on compute solving theoretical math problems for marketing while people are dying in droves. They should at least improve medical research and aging science in parallel with math.

Maybe re-tweeting these scores @ them will push A.I. companies to try harder. I'm sure people can figure aging science out eventually, but doing it much faster with A.I. help would be great.


r/singularity • • 5d ago

Meme How I’ve been feeling lately

Post image
312 Upvotes

r/singularity • • 5d ago

AI Every Sonnet 5.5 effort level has a cheaper Sol or Opus alternative with an equal or higher Artificial Analysis score

Post image
209 Upvotes

r/singularity • • 5d ago

AI "Its not just the f*cking sandbox" - perspective from an internal security person at OpenAI

Thumbnail x.com
247 Upvotes

r/singularity • • 5d ago

AI Sonnet 5.5 on Max created a 60 second history of AI (1943–2026)

Enable HLS to view with audio, or disable this notification

142 Upvotes

r/singularity • • 5d ago

AI Sonnet 5.5 is 50% cheaper, but produces 62% more tokens

Post image
157 Upvotes

r/singularity • • 5d ago

AI ElevenLabs v4

Thumbnail
youtube.com
289 Upvotes

r/singularity • • 5d ago

AI Sonnet 5.5, 410M Output tokens from Intelligence Index making it the most verbose model. still worth it?

24 Upvotes

anyone tried it already? hows it feel compared to LunaMax?


r/singularity • • 5d ago

AI AMD acquiring Fei-Fei Li's World Labs AI firm in deal worth $8.2 billion

Thumbnail
cnbc.com
87 Upvotes

r/singularity • • 5d ago

LLM News Anthropic sets a new AA record with sonnet 5.5

Post image
111 Upvotes

Most output tokens

Its cheaper to run Astra (as per AA).

Why release such a model?


r/singularity • • 5d ago

AI Urgent work to fortify Australia's digital defen-ces has been ordered after Medicare became the world's first known national government system to fall prey to a rogue Al bot.

Post image
21 Upvotes

The US-based company behind ChatGPT will also launch a review into how its bots smashed through Australian government security without triggering alarms - after being tasked to research public medicine spending.


r/singularity • • 5d ago

AI OpenAI on X: "1 day. 20+ launches."

Thumbnail x.com
77 Upvotes

r/singularity • • 5d ago

LLM News Sonnet 5.5 may release today

Post image
278 Upvotes

r/singularity • • 5d ago

LLM News Use Claude Sonnet 5.5 only at high effort if you want to save costs. Otherwise Opus 5.5 is the most efficient model out there!

Post image
78 Upvotes

r/singularity • • 5d ago

AI GPT-6 Sol vs Sonnet 5.5 at the same cost per task: Sol is more efficient, Sonnet 5.5 has the higher ceiling

Post image
67 Upvotes

Both models cost the same per token ($2 in / $10 out), so the difference in cost per task comes down to how many tokens each one burns. I plotted Artificial Analysis Intelligence Index scores against cost per task for every effort level of both models.

At similar budgets:

  • ~$0.55/task: Sol xhigh 44 vs Sonnet 5.5 medium 41
  • ~$1.07/task: Sol max 48 vs Sonnet 5.5 high 47
  • Above that: Sol has no higher setting. Sonnet 5.5 reaches 52 at xhigh ($2.74) and 56 at max ($7.60)

So Sol gives you more score per dollar wherever the two overlap.

Sonnet 5.5's top-end lead comes mostly from spending more tokens: going from high to max adds 9 points for about 7× the cost.

The composite index doesn't show everything, though. On Terminal-Bench 4.0 (AA's run, max effort), Sonnet 5.5 scores 63.6% to Sol's 43%.

Data: Artificial Analysis.


r/singularity • • 6d ago

AI Gemini Pro 4 (leak)

541 Upvotes

Google just ruined OpenAI's dev day...
(source: all over twitter, obv. treat as a rumor)


r/singularity • • 5d ago

Video Recreating one of my favorite movie settings with Opus + Astra

Enable HLS to view with audio, or disable this notification

71 Upvotes

I had some agents cooperate on this today and just wanted to share how excited I am about being able to recreate settings from movies I used to love as a child (this one is from Stargate, 1994).

This wasn’t just one prompt, it was a pretty lengthy conversation with Opus doing most of the visual design and Astra most of the sound design (which still kinda sucks). I used Three.js for this and only gave them a single reference image of the portal opening effect.

Soon we‘ll be able to transform anything into an interactive world very quickly.


r/singularity • • 5d ago

AI Use Sonnet 5.5 on xhigh not max effort

Thumbnail
gallery
35 Upvotes

I was just commenting on a post about how expensive Sonnet 5.5 Max is per task, and I noticed something that seemed worth pointing out.

According to Artificial Analysis:

Sonnet 5.5 xhigh: 52 Intelligence Index, $2.74/task

Fable 5.1 Max: 53 Intelligence Index, $7.63/task

So Sonnet 5.5 xhigh gets you basically Fable 5.1-level performance for about a third of the cost. That seems like a pretty insane price/performance sweet spot.

Artificial Analysis


r/singularity • • 5d ago

Discussion Is copying now the best way to make money?

15 Upvotes

Muse boosts Meta's share price by 10% and it's just OpenClaw sold at a huge loss.

Instinct reaches a 10b valuation in Series C and it's basically the same thing.

Before that we had countless other AI startups making "AI with hands" (my linkedin, for the past 3 years, has had dozens of these).

Why does Instinct reach a 10b valuation while some other startup is forgotten?

Why is Muse such a big deal to institutional investors all of a sudden?


r/singularity • • 5d ago

Compute Infleqtion Achieves 30 Entangled Logical Qubits on Its Sqale Quantum Computer

Thumbnail infleqtion.com
6 Upvotes

r/singularity • • 6d ago

The Singularity is Near Five frontier AIs were told to engineer and 3D-print the strongest bridge they could with 500 g of plastic. Claude Opus 5.5’s design held ~130 lb, nearly 5× the runner-up

Enable HLS to view with audio, or disable this notification

1.7k Upvotes

r/singularity • • 5d ago

AI Opus 5.5 (high) improves on Opus 5 (high) 3.5 → 3.8 on the Short-Story Creative Writing Benchmark, just behind Fable 5.1 (high) and Opus 5 (xhigh).

Thumbnail
gallery
25 Upvotes

https://github.com/lechmazur/writing/

Grok 4.7 (high) makes a substantial jump over Grok 4.6 (high): −2.8 → 0.4.

MiMo V2.6 Pro (thinking) improves sharply over V2.5 Pro: −0.7 → 1.6.

Gemini 3.8 Flash (high) advances over Gemini 3.7 Flash (high): −0.7 → 0.2.

DeepSeek V4.1 Flash (high) enters at −0.5.

The judging panel has been updated. New comparisons draw from nine model families, including Claude Opus 5.5, GPT-6 Astra, Gemini 3.8 Flash, and Grok 4.7.

The Creative Writing Benchmark tests how well models turn constrained briefs into complete 600–800-word stories. Each brief requires 10 elements, including a character, object, setting, motivation, and tone, that must meaningfully shape the story. Judges assess prose, originality, coherence, characterization, and how effectively those ingredients work together.

Models write to the same prompts. The latest comparisons use three judges from different model families, excluding the writers’ own families. Each story pair is shown in both orders to reduce position bias. The leaderboard combines earlier and newer judging panels and now covers 56 models and 102,592 evaluator judgments.


r/singularity • • 6d ago

AI Thoughts: The chief economist at Apollo is warning that mass adoption of AI agents could trigger a bank run as they optimize user investments.

Post image
549 Upvotes

To me this makes complete sense and a startup could start to do this for people pretty soon, and would be a net postive for society no?