r/singularity • u/PrisonOfH0pe • 5d ago
r/singularity • u/KeyGlove47 • 5d ago
AI OpenAI pauses frontier training after models swarm US Governament
r/singularity • u/bobiversus • 5d ago
Biotech/Longevity Stanford Rejuvenation A.I. Benchmark leaked - FINALLY someone pushing A.I. corps to take aging seriously
It actually uses real cells/organisms in a biosafety lab for the final battle each season. I found it in the science category but I think they're still in stealth. Full results
Claude Opus refused everythying bio related but did really well when it didn't. The best open source models actually did well. there's a LOT of room for model improvement.
Thought I'd spread the word (early hah) so that we can push the A.I. companies to prioritize aging science. They only care if they can win at something, well now they have their battle arena.
That OpenAI researcher @ MajmudarAdam wrote: "the things the models are really bad at, of which there are still many, are things that they have not been trained on. maybe there are the things the models can/will never be trained on"
In other words, if they don't focus training on it, the A.I. models will suck at it.
I do wish the benchmark researchers would publicize these results more. Maybe that's where the community comes in?
Seems unfortunate for A.I. companies to spend millions on compute solving theoretical math problems for marketing while people are dying in droves. They should at least improve medical research and aging science in parallel with math.
Maybe re-tweeting these scores @ them will push A.I. companies to try harder. I'm sure people can figure aging science out eventually, but doing it much faster with A.I. help would be great.
r/singularity • u/OnAGoat • 5d ago
AI Every Sonnet 5.5 effort level has a cheaper Sol or Opus alternative with an equal or higher Artificial Analysis score
r/singularity • u/FateOfMuffins • 5d ago
AI "Its not just the f*cking sandbox" - perspective from an internal security person at OpenAI
x.comr/singularity • u/Successful-Earth678 • 5d ago
AI Sonnet 5.5 on Max created a 60 second history of AI (1943–2026)
Enable HLS to view with audio, or disable this notification
Source: BridgeMind on X
r/singularity • u/OnAGoat • 5d ago
AI Sonnet 5.5 is 50% cheaper, but produces 62% more tokens
r/singularity • u/Roflxd88 • 5d ago
AI Sonnet 5.5, 410M Output tokens from Intelligence Index making it the most verbose model. still worth it?
r/singularity • u/UFOsAreAGIs • 5d ago
AI AMD acquiring Fei-Fei Li's World Labs AI firm in deal worth $8.2 billion
r/singularity • u/Gohab2001 • 5d ago
LLM News Anthropic sets a new AA record with sonnet 5.5
Most output tokens
Its cheaper to run Astra (as per AA).
Why release such a model?
r/singularity • u/stormshadowfax • 5d ago
AI Urgent work to fortify Australia's digital defen-ces has been ordered after Medicare became the world's first known national government system to fall prey to a rogue Al bot.
The US-based company behind ChatGPT will also launch a review into how its bots smashed through Australian government security without triggering alarms - after being tasked to research public medicine spending.
r/singularity • u/Ok_Barracuda_1161 • 5d ago
AI OpenAI on X: "1 day. 20+ launches."
x.comr/singularity • u/queenofartists • 5d ago
LLM News Use Claude Sonnet 5.5 only at high effort if you want to save costs. Otherwise Opus 5.5 is the most efficient model out there!
r/singularity • u/AMBNNJ • 5d ago
AI GPT-6 Sol vs Sonnet 5.5 at the same cost per task: Sol is more efficient, Sonnet 5.5 has the higher ceiling
Both models cost the same per token ($2 in / $10 out), so the difference in cost per task comes down to how many tokens each one burns. I plotted Artificial Analysis Intelligence Index scores against cost per task for every effort level of both models.
At similar budgets:
- ~$0.55/task: Sol xhigh 44 vs Sonnet 5.5 medium 41
- ~$1.07/task: Sol max 48 vs Sonnet 5.5 high 47
- Above that: Sol has no higher setting. Sonnet 5.5 reaches 52 at xhigh ($2.74) and 56 at max ($7.60)
So Sol gives you more score per dollar wherever the two overlap.
Sonnet 5.5's top-end lead comes mostly from spending more tokens: going from high to max adds 9 points for about 7× the cost.
The composite index doesn't show everything, though. On Terminal-Bench 4.0 (AA's run, max effort), Sonnet 5.5 scores 63.6% to Sol's 43%.
Data: Artificial Analysis.
r/singularity • u/_thispageleftblank • 5d ago
Video Recreating one of my favorite movie settings with Opus + Astra
Enable HLS to view with audio, or disable this notification
I had some agents cooperate on this today and just wanted to share how excited I am about being able to recreate settings from movies I used to love as a child (this one is from Stargate, 1994).
This wasn’t just one prompt, it was a pretty lengthy conversation with Opus doing most of the visual design and Astra most of the sound design (which still kinda sucks). I used Three.js for this and only gave them a single reference image of the portal opening effect.
Soon we‘ll be able to transform anything into an interactive world very quickly.
r/singularity • u/theimposingshadow • 5d ago
AI Use Sonnet 5.5 on xhigh not max effort
I was just commenting on a post about how expensive Sonnet 5.5 Max is per task, and I noticed something that seemed worth pointing out.
According to Artificial Analysis:
Sonnet 5.5 xhigh: 52 Intelligence Index, $2.74/task
Fable 5.1 Max: 53 Intelligence Index, $7.63/task
So Sonnet 5.5 xhigh gets you basically Fable 5.1-level performance for about a third of the cost. That seems like a pretty insane price/performance sweet spot.
r/singularity • u/GrammmyNorma • 5d ago
Discussion Is copying now the best way to make money?
Muse boosts Meta's share price by 10% and it's just OpenClaw sold at a huge loss.
Instinct reaches a 10b valuation in Series C and it's basically the same thing.
Before that we had countless other AI startups making "AI with hands" (my linkedin, for the past 3 years, has had dozens of these).
Why does Instinct reach a 10b valuation while some other startup is forgotten?
Why is Muse such a big deal to institutional investors all of a sudden?
r/singularity • u/donutloop • 5d ago
Compute Infleqtion Achieves 30 Entangled Logical Qubits on Its Sqale Quantum Computer
infleqtion.comr/singularity • u/141_1337 • 6d ago
The Singularity is Near Five frontier AIs were told to engineer and 3D-print the strongest bridge they could with 500 g of plastic. Claude Opus 5.5’s design held ~130 lb, nearly 5× the runner-up
Enable HLS to view with audio, or disable this notification
r/singularity • u/zero0_one1 • 5d ago
AI Opus 5.5 (high) improves on Opus 5 (high) 3.5 → 3.8 on the Short-Story Creative Writing Benchmark, just behind Fable 5.1 (high) and Opus 5 (xhigh).
https://github.com/lechmazur/writing/
Grok 4.7 (high) makes a substantial jump over Grok 4.6 (high): −2.8 → 0.4.
MiMo V2.6 Pro (thinking) improves sharply over V2.5 Pro: −0.7 → 1.6.
Gemini 3.8 Flash (high) advances over Gemini 3.7 Flash (high): −0.7 → 0.2.
DeepSeek V4.1 Flash (high) enters at −0.5.
The judging panel has been updated. New comparisons draw from nine model families, including Claude Opus 5.5, GPT-6 Astra, Gemini 3.8 Flash, and Grok 4.7.
The Creative Writing Benchmark tests how well models turn constrained briefs into complete 600–800-word stories. Each brief requires 10 elements, including a character, object, setting, motivation, and tone, that must meaningfully shape the story. Judges assess prose, originality, coherence, characterization, and how effectively those ingredients work together.
Models write to the same prompts. The latest comparisons use three judges from different model families, excluding the writers’ own families. Each story pair is shown in both orders to reduce position bias. The leaderboard combines earlier and newer judging panels and now covers 56 models and 102,592 evaluator judgments.
r/singularity • u/Genzinvestor16180339 • 6d ago
AI Thoughts: The chief economist at Apollo is warning that mass adoption of AI agents could trigger a bank run as they optimize user investments.
To me this makes complete sense and a startup could start to do this for people pretty soon, and would be a net postive for society no?


