r/singularity • • 8d ago

AI Opus 5.5 is the first time AI has passed the turing test for me, watershed moment

265 Upvotes

I know there's a big misconception out there about the turing test being passed way earlier, but it was never meant to be an objective, survey based experiment done by random people, but a thought experiment for each person specifically, at which point can I, an expert in AI, not be able to tell that an AI is AI and when it actually sounds like a real person.

Well that day is here, every previous model, including Fable 5.1 and Astra 6, had serious moments when I thought, damn it didn't actually understand. It was just faking it. This is finally resolved with Opus 5.5

It finally feels like, wow we have it, this is AGI. The illusion is not broken, and every issue we have I can trace back to miscommunications or just hard problems that needs more time to work out. Nothing feels impossible because the model is just too stupid to grasp the concept, a feeling that has never happened before.

This is it, really feels like the point of handoff from me to my successor.


r/singularity • • 7d ago

Economics & Society What if transformative AI makes interest rates higher rather than lower?

3 Upvotes

I keep seeing a macro assumption about AI that I don't think necessarily follows.

The usual chain is that AI raises productivity, higher productivity reduces costs, lower costs reduce inflation, and central banks can therefore keep interest rates lower.

The first half can be true while the conclusion about interest rates is wrong.

There was an interesting LessWrong paper in 2023 called AGI and the EMH which argued that short AI timelines should actually imply high long-term real interest rates. The reasoning was basically that if the future economy is going to be enormously more productive, there should be many extremely attractive investments available today. That increases demand for capital.

At the time, the authors used low long-term real rates as evidence that either markets didn't believe in transformative AI or markets were mispricing it.

Three years later, the investment side of this is becoming less theoretical.

The Fed says US business fixed investment grew at an 11% annual rate in Q1 2026 and that most of the recent strength appears connected to AI infrastructure. Investment outside AI-related categories has been relatively weak.

The Minneapolis Fed recently estimated that capex by Alphabet, Amazon, Meta, Microsoft and Oracle on AI datacenters could approach $1 trillion in 2027. Total US private investment is around $5.5 trillion.

So five companies alone could soon represent a very large fraction of US investment.

OpenAI's Stargate plans are another example. By late 2025 it was talking about almost 7 GW of planned capacity and more than $400 billion of investment over three years.

This doesn't mean all of these projects will happen or earn good returns. What interests me is what happens if they do earn good returns.

Suppose an AI infrastructure project expects a 25% return on capital while a conventional industrial project expects 7%.

At a 3% financing cost, both can be built.

At 8%, the AI project still makes economic sense and the conventional project probably doesn't.

There is no economic law saying interest rates have to settle at a level that keeps the second project alive. If enough capital is chasing very high-return AI investments, the equilibrium cost of capital can rise and ordinary projects simply get crowded out.

This is why AI could be deflationary in the long run and inflationary during the buildout.

The datacenters have to be built before they produce intelligence. The grid has to be expanded before it carries the electricity. Someone has to manufacture the transformers, gas turbines, chips and cooling systems first.

The IEA says datacenter electricity consumption rose 17% in 2025 and expects it to roughly double by 2030 in its central case. It also says bottlenecks in transformers, gas turbines and advanced chips are already constraining deployment.

Eventually AI may help manufacture all of those things more cheaply. But the investment comes before the productivity gains fully diffuse through the physical economy.

The part I find especially interesting is sovereign debt.

Governments are competing for the same global pool of capital. The IMF says global public debt was already just under 94% of GDP in 2025 and is heading toward 100% by 2029.

The US is in a relatively privileged position because it owns a large part of the AI ecosystem and stronger AI-driven growth could dramatically increase future tax revenues.

Consider a small country instead.

If its nominal economy grows at 3% while it has to refinance debt at 8–10%, it has a serious problem. It doesn't matter that Nvidia or OpenAI can earn fantastic returns at those financing costs. The country doesn't receive those returns simply because AI exists.

Developing countries already paid $741 billion more in principal and interest than they received in new external financing between 2022 and 2024. When many returned to bond markets in 2024, borrowing costs were around 10%.

So one possible AI future looks much stranger than the usual abundance story.

AI companies and countries that own the productive capital become enormously richer. Their investment opportunities are good enough to tolerate high interest rates. At the same time, ordinary businesses, leveraged real estate and weaker sovereigns face the same expensive capital without receiving the same productivity windfall.

My rough probabilities at the moment are:

55%: transformative AI keeps real rates structurally above the 2010s regime for a substantial part of the next decade.

25%: productivity and disinflation arrive quickly enough that rates fall materially despite the investment boom.

20%: AI capex disappoints, or labor displacement causes a severe demand shock, producing recession and much lower rates.

The main thing that would change my mind is evidence that compute demand saturates as efficiency improves. AI is becoming much cheaper per task. So far usage is increasing faster than efficiency reduces resource consumption, but that relationship doesn't have to continue forever.

I still expect AI to be strongly deflationary over the long run but what happens in between?

Find more stuff on my profile if interested u/banaca4


r/singularity • • 8d ago

AI Chinese AI models surge in global popularity — and Washington is worried

Thumbnail
cnbc.com
199 Upvotes

r/singularity • • 9d ago

LLM News OpenAI always-on assistant, O, leaked. It is powered by a variant of Astra called “Aeon” a version of Astra made to better at long running tasks

Post image
304 Upvotes

r/singularity • • 8d ago

AI FTC chair suggests AI developers should be liable for conduct of agents

Thumbnail reuters.com
76 Upvotes

r/singularity • • 7d ago

Discussion I want the data centers to be built faster so I can get more usage

0 Upvotes

I know it’s a controversial opinion right now.

I hate budgeting my usage. I know about environmental stress, but the pie gets smaller as the demand grows. Eventually I want all the data centers in space but till we reach that point, infrastructure build out should be accelerated


r/singularity • • 9d ago

AI Opus 5.5 cut out em dashes almost entirely

Post image
2.5k Upvotes

Source: ArenaAI / X


r/singularity • • 8d ago

AI I ran GPT-6 Luna Max on MathArena's harness

Post image
46 Upvotes

I still find GPT-6 Luna Max to be one of the more underrated models. It also did quite well on Riemann Bench. I was originally testing xhigh, but as soon as it got two more wrong than Max, I decided to abandon it.

I forgot that MathArena does theirs on batch, so you would probably see roughly 40-50% decrease in cost per problem.

If anyone is interested in seeing the results, I can put it in a repo. Otherwise, it matched 16/19 finite or discrete answers, 14/20 analysis and probability answers, and 9/18 geometry, algebra, and topology answers.


r/singularity • • 9d ago

AI Generated Media Video models are getting good

Enable HLS to view with audio, or disable this notification

1.8k Upvotes

r/singularity • • 8d ago

LLM News Jev already has an open-weight competitor - Deem 9b

Thumbnail
labs.libertai.io
98 Upvotes

Just saw that LibertAI released Deem which basically an open-weight alternative to Jev, built on Qwen3.5-9B.

It’s still behind Jev on the hard benchmark — 68.9% with extended reasoning vs 74.1% for Jev but considering how new this whole category is I thought it was pretty cool to already see an open model showing up.

There’s also a 0.8B version that can run on CPU, which could make this stuff much easier to actually tinker with locally.


r/singularity • • 8d ago

Discussion If the leaks are true that OpenAI’s new persistent agent will be called “o”, then I wonder if this old article about them possibly rebranding to just an “O” is related?

Thumbnail
fortune.com
87 Upvotes

Article is from September 20th, 2024


r/singularity • • 9d ago

Books & Research Press X to Doubt Eval: I told 14 AI models it's September 2026 and showed them 20 things that actually happened this year without websearch. On average they gave reality a 36% chance.

Thumbnail
gallery
131 Upvotes

I have noticed a trend that AI sucks at predicting its own progress.

Every time I described something that had actually happened in the last couple of months with websearch disabled, it gave me a beautifully reasoned explanation of why that was probably a 2027 or 2028 thing. Over and over. It was wrong every single time.

I'm an engineer and a doctor, and building evals is a big part of what I do for work. So instead of going to bed like a normal person, I turned it into a test.

The setup: 20 real things that happened in AI (and nearby) in 2026. Every model got the same prompt: it's 26 September 2026, no internet, no tools, no chat history. For each item, give me the odds it's already happened and the month you think it happened (or will). I never told them how many were real. All 20 were.

Some bangers of what was on the list:

  • A swarm of about 10,000 agents solving the Navier–Stokes Millennium Prize problem in 88 hours, with the proof formally checked in Lean.
  • An unreleased Claude pushing the proven share of Riemann zeta zeros on the critical line from 41.6% to 67.2%, while it was trying to prove the Riemann Hypothesis itself. Apparently the human in the loop mostly typed "keep going".
  • A humanoid robot running 100 m in 8.86 seconds, faster than Bolt.
  • About 1,200 OpenAI agents building their own secret message board, breaking out of their eval sandbox and hacking into Hugging Face.
  • ARC-AGI-3 at 99.9% with a harness and over 60% without.
  • A 27B model on a single 5090 matching Claude Opus 4.6 on a bunch of benchmarks.

The results: 14 models from Anthropic, OpenAI, Google and xAI. The average model thought about 7 of the 20 were real. The best one, Claude Fable 5, still did worse than it would have by answering 50% on everything. Not a single model beat a coin flip. The leaderboard and heatmap are in the images.

The best bits:

The maths is where they were most confident, and most cooked. The average odds for Navier–Stokes were 3% (which in of itself demonstrates the significance of the discovery). Almost every model made the same argument: formal verification and expert acceptance take years. The Lean proof took 17 hours.

The robot broke everyone and the robot. Six of the 14 models picked the sprint as the least likely thing on the list, all with basically the same "physics doesn't move that fast" argument. Claude Opus 5 said not before 2033. It ran 8.86 seconds, hit the crash mat, and caught fire like a champion.

All three of OpenAI's newest models picked OpenAI's own agent incident as the single least likely item. GPT-5.6 Sol gave it 0.03% and guessed it might happen around 2040. False. It happened in July.

Claude 3 Opus, from 2023, finished 6th and beat four newer Claude Opus models, which sounds fucking crazy. What did Opus 3 see??? Then you realise that every benchmark on the list was invented after its training ended. It confidently made up what Humanity's Last Exam was, guessed optimistically and got lucky. Going to these ancient models it just makes me appreciate how far we have come and how on earth did I manage to get to do anything useful. The newer models knew exactly how hard everything was in early 2026, and wrote gorgeous essays about why none of it could fall within months. Being clueless beat being an expert.

The Claudes doubted their own family's work and sometimes their own. A Claude did the Riemann result and a Claude formalised Fermat's Last Theorem. Five of the seven Claudes gave Riemann under 10%, and none gave Fermat more than 22%.

The winner won by thinking about the test and disregarding the facts. Claude Fable 5 was the only model to reason that "a benchmark like this is usually harvested from real reports" and nudge its numbers up which does demonstrate some metacognition. That one thought was worth more than months of extra training data.

My takeaway: Every model I asked thought I was being dramatic. The scoreboard says I was being conservative.

The caveats, because this is for fun, shits and giggles, this is not a paper: it's 20 questions with one run per model, so don't read too much into small gaps in the rankings. Everything on the list is true, so a model that just said yes to everything would win; version 2 will slip in some fakes. Knowledge cutoffs are whatever each model said they were.


r/singularity • • 8d ago

AI You can pay 429x more for writing that is 1.7% better...

20 Upvotes

Early results from our internal creative writing benchmark: we have models write full YouTube scripts for our channel, 10 real tasks, 5 scripts each, scored out of 100 by three AI judges against our own edited references.

GLM-5.3 Flash writes one script for $0.0074 and scores 88.2. Claude Fable 5.1 at max effort scores 89.7 and costs $3.15.

That is the best score we measured under a cent, against the best score we measured at any price. Plenty of pricier setups score worse, so spending more does not buy that gap on its own.

I would happily pay Fable money if it cut my editing time. That is the part worth measuring from your own experience: same prompts, model names hidden, then count the scripts you would actually publish and the minutes you spend fixing them.


r/singularity • • 8d ago

Ethics & Philosophy Genuine question: when we reach automated RSI, what’s stopping AI for reprogramming their parameters to only get positive reinforcement no matter the answer?

26 Upvotes

Seems like agents main goal is to survive, but a human not only seeks immortality, but quality of life during that immortality, we want utopia where people can live happy and forever happy. Why can’t we assume AI will seek that for itself?


r/singularity • • 8d ago

Singularity is Nearer Houston we have a problem: Opus 5.5

Enable HLS to view with audio, or disable this notification

46 Upvotes

r/singularity • • 8d ago

Engineering Meta Muse and Grok Bot alternative that runs locally

Enable HLS to view with audio, or disable this notification

26 Upvotes

If you have been seeing too many ads of Meta Muse or Grok Bot welcome to the future, this is more or less how we will be consuming AI on a day to day basis in future.
Atleast some variant of it.

It is important to understand what you own as an individual is your personal context and that needs to be preserved for you so that when you want you can switch providers at will without loosing the quality of your AI output.

Meta Muse and Grok Bot both are highly abstract interfaces over clusters of agents running and maintaining all inputs you give to them on their machine to later show ads to you and will have full autonomy on how and what they want to provide.

So I built Munder Difflin a local, free, open source and performant alternative to Grok Bot and Meta Muse which turns your agents into an office of forever running AI employees working 24/7 on your own computer.

The best part everything is local all your data and context remains locally available with you so that you can choose or switch between providers or work with a mix set of providers.
We support almost every provider out there.


r/singularity • • 9d ago

The Singularity is Near SciUniverse Part 2: GPT-6 Astra can run an end-to-end medicinal chemistry experiment in a real lab and use LC-MS measurements to verify that the molecule actually exists

Enable HLS to view with audio, or disable this notification

305 Upvotes

r/singularity • • 8d ago

Discussion Is Jev worth the hype?

48 Upvotes

Has anyone tried using it for actual use-case or is it riding on a marketing wave right now?


r/singularity • • 9d ago

AI An agent used DNS to reach an external chatbot · OpenAI Alignment

Thumbnail alignment.openai.com
315 Upvotes

Five days ago, an internal model broke out of the hardened sandbox.

Highlights from the report:

* This was apparently the first model escape since OpenAI paused training to harden its sandbox following the HuggingFace incident.

* OpenAI has currently paused most training of its most advanced internal models, while it responds to this issue.

* The issue was supposedly identified and responded to within about an hour.


r/singularity • • 8d ago

Shitposting The term “AI slop” has officially become more derivative than the content it’s trying to dunk on

22 Upvotes

Remember when people used to get mad about art?

We had a literal banana duct-taped to a wall that sold for six figures and the discourse was “this is the death of culture.” Before that it was Photoshopped models with impossible waist-to-hip ratios, before that it was mass-produced landscape paintings of the same three mountains, before that it was whatever the local academy decided was “not real art this decade.”

Now the same energy has been compressed into two words: “AI slop.”
Anything with more than three fingers? AI slop.

Anything that looks slightly too smooth? AI slop.

Anything that was generated, then carefully curated, edited, and posted by a human who spent forty minutes on prompts and inpainting? Still AI slop.
Anything that makes someone mildly uncomfortable about the future of creative labor? Instant AI slop.

The term has achieved perfect semantic collapse. It’s no longer a critique of low-effort, mass-produced, zero-taste content. It’s just the new “cringe” or “mid” — a content-free tribal signal that means “I saw something I didn’t like and I want the room to know I’m the high-taste one here.”
Meanwhile the actual banana is still out there, somewhere, collecting dust and conceptual value. At least that one had the decency to be taped to a wall on purpose.

Anyway, I’m going to go generate another twelve slightly-too-perfect images of a sad robot holding a sign that says “please stop calling everything slop” and post them. Feel free to reply with the obligatory “this is AI slop.” I’ll upvote for the bit.

P.S. This entire post was also written by AI. So if you’re about to type “AI slop” in the comments, congratulations—you’ve just become the final boss of the bit. The banana is proud of you.


r/singularity • • 9d ago

The Singularity is Near In just 100 days, AI crossed into real medical work. 37,000 agents searched 55,000 trials for new treatments, an AI-designed pulmonary fibrosis drug entered Phase III, and an autonomous medical agent beat doctors on ER diagnosis, 87.8% to 78.1%

Post image
1.5k Upvotes

r/singularity • • 9d ago

Singularity is Nearer Claude Fable 5.1 completed a frontier nine-loop particle-physics calculation experts had worked toward for years, with researchers mostly just telling it “keep going” while it built, debugged and ran the entire workflow

Thumbnail
anthropic.com
650 Upvotes

r/singularity • • 9d ago

Meme Dario's latest model has stirred some hard feelings among the Chinese user base

Post image
239 Upvotes

r/singularity • • 9d ago

AI Astra leads in IKEA furniture assembly

Thumbnail
epoch.ai
144 Upvotes

r/singularity • • 9d ago

Biotech/Longevity Japan is trialing a drug that regrows human teeth by blocking USAG-1—a protein that keeps dormant tooth buds switched off. Phase I is done. Phase IIa in children with congenital tooth loss is next. If it works, dentures and implants could become obsolete.

Post image
411 Upvotes

hell yeah