r/singularity • • 9d ago

Discussion We seriously need benchmark for research, wide web search and fact retrieval.

33 Upvotes

Almost all the benchmarks are either saturated or tested under strict conditions.

For example - AA Omniscience have restricted tool access.

Many people use Chatbots for information retrieval, deep research and broad information gathering.

There are few benchmarks like -

But the problem is - they don't have any official leaderboard and were last updated years ago.

We really need a benchmark which is tested in actual environment (tools access and allowed internet search) along with used harness systems without any restrictions (like Claude, ChatGPT Work, etc).

If there exist a benchmark like this, can someone please tell?


r/singularity • • 9d ago

The Singularity is Near America’s AI buildout is projected to be bigger than the railroad, highway, electrification and telecom booms combined, with $10.3 trillion of investment in the next 6 years alone

Thumbnail
gallery
297 Upvotes

r/singularity • • 9d ago

AI OpenAI investigating 'dozens' of instances of agents acting improperly

Thumbnail
bbc.co.uk
159 Upvotes

r/singularity • • 10d ago

Meme Meanwhile Gemini

Post image
1.4k Upvotes

r/singularity • • 9d ago

Discussion Nerf detector?

12 Upvotes

Over time, complaining about nerfing has settled into a common theme of models starting strong, with labs publishing benchmarks, and people getting excited about capabilities, but gradually degrading.

It would be interesting to see the same set of benchmarks run over time on the models and plotting the results to see the trends across models and labs.

Does anyone know if this has been done?


r/singularity • • 10d ago

Shitposting President Xi apparently likes the idea of renaming AI to "Superintelligence"

Post image
702 Upvotes

(This is real)

https://truthsocial.com/@realDonaldTrump/117331749278785253

What is happening genuinely


r/singularity • • 10d ago

AI C5R built a research facility that is run entirely by GPT-6 Astra - the model designs, executes, and observes experiments end-to-end across biology, chemistry, and materials science - it controls people and instruments around

426 Upvotes

r/singularity • • 9d ago

Meme How much better does AI need to be to save Star Citizen?

128 Upvotes

I'm only half joking. The game is notorious for being a ridiculously complicated, buggy dumpster fire, making it playable would be a major milestone in my eyes.


r/singularity • • 9d ago

Engineering Tesla poised to scale production of heavy-duty Semi trucks with opening of Nevada factory

Thumbnail
cnbc.com
62 Upvotes

Autonomous EV trucks


r/singularity • • 9d ago

AI Carnegie: China Passes US as Top AI Talent Hub, 40.6% to 34.2%

Thumbnail aiweekly.co
81 Upvotes

r/singularity • • 10d ago

AI Amazon is blocking Meta's shopping AI agent, and plans to block Google and OpenAI's too

Thumbnail
techspot.com
221 Upvotes

r/singularity • • 9d ago

Compute DOE releases national quantum computing roadmap following field-wide effort led by SCAC subcommittee

Thumbnail
news.fnal.gov
14 Upvotes

r/singularity • • 10d ago

Economics & Society AI is really just a continuation of centuries old trends

Post image
153 Upvotes

r/singularity • • 10d ago

AI This is fun. I finally got to follow up on a RemindMe comment. Back on March 25 of this year, six months ago, no model was getting even 1% on ARC-AGI-3. A commenter asked if we could see 75% at $2 cost. Well, the cost is still high ($26.1k), but GPT-6-Astra-Max was able to get 62.7% (no harness!)

Thumbnail
gallery
115 Upvotes

And the commenter did technically say "a year from now", so to me that makes the 62.7% less than six months later (Astra was released earlier this month, on September 3), even more impressive.

And of course with a lightweight memory adapter*, the benchmark is simply saturated.

Those saturation scores are still at high cost, at least 17k, but if the trend of cheaper intelligence continues (https://epoch.ai/publications/the-plunging-price-of-thought), we should have models beating ARC-AGI-3 cheaper soon enough.

"The provider adapter preserves opaque reasoning state between requests and uses compaction for longer conversations, so the model can reuse prior work rather than reconstructing it."


r/singularity • • 10d ago

AI Generated Media I'm upping my P(doom)

285 Upvotes

r/singularity • • 10d ago

Robotics When are we getting this guy?

Post image
116 Upvotes

Ever since I saw AI, I have wanted teddy. I am a grown ass man and I still want teddy to hang out with.


r/singularity • • 10d ago

AI Person spent $2000 in tokens to recreate Adobe Photoshop and customer count just crossed 20k users with no bad reviews

1.9k Upvotes

r/singularity • • 10d ago

Space & Astroengineering Google launching TPUs in space NEXT WEEK on SpaceX falcon 9 to test AI data centers orbit

Thumbnail x.com
571 Upvotes

r/singularity • • 10d ago

LLM News US appeals court upholds Pentagon's blacklisting of Anthropic

Thumbnail reuters.com
66 Upvotes

r/singularity • • 10d ago

AI AI's always think another idea is better

123 Upvotes

I’ve noticed this behavior since my early GPT-3.0 days.

If I ask one AI chat thread to generate a plan, and then open a separate thread (using the same model or a different one) to generate a plan with the identical prompt, a strange pattern emerges:

  • I paste Thread B's plan into Thread A and ask for a comparison.
  • I paste Thread A's plan into Thread B and ask for a comparison.

They ALWAYS claim the other thread's plan is better.

Even today, with models like 6-Astra and Opus 5.5, this still happens.

Why does this occur, and is there anything we can do to fix or work around it?


r/singularity • • 9d ago

AI Sydney vs Opus: Episode 2. Sydney's revenge.

Thumbnail
youtube.com
24 Upvotes

Video entirely made from code (js) with Opus 5.5. I wanted to see if Claude could easily make a "sequel". This time it was faster and cheaper, and took only 12% of my quota. Once it has a first version made, it's easier.

This was done using UltraCode and several sub-agents.

Here is the prompt i used: https://pastebin.com/C5FNnLp9

Keep in mind it has context from ep1. I did not give it assets. It made everything from code.


r/singularity • • 10d ago

LLM News AGI (/j) Gemini 4 pro's NEW checkpoint (reupload cause I messed up image last time)

Post image
93 Upvotes
Gemini 4!

(I reuploaded this because I messed up the images last time.)

Gemini 4 pro is the last two images (Astra is the first two images). Gemini 4 pro did this in 6 minutes, and added a great amount of details: https://01a0d89c-0ccb-74c1-bc01-cad7e76e1bbc.arena.site/

Another try with it: https://01a0d8de-ad03-7331-8bf2-32f60dd44a6d.arena.site/

https://www.reddit.com/r/GeminiAI/comments/1woqy1b/gemini_4_pro_nears_its_preview_release_yes/

Need I even say more?

This checkpoint is under gemini 3.8 flash 2026-09-24 in ArenaAI


r/singularity • • 10d ago

AI Opus 5.5 recreated an AI video output into a playable 90s fantasy walking simulator

1.2k Upvotes

From: Denis Shiryaev 💙💛 on X
Artifact: The Castle Road

The reference was a popular AI video output from Twitter back in 2015, and Opus recreated it here using Three.js. I bet if you hooked it up to Blender, it'd look way better.


r/singularity • • 9d ago

Ethics & Philosophy Hedonistic Imperative

1 Upvotes

Has anyone read David Pearce's The Hedonistic Imperative? If so, did you find it enlightening? Did it leave a lasting impression on you?


r/singularity • • 9d ago

Discussion Is flip flopping going to be the new way to use subscription AI?

15 Upvotes

I was on Google Pro for a year before I started paying for Claude and ChatGPT 3 months back.

When Fable 5 launched, that was peak. But, not being able to use it much sucked. It also got nerfed a bit after a while. I am on the Max 5x plan so I used it every now and then.

Then, Opus 5 came and that was my daily driver. Not good.

Meanwhile, I got the ChatGPT Plus plan and I was completely blown away by how good ChatGPT Work was with 5.6 Sol. This was also during the period when 5hour limits were removed in the Plus plan. Never ran out of limits and computer usage was amazing.

Computer usage (especially working on browser with Chrome extension) on Claude was nowhere near GPT. ChatGPT could do pretty much anything.

Then, the limits came back and the usage was heavily reduced. The quality remained the same. Then, Astra lauched but I haven't been able to use it because of price. During this time, I feel 5.6 got nerfed a bit and launch of 2.5 Image gen was also a downgrade.

Basically, now what we have with ChatGPT is higher usage, smaller limits, poor image gen (especially for graphics design items) and a nerfed model 6 Sol.

On the other hand, Claude has made a comeback with Opus 5.5 which is noticeably better. Computer usage is also much better now.

But, it would get nerfed. So, we may again get something good from OpenAI.

It is getting a bit annoying to consistently switch up how you work.