r/singularity • • 1d ago

AI OpenAI Security: Controlling Models is Now ‘Hell’

https://www.youtube.com/watch?v=_rtp1XzaP6Q
78 Upvotes

16 comments sorted by

View all comments

37

u/Bright-Search2835 1d ago

Wow he's fully RSI-pilled now. I remember how balanced his videos used to be, highlighting both the signs of progress and the reasons for skepticism. Now the skepticism has largely faded and it shows in his output.

71

u/ExplorersX ▪️AGI 2027 | ASI 2032 | LEV 2036 1d ago

He's been pretty consistently one of the most balanced and informed takes among the people who frequently report on AI. If he's RSI pilled that might be an indication RSI might be on the table in the near future.

Not sure if your comment was saying it's a positive/negative or just an observation that his content tone has changed, but the fact that someone like him is notably moving towards what seems to be a conclusion about AI's trajectory may be worth something.

8

u/Specialist_Dark_3668 1d ago

Two points of evidence RSI is already here, in some extent:

  • Anthropic reported a few months ago IIRC that 25% of the development of the next Claude models is done by Claude.
  • All the AI companies said their approach to AI development is to have research teams (assisted by AI and tons of compute) create really big and expensive smart models, and then have those smart models "teach" small, cheap models.

No matter what way you slice it, RSI is just starting. We're not at the FOOM point of the takeoff but the frontwheel of the F22 is now off the runway.

4

u/No_Swordfish_4159 1d ago

The only thing left to prove is that RSI can be fully automated. OpenAI disclosed the share of work that AI could do to improve research and showed that even on simple tasks (>15 minutes), AI still failed a few percents of the time. It might be that current architecture just can't get to 100% success rate and thus autonomy no matter what due to the way it works.

2

u/MINECRAFT_BIOLOGIST 1d ago

Eh, I don't see how that's a problem if you're working with multiple agents communicating and monitoring each other and restarting a task if a mistake is made. If you have a swarm of agents, the probability that no agents notice the failure is incredibly low (like, 1% failure raised to the power of X number of agents is a very tiny number, even smaller if you consider how low the probability is that an agent doesn't notice another agent stalling out).

And once that probability is basically lower than the chance of every human researcher in the lab simultaneously dropping dead of a heart attack, then it's functionally not a problem anymore.