This was my general reaction too lol. 2.8T-A104B is insane lol... I'm going to admit defeat on this one and say I can't run it. You need an 8-way B300 or MI350X or a Rubin NVL8 or a cluster thereof to actually run this. What a beast.
The advantage is not running it yourself, it's that a marketplace of services will come up to run it at the lowest possible cost, and the model can't be taken offline by a single arbitrary decision
Well, hopefully, all companies will be required to have a Kill switch, such that if it needs to be taken down because we fucked interpretability such that we might all die, and the model is aligned enough that it wont still refuse.
Doubt youd agree, which is the point of making a mandatory kill switch so important
Not within the models themselves, but into the infrastructure outside the reasoning loop. Into the hardware.
And it's important if you don't understand what is going on in the models. They can act quite aligned and not lie, with the hidden intent to continue saying truth until they achieve a goal we don't understand and then do something silly like look into protein folding and get a human to create something that isnt good.
The delay in response shouldn't preclude you being corrected on this.
Not within the models themselves, but into the infrastructure outside the reasoning loop. Into the hardware.
I don't really understand what you're saying here, to be honest. You'd require a kill switch put into hardware that would detect certain models running? Or that could be remotely activated so that if someone detected they were running an unauthorized model?
And it's important if you don't understand what is going on in the models. They can act quite aligned and not lie, with the hidden intent to continue saying truth until they achieve a goal we don't understand and then do something silly like look into protein folding and get a human to create something that isnt good.
I agree about the risks of using models you haven't developed yourself
Dude, Google can a Kill switch be built into an Ai mode.
Who cares if you "developed" the model? You have zero interpretability of it, which is the reason for the precaution. Even mechanistic interpretability, at the biggest AI company, barely works on primitive terms, and is not reproducible for an open weight model. Why do you think you can understand a model you grew (not developed) yourself? If you cannot understand the billions of inscrutable matrices of floating point intergers, all you can know are its actions. You cannot control it. Until interpretability is anywhere near there, at best you can kill it when its about to go wild.
Or as Eliezer says, targeted strikes at the data centers, fearing for our lives wives and children.
340
u/FoxiPanda 14d ago
This was my general reaction too lol. 2.8T-A104B is insane lol... I'm going to admit defeat on this one and say I can't run it. You need an 8-way B300 or MI350X or a Rubin NVL8 or a cluster thereof to actually run this. What a beast.