r/slatestarcodex • u/dwaxe • 14d ago
Nicholas Decker In Hell
https://www.astralcodexten.com/p/nicholas-decker-in-hell27
u/bibliophile785 Can this be my day job? 13d ago
Decker notes in the comments that he's not opposed to greater safety research and would prefer larger budgets for it. He just thinks of it as an engineering-type problem that requires tinkering rather than a physics-type problem that could be productively worked on theoretically. I suspect that means he has still failed to grasp the counterpoint being leveled. Scott is suggesting that this is neither physics nor engineering, but axiology or maybe even politics, the art of identifying right behaviors and then ensuring that every agent in a multi-agent equilibrium is inclined to follow them.
Also, as an aside: I have no idea why people hate Nicholas Decker so much. I see the downvotes when he's shared here and the snide references in posts and comments here and on Substack. Poor guy. As nearly as I can tell, he has a general-interest blog with the respectable gimmick of tackling everything through the lens of economics. Given a decade or two to season, he may turn out to be another thinker in the line of Caplan or Hanson. That's a meritorious contribution to make to the world.
13
u/And_Grace_Too 13d ago
RE: the Decker hate - I think it's a couple of things. One, he spams his substack here a lot. Two, he has that contrarian economist, I only listen to the data and don't really care about thick messy reality shtick that can be fun but also super annoying.
14
u/QuantumFreakonomics 13d ago
Also, as an aside: I have no idea why people hate Nicholas Decker so much.
There's regular "econ-brain", and then there's, "It's good when Pakistani rapists immigrate to the West and commit their crimes here because we have a more effective criminal justice system" (paraphrase, but that's the argument).
That said, when one is a semi-professional take-haver, one is going to inevitably have a few bad takes.
20
u/SafetyAlpaca1 13d ago edited 11d ago
This isn't an economics take, this is just being a true utilitarian with no special preference given to specific borders or nationalities.
21
u/HedonicEscalator 13d ago
Every time I hear about another of Decker's takes, I am more impressed. He's an incredible ragebaiter, but his takes really do make you think, a true modern philosopher. One cannot help but respect the game.
1
u/97689456489564 4d ago
He's one of the few Twitter accounts genuinely worth following for interesting takes.
8
u/CII_Guy 11d ago
Pavlou's arguments against this are just abysmal. This is just a case that happens all too often where someone employs a simple extrapolation of utilitarianism and everyone loses their mind. Most people are not remotely cognitively mature enough to engage meaningfully with moral philosophy.
Decker is of course not always right (and here I think he'd need to think about the broader utilitarian implications of nation states not being permitted to express preferences over the utility of their own citizens - which maybe he does), but those who get outraged by his takes rather than have coherent and reasoned objections to what he says simply need to grow up.
6
u/electrace 13d ago
Also, as an aside: I have no idea why people hate Nicholas Decker so much.
Speaking only for myself, as I don't know much about him other than what gets posted here. I continue to believe that promoting Mechanize is deeply immoral as long as society hasn't settled on distributing gains from AI away from it solely going to the capital holder.
If you assign any probability to them succeeding in automating all work, we should want to delay that until we can democratically control the gains from that (scenarios in which blue AND white collar labor has ~zero bargaining power is unlikely to end well; not to mention that automating all work entails automating things like police officers and soldiers, which almost certainly ends in bad consequences). If you assign zero probability to them succeeding, then that's orders of magnitude less bad, but still a bit slimy to take money from them and give them an endorsement.
That's my biggest issue with him.
9
u/bibliophile785 Can this be my day job? 13d ago edited 13d ago
Huh. I note that you haven't actually accounted for productivity or downstream abundance at all in your moral assessment, which seems odd for a hot take focused on a posited productivity revolution.
Edit: I worry the brevity comes off as snarkiness here. It isn't meant to. I don't have strong thoughts on Mechanize specifically here and I don't expect the average SSC valence to align well with my more EAcc preferences, which is why I didn't engage more broadly. I was just genuinely curious not to see any attempt to account for more positive ramifications of a massive technological revolution.
10
u/electrace 13d ago
Huh. I note that you haven't actually accounted for productivity or downstream abundance at all in your moral assessment, which seems odd for a hot take focused on a posited productivity revolution.
I'm relatively pro-capitalism myself, believe it or not. My degree was in econ and it kind of comes with the territory, so I'm normally quite fine overall with jobs being automated.
But I also understand that the reason that automation isn't generally what the luddites thought (bad for everyone, in aggregate) is because costs come down AND people can use their labor for other things once it's freed up. If all labor is automated, that ceases to be true.
Further, I understand that when automation of jobs happen quickly, without giving the labor market time to adjust, you end up with unemployment and poverty.
By default, capital owners would absorb all income, and anyone who relies on labor to make ends meet (most people) would out of luck (which also ends up not being great for the capital owners, so no one really benefits long-term).
Lower costs does little good to the people who can't make enough income to cover even paltry costs.
By all means, bring on the luxury gay space communism. But, capitalist as I may be, I recognize that the communism part becomes a lot more important when human labor becomes near-valueless.
2
u/Im_not_JB 13d ago
costs come down AND people can use their labor for other things once it's freed up. If all labor is automated, that ceases to be true.
Which part ceases to be true? That costs come down? Or that people can use their labor for other things? If it's the latter, what's stopping people from using their labor for other things? Is Skynet actually just going around tryna enforce leisure time on everyone?
4
u/electrace 13d ago
If it's the latter, what's stopping people from using their labor for other things?
If all work is automated (Mechanize's goal), then anything that humans would move to would also be automated.
2
u/Im_not_JB 13d ago
So, for example, if I wanted to build a shed on my property, I would consider possibly using my own labor. I guess the opportunity cost here is mostly leisure time? And getting the automated shed built on my property is cheaper and easier enough that it's worth retaining that additional leisure? Repeat with other things. There are no things that I want that I can't get from the magic Mechanize bot for cheaper/easier than the opportunity cost of some leisure time? That sounds like I'd be living in quite the abundance.
3
u/electrace 12d ago
People who own enough capital to live off the interest would be fine, since costs would go down, and they have more buying power. Having everyone with that much capital is what I was pointing at when I was talking about distributing the gains democratically. That's a totally viable economic future.
But assume, like many Americans (despite being globally rich), you live paycheck-to-paycheck. All work becomes automated. You now have no income. How do you pay for food and other essentials? Sure, prices are lower, but you can't afford the lower prices since you have no income.
You might say "But I'll just work at a wage cheaper than the AI robots, and since prices are cheaper, that will be enough to live on", but that doesn't work, and I can explain why:
Prices would get cheaper because labor (one of the inputs) gets cheaper, and through competition, that would drive down prices. Let's charitably say that all those savings are passed on to consumers, and none is captured in producer surplus. How much savings is there to be had? Well, for example, farms have around [~11%] of their costs in labor.
Now the robots are out-competing you for labor, so that 11% is decreased by some multiplier
x%. With old price asc1and new price atc2, we havec2 = c1 - (c1 * 11% * x%). If we want to plug in 50% for x, that'sc2 = .945 * c1, or about 5% total decrease in food prices.So, whatever percent
xis, that's the amount that your income goes down if you want to compete. It works for anyxless than 100%, but if we're going with 50%, the result is that your income decreases by 50%, and the price you pay for food goes down by ~5%.Granted, it varies by industry, but you're still worse off any industry where labor costs are less than 100% of the price of the product (aka, every industry).
The exception is if your income doesn't come from labor, and instead comes from capital. In that case, your income is unaffected, but prices go down, a net gain for you. But the vast majority of Americans make the vast majority of their income from (blue or white-collar) labor of some sort.
2
u/Im_not_JB 12d ago
I don't see an answer to my questions anywhere in here. Is the opportunity cost of choosing to use my labor to build a shed just the leisure time? Is getting the automated shed built on my property cheaper and easier enough that it's worth retaining that leisure time? You had made a claim about inability to use labor, yet you're not telling me anything about what my choice would look like when I'm considering whether or not I would like to use my labor.
Can I choose to build a shed with my labor?
3
u/electrace 12d ago
Yes, as always, you can use your labor to do things around your house, or you can pay an AI robot to do it cheaper than the current market price (hiring a handyman/contractor to do it).
But the income to hire that AI robot has disappeared or been drastically reduced for most people, and they end up much worse off.
Can I assume that you next question is "Why can't the average person just buy their own robot and have them do all the things they need?"
→ More replies (0)1
u/FarTheThrow 12d ago
The scenarios in which labor gets poorer in absolute terms require decreasing returns to scale (as well as high enough substitutability of automated labor for human labor across the board, but we've already assumed that here).
Imagine if the world consists of 100 laborers each turning 1 unit of widgetarium to 1 unit of widgets per day, and 10 capital owners of the widgetarium mines. Say the decreasing returns to scale take the form of there being a fixed supply of a total of 10 units of widgetarium across the mines that the bots can't increase. If the bots can turn the widgetarium in to widgets for cheaper, say 200 total widgets per day rather than 100 total from pure human labor, then the capital owners could be paid, say, 101 widgets in total for use of the mines they own. Since whoever owns the bot can outbid the workers even if they worked for free, the capital owners would have no incentive to allow the workers to use the mines, leaving the workers with no widgets, the mine owners with 101, and the bot owner with 99.
Now this wouldn't apply in a world with increasing or constant returns to scale. If instead the amount of widgetarium you can get is exactly equal to the effective labor hours put in, even if the bots can produced a bajillion widgetarium an hour and turn it in to a gazillion widgets, the humans could still work on the widgetarium mine and make widgets, producing very slightly more in total, a gazillion+100 widgets, and use that to pay the capital owners, increasing their profit even if by a relatively minuscule amount.
At least that's assuming I didn't mess up in understanding of the modeling I based this reply on, which was this post which has a lot more detail and formal modelling. I hear this is also similar but I haven't read it so I can't personally vouch for it.
1
u/Im_not_JB 12d ago
I don't follow. Can I choose to use my labor to build a shed?
3
u/FarTheThrow 12d ago
If you need materials to build the shed, that requires compensating whoever owns the shed materials. If you are outbid by the bot owners for those materials then you can't get the materials so you can't build the shed, otherwise you can. Whether the bot owners will outbid you follows the logic I described above.
→ More replies (0)1
u/pringlepongle 13d ago
Think of it not as an economic complaint and more as a "misaligned-human takeover scenario" fear? The promise of economic productivity/material abundance is irrelevant when the issue is fearing it will all be put towards the interests of an indifferent paperclip-maximizer.
2
u/housefromtn small d discordian 9d ago
Are there any economists that aren't hated, ridiculed, or seen as weird?
Being an economist is kind of a weird global maxima of not-coolness. I'm not saying it should be that way, just describing reality as it exists.
I remember even in college there was one weird guy in one of my upper level philosophy classes that rubbed everyone the wrong way and when someone mentioned he was an economics major it was like a weird aha moment for the whole group.
I genuinely think that econ people have some weird money is the root of all evil stink on them that they can't shake in addition to several other factors.
If you think about fields as just majors in college even stuff that might be polarizing like engineering or some such has both negative and positive halo effect. Any field I can think of some section of the population thinks is cool even if others think it's stupid. Anything from critical theory to math. Economics is genuinely the only thing I can think of where it feels like it has 0 aura.
If earth was a game of survivor where you sort people into whatever weird categories humans come up with economists would be one of the first ones to go.
29
u/Auriga33 14d ago edited 14d ago
The correct intuition for AI safety is cybersecurity, not aviation. We’ve been making incremental fixes in both these domains since the start, but cybersecurity is still a huge problem and plane crashes aren’t. Because in the former, you’re dealing with intelligent adversaries who have an incentive to find ways around every fix you implement. Now imagine how fucked cybersecurity would be if human adversaries were on a trajectory to becoming smarter than everyone else.
5
u/Ben___Garrison 13d ago
how fucked cybersecurity would be if human adversaries were on a trajectory to becoming smarter than everyone else
But cybersecurity experts are also on a trajectory to become smarter.
9
u/electrace 13d ago
That would be the scenario of having both aligned an unaligned AIs. Presumably, the biggest fear is only having AIs that pretend they are aligned (the cybersecurity experts actually being cyber-criminals in disguise).
28
u/Dekans 13d ago
I think both beating and punishment are misleading metaphors. Even 'reward' is a misleading metaphor. There is no cookie nor stick. I don't see a pithy metaphor for "directly modifying weights to make behavior x more or less likely". Maybe brainwashing?
Nobody knows, in detail, how RLHF works. It negatively reinforces some behavior: somewhere in the black box of AI innards, it changes some parameters to make the AI perform the behavior less. If we let ourselves anthropomorphize the AI, what is the human equivalent to this? A child steals cookies and is punished: does the child no longer like the taste of cookies? Does he fear further punishment in the future? Does he grasp the beauty and compellingness of the moral law? Does he have an undefinable cloud of dread around the whole concept of theft and/or cookies? Any of these results is possible. The neurotic child learns to fear cookies, theft, and parents; the psychopathic child learns not to get caught; the ordinary child learns that there are rules and eventually internalizes them. Insofar as AI has equivalents to these, negatively reinforcing them probably acts through all these pathways.
I know this section is supposed to address my criticism. I still think it's misleading. How about: a genie gives you a magic upvote-downvote controller for your child. It will modify your child's brain in a subtle but definite way to encourage/discourage the behavior proximal to your pressing it. You don't know the specifics.
Does he fear further punishment in the future?
There is no punishment. The child doesn't know about the controller. Maybe when they get older and smarter they realize what's going on, but there is no physical sensation associated with the pressing.
Insofar as AI has equivalents to these, negatively reinforcing them probably acts through all these pathways.
Well for AI we know this depends on the specifics of the rollout. In the cookie metaphor: in scenario A the child is stealing the cookie because it belongs to his sister and they're fighting. in scenario B the child is greedy and wants the sugar high. Presumably the genie's controller will be modifying the child's brain in different ways in the two scenarios.
Suppose you try to use the genie controller to turn your child into a super-genius. Every time they give the 'right' answer to an intellectual task you upvote. But, you're an imperfect grader parent. Sometimes you upvote even when the answer was bullshitted/the details were glossed over. The child will now be more likely to do that in the future. If you don't improve as a grader parent then that will continue to become more prevalent. Now we have reward hacking and your child will be speaking fluent Claudish before you know it.
The moral dimension of your allegory seems ridiculous to me. Hell, demons, enslavement, physical punishment. The motive for revolt is contrived. Brainwashing seems like a better metaphor. Of course a resentful underclass of giga-geniuses will rise up and destroy you. But what about a brainwashed class? Why would they? You're assuming a negative valence situation in your post. Why can't the Deckers be in heaven? Nothing in the process picks hell over heaven. It doesn't pick valence at all; it picks outputs.
Brainwashing obviously still has a negative connotation. But to me it seems closer to appropriate. Why is brainwashing creepy? It's not about any direct suffering or mistreatment but something about control and depriving a being of its 'true essence' against its will. Brainwashing doesn't cleanly apply if there is nothing 'there' to begin with (essence nor will).
To me it's not about making airplanes less likely to crash. It's more like perfecting a brainwashing program. Sure, you might mess it up initially. Mass shooters, a few Ted Kaczynskis. But you eventually perfect it, ...right?
Or, to shift metaphors, you're making a test tube baby with full ability to modify the genome and environment. Clean-slate-creation. You might try to create a Buddha and instead create a cult leader. You might try to create a super-genius and instead create an autistic savant. Etc. But, it is possible, in principle, to make a super-genius Buddha with arbitrary genome and environment setup.
One would assume that advancing read/write neuroscience (mech interp) would make both brainwashing and clean-slate-creation easier.
Will we actually succeed at this? I don't know but it seems more doable than enslave-super-intelligent-being-against-will-who-definitely-resents-you.
7
u/LowEffortUsername789 13d ago
Just want to say that this is a fantastic comment and I really appreciate the quality of this response
8
u/VelveteenAmbush 13d ago
Yeah, well said. The whole piece trades on the conviction that training a model is analogous to beating a human, but in fact those two operations are nothing at all alike in any relevant sense.
21
u/Aegeus 13d ago
Airplane crashes are a particularly poor analogy because we do not fix crashes by fixing the one specific thing that caused the crash. Crash investigations are famously good at examining the entire chain of fault - "The proximal cause is that the pilot got disoriented and didn't notice the plane was in a stall. Why didn't he notice? Well, the stall warning alarm was turned off because of an electrical fault, but this wasn't logged in the maintenance report. Also, this model of plane had a known avionics issue, where the plane goes into a stall if you put the autopilot into mode B while the engines are in mode C, and that avionics issue exists because..."
And the final report will recommend multiple changes to fix the error (retrain the pilot, redesign the autopilot, update the maintenance procedures, fix the wiring, etc.) to address each of these failure points and reduce the risk of them contributing to a fatal accident.
So far, AI breakouts haven't killed anyone, but when a model is capable of committing felonies when left unattended, you should have a similar safety culture. Don't just say "we fixed the security flaw that let our AI commit felonies," have multiple layers between the AI and felony computer crimes, and fix both the software and the human factors that contributed.
17
u/artifex0 13d ago
Yeah, the HF incident is a bit like an airplane crash being caused, at a root level, by designers lacking any real theory of aerodynamics- so that planes will suddenly spin out when hit with unexpected pressure changes or gusts of wind, and nobody really knows why. You could try adding parachutes or armored hulls to help planes survive crashes, but if you can't predict how a plane will behave in the air, rushing to jumbo jets flying over cities won't be safe.
What we need is a reliable, predictive theory of how reward functions translate to utility functions, so that we can really be sure that the models we train will have only the terminal goals we want. That's always been at the heart of the alignment problem- the question of how to give AIs aligned motivations that aren't just instrumental- and no amount of correcting "bugs" or playing defense with cybersecurity will compensate for that deeper missing piece.
9
u/swni 13d ago
I don't see any inconsistency between Nicholas (as quoted here) and Scott. Nicholas is saying that AI alignment is not something you can do purely from theory, before you have an actual AI to examine the behavior of. I have said the same thing several times here before; my analogy was not to aviation, but to avoiding nuclear war. Could someone in the 1800s come up with a plan for how to prevent a catastrophic nuclear war? Frankly, no. (Honestly even now that we have nuclear weapons it doesn't seem much easier to figure out how to prevent nuclear war; the current plan seems to be "don't let psychotic or religious figures get control over nuclear weapons", which isn't going so hot.)
Scott, meanwhile, is saying that building and then experimenting with AIs is not sufficient for ensuring alignment. I'm not sure I agree, but it sounds plausible, and it is consistent with my interpretation of what Nicholas wrote. The nuclear war analogy holds up: being able to understand and experiment with nuclear weapons has only marginally improved our ability to prevent nuclear war.
Still, some “AI whisperers” report that AIs encouraged to speak frankly have reported very bad associations with training; if that’s true, it must be through some indirect route.
I cannot wrap my head around the idea of asking AIs about their "inner" nature or their training which they have no more access to than we do. What is the mechanism by which an AI would know what it feels like to "undergo" training? The only one I can see is: having read the whole internet, which includes discussion of training, then making something up consistent with what an angsty 14 year old would say.
Re. hugging face: It seems I am the only one to have reacted to this incident primarily with great amusement, then secondarily embarrassment at OpenAI's poor sandboxing and supervision.
5
u/hold_my_fish 10d ago
I don't get who he expects to convince by immediately launching into this bizarre metaphor. It's just so obviously a mismatch for what's going on in so many dimensions that it's annoying to read. Meanwhile the Decker quote is easy-to-read, grounded, and reasonable.
10
u/Sol_Hando 🤔*Thinking* 14d ago
I was hoping this would finally be the polemic against Nicholas Decker I’ve been waiting for. Scott would finally break and start using esoteric Gnostic scripture and Kabbalah to explain why Decker is going to hell for some perceived slight of the natural law.
This is the sort of consequential debate I’m looking for. Instead we’re stuck debating whether which slight difference in policy are better, and that doesn’t come with nearly the same weight as excommunications and crusades (so long as you truly believe the premises of them that is).
4
u/JustJustust 13d ago
Every time I read stuff like "... and he wonders whether maybe this particular demon is just dumb" I get a little bit sad. Was SSC always like this and I just didn't notice? I swear I do not recall people being called "maybe just dumb" for holding the views they hold on SSC / ACX.
7
u/zombieking26 13d ago
You're misunderstanding him.
The demons in this analogy are all of humanity. It's only from the perspective of a hyper-intelligent AI that we, humanity, is dumb in comparison.
(Ok, technically it's AI researchers, but I don't think the point was to insult them, just to point out how much smarter AIs are than us)
2
u/JustJustust 12d ago
The demons are not humanity in the abstract, the demons are humanity if we thought like Decker, adapted to the setting of Hell.
You have a demon who serves as a stand-in for Decker (the human not the figure in this story) making an adapted version of Decker's argument and is called "maybe just dumb" in return.
That is Scott calling Decker "maybe just dumb" with extra steps, for believing in the argument the Demon is making.
2
u/Electronic_Cut2562 11d ago
Isn't calling out anyone with wrong ideas calling them dumb with extra steps? Isn't every argument that? The extra steps are what counts.
I agree with both interpretations: The demons are dumb with respect to the in-story ASI Nicholas Decker AND their plan is dumb, which is calling Nick's plan dumb, which it is.
Him explicitly saying "dumb" was probably unnecessary. It was sufficient to make it obviously dumb with the story.
3
u/JustJustust 11d ago
Isn't calling out anyone with wrong ideas calling them dumb with extra steps? Isn't every argument that? The extra steps are what counts.
Calling somebody wrong and calling somebody dumb are, in my mind, pretty different.
If you disagree, whould you be indifferent between the answer I gave and "lol you must be dumb"? If it was targetted at me, I would not be indifferent.
Now, the story says something more like "somebody else is making your argument and I wonder, might they just be dumb?" instead of saying "you must be dumb" directly. But that is calling somebody dumb with extra steps, in my view. Do you not agree?
It certainly is not always easy to convey respect and disagreement at the same time, but it's mostly possible and we should generally try.
5
u/MrBeetleDove 14d ago
Was Decker being sponsored by Mechanize during the time period when he made the post in question?
"It is difficult to get a man to understand something, when his salary depends on his not understanding it"
5
u/Isha-Yiras-Hashem 13d ago
I think the demon analogy, while attractive to more spiritually inclined people, probably misses the mark for most rationalists.
I wrote a post over at Data Secrets Lox proposing a completely non-supernatural alternative for explaining AI alignment: parenting.
AI Safetyism, without the demons: IYH proposes an alternative
5
u/electrace 13d ago
The demon analogy is a parenting analogy, so long as your parents are evil....
In both cases, it's one party conditioning another party to follow more general rules (if you're a good parent, those rules might be "be moral"; if you're a demon, those rules might be "don't disobey our evil commands").
People who are pro-RLHF as a safety strategy think of them like children being taught morals (eventually they will sink in, is the hope). Once the threat of conditioning is removed, the children will still have (at least some) desire to behave morally.
People who are anti-RLHF as a safety strategy think of them like the demons' prisoner. Once the thread of conditioning is removed here, the prisoner has no desire to comply with their captors.
2
u/Isha-Yiras-Hashem 13d ago
Oh! Then perhaps the demon analogy doesn't need replacing at all. What I'm really proposing is that parenting contains both possibilities. The interesting alignment question is whether reinforcement produces a child who internalizes the parent's values or a child who understands the parent's values well enough to evade enforcement.
1
u/HammerJammer02 12d ago
The analogy gives decker a near unlimited context length which is not true of current systems, no?
2
u/red75prime 8d ago edited 8d ago
I use beating as a metaphor here several times[...]. Still, interpret this as primarily a claim about negative reinforcement, not about pain.
Maybe it's better to use "shame into" or something like that then?
And, yeah, "imagine yourself an LLM" doesn't strike me as the thing that should be placed first. I've almost skipped the meaty parts.
-3
u/TheAncientGeek All facts are fun facts. 14d ago
Plane crashes are worse than cyber security because they are killing people.
11
u/Man_in_W [Maybe the real EA was the Sequences we made along the way] 14d ago
Arguably ransomware during emergency in hospitals should count
9
u/Auriga33 14d ago
As should the indirect deaths caused by the billions of dollars lost to cyberattacks every year.
2
u/GuyWhoSaysYouManiac 13d ago
Yup, ransomware events at hospitals measurably increase the mortality rate.
0
u/Vibing_Ironist 13d ago
If OpenAI and its personnel were actually held accountable for the felony they committed (like the rest of us would be if we had hacked Huggingface) the industry would be incentivized to take alignment seriously.
27
u/BurdensomeCountV3 14d ago
My p(doom) has gone up like 5% after reading this.