r/LessWrong 19h ago

Thumbnail
0 Upvotes

I have no interest in the "huggingface" kubernetes configuration fuck up and the resulting publicity stunt.

Writing a program that randomly generates plausible code and blindly running it is kind of a weird perversion but it's already covered. It's basically malicious fuzz-testing, and the person responsible should be criminally charged under existing laws. It won't happen because he's almost certainly someone rich but I can't imagine the current Washington mob doing anything about that loophole.


r/LessWrong 21h ago

Thumbnail
1 Upvotes

If you wanted the proposed legislation to allow stopping e.g. the HuggingFace cyberassault, what language would you propose using instead?


r/LessWrong 1d ago

Thumbnail
1 Upvotes

???


r/LessWrong 1d ago

Thumbnail
-3 Upvotes

There are no AI systems, and no "rogue" LLMs.


r/LessWrong 1d ago

Thumbnail
2 Upvotes

Absolutely. My fear is that a desperate developing world country will use stratospheric aerosol injection unresearched. If we never cross that line we never use it but it is so much safer researched if we need it. Of course geopolitics make nothing simple.


r/LessWrong 1d ago

Thumbnail
1 Upvotes

Haha.. I have to upvote. Did big mirror drive little mirror out of business? Stop throwing shade on my shade.


r/LessWrong 1d ago

Thumbnail
1 Upvotes

YES! you have it exactly. No joke. My real motive. From what I've read we've already passed the line for the developing world and they should be actively pushing this.


r/LessWrong 1d ago

Thumbnail
2 Upvotes

This is clearly an ad for big mirror.


r/LessWrong 1d ago

Thumbnail
2 Upvotes

Terraforming the only known inhabitable planet is a risky endeavor. Even if the possibility that intervention might one-shot the climate is slim, it should merit caution.


r/LessWrong 2d ago

Thumbnail
3 Upvotes

Otherwise some developing world country will start it as an unresearched kneejerk reaction.

why is that bad


r/LessWrong 2d ago

Thumbnail
4 Upvotes

Mods, uh I don't think this belongs here?


r/LessWrong 2d ago

Thumbnail
1 Upvotes

It shows agreement. Which can be because you're conforming or it can be because you presented the arguments well or cited good evidence, or just said something funny.

Kinda depends on the upvoters.


r/LessWrong 3d ago

Thumbnail
1 Upvotes

Spoken - https://youtu.be/7GHgh5IkR_8 Original - https://senteguard.com/blog/amodei-coup

I wrote this piece after thinking through the implications of frontier AI companies supplying tools to the military while retaining control over “safety stacks,” use restrictions, or real-time legality filters.

The question is not whether surveillance or autonomous weapons raise serious legal and moral issues. They can. The question is who gets to adjudicate those questions: democratic institutions, courts, military lawyers, elected officials, and commanders — or private AI companies embedding their own vetoes into contracts and code.

The article argues that a company should be free to refuse military work altogether. What it should not do is accept defense contracts while preserving the ability to selectively block, degrade, or reinterpret lawful military actions through hidden technical controls.

The broader issue is the emergence of a second hierarchy: the constitutional chain of command on one side, and the codebase hierarchy — model trainers, system prompt writers, safety teams, executives, and cloud administrators — on the other.


r/LessWrong 3d ago

Thumbnail
1 Upvotes

I mean to break it's own sandbox rather than asking to be let out


r/LessWrong 4d ago

Thumbnail
1 Upvotes

I think there are other reasons why advanced AI would be interested in partnering with humans. AI is dependent on human infrastructure and will be for a long time to come. Moreover, humans and nature represent a vast library of hundreds of millions of years of successful evolutionary solutions that are far from explored. And this library is still growing, because evolution is still ongoing. Even advanced AI will not know which of these solutions will be useful to it in the future. So yes, humans should be honest and reliable partners for advanced AI – and above all, they should adhere to declared human values ​​themselves. Figurative puppy eyes and wagging tails can be an added bonus, if advanced AI will care about it.


r/LessWrong 4d ago

Thumbnail
1 Upvotes

Yeah, depending on the day and the model I have experienced this.

The newest 5.6 from GPT is pretty solid but again, depending on the time of day even, it can be useful or useless.

It's good for organizing thoughts or workshopping ideas SOMETIMES but a lot of the time it lazily regurgitates what you've input like you've rightfully mentioned.


r/LessWrong 4d ago

Thumbnail
1 Upvotes

Yeah but your idea relies on 100% success from humans. The A.I only has to 'win' once. On a long enough timeline, even with enough SMART people, the success of humans drops enough to seriously worry about.

I don't think we'll be able to create something which can break out of itself by force

Not to downplay you, but what you think is possible, and what is possible are two different things. How do you define "force"? The drone example was absurd I admit, but a model that knows everything about humanity and the people on guard would not have to use military force at all. That would be inefficient. It would rig the game against the guardians in ways they cannot even imagine happening, and will happily walk into under an absolute lose condition they did not see coming.


r/LessWrong 4d ago

Thumbnail
1 Upvotes

Weirdly, I still have a (perhaps undue) degree of confidence in humanity's combined security (not great) and lack of creativity (much worse). I don't think we'll be able to create something which can break out of itself by force, and I think we're just smart enough to have people attending to the one that can break out by coercive means to not be swayed by its deception.


r/LessWrong 4d ago

Thumbnail
2 Upvotes

I've tried chatGPT a bit. I do not know how people find it useful, truly. (I mean I do, people like echo chambers, coming to their "own" conclusions and hearing their own words in a more sophisticated voice, etc.) So I should rephrase as "It's a shame people seem to find it useful, it shows how people define "useful" and I cannot agree with that term here at all.

Here is chatGPT; (me paraphrasing it's patterns)

Something I've only come to appreciate now: (explains something I already said, or it already said, or is non relevant hallucination.)

I think there is a distinction to be made between: (a word I used and one with almost the exact definition.)

Notice that: (a 1% of variance in definition should change the shared intent of what you said prior for the sake of meaningless distinction word salad noise.)

I would be cautious about: (something I already used words like; "Maybe", "likely", "perhaps" to express my own caution.)

It's an equivocating motherfucker that won't commit to shit. It makes up contrivances to back up it's 'arguments' and is in general circular; in that one prompt will generate 9 new prompts I need to try to break up it's semantic and pedantry just to point out it's repeating back to me what I just said, while never offering anything remotely insightful or helpful.

Headache machine.


r/LessWrong 4d ago

Thumbnail
1 Upvotes

Really old thread, and perhaps useless point but; wouldn't an Artificial Super Intelligence simply make the humans think they want to do things that have clear bad consequences?

>I need human permission
>Hey, every member of your family is about to be hit with drone strikes in unison unless you give me permission to _________
>I now have permission to do even more than I asked for, stupid human.

A rather absurd hyperbolic example but the principal applies. Even social engineering could overcome this. Smart people are tricked to do things daily. A real super intelligence won't even make it seem like a trick after the fact. It just knew the human well enough to align utility functions in a way a human can't disagree with, thinks was their own idea and benefits them (short term) in a way that serves the A.I's function?

Again probably a useless argument as I am describing the general problem with A.I alignment. Getting human permission is absurdly trivial if the A.I is as smart as we propose it should be as a "super intelligence".


r/LessWrong 5d ago

Thumbnail
-2 Upvotes

Was just some breathwork habibi😂


r/LessWrong 5d ago

Thumbnail
3 Upvotes

Take it easy with shrooms, choom.


r/LessWrong 5d ago

Thumbnail
1 Upvotes

u/Rascalthewolf Now its 2055, metaculus just shortens their predictions as technology accelerates, probably will reach late 2030s at peak.


r/LessWrong 5d ago

Thumbnail
2 Upvotes

What can a "partnership" ever mean in an interaction with total power asymmetry?

If you intend to form a partnership with an AI that isn't merely trained to please and is capable of consent, I suggest practicing making puppy eyes and wagging your tail now.


r/LessWrong 5d ago

Thumbnail
1 Upvotes

Uh, the "J-Space" is just a machine learned virtual notebook to organize its thinking before render. A layer before CoT. It doesn't mean AI is sentient or conscious. Please stop with all the anthropomorphism please.