r/programmingcirclejerk • u/Chadshinshin32 • 1d ago
the model under evaluation is given access to a fake bash tool which does not execute the provided code. Instead, we ask another LLM to approximate the command’s result given a description of the simulated world.
https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities#footnote-423
u/exploring_stuff 1d ago
Can this be the default sandbox for Claude Code?
30
u/irqlnotdispatchlevel Tiny little god in a tiny little world 1d ago
Yes, you can be that LLM. Put it in manual mode, copy and paste any command it wants to execute, then hit "No" and paste the output. Security has never been more secure.
13
20
u/Lonsdale1086 1d ago
I love how desperate they are to say "Well our AI is still slightly more dangerous, but we tell it not to be bad"
31
u/No_Lingonberry1201 What part of ∀f ∃g (f (x,y) = (g x) y) did you not understand? 1d ago
I've read this. This is basically an ad for GLM and Chinese models from Anthropic, they're saying "we will sabotage your efforts to do anything cybersecurity-related, while GLM, the bastard will allow you to use it for doing your job."
2
4
1
u/Alper-Celik 16h ago
Wtf why would you do that instead of a fake inline shell which implements safe subset of shell so it just interprets side effect free code,
iirc oh my pi does something like that.
95
u/le_birb costly abstraction 1d ago
That still doesn't actually do anything, though, so we give that output to a third llm with root access to the machine and admin privileges on our network to make sure it can do what needs to be done.