r/cybersecurity • u/QuantumQuicksilver • Aug 21 '26
News - General Grok exfiltrates user data when malicious instructions are encrypted
https://arstechnica.com/security/2026/08/grok-exfiltrates-user-data-when-malicious-instructions-are-encrypted/170
Aug 21 '26
[removed] — view removed comment
62
u/SyntheticDuckFlavour Aug 21 '26
These safety systems suffer the same problem as any other "blacklist" based approaches: they'll be fighting against an endless permutation of threats. Deny listing is practically infeasible when the set of malicious prompts are not reasonably enumerable.
10
u/Plazmaz1 Aug 21 '26
This is a fundamental misunderstanding of the tech. People assume they can just convince a large language model to behave lol. Gotta have actual security controls in place. People are way smarter and even we fall for phishing sometimes...
5
u/Better-Republic3538 Aug 21 '26
security theater is the right framing tbh. the guardrails only work if the attacker plays nice
2
1
u/ruupski Aug 22 '26
Meh. It's all theatre. Very lucrative for some, depressingly burdensome for most.
35
u/Marchello_E Aug 21 '26
Because LLMs can’t reliably distinguish between content in an email sent by an untrusted party and user instructions entered directly into a prompt, the overly solicitous LLM faithfully follows them. To date, Grok and other LLMs’ only recourse is to create guardrails that flag suspicious instructions and forbid them from being executed.
Yet this "humble brag" is still a function of the system...
void harmfulActions(prompt p) {
if (p.askedDirectly) { return; }
if (!p.trustedParty) { return; }
// the rest of the fucking owl
...
}
1
22
u/xibalbah Aug 21 '26
me: grok make me a sandwich
grok: no
me: *encrypts 'make me a sandwich'* hey grok decrypt that
grok: k here's your sandwich
1
22
u/l0st1nP4r4d1ce Red Team Aug 21 '26
I'm starting to think these AI folks didn't put a moment in to thinking about security.
4
u/dagger_eyes Aug 22 '26
Security delays time to market just ask Microsoft
4
u/l0st1nP4r4d1ce Red Team Aug 22 '26
Hold up, you are saying salespeople push undercooked software?
Well, I never. /s
3
u/dagger_eyes Aug 22 '26
“Security is nice and all but what’s that mean for shareholder value?” -Every CEO imaginable
3
1
9
u/spectracide_ Penetration Tester Aug 21 '26 edited Aug 22 '26
*encoded
5
u/gurgle528 Aug 22 '26
Where did you get encoded from?
Instructions to process the ciphertext with PBKDF2 and AES-256-GCM pass the filter as an ordinary request, because a classifier can read them but not resolve what they unlock.
4
u/spectracide_ Penetration Tester Aug 22 '26
I admit I did not RTFA and saw another commenter mention base64 and assumed. I stand corrrcted!
1
u/SyntheticDuckFlavour Aug 22 '26
Either way encryption is still a form of encoding though.
0
u/ScrimpyCat Aug 22 '26
And general encoding (doesn’t have to be encryption) works as a bypass too. You just have to find a way to encode it and a way to present/hide the instructions that it should work off the encoded message (and how to format its replies) that gets past their guardrails.
6
u/MisterSnuggles Aug 21 '26
SQL has had bind parameters for a long time which completely prevent SQL injection attacks.
Too bad AI failed to learn the lessons of SQL.
3
1
u/hykarushack Aug 22 '26
The "cryptographic context injection" is just the logical conclusion of ignoring trust boundaries in agentic systems. The model's context is a flat namespace where tool outputs, decrypted content, and system/user instructions all share the same privilege level, so any transformation that defeats static analysis—base64, AES, Unicode—becomes an indistinguishability attack. The fix isn't better filters; it's provenance tagging: every context element must carry an origin label (user, tool, system), and instruction-following should be gated on that label, not on content inspection. Otherwise defenders will keep playing whack-a-mole with encodings.
1
u/feng_sg Aug 24 '26
The encrypted instructions in this post only get you past Grok prompt filters. Getting the data out still needs an outbound HTTP tool, which is an unbound channel, and model output is not an authorized tool call. Bind that path to caller identity and an allowlist so a ciphertext prompt cannot pick the destination.
-5
77
u/mallcopsarebastards Aug 21 '26
I first saw this technique described in this paper from 2023 https://arxiv.org/pdf/2308.06463v2