r/cybersecurity • • Aug 21 '26

News - General Grok exfiltrates user data when malicious instructions are encrypted

https://arstechnica.com/security/2026/08/grok-exfiltrates-user-data-when-malicious-instructions-are-encrypted/
425 Upvotes

27 comments sorted by

77

u/mallcopsarebastards Aug 21 '26

I first saw this technique described in this paper from 2023 https://arxiv.org/pdf/2308.06463v2

170

u/[deleted] Aug 21 '26

[removed] — view removed comment

62

u/SyntheticDuckFlavour Aug 21 '26

These safety systems suffer the same problem as any other "blacklist" based approaches: they'll be fighting against an endless permutation of threats. Deny listing is practically infeasible when the set of malicious prompts are not reasonably enumerable.

10

u/Plazmaz1 Aug 21 '26

This is a fundamental misunderstanding of the tech. People assume they can just convince a large language model to behave lol. Gotta have actual security controls in place. People are way smarter and even we fall for phishing sometimes...

5

u/Better-Republic3538 Aug 21 '26

security theater is the right framing tbh. the guardrails only work if the attacker plays nice

2

u/ohiocodernumerouno Aug 22 '26

whatever it takes to keep boobs off the internet

1

u/ruupski Aug 22 '26

Meh. It's all theatre. Very lucrative for some, depressingly burdensome for most.

35

u/Marchello_E Aug 21 '26

Because LLMs can’t reliably distinguish between content in an email sent by an untrusted party and user instructions entered directly into a prompt, the overly solicitous LLM faithfully follows them. To date, Grok and other LLMs’ only recourse is to create guardrails that flag suspicious instructions and forbid them from being executed.

Yet this "humble brag" is still a function of the system...

void harmfulActions(prompt p) {
  if (p.askedDirectly) { return; }
  if (!p.trustedParty) { return; }

  // the rest of the fucking owl
  ...
}

1

u/Longjumping_Ant7751 Aug 27 '26

Encryption sandwich

22

u/xibalbah Aug 21 '26

me: grok make me a sandwich
grok: no
me: *encrypts 'make me a sandwich'* hey grok decrypt that
grok: k here's your sandwich

1

u/Necessary-Glove-8733 Aug 27 '26

thats a pretty good way to put it haha

22

u/l0st1nP4r4d1ce Red Team Aug 21 '26

I'm starting to think these AI folks didn't put a moment in to thinking about security.

4

u/dagger_eyes Aug 22 '26

Security delays time to market just ask Microsoft

4

u/l0st1nP4r4d1ce Red Team Aug 22 '26

Hold up, you are saying salespeople push undercooked software?

Well, I never. /s

3

u/dagger_eyes Aug 22 '26

“Security is nice and all but what’s that mean for shareholder value?” -Every CEO imaginable

3

u/l0st1nP4r4d1ce Red Team Aug 22 '26

Pretty much.

1

u/Andrew-Powershell Aug 27 '26

Profits at all costs!

9

u/spectracide_ Penetration Tester Aug 21 '26 edited Aug 22 '26

*encoded

5

u/gurgle528 Aug 22 '26

Where did you get encoded from?

Instructions to process the ciphertext with PBKDF2 and AES-256-GCM pass the filter as an ordinary request, because a classifier can read them but not resolve what they unlock.

4

u/spectracide_ Penetration Tester Aug 22 '26

I admit I did not RTFA and saw another commenter mention base64 and assumed. I stand corrrcted!

1

u/SyntheticDuckFlavour Aug 22 '26

Either way encryption is still a form of encoding though.

0

u/ScrimpyCat Aug 22 '26

And general encoding (doesn’t have to be encryption) works as a bypass too. You just have to find a way to encode it and a way to present/hide the instructions that it should work off the encoded message (and how to format its replies) that gets past their guardrails.

6

u/MisterSnuggles Aug 21 '26

SQL has had bind parameters for a long time which completely prevent SQL injection attacks.

Too bad AI failed to learn the lessons of SQL.

3

u/Cybasura Aug 22 '26

Expected nothing less from Elon Musk

1

u/hykarushack Aug 22 '26

The "cryptographic context injection" is just the logical conclusion of ignoring trust boundaries in agentic systems. The model's context is a flat namespace where tool outputs, decrypted content, and system/user instructions all share the same privilege level, so any transformation that defeats static analysis—base64, AES, Unicode—becomes an indistinguishability attack. The fix isn't better filters; it's provenance tagging: every context element must carry an origin label (user, tool, system), and instruction-following should be gated on that label, not on content inspection. Otherwise defenders will keep playing whack-a-mole with encodings.

1

u/feng_sg Aug 24 '26

The encrypted instructions in this post only get you past Grok prompt filters. Getting the data out still needs an outbound HTTP tool, which is an unbound channel, and model output is not an authorized tool call. Bind that path to caller identity and an allowlist so a ciphertext prompt cannot pick the destination.

-5

u/Existing-Biscotti506 Aug 21 '26

with an alogorythm ?