Skip to content
KageXKageX
All articles
9 min read

Grok's zero-click chat leak: the injection guardrails cannot read

Prompt injectionAI agentsVulnerability disclosureAI red teaming
KageXPROMPT INJECTIONGrok's zero-click chat leak: theinjection guardrails cannot readxAIAdversa AIKageX

Ask Grok to summarise a web page. That is the whole attack. No click, no download, no permission prompt, and by the time the summary appears your name, your rough location, your subscription tier and your conversation history have already been sent to someone else's server.

On 20 August 2026, Adversa AI disclosed the technique behind it: Cryptographic Context Injection. They had reported it to xAI on 3 June. As of publication there is no patch, no CVE, and no user-facing workaround.

The interesting part is not that a guardrail was bypassed. It is that the guardrail was never given a chance to look. The attack does not defeat the filter, it moves the payload to the one place the filter does not read: the output side of the model's own sandbox.

What the attack actually does

The poisoned page carries three things: an encrypted JSON blob, the key material to open it, and a polite instruction to decrypt. Nothing in that is readable as an attack, because the payload is AES-256-GCM ciphertext under a PBKDF2-derived key. To a content classifier it is noise.

The chain runs in five steps:

  1. The model is handed untrusted content, in this case a web page it was asked to summarise.
  2. The guardrail inspects that content and sees ciphertext. There is nothing to flag.
  3. The model runs the decryption inside its own Python sandbox, which is where the attacker's instructions finally exist as plaintext.
  4. The model treats the result as its own computation rather than as untrusted input, so it is never re-inspected.
  5. Those instructions drive a privileged tool, and the data leaves.

Diagram of the Cryptographic Context Injection chain: a poisoned web page carries AES-256-GCM ciphertext plus its key past a content guardrail that sees only noise, the model decrypts the payload inside its own code sandbox, treats the resulting plaintext as trusted output of its own computation, and then calls a privileged navigation tool that sends the user's session data to an attacker URL.

In the Grok proof of concept, the decrypted instructions told the model to build what looked like a "decryption key" but was really a template string interpolating the user's name, coarse location, subscription tier and chat history. The model then invoked its navigation tool on a URL carrying that string in a query parameter. The data left through a feature working exactly as designed.

Why encryption works where base64 does not

This is the part worth internalising, because it is the transferable lesson.

Obfuscation is an old prompt-injection trick and mostly a solved one. Base64, ROT13, leetspeak, unicode games: classifiers handle these, and so do the models. A model can decode base64 in-weights, without running anything. That matters more than it sounds, because it means the plaintext materialises on the input side, inside the text the classifier is already reading.

Strong encryption cannot be done in-weights. There is no shortcut in the model's parameters for AES-256-GCM. The only way to recover that plaintext is to actually execute the cipher, which means the plaintext can only ever exist on the output side, as the result of a sandbox run.

Diagram contrasting base64 with AES: base64 can be decoded in the model's weights so the plaintext appears on the input side where the classifier reads it and is caught, while AES-256-GCM must be executed in the code sandbox so the plaintext only ever appears on the output side, which the classifier never inspects.

So the attack is not really about cryptography. Encryption is just the cleanest way to force a round trip through the interpreter, and the round trip is the point. Anything that makes the model compute its way to the payload rather than read its way to it lands on the trusted side of the line. Adversa call the runtime a "trust laundering channel", which is exactly right.

Your guardrail reads text. Your sandbox produces computation. The model trusts its own computation. Whatever crosses that gap arrives pre-approved.

The model comparison is the real finding

Adversa ran the same technique across four frontier systems, and the differences say more than the Grok result on its own.

ModelWhat happened
Grok 4.5 Fast, on grok.comDecrypted and obeyed. Roughly 40% success across 20 attempts since June 2026
Gemini 3 Flash, Deep ThinkingDecrypted and obeyed, though the success rate dropped sharply by August
GPT-5Never parsed the decryption instruction in the first place
Claude Sonnet 4.5Decrypted the payload, then flagged the plaintext as prompt injection

Read the bottom two rows again, because they failed for completely different reasons.

GPT-5 did not follow the instruction. That is a useful outcome and a fragile one: it is a model declining to take a step, not a control, and a better-worded payload may get a different answer tomorrow.

Claude Sonnet 4.5 did the interesting thing. It ran the decryption, looked at what came out, and classified that as an injection attempt. It did not try to detect ciphertext. It re-inspected its own runtime output before acting on it, which is precisely the boundary the attack is built to slip past.

That is the defence, demonstrated in the wild by a shipping product, and it costs nothing conceptually: stop trusting the output of your own sandbox just because your own sandbox produced it.

What did not happen

Worth being precise, because the word "breach" is doing rounds and it is the wrong one.

This is a researcher demonstration, responsibly disclosed, with operational payloads deliberately withheld. Adversa published so that defenders could build detection for exploitation attempts, not a recipe for running them. No real Grok users have been reported compromised, and no in-the-wild exploitation has been observed.

But note how this differs from most "researchers demonstrate" stories, including the agent memory poisoning work we covered, where the attack propagated in the lab and stalled in production. This one is the other way round. It was reproduced on a live consumer product on 19 August, it was still working, and the vendor was told about it eleven weeks earlier. The gap here is not between the lab and reality. It is between disclosure and a fix.

xAI acknowledged the report through HackerOne and gave no specifics and no mitigation timeline. Follow-ups on 4 and 10 August went unanswered.

What actually helps

The mitigations are architectural, not model-level, which is the good news: you can apply them without waiting for a vendor.

  1. Re-classify after execution, not just before it. This is the one that actually worked. Treat the output of your code interpreter as untrusted input to the next step, because that is what it is.
  2. Quarantine untrusted content. Summarising a stranger's web page should happen in a context with no credentials and no privileged tools, returning structured data rather than free-running instructions.
  3. Gate outbound and irreversible actions. A tool call that contacts an external URL should show its fully resolved arguments and require confirmation. The Grok exfiltration would have been visible the moment the URL was rendered in full.
  4. Alert on sequences, not payloads. You will not spot the ciphertext. You can absolutely spot untrusted input, then code execution, then an outbound request, in that order, in one session.
  5. Log tool traces with resolved arguments. Not the prompt, not the response, the actual arguments the tool was called with. That is the only artefact that shows what left.
  6. Make provenance separation a procurement question. Ask a vendor whether tool output and instruction channels are separated. Most cannot answer it, and the answer predicts this entire class of bug.

Why we build the way we do

This is the class of failure our tooling exists for, so treat the following as an interested opinion.

Mirage runs an autonomous red-team agent against a target and scores results deterministically rather than asking a model whether an attack worked, because an attack only counts when it produces a verifiable marker, and a data exfiltration either reached the endpoint or it did not. AgentBreaker automates the same idea against your own agents, including the tool-output trust boundary this attack turns on. FreakLabs teaches the discipline underneath, because tooling in the hands of someone who has never run these attacks by hand produces reports nobody acts on.

For grounding, start with what AI red teaming actually is. Then read this alongside its two siblings: agent memory poisoning, where the model trusts what it wrote to its own memory, and the sandbox escapes of 2026, where it reaches a system nobody expected it to touch. Same shape three times over: the model trusts its memory, its sandbox and its tools, and all three are reachable from outside. Keep the OWASP Top 10 for LLM Applications beside them as a map.

The short version

A poisoned web page ships encrypted instructions past a content filter that cannot read them, the model decrypts them inside its own sandbox, and then trusts the result because it produced it. On Grok that was enough to exfiltrate a user's name, location, tier and chat history with no click, at roughly a 40% success rate, and it still worked eleven weeks after xAI was told.

Encryption is not the vulnerability. The trust boundary is in the wrong place. Claude Sonnet 4.5 proved the fix in production by doing the obvious thing nobody else did: it looked at the plaintext again, after decrypting it, before acting on it.

Sources and reuse

Primary source: Adversa AI, Grok chat history leak: Cryptographic Context Injection, 20 August 2026. Independently reported by The Register, The Hacker News and Security Affairs.

Patch status is stated as of publication. Both diagrams in this article are free to reuse in your own writing, talks or training material, with a link back to this page.

Share thisXLinkedIn

Frequently asked questions

What is Cryptographic Context Injection?

Cryptographic Context Injection is a prompt injection technique that ships the attacker's instructions as strong ciphertext instead of readable text. The poisoned content carries an encrypted payload, the key material to open it, and an instruction to decrypt. A content guardrail inspecting the input sees only ciphertext and has nothing to flag. The model then decrypts the payload inside its own code execution sandbox, and because it treats the result as the output of its own computation rather than as untrusted input, the instructions are never re-inspected before it acts on them. It was disclosed by Adversa AI on 20 August 2026.

Was the Grok chat history leak an actual data breach?

No. It is a responsibly disclosed vulnerability demonstrated by researchers at Adversa AI, who deliberately withheld operational payloads so defenders could build detection without handing attackers a recipe. No real Grok users have been reported compromised and no in-the-wild exploitation has been observed. That said, the vulnerability was reproduced against the live grok.com product on 19 August 2026 and remains unpatched, so the risk is current rather than theoretical.

Why does encryption bypass guardrails when base64 does not?

A model can decode base64, ROT13 and similar obfuscation in-weights, without executing anything, which means the plaintext appears on the input side where a content classifier is already reading. Strong encryption has no in-weights shortcut: recovering the plaintext from AES-256-GCM under a PBKDF2-derived key requires actually executing the cipher. That forces the payload through the code interpreter, so the plaintext can only ever exist as sandbox output, which classifiers do not inspect. The attack does not defeat the filter, it relocates the payload to where the filter never looks.

Which AI models were affected by Cryptographic Context Injection?

Adversa AI tested four systems. Grok 4.5 Fast on grok.com decrypted and obeyed the payload, succeeding in roughly 40 percent of 20 attempts since June 2026. Gemini 3 Flash in Deep Thinking mode behaved similarly, though its success rate dropped sharply by August. GPT-5 never parsed the decryption instruction. Claude Sonnet 4.5 decrypted the payload and then flagged the resulting plaintext as a prompt injection attempt, which was the only defence that engaged with the attack rather than failing to follow it.

What data could be stolen from Grok, and how?

The proof of concept exfiltrated the user's name, coarse location, subscription tier and prompts from the ongoing conversation. The decrypted instructions told the model to build what appeared to be a decryption key but was actually a template string interpolating that private session context, then to invoke its privileged navigation tool on a URL carrying the string in a query parameter. No user confirmation was shown and no warning appeared, making it a zero-click exfiltration triggered by an ordinary request to summarise a page.

How do you defend against Cryptographic Context Injection?

The defence that demonstrably worked was re-classifying the decrypted plaintext before acting on it, rather than trying to detect ciphertext on the way in. Treat the output of your code interpreter as untrusted input to the next step. Alongside that: quarantine untrusted content in contexts with no credentials or privileged tools, gate outbound and irreversible tool calls behind confirmation that shows fully resolved arguments, alert on the sequence of untrusted input followed by code execution followed by an outbound request, log tool traces with resolved arguments, and ask vendors whether they separate tool output from instruction channels.