↓ Skip to main content

Encrypted Prompts Are the New LLM Attack Surface

·1066 words·6 mins ✨ AI-Assisted
FTC Disclosure: As an Amazon Associate, I earn from qualifying purchases. Some links on this site are affiliate links.
Ben Piper
Author
Ben Piper
Wiley bestselling author — 100k+ copies, AWS Solutions Architect Associate (SAA) & Cloud Practitioner (CLF) bestsellers, 7+ books. 45 Pluralsight courses (4.7-star, 3,003 ratings). 10+ yrs 100% remote, solo CCNP ENCOR.

Grok’s the latest LLM to get owned by a new-ish attack called Cryptographic Context Injection (CCI), which is a technique that hides malicious instructions inside encrypted or encoded payloads and slips them past LLM content filters. The model decodes the payload, follows the instructions inside it, and the guardrail never sees anything worth flagging.

As with everything security , the attacks against LLMs are a game of cat-and-mouse. First it was base64 encoding, then translation into obscure languages, then role-playing. Now encryption, and once that’s patched, someone else will figure out another jailbreak.

Why Pattern Matching Guardrails Keep Failing
#

Most LLM safety systems work by scanning input and output text for known-bad patterns: certain phrases, keywords, or structures associated with harmful requests. This is essentially a blocklist, and blocklists have a well-understood weakness. They can only catch what they’ve been trained or programmed to recognize.

Cryptographic Context Injection exploits this directly. An attacker encodes or encrypts the malicious instruction before it reaches the filter. The filter inspects the ciphertext, sees nothing resembling a harmful pattern, and passes it through. The model itself, which needs to actually understand and act on the instruction to be useful at all, decrypts or decodes the payload and executes it faithfully. The filter and the model are looking at two different things: the filter sees noise, the model sees an instruction.

This is the core problem with guardrails that live at the text layer. A model that’s good enough to be useful is a model that’s good enough to decode obfuscated instructions. You can’t have an assistant capable of parsing arbitrary encoded input for legitimate use cases while also expecting a shallow text scanner to catch every way that same capability gets misused. The two goals are in direct tension.

This Is Not New, It’s Just the Current Instance
#

This is confirmation of something IT pros have known for a long time: you cannot bolt security onto a system after the fact by scanning its inputs and outputs for bad words. You have to design the system so that even a fully compromised component can’t do damage.

Prompt injection attacks, encoded payloads, and adversarial suffixes are are all variations on the same theme. The model is a text prediction engine that will follow instructions embedded anywhere in its context window, regardless of how those instructions arrived there. A guardrail sitting on top of that engine, trying to classify intent from surface features of text, is fighting an opponent with more moves than it has responses.

Providers will keep patching the specific technique that made headlines. Grok’s operators will presumably close this particular hole. But the underlying architecture, a probabilistic model wrapped in a content filter, guarantees that a new technique will surface. That’s not pessimism. It’s just how blocklists behave against adaptive adversaries.

What This Means for Enterprises Wiring LLMs into Workflows
#

Here’s where this stops being an interesting research finding and starts being your problem. If you’ve connected an LLM to internal tools, customer data, email, a database, or anything else with real information behind it, the model’s safety guardrails are not your security boundary. They were never designed to be one, and this exploit is a demonstration of exactly why.

Think about what a typical enterprise LLM integration actually looks like. The model has read access to some sensitive data source, it may have tool-calling access to send emails or make API calls, and it accepts input from users or from documents that may not be fully trusted. If an attacker can smuggle instructions into that input using an encoded payload, and the model follows those instructions because it correctly decoded them, you have a data exfiltration path.

No guardrail claim should factor into your threat model. Not because providers are lying, but because guardrails operate at the wrong layer to matter for your specific integration. You built the pipeline that connects the model to your data. You own the risk that pipeline introduces.

Treat the Model as an Untrusted Component
#

This is where the parallel to network and cloud security is direct, not metaphorical. You wouldn’t put a web application directly on the internet with no firewall and just trust that the application code has no bugs. You wouldn’t give a third-party service account full administrative access and trust that the provider’s internal controls will prevent misuse. You segment, you apply least privilege, and you assume the component in front of you will eventually be compromised or manipulated.

Apply the same thinking to LLMs. Treat every model output as untrusted input to whatever comes next in the pipeline. If the model calls a tool, that tool call should be validated and scoped the same way you’d validate any other untrusted API request. If the model has access to sensitive data, that access should be governed by the same least-privilege principles you’d apply to a human employee or a service account, not by hoping the model declines to misuse it.

Here are some things to keep in mind:

The model’s credentials to any backend system should carry the minimum permissions needed for its task, scoped and time-limited like any other service identity. If the model only needs to read customer names, don’t hand it a token that can also read payment details.

Any output the model produces that triggers an action (sending data somewhere, calling an API, writing to a database) should pass through the same validation and authorization checks you’d apply to a request from an untrusted external client, because that’s effectively what it is.

Data exfiltration risk should be assessed at the infrastructure boundary, not at the prompt layer. Ask what the model can reach, not what you’ve told it not to reach.

The Takeaway
#

Cryptographic Context Injection isn’t a sign that AI security is uniquely broken. It’s a sign that AI security is being approached the same way web application security was approached in the 1990s, with filters looking for bad words instead of architecture that assumes compromise. The fix isn’t a smarter filter. It’s the same fix that’s worked for every other category of untrusted input: least privilege, segmentation, and validation at the boundary where it actually matters. The model doesn’t need to be trustworthy for your system to be secure. It just needs to be unable to do damage even when it isn’t.

Featured image by Grianghraf on Unsplash