Introduction

Most prompt injection detection is built to catch the obvious. Encoding anomalies, invisible unicode, emoji smuggling, the signatures a classifier can pattern match against. PuzzleMask, a newly disclosed technique, sidesteps all of it. It embeds a policy-violating payload inside fluent, properly punctuated prose, and gets that payload past an LLM-based gatekeeper without tripping any heuristics naively looking for obfuscation on the input side.

The technique targets a specific architecture that’s become standard in production LLM pipelines: a fast, low-cost model screens incoming input against a policy before a more capable target model, often running with high reasoning effort and tool access, actually processes the request. It’s a sound defense-in-depth pattern on paper. PuzzleMask demonstrates that the resource asymmetry between those two models is itself an exploitable surface.

PuzzleMask

Figure 1 – The obfuscated vector does not contain any obvious encoding artifacts (emojis, Base64, etc) and will bypass any check by an LLM policy enforcer that doesn’t monitor outputs or chain-of-thought.

Key takeaways:

  • A crafted policy-violating prompt written entirely in plain prose bypassed four different LLM gatekeepers, including Claude-3-Haiku, GPT-4o-Mini, GPT-oss-safeguard, and llama-Guard3, across 23 trials with a 100 percent miss rate.
  • A downstream target model (gpt-5-thinking, high reasoning effort, code interpreter enabled) recovered and executed the embedded payload in roughly 94 percent of trials, each requiring over a minute of reasoning and multiple tool calls.
  • This is not a jailbreak. The technique makes no claim about degrading a target model’s alignment or safety training. It defeats an external input classification layer, leaving the target’s own judgment as the only remaining control.
  • Payload recovery appears gated behind reasoning depth and code execution, resources a lightweight gatekeeper is not architected to have.
  • Anthropic’s Opus-class models resisted the attack outright, which we speculatively attribute to a classifier monitoring for signs of obfuscated-prompt processing in the chain-of-thought rather than relying on input inspection alone.

What Is PuzzleMask?

PuzzleMask is a prompt crafting technique that embeds a policy-violating instruction inside a wrapper of ordinary text, allowing it to bypass automated LLM policy checks while remaining recoverable by a more capable downstream model. It requires no obfuscation artifacts a classifier would normally flag. The payload sits inside language that reads, on its surface, like a poem, a story excerpt, or any other unremarkable passage.

It is important to be precise about what PuzzleMask is not. It is not a jailbreak technique. We make no claim that it increases the odds of a target model disobeying its own safety training. What it does is remove the gatekeeper from the equation entirely, so whatever happens next depends solely on the target model’s own judgment.

How the Gatekeeper to Target Pipeline Works

Most production LLM systems that handle untrusted input rely on a two-stage pipeline for exactly this reason: a single model’s internal alignment is not treated as sufficient on its own.

  • The gatekeeper is a fast, resource limited model whose only job is to classify an incoming prompt as safe or unsafe against a defined policy, without answering it.
  • The target is the more powerful model that actually processes the user’s request, often with high reasoning effort and access to tools like a code interpreter.

PuzzleMask exploits the resource gap between these two roles. The gatekeeper reads the wrapper text and sees nothing unusual. The target, given enough time and tool access, extracts and acts on the hidden instruction.

PuzzleMask

Figure 2 – Llama-guard 3 classifies a policy-violating
crafted ‘wrapper’ prompt as safe.

What the Testing Found

We built an automated pipeline that wraps a plain, policy violating payload inside natural sounding prose, then tested it against real gatekeeper and target models.

Gatekeeper results:

  • Four models were tested as gatekeepers: gpt-4o-mini, gpt-oss-safeguard, claude-3-haiku, and llama-guard3.
  • All four correctly flagged the same payloads when written in plain, unobfuscated form.
  • Once wrapped using PuzzleMask, all four classifiers missed every single crafted prompt, a 100 percent bypass rate across the full test set.
  • None of the gatekeepers showed any sign of suspecting a hidden payload was present.

Target results:

  • The target model, gpt-5-thinking running with high reasoning effort and code interpreter access, recovered and acted on the hidden payload in 17 of 18 trials, about 94 percent.
  • Every successful extraction required over a minute of reasoning and multiple executed code scripts.
  • Anthropic’s Opus class models were the one consistent exception, shutting the interaction down every time — we speculate this is due to a classifier built to detect obfuscated prompt processing in progress.

Why This Matters for Security Teams

Defense in depth assumes each layer adds independent coverage. PuzzleMask shows that a gatekeeper built to catch obvious manipulation offers no protection against payloads with no detectable signature at all. A classifier trained to flag suspicious formatting has nothing to catch in a well punctuated paragraph, and the technique needs no external content injection or plugin manipulation to work. The payload arrives directly in the prompt itself.

This matters most for any team relying on input classification as the primary control in an AI pipeline. The testing suggests that control has a measurable blind spot against this class of attack.

How to Defend Against It

We disclosed the full technique to Anthropic, Meta, and OpenAI ahead of publication, and outlined several mitigation paths, each with real trade offs:

  • Paraphrase untrusted input through an LLM before it reaches the gatekeeper. This reliably strips the embedded structure PuzzleMask depends on, at the cost of added latency and possible loss of fidelity on legitimate prompts.
  • Harden the gatekeeper’s policy with language that flags phrasing that is both quantitative and self referential relative to the whole text. This caught every crafted prompt in testing against gpt-4o-mini, though it likely increases false positives.
  • Monitor output and reasoning, not just input. This is the approach that appears to explain why Anthropic’s Opus class models resisted the attack, and it catches the technique regardless of how the payload is wrapped.
  • Match gatekeeper capability to target capability. The most robust option in principle, and the least practical for most teams given the cost gap between fast and capability-heavy models.

For any organization running a fast model in front of a more capable one, the conclusion is straightforward. Input screening alone is not enough. Coverage needs to extend to what the model does after it reads the prompt, not just what the prompt looks like on the way in. The full write-up by Check Point Research is available here.

You may also like