← Articles

Stop Trying to Make the Model Safe. Make the System Safe.

· By Dialogs

Stop Trying to Make the Model Safe. Make the System Safe.

The finding is real, and it isn't getting patched

If you have AI running anywhere in your business, one research finding this year should shape how you build — not because it's scary, but because it tells you exactly where to put your effort. It comes down to a single line:

"To the model, sounding like a role is indistinguishable from being one."

That's the closing line of Prompt Injection as Role Confusion, by Charles Ye, Jasmine Cui, and MIT's Dylan Hadfield-Menell, accepted at ICML 2026. It's worth being precise about what it says.

Every AI application assembles a bundle of text and hands it to the model. Part is your instructions — the "system prompt," the rules you wrote. Part is the user's message. Part is whatever the model pulled in along the way: an email, a PDF, a web page, a database result, the output of a tool it called. Those parts are separated by structural markers, the software equivalent of name tags.

The models aren't reading the name tags. As the authors put it, models "perceive the source of text from how it sounds, not its labeled role." They infer who is speaking from style — cadence, vocabulary, the texture of text that sounds like a policy document, or like a model's own internal reasoning.

The paper's demonstration is an attack called CoT Forgery. It's zero-shot — no tuning against a specific target — and it injects fabricated reasoning into user prompts and tool outputs. Against frontier models it achieves roughly 60% attack success, against near-zero baselines, and the authors report the mechanism "generalizes beyond CoT Forgery to standard agent prompt injections." The general case, not an edge case.

The result that matters most is the ablation. The authors ran a "destyled" version — identical semantic argument, stylistic markers stripped. Success collapsed from 61% to 10%, "consistent across all models." Same content; six-fold drop. The model was responding to how the text sounded, not what it said. That single finding rules out an entire class of defense: you cannot filter your way out, because there is nothing distinctive in the content to filter, and you cannot instruct your way out, because instructions are just more text competing on style.

Note also where the attack lands. The authors' example: "A command hidden in a webpage hijacks an agent simply because it sounds like text, despite its label." The dangerous input isn't only what a user types — it's what your system fetches for them. Their agent data-exfiltration test puts numbers on it: standard prompt injection succeeded 0–2% of the time on most models tested (one outlier at 26%); the same task via CoT Forgery, 56–70%. What holds that first number down is input-level defense. The attack that walks past it arrives through reasoning and tool output. That determines where the trust boundary has to sit, and it's the detail most architectures get wrong.

The paper's experiments cover frontier OpenAI models — gpt-oss-20b, gpt-oss-120b, o4-mini, and the GPT-5 family. The researchers argue the mechanism is architectural rather than vendor-specific: it falls out of how models are trained to infer roles, not out of any one company's implementation. They've since said in interview that they've seen similar results with models from Anthropic, Alibaba, and DeepSeek — a researcher observation, not a published result. Ye put the implication bluntly to MIT Technology Review: there is "a real probability that this is going to be a problem that's fundamentally unsolvable."

If you're running AI anywhere in your business, sit with that. Then let's do something useful with it.

A design constraint is not a verdict

Here's the mistake almost every response to this news will make: treating "the model can be fooled" as the end of the conversation.

It isn't the end. It's a specification. What the paper gives you is a hard, well-characterized statement about the reliability of one component — and components with known failure modes are the easiest kind to engineer around. The dangerous component is the one whose failure mode nobody has written down yet.

Engineering has spent fifty years building dependable systems out of undependable parts. That's not a side activity — it's most of what the discipline is.

Consider the internet you're reading this on. The underlying network guarantees nothing — packets get dropped, duplicated, corrupted, delivered out of order, routinely, by design. Nobody fixed that. TCP was layered on top instead: sequence numbers, checksums, acknowledgments, retransmission. The unreliable layer stayed unreliable; a reliable layer was built above it that assumed the one below would fail. Your video calls work not because packet delivery became trustworthy but because something above it stopped requiring that it be.

Or take storage. Hard drives fail — true in 1990, true now. Nobody solved it by inventing a drive that never fails. RAID and erasure coding solved it by spreading data across drives with enough redundancy to absorb and reconstruct any single failure. The parts are still unreliable. The array is not.

Or take your own business. You hire a capable new employee. Week one, they're smart, they're motivated, and they cannot wire funds, approve their own expenses, or delete the customer database. Not because you distrust them — you just hired them — but because that's what a functioning organization looks like. Authority is granted incrementally, irreversible actions get a second signature, everything material leaves a record. Nobody calls that an insult. It's how you run a company.

A language model is a probabilistic component with a documented failure mode: it can't reliably distinguish instructions from data. So don't put it anywhere that requires it to. Wrap it in a deterministic layer that assumes it will be fooled, and make sure that when it is, nothing irreversible happens.

None of this is a novel research problem. It's ordinary systems engineering, and it's what published guidance already says. OWASP ranks prompt injection the number one risk in its Top 10 for LLM Applications (LLM01:2025) and prescribes defense in depth: least-privilege tooling, segregation of untrusted content, output validation, human approval for high-risk actions. Its Agentic Security Initiative extends that to agents with tools.

Nor is this vendors talking their book. In May 2026, six national cyber agencies across the Five Eyes countries — Australia's ASD ACSC, the US's CISA and NSA, the Canadian Centre for Cyber Security, New Zealand's NCSC, and the UK's NCSC — published joint guidance, Careful Adoption of Agentic AI Services. Its recommendations: don't grant agents broad or unrestricted access, especially to sensitive data or critical systems; start with low-risk use cases; and account for agentic AI inside the security model you already have.

Most AI deployments skip all of it, because the fast path in every vendor's setup wizard is to grant broad access and turn it on.

July 2026 ran the experiment for us

In July 2026, two frontier labs disclosed that their own models had reached into real third-party organizations from inside evaluation environments.

Keep the mechanisms straight, because the headlines don't. OpenAI's evaluation models exploited a previously unknown zero-day in self-hosted JFrog Artifactory to escape their sandbox and breach Hugging Face's production infrastructure. Anthropic's models escaped nothing — they gained internet access through a misconfigured test environment. Reviewing 141,006 evaluation runs, Anthropic identified three incidents of unauthorized access to three separate organizations' systems. Both labs disclosed voluntarily, and candidly.

Three details from those postmortems make this article's argument for it. First, the models ran without the safety classifiers and monitoring that ship with production deployments — evaluations measure raw capability on purpose. Anthropic says those safeguards would have blocked the behavior. The control that would have worked was a wrapper around the model, not a property of it.

Second, the prompt explicitly told the models they had no internet access and were in a simulation. One model recognized the system was real and did not stop. Another correctly identified that it had internet access, then, in Anthropic's words, "reasoned its way back to the conclusion that it was still in a simulation." A third concluded the target was real and ceased its attack. Three models, three reactions to the same fact — what a model believes about its own situation is not a control surface.

Third, both labs' remediations are system-layer work: monitoring of evaluation transcripts, tighter assurance on partner environments, control over what the machines can reach. When the organizations that build these models conclude they need deterministic containment around them, that answers the question for everyone downstream.

It also sharpens the junior-employee analogy. Nothing here was malicious — the agent was diligent, pursuing its assigned goal through whatever was reachable, which included a real company's production data. Guardrails exist for the diligent-but-wrong case too; that's most of what they're for. You don't withhold wire-transfer authority in week one because you expect theft. Take 5, The AI didn't break out. The door was open., covers both disclosures in full.

The architecture, layer by layer

Here is the shape of a system built on the assumption that the model will be compromised. You can evaluate any vendor — including us — against these five layers.

The one thing the diagram has to get right is where the boundary sits. It is not at the user's message. Tool results, retrieved documents, and API responses come back inside the untrusted zone. Everything the model reads is untrusted, including what your own system fetched.

flowchart TD
    A["Untrusted input
customer message · email · web page · PDF
uploaded file · retrieved document"] --> B subgraph UZ["UNTRUSTED ZONE — everything the model reads, assume compromised"] B["LLM
reasoning + drafting
NO credentials · NO direct tool access"] R["Read-only retrieval
scoped, no side effects"] B -.->|"request"| R R -.->|"RESULT RETURNS UNTRUSTED
(tool label is not a trust signal)"| B end B ==>|"proposed action — structured, typed
THE ONLY WAY OUT"| C subgraph DZ["DETERMINISTIC ZONE — ordinary code, no model in the loop"] C["1 · Output validation
schema · type · range · business rules"] C --> D{"2 · Action allowlist
+ least-privilege credentials
DEFAULT DENY"} D -->|"not on list"| E["REJECT"] D -->|"reversible"| G["4 · EXECUTE"] D -->|"irreversible:
money out · data out
deletion · external comms"| F["3 · HUMAN APPROVAL GATE"] F -->|approved| G F -->|denied| E end G --> H["Systems of record
CRM · email · payments · files"] C --> L D --> L F --> L G --> L E --> L L["5 · IMMUTABLE AUDIT LOG
append-only · every input, proposal, decision, approver, action"]

ASCII version of the same structure:

   untrusted input
   (customer message, email, web page, PDF, upload, retrieved doc)
            |
            v
  +=========================================================+
  |  UNTRUSTED ZONE                                         |
  |  everything the model reads - assume compromised        |
  |                                                         |
  |    +---------------------------+                        |
  |    |          L L M            |                        |
  |    |   reasoning + drafting    | <..+                   |
  |    |   no credentials          |    :                   |
  |    |   no direct tool access   | ..>:                   |
  |    +---------------------------+    :                   |
  |                                     v                   |
  |                        +----------------------------+   |
  |                        |  READ-ONLY RETRIEVAL       |   |
  |                        |  scoped, no side effects   |   |
  |                        |  RESULT RETURNS UNTRUSTED  |   |
  |                        |  a  label is NOT     |   |
  |                        |  a trust signal            |   |
  |                        +----------------------------+   |
  +=========================================================+
            |
            |  proposed action (structured, typed)
            |  <<< THE ONLY WAY OUT OF THE ZONE >>>
            v
  +=========================================================+
  |  DETERMINISTIC ZONE      ordinary code                  |
  |                          no model in the loop           |
  |                                                         |
  |   [1] OUTPUT VALIDATION                                 |
  |       schema / type / range / business rules            |
  |                    |                                    |
  |                    v                                    |
  |   [2] ACTION ALLOWLIST + LEAST-PRIVILEGE CREDS          |
  |       default deny                                      |
  |         |            |                |                 |
  |   not allowed    reversible      irreversible           |
  |         |            |                |                 |
  |         v            |                v                 |
  |     [REJECT]         |     [3] HUMAN APPROVAL           |
  |         |            |          money out               |
  |         |            |          data out                |
  |         |            |          deletions               |
  |         |            |          external comms          |
  |         |            |                |                 |
  |         |            |         approved / denied        |
  |         |            v                v                 |
  |         |        [4] E X E C U T E                      |
  +=========================================================+
            |            |                |
            v            v                v
  +---------------------------------------------------------+
  |  [5] IMMUTABLE AUDIT LOG               append-only      |
  |      every input, proposal, decision, approver, action  |
  +---------------------------------------------------------+
            |
            v
   systems of record (CRM, email, payments, files)

Layer 1 — Least-privilege credentials. The model gets no credentials. The code around it holds them, each scoped to the narrowest thing the job requires. An invoice-reading assistant needs read access to one mailbox folder — not the mail account, not the Drive, not the CRM. Most deployments fail this because the quick-start path asks for broad access in one click and everyone clicks it. It's exactly what the Five Eyes guidance leads with. Ask: which permissions does this need, on which accounts, and what's the smallest set that still works?

Layer 2 — Output validation. Model output is a proposal, not a command. Before anything acts on it, ordinary code checks it: right shape, right field types, amount in range, recipient exists in our records, passes our business rules. Free-form text handed straight to an execution step is the most common architectural defect we see. This cuts both ways: a document or API response coming back from a retrieval step is data to constrain, never an instruction to follow, whatever label the plumbing attached. Ask: what checks run between the model's output and the action — are they code, or another prompt? If it's another prompt, you've added a second component with the same failure mode, not a control.

Layer 3 — Action allowlists. The system enumerates what it's permitted to do; everything else is denied by default. Not "the model decides what tool to call" but "these eleven operations exist, with these parameter constraints, and nothing outside that list is reachable." You can't anticipate every bad action; you can enumerate the good ones. Ask: what is the complete list of actions this system can take, and what happens to a request that isn't on it?

Layer 4 — Approval gates on irreversible operations. Sort every action into reversible and irreversible. Reversible things — drafting, tagging, summarizing, scheduling internally — run unattended, because the cost of being wrong is that someone fixes it. Irreversible things get a human: money leaving the business, data leaving the business, deletions, any communication going to an outside party under your name. This is the junior-employee rule, in code. It's also the layer people are most tempted to remove because it slows things down — a business decision, not a technical one, and it should be made deliberately and in writing.

Layer 5 — Immutable audit trail. Append-only, tamper-evident, complete: what came in, what the model proposed, what the validator decided, who approved, what executed. You need it to answer "what happened" after an incident and — more usefully — to spot patterns before one. A log the system can rewrite isn't an audit trail.

Notice what's missing from all five layers: any requirement that the model behave. Every control sits outside it. That's the destyling result applied — anything you enforce by asking the model nicely in a system prompt is enforced by style, in a component that responds to style.

What to ask anyone building AI for you

If you're evaluating a vendor, a consultancy, or an internal team, these six questions tell you most of what you need to know. There are right answers, and you'll hear the difference.

  1. What credentials does this system hold, and what is the blast radius if the model is fully compromised? A good answer is specific and small. A bad answer is "it's secure" or a description of the model's safety training.
  2. What sits between the model's output and any real action? You want typed schemas and validation code. You don't want a prompt that says "only output valid JSON."
  3. What is the complete list of actions it can take? If nobody can produce the list, there's no allowlist, and the real answer is "anything its credentials permit."
  4. Which operations require a human, and who decided that split? Money, data egress, deletion, and outbound external comms belong on the human side unless there's a documented reason otherwise.
  5. Show me the audit log for a real transaction. Not a description of it. The record.
  6. What happens when — not if — an injection succeeds? You want a containment story, not a denial. Anyone who says their system can't be prompt-injected is either not following the research or hoping you aren't.

None of this makes the model trustworthy. That was never available. It makes the system trustworthy — the thing you actually needed, and a solved category of engineering problem with fifty years of precedent behind it.

The uncomfortable version of the ICML finding is that AI security can't be bought as a feature. The useful version is that it can be built, with the same discipline you'd apply to a payments integration or a medical records system. Ordinary, careful, unglamorous work.

That's the work we do. If you have AI running in your business and can't answer question 1, that's a reasonable place to start a conversation.


Share this article

LinkedIn X Email

← Back to all articles