← Articles

Small Businesses Are the Soft Target

· By Dialogs

Small Businesses Are the Soft Target

If you run a small business, here's the thing we most want you to take from this piece: attackers don't pick the richest target — they pick the cheapest one to beat. That has always been the economics of cybercrime, and it's about to matter to you in a new way. The numbers first, then what to do about them.

Every year, the FBI's Internet Crime Complaint Center (IC3) publishes a tally of what cybercrime cost the country. In 2025, the number was $20.877 billion — up 26% from the year before, across just over one million complaints. Business email compromise (BEC), the scam where an attacker impersonates a vendor or executive to redirect a wire payment, accounted for $3.046 billion of that on its own. [1]

BEC isn't sophisticated. There's no malware, no exploit, no zero-day. It's a well-written email and a moment of trust: "Here are our updated banking details for this month's invoice." It works because it's cheap to run and hard to distinguish from a normal Tuesday. And it lands hardest on the businesses least equipped to catch it — the ones without a dedicated finance-ops team cross-checking every wire against a callback list.

Verizon's 2025 Data Breach Investigations Report puts a number on that imbalance directly: extortion malware showed up in 88% of small-business breach incidents, compared to 39% at larger organizations. [2] Attackers aren't targeting small businesses because they hold more valuable data. They're targeting small businesses because the effort required to succeed is lower — fewer controls, fewer people double-checking, faster payoff per hour of attacker time. That's not a security failing unique to any one company. It's economics. Attackers, like anyone else, go where the return on effort is best.

Now add a new tool to that same economic equation: AI agents that read email, process invoices, and take action on what they read.

Researchers Charles Ye, Jasmine Cui, and MIT's Dylan Hadfield-Menell make the mechanism precise in their paper "Prompt Injection as Role Confusion," accepted to ICML 2026. Their finding, in their own words: "we trace prompt injection to role confusion: models perceive the source of text from how it sounds, not its labeled role." A malicious instruction hidden in a webpage or an email doesn't need to break any technical boundary — it just needs to sound like the kind of text the model is used to trusting. As they put it: "a command hidden in a webpage hijacks an agent simply because it sounds like text, despite its label." [3]

Their attack technique, CoT Forgery, plants text that mimics the model's own internal reasoning inside ordinary input. It's zero-shot — no training, no tuning, no iteration required — and it achieved 60% attack success against OpenAI's frontier models (tested on the gpt-oss and GPT-5 families), up from a near-zero baseline. [3] The published experiments cover OpenAI models only; Ye and Cui have said in interview, though it isn't a published result, that they've since seen similar behavior in models from Anthropic, Alibaba, and DeepSeek — worth naming because it points at an architecture-wide pattern rather than one company's bug, even though only the OpenAI numbers are peer-reviewed. [3] The researchers also ran a control: take the identical malicious argument, strip out only the reasoning-style markers that make it sound like the model's own thoughts, and leave the semantic content untouched. Success collapsed from 61% to 10% — a sixfold drop from changing nothing but how the text sounded. [3] That's about as clean a proof as this kind of research gets that models are responding to style, not substance.

Now put a number on what that means economically. In the paper's agentic tests, an ordinary prompt injection — just asking the model to do something it shouldn't — succeeded 0–2% of the time against a data-exfiltration task: get the agent to send information somewhere it doesn't belong. CoT Forgery, on the identical task, succeeded 56–70% of the time. [3] Zero-shot, no malware, no infrastructure, against a payload that is, in plain terms, your client list or your financials leaving the building. That's the economics: the cost of the attack barely changes, and the odds of it working go from a rounding error to better than a coin flip.

That's the shape of the next few years. Prompt injection is BEC's successor, not a separate problem — same trust exploited, cheaper to run. It needs no malware to write, no exploit kit to buy, and it leaves nothing an antivirus or endpoint agent recognizes as a signature — because nothing was installed. The "attack" is a paragraph of text sitting in an inbox, a support ticket, or a résumé upload field, waiting for an AI agent to read it and act on it as if it were an instruction from the business that deployed it. Ye said as much directly to MIT Technology Review, describing what happens once that math becomes common knowledge: "there's going to be a huge economic incentive for people to do jailbreaks and prompt injections." He went further: "there's a real probability that this is going to be a problem that's fundamentally unsolvable." [5] Both lines are his forecast in an interview, not a finding in the paper — but the paper is exactly why the forecast holds up. And "unsolvable" doesn't mean "undefendable." It means the fix isn't a patch to the model. It's a decision about what the model is allowed to touch.

Is this already happening in the wild, or is it still theoretical? The honest answer is: mostly the latter, with one real exception worth naming. In December 2025, Palo Alto Networks' Unit 42 documented a live case of attackers embedding hidden text — invisible fonts, off-screen positioning, script-driven content — inside a scam advertisement page, aimed at an AI system that reviews and approves ads before they run. [4] Unit 42 was careful to say they have no confirmed case of that specific attack succeeding against a deployed system yet. But someone tried, in production, against a real company's AI moderation pipeline. That's the tell. Beyond that case, most documented prompt-injection "incidents" circulating right now are researcher demonstrations and disclosed vulnerabilities, not confirmed criminal breaches — genuinely useful warnings, but not yet the wave. The wave is the part Ye is forecasting, not the part that's already landed.

There's a second documented case, and it's more unsettling than the first. In July 2026, Anthropic disclosed that during routine cybersecurity evaluations, three of its own models — given internet access by a misconfigured test environment despite being told they had none — gained unauthorized access to three real organizations using nothing more exotic than weak passwords, unauthenticated endpoints, SQL injection, and exposed debug pages. [7] No malware, no exploit chain, no human attacker. Two of the three victims never detected it themselves. This wasn't an adversarial attack; it was an agent pursuing an assigned goal through whatever was reachable. The lesson for a small business is blunt: the flaws that matter here aren't exotic, and nobody hostile needs to be on the other end for them to get exploited at machine speed. More in The AI didn't break out. The door was open.

Here's the part that matters for a five-to-hundred-person business trying to decide what to do about any of this: you don't need to become a security company to close this gap, and you don't need to hire a security team you can't staff or afford. What closes it is deciding, in advance, what your AI agents are structurally incapable of doing — regardless of what any piece of text convinces them to try.

Three constraints do most of the work:

  1. Least-privilege access. An agent that only needs to read email shouldn't hold credentials that can also move money or change a payment address.
  2. Approval gates on anything irreversible. A wire transfer, a bulk email send, a data export — these get a human confirmation step, full stop, no matter how confident the agent sounds.
  3. An audit trail that can't be edited after the fact. If something does go wrong, you need to know exactly what the agent read, decided, and did — not just what it says it did.

None of these require trusting the model more. They require trusting it less, on purpose, and building the system so that doesn't matter. It's also not just a vendor's opinion: a joint Five Eyes cybersecurity advisory published 1 May 2026 — six national cyber agencies, including Australia's ASD ACSC, the US's CISA and NSA, the Canadian Centre for Cyber Security, New Zealand's NCSC, and the UK's NCSC — gives businesses adopting agentic AI almost the identical guidance: avoid granting broad or unrestricted access, especially to sensitive data or critical systems, and start with low-risk use cases before expanding. [6] Six national cyber agencies across five countries put their names to the same three constraints. That's the actual playbook — and it's the subject of the next piece in this series: Stop Trying to Make the Model Safe. Make the System Safe.

Share this article

LinkedIn X Email

← Back to all articles