The question nobody asks before turning on an AI agent
When a small business connects an AI tool — a chatbot, an email assistant, an agent that drafts invoices or triages support tickets — the setup conversation is almost always about capability. Can it read the inbox? Can it pull customer records? Can it draft and send? The answer, for the sake of getting the thing working by Friday, is usually yes to all three.
The question that doesn't get asked is: what happens when it's wrong?
Not "if." When. Every LLM-based system, including the best ones available in 2026, can be pushed off the rails — through a malformed input, a bad prompt, or a document quietly containing tampered instructions. That's not a hypothetical; it's a structural property of how these models process text. Researchers at MIT and collaborators demonstrated it directly on frontier OpenAI models: a model doesn't judge whether a piece of text is a trusted instruction or hijacked content by the label on it — it judges by how the text sounds. A command hidden in a scraped webpage, a vendor API response, or a note left in a shared CRM record can hijack an agent simply because it reads like an authoritative instruction, regardless of which channel it arrived through. The researchers argue the underlying mechanism is architectural, not specific to one vendor's models — though the published results themselves cover OpenAI's model family, not every provider.
The number that should get your attention: in the same study's data-exfiltration test — an agent tricked into leaking information it was trusted to handle — a standard injected instruction mostly failed, succeeding 0–2% of the time on most models tested (one model spiked to 26%). Dress that same instruction up to sound like the model's own reasoning, and it worked 56–70% of the time, across every model tested. The content of the attack didn't change. Only how it sounded did. That's the whole argument for auditing what an agent can reach, instead of trusting that it will correctly filter what it's told: the filtering approach is the one that failed 98–100% of the time it was tested against a dressed-up attack.
That means every integration your agent reads from — not just the inbox it's watching — is a potential instruction channel. (Full technical breakdown in the companion piece, "There is no inside voice.")
Given that the model will eventually do something you didn't intend, the only question that matters operationally is: what could it reach when it did?
That's blast radius. And for most small businesses, the honest answer is "more than anyone checked." Not because anyone was careless — because the default setup path for nearly every AI integration is "grant broad access once, move on." Nobody sits down deliberately to overprovision an AI agent. It happens by default, one OAuth "Allow" click at a time, during onboarding, when the goal was to get the demo working.
This is a checklist to go find out what you actually granted. It takes about 20 minutes, needs no security background, and ends with a number you can act on.
Why this is a permissions problem, not a "is the AI good enough" problem
Every AI vendor will tell you their model is safe, well-tested, aligned, whatever term is current. Take that as roughly true and it still doesn't answer the question you care about. Safety testing describes how the model behaves in the scenarios someone thought to test. Permissions describe the actual ceiling on damage when it encounters one nobody tested.
An AI agent's authority — what accounts it can touch, what actions it can take without asking, what it can send outside your organization — is your security policy for that system, whether you wrote it down or not. If you never reviewed it, you have a policy anyway; you just don't know what it says.
Where to actually look
Most small businesses have three or four places this access lives. Each takes a few minutes.
1. Google Workspace — API controls If your business runs on Google Workspace, go to the Admin console: Menu → Security → Access and data control → API controls, then Manage App Access. This lists every internal and third-party app with OAuth access to your organization's data, and shows the access level assigned to each: Trusted (access to everything, including restricted services), Limited (unrestricted services only), Specific Google data (scoped to exact permissions an admin chose), or Blocked. Look for any AI tool — chatbot, agent, browser extension, "AI assistant" plugin — sitting at Trusted when it only needs read access to one shared drive.
2. Microsoft 365 / Entra — Enterprise applications If you're on Microsoft 365, sign in to the Microsoft Entra admin center (entra.microsoft.com) as a Global Administrator and go to Identity → Applications → Enterprise apps → Consent and permissions → User consent settings. This shows whether your organization allows any user to grant an app access on their own (a common default), and lets you restrict it to apps from verified publishers only, or to admin-approved apps exclusively. From Enterprise apps, click into any individual application to see the exact permissions it was granted — mailbox read/write, files, contacts, directory data — and revoke what's unused.
3. Your CRM's integrations or connected-apps panel Every CRM (GoHighLevel, HubSpot, Salesforce, and the rest) has a settings page listing connected apps and their granted scopes — usually under Settings → Integrations, Connected Apps, or API Keys, naming varies by vendor. Check what any AI or automation tool can do here: read contacts, read and write contacts, trigger workflows, export data, send on your behalf. Export is the one people miss — an agent that can read your entire contact list and email it somewhere is functionally the same risk as a breach, whether or not anything malicious ever happens.
4. The AI vendor's own admin settings Whatever chatbot, copilot, or agent platform you use — its own admin panel is the fourth place. Look for anything labeled permissions, scopes, connected accounts, or data access. Many agent platforms let you scope a connection down to specific mailboxes, folders, or objects instead of "all of Gmail" or "all of Salesforce." If that option exists and you skipped it, that's the fastest fix on this list.
Bonus question: What can your AI actually reach? In July 2026, Anthropic disclosed that its own evaluation sandboxes — built by one of the most safety-conscious labs in the industry — turned out to have live internet access nobody intended, despite the eval prompts explicitly stating there was none. If that lab can misjudge what an environment can reach, assume your own integrations reach further than you think until you've checked egress paths, not just the permission grants above. More on what happened next: The AI didn't break out. The door was open.
The Dialogs AI Blast Radius Score
Score each category 0–2. Be honest — this is diagnostic, not a report card.
| Category | 0 points | 1 point | 2 points |
|---|---|---|---|
| Credential scope | Agent has org-wide or "Trusted" access | Scoped to a department or shared resource | Scoped to the specific mailbox/folder/object it needs |
| Irreversible-action gating | Sends, deletes, payments, or external emails fire with no review | Some actions gated, others not | Every irreversible action (send, pay, delete, share externally) requires a human click |
| Data-egress paths | Agent can email, export, or post data to any external address/destination | Egress limited to a known allowlist | No unreviewed path for data to leave your systems |
| Logging & auditability | No record of what the agent did or why | Basic activity log, not reviewed | Full action log, reviewed on a schedule |
| Revocation readiness | Nobody knows how to cut the agent's access quickly | Access can be revoked, but it'd take research under pressure | You could revoke this agent's access in under 5 minutes, today, and you know exactly how |
8–10: Solid. Keep the habit of re-checking after every integration change. 4–7: Typical, and that's the problem — typical is where the exposure lives. Fix credential scope and irreversible-action gating first; they carry the most risk per point. 0–3: The agent currently has more reach than anyone has verified is safe. This isn't a five-alarm situation, but it's the first thing to fix this week, not next quarter.
(The Blast Radius Score is a Dialogs framework, not a published industry standard — built from the same principles behind OWASP's guidance on excessive agency in AI systems and the joint government guidance below. Use it as a diagnostic starting point, not a compliance benchmark.)
The three constraints that matter most
This isn't just a Dialogs opinion — it's Five Eyes joint guidance. In May 2026, six national cybersecurity agencies across five countries — Australia's ASD ACSC, the US's CISA and NSA, the Canadian Centre for Cyber Security, New Zealand's NCSC, and the UK's NCSC — jointly published the same lead recommendation this audit is built around: don't grant agentic AI broad or unrestricted access, particularly to sensitive data and critical systems. The same guidance urges organizations to deploy incrementally and limit agents to low-risk tasks at first, enforce strict privilege controls, monitor continuously, keep a human in the loop on oversight, and fold agentic AI into the security framework they already run rather than treating it as a separate experiment outside normal governance. That last cluster maps almost one-to-one onto the five categories in the score above — this rubric tracks what six governments told every organization to check, not just a Dialogs preference.
If you only fix three things, fix these:
- Scope every credential to the smallest thing that works. Not "read/write all mail," but "read this one shared inbox." Most platforms support this; most setups skip it.
- Put a human between the agent and anything irreversible. Sending money, deleting records, emailing outside the company, changing a customer's account — one approval click costs a few seconds and eliminates the failure mode that actually matters.
- Make sure you can pull the plug in minutes, not hours. Know where the off switch is before you need it. If you had to go find out how to revoke this agent's access right now, how long would it take?
None of this requires a security team or a large budget. It requires twenty minutes and a willingness to look.
Where this goes next
A permissions audit tells you what your AI can do. It doesn't tell you what to do about the gap once you've found it — how to build the approval gates, the audit trail, and the allowlists so the system stays safe as it scales past one agent and one integration. That's the subject of the pillar piece in this series: "Stop trying to make the model safe. Make the system safe."