Most businesses running AI agents can answer one question only vaguely: what can this agent reach, and what happens if someone tricks it? An agent that only drafts text is a risk to your reputation. An agent that reads your inbox, edits your files and sends messages is a risk to your operations, your customers and your money. This guide covers how to secure the second kind, in the order we would tackle it with a client.
What AI Agent Security Actually Covers
Chatbot security is mostly about what the model says. Agent security is about what the model does. An agent is a language model connected to tools, such as an email account, a CRM, a file store, a payment API or a database, and it decides which tool to call and with what inputs. That creates four distinct risks, and most incidents involve more than one of them:
- Prompt injection. Text inside an email, support ticket, web page or document tells the agent to do something its owner never asked for, such as forward a file or change a bank detail. The agent cannot reliably tell instructions from content.
- Excessive permissions. The agent runs on a person's login or a shared admin account, so any mistake or hijack has the full reach of that account. The OWASP Top 10 for LLM Applications lists both prompt injection and excessive agency among the top risks for this reason.
- Irreversible actions. Deletes, payments, outbound email and contract changes often cannot be undone with a click. The damage lands before anyone reviews the output.
- Silent drift. The agent keeps running after its job has changed, or its behavior shifts after a model or prompt update, and nobody notices until a customer does.
The controls below map to these four risks. None of them requires an enterprise security team. They require deciding, in writing, what each agent is for.
Start With an Agent Inventory
You cannot secure agents you have not listed. Most small businesses discover at least one agent they forgot about during this step, often a trial that was never switched off or a browser extension that quietly acts on someone's behalf. For every agent, record these fields in one spreadsheet:
- Name, business purpose and a named human owner.
- Every system it can read, and every system it can write to.
- Which identity it uses: a personal login, a shared account, or its own account.
- Which actions need approval, and who gives it.
- The date it was last reviewed.
An inventory that takes an afternoon to build closes the largest gap most businesses have. The inventory also feeds every later step, because you cannot set a permission for an agent you have not written down. For the shadow-agent side of this problem, where employees connect personal AI tools to work accounts, see wizeb.com/blog/shadow-ai-agents-dots-muse-2026.
Give Every Agent Its Own Identity
The fastest way to get an agent running is to connect it to someone's account. The fastest way to create a breach is to leave it there. A dedicated identity lets you scope access, audit actions and revoke one agent without disrupting staff. The core rules:
- One identity per agent. Never a person's login and never a shared admin account.
- Read access by default. Grant write access only where the job requires it, and only for the specific action.
- Narrow the data surface. A maintenance agent needs a maintenance inbox, not the office manager's full mailbox.
- Make credentials revocable. You should be able to cut one agent off in minutes without touching the others.
Permissions also accumulate. Each agent is usually provisioned with whatever access made it work on day one, and few businesses schedule a review six months later. Put a recurring access review on the calendar. We go deeper on this in wizeb.com/blog/ai-agent-permissions-least-privilege-2026 and on separate accounts for agents in wizeb.com/blog/ai-agent-delegated-spending-accounts-2026.
Separate Reading Untrusted Content From Taking Action
Prompt injection enters through data the agent reads. An inbound email, a web page the agent browses or a support ticket written by a stranger can all carry instructions. You cannot fully solve this by asking the model to ignore them, because the model is designed to follow text that looks like instructions. The defense is structural.
In practice, that means three design choices. First, an agent that reads untrusted content should not also hold broad write power in the same step. A common pattern is a reader that summarizes and classifies, and a separate action step that only accepts a short, validated set of commands. Second, tool arguments that come from untrusted text should be checked against an allowlist before they run, such as an approved recipient domain or a known account number format. Third, outbound actions to new destinations should require confirmation. Each of these turns a successful injection from a silent breach into a blocked request.
Put Limits in the System, Not the Prompt
A line in the system prompt that says "never spend more than $500" is a suggestion. A payment rule that rejects any transaction above $500 is a control. The same logic applies to every limit that matters:
- Spending caps per transaction, per day and per month, enforced by the payment system or the API gateway.
- Recipient allowlists for outbound email and messages.
- Row and record thresholds, such as refusing to bulk-update more than 50 contacts without approval.
- Rate limits so a runaway loop stops itself after a fixed number of calls.
Write the prompt as guidance and the tool configuration as the actual boundary. If the prompt and the configuration disagree, the configuration wins, which is the behavior you want.
Decide in Advance What Needs a Human
Approval rules work best as three tiers, defined before the agent goes live:
- 1Automatic. Reversible, low-value and inside a known pattern. Example: tagging a support ticket or filing a receipt into a folder the agent owns.
- 2Queued for review. Visible but not yet final. Example: a drafted customer reply that a person sends, or a vendor invoice held for a same-day check.
- 3Approval required. Irreversible, high-value or new. Example: a payment, a deletion, a contract change or an email to a recipient outside the approved list.
The tiers matter more than the tooling. Agents that ask for approval on everything get ignored within a week, and agents that never ask get blamed when something goes wrong. The cutoff points should be reviewed after the first month of real use. For a worked pattern, read wizeb.com/blog/human-in-the-loop-ai-agents-checkin-pattern-2026.
Make Destructive Actions Reversible
Assume every agent will eventually make a wrong, destructive change, and design so the blast radius is small and the recovery is fast. Three practices cover most of it:
- Snapshot before writes. Keep a point-in-time copy of any folder or table the agent can modify, retained long enough to matter.
- Soft delete by default. Agents move files to a recycle area or mark records inactive. Permanent deletion stays with a person.
- Test the restore path. A backup you have never restored is a hope, not a plan. Run one restore drill per quarter.
Pair these with a kill switch. You should be able to stop one agent, or all of them, within minutes, and the person who can do it should be named in the runbook. Most businesses discover during their first drill that nobody knows where the agent's credentials are stored. Wizeb covers the operational side in wizeb.com/blog/ai-agent-kill-switch-incident-response-2026 and file recovery in wizeb.com/blog/ai-agent-file-deletion-backup-safety-2026.
Log What Each Agent Did, and Why
Logs answer the questions that come after an incident: what did the agent read, what did it change, who approved it, and did it behave differently last week? A useful log entry records the trigger (which email or ticket started the run), each tool call with its arguments, the outcome and any approval. Store logs somewhere the agent cannot edit.
Logs alone do not catch drift. Set a baseline for each agent, such as the typical number of records changed per day, the share of tasks escalated for approval and the error rate. Review a weekly sample of actions and alert when a metric moves well outside its range. Silent failures are usually visible in these numbers weeks before anyone complains. The monitoring side is covered in wizeb.com/blog/ai-agent-observability-silent-failures-2026.
Patch the Agent Stack Like Any Production Software
Agent frameworks, connectors and plugins are software, and they carry vulnerabilities like any other software. The difference is that a flaw in an agent layer often hands an attacker the same tool access the agent has. Track the exact framework and connector versions for every agent, name an owner for advisories, and agree on a same-day patch process for critical issues. Disclosed exploits against agent tooling have been seen within hours of publication, so a monthly patch cycle is too slow for those.
The same discipline applies to the SaaS vendors you connect. For each one, ask where data is stored, how long it is retained, whether it is used to train models, how access is revoked, and what happens to your data if you leave. The lessons from a recent vendor breach are in wizeb.com/blog/ai-vendor-security-tldv-breach-lessons-2026.
What This Looks Like in a Real Small Business
Picture a twelve-person accounting firm, which we will call the firm for this example. It runs an agent that drafts replies to client emails and files receipts. Set up in a hurry, the agent was connected to a partner's full mailbox and to a shared drive with admin rights. An audit found that the agent could read three years of client correspondence, and that a single malicious invoice email with hidden instructions could have pulled files out through an outbound message.
The fix took about two days of setup. The agent got a dedicated drafting mailbox with no send rights to external addresses, and a person sends every reply. Its drive access was narrowed to the receipts folder. It has no payment authority. Any outbound message to a new domain now goes to the partner for approval. The agent's job did not change. What changed is that the worst successful attack now reaches a receipts folder and a drafts queue, instead of the firm's client history. This example is hypothetical, but the pattern is one we see often.
Keep Compliance in the Same Inventory
If an agent touches personal data, customer records or regulated decisions, you may have obligations under data protection law and, in some cases, AI-specific rules. The inventory, logs and approval tiers above are also the evidence you would need to show what the agent does and who oversees it. Whether a particular obligation applies depends on where you operate and what the agent does, so confirm it with qualified counsel. For the EU timeline and what it means for agent deployments, see wizeb.com/blog/eu-ai-act-august-2026-ai-agent-compliance.
A 30-Day Plan to Secure Your Agents
- 1Week one: build the inventory. List every agent, its owner, its systems and its identity. Switch off anything nobody owns.
- 2Week two: separate identities and cut permissions. Move each agent to its own account and remove write access it does not need.
- 3Week three: set limits and approval tiers. Add spending caps, recipient allowlists and the three-tier approval rules, enforced in the systems themselves.
- 4Week four: prove it works. Test the kill switch, restore one snapshot, review a week of logs and set the drift baseline for each agent.
After month one, the agents you run will be fewer, narrower and easier to explain. That is the goal. Security for agents is less about one perfect control and more about making every agent's reach small, visible and reversible.
How Wizeb Builds Secure Agents
Wizeb builds agents with the controls already in place. Every agent we deliver ships with its own identity, task-scoped permissions, limits enforced by the system, approval tiers agreed with the owner, logging and a documented kill switch. We review existing agents too, mapping what each one can actually reach and rebuilding the access around the job it does. Our AI agent work is described at wizeb.com/services/ai-agents, and the scheduled, rule-based parts of many workflows are handled through wizeb.com/services/automation.
Review one agent this month
Pick the agent with the broadest access in your business. Wizeb will map its real permissions, show you the one write action it should not be able to take, and scope a fix. Start at wizeb.com/services/ai-agents.
