How to let an AI agent work on your accounts without letting it send, spend or delete by mistake
Five controls to set before an agent touches your inbox, files or payment accounts, with the order to set them in.
Give the agent its own limited access, require a person to approve anything that sends, pays, posts or deletes, log every action, test the workflow on sample data against a written pass/fail check, and keep a way to switch it off. Set the controls in that order, before the agent touches real email, files or money.
An agent is software that takes actions on your behalf, so a mistake is an action, not a bad sentence. A wrong draft costs you a minute. A wrong send, payment or deletion can cost far more, and some cannot be reversed. These five controls keep the cheap mistakes cheap.
The five controls
1. Give the agent its own access
Do not sign the agent in as you. Create a separate account, a shared folder or a scoped token that reaches only what the task needs. Start read-only. If the agent can only read your inbox, the worst result is a wrong summary.
2. Put a person in front of irreversible actions
The agent prepares the action and a person approves it. A draft email waits in a folder until someone presses send. A purchase waits in a queue. A deletion waits in a list of proposed deletions. The agent does the typing and the person owns the decision.
3. Log every action
Keep a record of what the agent read, what it changed and what it proposed, with times. A log lets you answer what happened last Tuesday without guessing, and it shows you where the agent keeps getting things wrong.
4. Write a pass/fail test first
Before building the workflow, write down what correct looks like in a form you can check. For inbox triage: messages from these ten senders land in these folders, nothing is sent, nothing is deleted. Run it on sample data. A workflow without a written check will pass because nobody defined failing.
5. Keep a stop switch
You need two. One stops the agent process. The other revokes the access it holds, such as the token or the account share. The second one matters more, because it works even when the process does not respond.
What to gate and what to leave alone
Gating everything makes the agent useless and teaches people to click approve without reading. Gate by consequence.
| Action | Approval | Why |
|---|---|---|
| Read files, mail or records | No | Nothing changes |
| Draft text, summaries, reports | No | A person reads it before it goes anywhere |
| Send a message or share a file | Yes | It leaves your control |
| Spend or move money | Yes | Often irreversible |
| Post publicly | Yes | Others see it at once |
| Delete or overwrite data | Yes | Recovery may not exist |
| Change permissions or settings | Yes | It widens what the agent can do |
Instructions are not enforcement
Writing ask before sending in the agent instructions helps, and the model usually follows it. It is still a request. The real gate is a setting the harness enforces or an account that cannot do the thing at all.
Claude Code asks before it runs most commands and edits files, unless you change its permission mode. Codex has sandbox and approval settings, such as a read-only sandbox, that you choose when you run it. Use those settings, and scope the accounts so that even a model that ignores its instructions cannot send, spend or delete. Markdown skills that say ask first are useful reminders and should not be the only barrier.
The same holds for local models. A small model follows long instructions less reliably than a large one, so enforcement matters more, not less. The agent comparison lists the harnesses and what each one controls.
Mistakes that defeat the controls
- Signing the agent in as you. Every action then carries your full permissions and your name, and revoking the agent means changing your own password.
- Approving without reading. If every request looks the same, people click through. Show the exact message, amount or file list in the approval request.
- Keeping the log where the agent can edit it. Write logs somewhere the agent can only append to, or somewhere it cannot reach.
- Testing on real data first. A mistake on sample data is a bug report. The same mistake on a live inbox is an incident.
- Never testing the stop switch. Revoke the token once on purpose, so you know it works and how long it takes.
The order to set them
- Create the agent's own account or token, read-only.
- Turn on the harness approval settings and confirm that a gated action actually stops and asks.
- Turn on logging and check that it records a test action.
- Write the pass/fail check and run it on sample data.
- Test the stop switch by revoking the token and watching the agent fail.
- Only then move to real data, one task at a time.
I hold software to this standard because my day job is as a firefighter-paramedic, where a checklist that fails at the wrong moment hurts someone. The Agent Workflow package at $4,000 builds one workflow this way: a written scope and pass/fail test first, a person approving anything sent or changed, and a log of every action. About has more on how I work, and the contact page is the place to start a conversation.
Questions
What should an AI agent never do without approval?
Anything you cannot undo or that leaves your control: sending a message, spending money, posting publicly, deleting data, changing permissions and sharing files. Reading and drafting are safe to leave ungated.
Is a rule in the prompt enough to keep an agent safe?
No. An instruction such as ask before sending is guidance the model usually follows. Enforce the rule with the harness permission settings and with account access that cannot send or delete in the first place.
Can an agent have read-only access to my inbox?
Yes, and it is the best place to start. Give the agent a read-only token or a separate account, let it draft replies as text, and have a person send them until the drafts have earned trust.
How do I stop an agent that is doing the wrong thing?
Keep two switches within reach: stop the agent process, and revoke the access it holds. Revoking the token or the account share stops it even if the process keeps running.
What is a pass/fail test for an agent workflow?
A written check that says what correct output looks like, run on sample data before real data. For inbox triage it might say every message from a named sender lands in the right folder and nothing is sent.