Work in progress · updated 30 Sep 2026 · notes, not conclusions
Letting an agent act alone
When should a security lead let an AI agent act without asking first, and how do they take that permission back?
For exampleThe agent has cut 200 infected laptops off the network correctly this quarter, and wrongly once. Should it do the next one alone, at 3 am, without waking anyone?
- Ask every timewhere every action starts
- Ask if unsureafter a good track record
- Act, then tella person reviews after
- Act alonefor this one action only
If its accuracy on that action drops, it steps back down and says why. Set separately for each action: isolate a laptop, block a site, reset a password.
Why it matters
Agents are fast, and asking a person every time makes them slow again. But a wrong automatic action stops someone's work, or worse. Today that trade-off is usually a hidden setting, not a decision anyone can see or review.
What I'm trying
- A track record per action: how often the agent was right, what it cost when it wasn't, and who fixed it.
- A switch per action type (not one global "autonomy" dial), with the reason for each choice written down.
- Automatic step-down: if the agent's accuracy on an action drops, it goes back to asking, and says why.
What would show it works
Security leads who use the screen set permissions they don't later reverse, and can explain each one in a sentence.
Next step
Sketches of the track-record screen, then three conversations with SOC leads.