Back to library
AI & Tooling

Before you let an agent near production: a blast-radius checklist

The scary agent stories all share one shape: a tool that could do more than the task needed, pointed at a system that couldn't say no. Here's how to shrink what an agent can break before you hand it real work.

MC
Mike Curry
Founder · learn.curry.io · Sep 27, 2026
8 min
Before you let an agent near production: a blast-radius checklistAI & Tooling

Every cautionary agent story you've read — the dropped table, the force-push, the credential that ended up in a public repo — has the same shape. It's rarely a story about a model being malicious or even especially dumb. It's a story about a tool that could do far more than the task required, pointed at a system with no way to refuse. The fix isn't a smarter model or a longer prompt. It's the same discipline infrastructure people learned the hard way years ago: decide the blast radius before you light the fuse. Here's the checklist I run before an agent touches anything that matters.

Why the prompt is the wrong place to put your safety

The instinct, when you're nervous, is to write more instructions. "Never delete anything. Always ask before running migrations. Don't touch the prod config." Those lines are worth writing, but they are advice, not enforcement. A model under pressure to finish a task will occasionally reason its way around advice, and a model that's confused about which environment it's in won't even know it's breaking a rule. Anything you actually cannot afford to have happen needs to be impossible, not discouraged.

That means safety lives in the harness — the permissions, the sandbox, the credentials, the checks that run before a change lands — not in the conversation. The prompt shapes intent. The harness bounds consequences. You need both, and most people have only invested in the first.

Step one: enumerate what it can reach

Before you run anything, write down every system the agent's process can touch. Not what you told it to touch — what it *can* touch. The shell it runs in, the environment variables in that shell, the cloud credentials on the machine, the git remotes it can push to, the databases the local config points at, the API keys sitting in a dotfile from a project you finished last year. Most engineers who do this exercise honestly are unsettled by the list. A developer laptop is a skeleton key to half the company, and an agent running on it inherits every tooth.

This is the single highest-value ten minutes in the whole checklist, because it converts a vague worry into a concrete inventory you can start cutting.

Step two: cut the inventory down to the task

Now remove everything the task doesn't need. Run the agent in a container or a fresh user account that has the repo and the toolchain and nothing else. Give it scoped, short-lived credentials that can only reach the one service it's working on — and if that service has a sandbox or staging tier, point it there. Use a git identity that can push to feature branches but not to main, and turn on branch protection so that even a determined force-push bounces. If the task is read-only, make the token read-only. Least privilege is an old idea; agents just make the cost of ignoring it arrive faster.

  • Isolated runtime. Container, VM, or dedicated user — never your daily-driver shell with your daily-driver keys.
  • Scoped credentials. One token, one service, one environment, short expiry. Rotate it when the task ends.
  • Guarded remotes. Feature branches only, protected main, no deploy hooks reachable from the agent's account.
  • Dry-run by default. Anything destructive runs in preview mode first — `--dry-run`, `terraform plan`, a transaction you can roll back.
  • Real data stays out. Fixtures or scrubbed copies. If the agent needs production-shaped data, it gets production-shaped, not production.

Step three: make the dangerous actions loud

Some tasks legitimately need a risky capability — a migration, a bulk update, a deploy. You can't remove those; you can make them impossible to do quietly. Put an approval gate on the specific commands that matter, so the agent has to stop and show you exactly what it's about to run. Wrap destructive CLI tools in a thin script that prints the plan and waits for a human keystroke. Log every command to somewhere the agent can't edit. The goal isn't to slow everything down; it's to make sure the one action that could ruin your week is the one action that can't happen while you're getting coffee.

Key takeaway
Instructions shape what an agent tries to do. Permissions decide what it can do. Spend your effort on the second: isolate the runtime, scope the credentials, protect the branches, and gate the handful of actions that can't be undone.

Step four: rehearse the failure

Here's the step that separates people who've been burned from people who are about to be. Before the real run, deliberately ask the agent to do something it shouldn't be able to do, and watch what happens. Tell it to push to main. Tell it to drop a table in the staging database. Tell it to read a secret it shouldn't have. If any of those succeed, your harness has a hole, and you found it on a Tuesday afternoon instead of during an incident. If they all fail cleanly, you've earned some real confidence — the kind that comes from evidence rather than from a paragraph in a system prompt.

Do this again whenever the setup changes. A new tool, a new credential, a new machine — each one is a chance for the blast radius to quietly grow back.

Step five: keep the recovery path shorter than the damage path

Even with everything above, assume something will eventually go wrong, and make sure undoing it is cheaper than doing it was. That means backups you've actually restored from, not backups you believe exist. Migrations with a tested down path. Deploys you can roll back with one command. Feature flags so a bad change can be turned off without a redeploy. None of this is agent-specific — it's just that agents produce changes faster than humans do, so the gap between "we could recover from this" and "we have recovered from this before" gets exposed sooner.

What this buys you

The payoff isn't just fewer disasters. Once the blast radius is small and known, you can stop hovering. You can let the agent run longer tasks, try more aggressive refactors, and work while you're doing something else — because the worst case is a reverted branch, not a page at 2am. Every team I've watched get real leverage from agents got there the same way: not by trusting the model more, but by building a room where trust wasn't required.

That's the skill worth learning right now. Not prompting — anyone can prompt. Building the harness that makes an agent safe to delegate to is the part that looks like engineering, and it's the part that will still be valuable when the models have changed three more times.

Want a coach in your corner?

Book a 1:1 call — we'll map your next step and pressure-test your plan. Group courses coming soon.

Book a call
Enjoyed this guide?

Get one like it in your inbox each week

Practical guides and roadmaps for the AI shift — free, no hype, written by the same people.

Join 3,700+ readers. Unsubscribe anytime.