Guide

AI agent kill switch: how to stop or pause an agent in production

An AI agent kill switch is a control outside the agent’s prompt and reasoning loop that lets a person, or an automatic rule, stop or pause the agent’s next actions immediately, without shutting down or redeploying the application it runs in.

Last updated:

Why production agents need one

An agent calls a model, reads the answer and decides what to do next, often in a loop. When something goes wrong, it keeps going: a retry loop burns tokens, a bad prompt reaches every user, a provider degrades, or a tool call does something it should not. Developers have written about single agents running up tens of thousands of dollars before anyone noticed (one example on DEV). A redeploy takes minutes and needs an engineer; a kill switch takes seconds and can be pressed by whoever is on call.

The controls a real kill switch includes

  • Pause and resume: stop new model calls, keep the app running, continue later.
  • Stop: the same effect as pause, used when the agent should not come back without a decision.
  • Model swap: move an agent to another model (cheaper, or from a provider that is healthy) without a deploy.
  • Automatic pause on thresholds: a daily spend cap, an error rate or a spend anomaly pauses the agent without waiting for a human.
  • Scopes and approvals: decide which commands an app accepts, and require an owner or admin to approve risky changes such as a new prompt.
  • Audit trail: who sent which command, when, and whether it was delivered.

What makes it trustworthy

  • Out of band. The switch must not depend on the model agreeing to stop. It lives in your code, around the model call.
  • Authenticated. Commands are signed (HMAC-SHA256 with a timestamp and a single-use nonce), so nobody else can pause your agent or replay an old command.
  • Honest about timing. A pause prevents the next model call; it does not cancel a request already in flight. Plan for that.
  • Safe when the control plane is down. Reporting should fail open, so an outage of the control plane never takes your agents down with it. The flip side: a switch cannot be pressed while the control plane is unreachable, so keep a local override (an environment variable) as well.

How to add one with Agent Control Panel

Agent Control Panel (ACP) delivers commands (pause, resume, stop, restart and model updates) as signed webhooks to your app. The SDK keeps the agent’s state locally; while an agent is paused or stopped, the wrapped call throws AgentPausedError without calling your model. Alert rules can attempt an automatic pause on a daily spend cap, error rate, spend anomaly or lost heartbeat, within the app’s control scope. Every command is recorded.

Node.js

import { createClient, AgentPausedError } from "agent-control-panel";

const acp = createClient({
  url: process.env.ACP_URL,
  apiKey: process.env.ACP_API_KEY,
  agentId: process.env.ACP_AGENT_ID,
  webhookSecret: process.env.ACP_WEBHOOK_SECRET,
});
acp.start();
app.post("/api/acp/webhook", acp.expressWebhookHandler()); // before express.json()

try {
  const answer = await acp.wrapAgentRun(() => llm.chat(messages), {
    modelUsed: acp.getEffectiveModel("gpt-4o-mini") ?? "gpt-4o-mini",
  });
} catch (err) {
  if (err instanceof AgentPausedError) { /* paused from the dashboard: skip the call */ }
  else throw err;
}

Python

from agent_control_panel import AgentPausedError, create_client

acp = create_client(url=ACP_URL, api_key=ACP_API_KEY,
                    agent_id=ACP_AGENT_ID, webhook_secret=ACP_WEBHOOK_SECRET)
acp.start()

try:
    answer = acp.wrap_agent_run(lambda: my_agent(question),
                                model_used=acp.get_effective_model("gpt-4o-mini"))
except AgentPausedError:
    answer = "This assistant is paused. Please try again later."

Then press Pause on the agent in the dashboard, or create an alert rule with automatic pause. Start with the app in monitor scope, confirm heartbeats, then allow the control scope. Details: controlling agents, control scopes, alerts and automatic pause.

Frequently asked questions

What is an AI agent kill switch?

A control outside the agent’s prompt and reasoning that stops or pauses its next actions immediately, without shutting down or redeploying the application it runs in.

Does pausing cancel a model request that is already running?

No. A pause prevents the next model call. A request already in flight finishes, so design long-running tools with that in mind.

Can an agent be paused automatically?

Yes. In ACP, alert rules on a daily spend cap, error rate, spend anomaly or lost heartbeat can attempt an automatic pause, within the app’s control scope.

What happens to my app if the control panel is down?

The SDK fails open: your agents keep running and reporting resumes later. For that reason, also keep a local way to disable an agent, such as an environment variable.

Do I need to change my model provider or send my API keys?

No. Your agent keeps calling its model provider directly with your own keys; ACP only receives telemetry and sends signed commands.

Next: connect an existing app or request early access.