How to switch an AI agent’s model without redeploying
To switch an AI agent’s model without redeploying, the model name must come from somewhere you can change at run time, not from the code you shipped. Your agent reads the current model just before each call, and a person (or a rule) changes it from outside the app. The next call uses the new model, and nothing is rebuilt or restarted.
Why teams need to switch models quickly
- A provider is down or rate-limiting you. Every minute on a failing model is a minute of failed answers.
- Cost. An agent turns out to be fine on a cheaper model, or a loop is burning tokens on an expensive one.
- Quality. A new model version answers worse for your use case, and you want the previous one back now.
- Trying something new. You want one agent on a new model while the others stay where they are.
Three ways to do it
| Approach | How fast | Granularity | Trade-off |
|---|---|---|---|
| Environment variable or config file | Minutes: needs a redeploy or restart | Usually the whole service | Simple, but needs an engineer and a deploy at the worst moment |
| AI gateway or proxy (routing and fallbacks) | Automatic on errors, or a config change | Per route | Every request goes through another service, which usually holds your provider keys |
| Control plane command to the agent | Seconds | Per agent | Your code must read the current model before each call (one line) |
They can be combined: a gateway that falls back automatically on provider errors, plus a per-agent switch that a person can use for decisions a rule cannot make, such as cost or quality.
How it works with Agent Control Panel
- An owner or admin chooses a new model for the agent in the dashboard. The app must allow at least the control scope; new apps start in monitor scope, which allows no commands.
- Agent Control Panel (ACP) sends an
update_modelcommand to your app as a signed webhook (HMAC-SHA256, timestamp, single-use nonce). The SDK verifies it and stores the new model. - Your code asks the SDK for the model before each call:
getEffectiveModel(default)in Node.js,get_effective_model(default)in Python. The next call uses the new model. - The SDK reports its model on the next heartbeat (every 30 seconds), and that is how ACP confirms the swap. If the webhook never arrived, for example because the app has no public URL, the heartbeat reply carries the model instead, so the swap still lands within about 30 seconds and survives restarts and extra replicas.
Switching within one provider (say, one OpenAI model to another) needs only the model name. Switching across providers also needs your code to pick the matching client, as in the examples below.
Node.js
import OpenAI from "openai";
import Anthropic from "@anthropic-ai/sdk";
import { acp } from "./acp"; // createClient(...) + acp.start(), see the quickstart
const openai = new OpenAI();
const anthropic = new Anthropic();
export async function answer(question: string, system: string) {
// The model ACP asked for after a swap, otherwise your default.
const model = acp.getEffectiveModel("gpt-4o-mini") ?? "gpt-4o-mini";
return acp.wrapAgentRun(
() =>
model.startsWith("claude")
? anthropic.messages.create({ model, max_tokens: 1024, system, messages: [{ role: "user", content: question }] })
: openai.chat.completions.create({ model, messages: [{ role: "system", content: system }, { role: "user", content: question }] }),
{ userInput: question, modelUsed: model },
);
}
Python
from anthropic import Anthropic
from openai import OpenAI
openai, claude = OpenAI(), Anthropic()
def answer(question: str, system: str):
# The model ACP asked for after a swap, otherwise your default.
model = acp.get_effective_model("gpt-4o-mini")
if model.startswith("claude"):
call = lambda: claude.messages.create(model=model, max_tokens=1024, system=system,
messages=[{"role": "user", "content": question}])
else:
call = lambda: openai.chat.completions.create(model=model, messages=[
{"role": "system", "content": system}, {"role": "user", "content": question}])
return acp.wrap_agent_run(call, user_input=question, model_used=model)
Both SDKs read token counts and the answer text from OpenAI, Anthropic and Google Gemini responses, so the dashboard keeps comparing cost and errors across the switch. Setup: connect an existing app. Command details: controlling agents and control scopes.
What to check before you switch
- The prompt. A prompt tuned for one model can behave differently on another. Try it on a few real inputs first.
- Tool calls and output format. Providers format tool calls and structured output differently. If your agent parses either, test that path.
- Price and limits. Check the new model’s price per token, context window and your rate limits.
- Calls already running. A switch affects the next call; a request already in flight finishes on the old model.
- A way back. Note the previous model so you can switch back the same way.
Frequently asked questions
Can I switch an AI agent’s model without redeploying?
Yes, if your code reads the model name at run time. With Agent Control Panel, an owner or admin sends a model change from the dashboard and the SDK applies it before the agent’s next call.
Can I switch from OpenAI to Anthropic the same way?
Yes. ACP changes the model name; your code picks the matching provider client, for example by checking whether the name starts with “claude”.
Does ACP need my provider API keys to switch models?
No. Your agent keeps calling its providers directly with your own keys. ACP only sends the model name in a signed command.
What if my app has no public webhook URL?
The swap still arrives: the SDK’s heartbeat reply carries the model ACP holds, so the agent adopts it within about 30 seconds.
Who can switch a model?
Owners and admins of the organization, on apps whose control scope allows it. Every command is recorded in the audit log.
Related guides
- What is an AI agent control panel?: How it differs from a control plane, observability, gateways and orchestration frameworks.
- How to monitor AI agents in production: The signals that matter (liveness, errors, cost, quality) and what to do when one moves.
- Agent Control Panel vs Langfuse and LangSmith: Observing agents versus operating them, and when to use both.
- Roll out a new AI agent prompt safely: Versioning, review, approval, testing and rollback for prompts in production.
- Daily spend caps and automatic pause for AI agents: Stop runaway token spend before the invoice: caps, loop detection and safe defaults.
- What is an agent control plane?: How it differs from observability and gateways, and when you need one.
- AI agent kill switch: How to stop or pause an agent in production, and what a trustworthy switch needs.
Next: connect an existing app, read the documentation, or request early access.