Model Swap and Tool Calls: What Changes Under an Agent

Vendors retire models on schedules you do not set, and the replacement may speak a different tool-call dialect. What the primary docs say differs, and a contract test to run before any swap.

By Harinderpal Hanspal on September 2026. Updated October 2026

A model swap under an agent can change how tool calls are encoded, identified and forced, so an agent can start answering in prose with no error. OpenAI returns arguments as a JSON string with a call id, Anthropic as an object, and Ollama documents neither an id nor a tool_choice field.

Swapping the model under an agent can change the tool-call format and the agent goes quietSketch of an agent loop on the right expecting tool call arguments as a string with a call id, and a swapped-in model on the left that sends arguments as an object with no id. A dashed arrow between them is crossed out in red, labelled dropped. Below, a note says the agent answers in prose with no error.new model: argumentsas an object, no callidagent loop expects astring and an idcall droppedsame agent, same prompt, different modelresult: the agent answers in proseno error, no tool run
Swapping the model under an agent can change the tool-call format and the agent goes quiet

Anthropic notified developers on 5 June 2026 that Claude Opus 4.1 would retire, and retired it on 5 August (Anthropic, model deprecations). Its stated policy is at least 60 days' notice for publicly released models. OpenAI shut down its Assistants API on 26 August 2026, a year after announcing it, and told developers to move to the Responses API (OpenAI, deprecations). Both are as of 6 October 2026. Anthropic's page adds that applications relying on its models "may need occasional updates".

So a model under your agent will be swapped. The question for a plant or an OEM is what else changes when it is.

Four places the tool-call contract differs

An agent loop depends on a small contract: the request carries tool definitions, the response carries a call, the result goes back tied to that call. Each piece is documented differently.

Argument encoding. OpenAI's function-calling guide shows arguments as a JSON-encoded string, for example "arguments": "{\"location\":\"Paris, France\"}". Anthropic's tool-use documentation shows "input": { "location": "San Francisco, CA" }, an object. Ollama's tool-calling page also shows arguments as an object (OpenAI, Anthropic, Ollama). Code that calls a JSON parser on every argument field breaks on an object, and code that reads fields directly breaks on a string.

Call identity. OpenAI requires each tool output to reference the call's call_id. Anthropic's example carries an id beginning toolu_. The Ollama page we read shows an index and a name, with no id field. A loop that pairs results to calls by id has nothing to pair on.

Forcing a call. OpenAI's tool_choice accepts auto, required, a named function or allowed_tools. Anthropic's has auto, any, tool and none, and on its page any and tool return a 400 error on Claude Opus 5.5, Sonnet 5.5, Fable 5.1 and Mythos 5.1, which should use strict tool use instead. The Ollama page we read documents no tool_choice parameter at all. A rule like "the agent must call a tool before it answers" is a provider feature, not a universal one.

Parsing on a self-hosted engine. vLLM states that different model families need their own tool-call parser and often their own chat template, and lists parsers such as llama3_json for Llama 3.1 and 3.2, llama4_pythonic for Llama 4, hermes, mistral and qwen3_xml (vLLM). Its page puts the formats at JSON, Python syntax or XML tags. With the wrong parser, a call the model did emit can reach the agent as plain text. That is our inference from the parser requirement, not a vLLM statement.

Request parameters move too. Anthropic's page says temperature, top_p and top_k return a 400 on Claude Opus 4.7 and later when set to non-default values. That one fails loudly. The dropped tool call does not.

Why it fails quietly

A refusal is easy to catch. A missing tool call is not an error to the loop: the model produced text, the loop saw no call, the run ended. The output looks like a model that declined. A team debugging it tunes prompts, when the cause is a definition that never reached the model, an id that was never returned or a parser that did not match.

A contract test to run before any swap

Send one fixed tool-calling request to the candidate through the same path production uses, and assert four things:

  1. The response contains a tool call, not prose.
  2. Its arguments parse into the type your loop expects.
  3. It carries an id your result message can reference, or your code supplies one.
  4. A request your agent relies on, such as a forced call, is either honored or refused with an explicit error.

It takes minutes. It tells you the plumbing works. It does not tell you the model picks the right tool or fills the values correctly, and we found no published, quantified comparison of agent failure rates before and after a swap.

What to ask a vendor that swaps models for you

The governance case for putting these checks below the agent, where a model change cannot remove them, is in our paper on governing agents in production.

Drawn from Anthropic's model deprecations and tool-use documentation, OpenAI's deprecations and function-calling guide, Ollama's tool-calling docs and vLLM's tool-calling docs, all read on 6 October 2026. No study of agent failure rates after a model swap was found, so none is cited.

Related notes

Related paper: Governing agents in production: what to ask before an agent acts