Skip to content

Decision guide / Reliability

What to measure before you switch LLM models

A model change touches quality, speed, cost, control and drift. What to baseline, compare and gate before you upgrade or migrate the model behind an agent.

A practical framework for an engineering decision.

The decision

A new model can be better on a benchmark and worse for your workflow. Score the current model on your own cases first, run the candidates the same way, check the checkers, and let a regression gate decide what ships. Then remove whatever the new model no longer needs.

  • Baseline the current model on your own cases before comparing candidates.
  • Compare all five measures, not only quality.
  • Write the gate and the decision down before the rollout starts.

Why a model change breaks agents that worked yesterday

Providers release new models every few months and retire old ones on fixed dates. A newer model changes how the agent works, not only how well: responses get longer, tool calls multiply, the wait between steps moves, and instructions that held for the old model stop holding.

Swapping the model name is the easy part. Knowing that the new model is better for your workflow, what it costs, and what it is now able to do is the work. That knowledge comes from your own cases, scored the same way before and after.

Baseline the current model first

  • Score quality on your own cases, counting every attempt, including retries and abandoned runs.
  • Measure speed where the user waits, including the time between tool calls.
  • Record the cost per accepted task, including the people who review the output.
  • List every prompt, tool description and setting tied to the current model.

Without the baseline, a comparison only tells you which candidate looks better in a demo. With it, every candidate is judged against the configuration your users rely on today.

Compare all five measures

What to compare across candidate models
MeasureWhat to compareWhat a single score hides
QualityTask success on your cases, over every attemptWins on easy cases that mask losses on the hard ones
SpeedTime to a useful result for the person waitingLonger outputs and extra tool calls that add seconds
CostTotal cost per accepted taskA cheaper price per token spent on more tokens
ControlAllowed and denied actions, prompt-injection and data-access checksA more capable model that reaches further than intended
DriftChanges against the last accepted configurationFormat, refusal and instruction changes on edge cases

Check the checker

Automated judges and graders change with the models too. Before trusting a judge's score on a new candidate, compare its verdicts with human decisions on a sample, and plant known failures to confirm the checks still catch them. A green result is a claim about the test, not proof about the agent.

Gate the change and write the decision down

Decide what blocks the change before the results arrive: the named cases, the thresholds for each measure, and the rule that one critical failure stops the release. Wire that gate into the pipeline so it runs on every model, prompt or setting change.

Then record the decision as adopt, hold or switch, with the evidence behind it. A hold is a useful outcome: it means the next release is judged against the same cases, and nobody has to rediscover why the last candidate was rejected.

Roll out with a way back, then clean up

  • Stage the rollout and watch the scorecard at each stage.
  • Test the rollback to the last accepted configuration before the first stage.
  • Plan reconciliation for actions a rollback cannot undo.
  • Remove the prompts, retries and workarounds the new model no longer needs.

Method and scope

The decision frameworks and synthetic examples are TeqEngine’s editorial guidance.

Have a system like this in front of you?

We can scope a platform engagement directly, or begin with an architecture review when the next decision needs more evidence.