# What to measure before you switch LLM models

A model change touches quality, speed, cost, control and drift. What to baseline, compare and gate before you upgrade or migrate the model behind an agent.

Author: TeqEngine
Published: 2026-10-05
Canonical: https://teqengine.ai/insights/what-to-measure-before-switching-models

## The decision

A new model can be better on a benchmark and worse for your workflow. Score the current model on your own cases first, run the candidates the same way, check the checkers, and let a regression gate decide what ships. Then remove whatever the new model no longer needs.

- Baseline the current model on your own cases before comparing candidates.
- Compare all five measures, not only quality.
- Write the gate and the decision down before the rollout starts.

## Why a model change breaks agents that worked yesterday

Providers release new models every few months and retire old ones on fixed dates. A newer model changes how the agent works, not only how well: responses get longer, tool calls multiply, the wait between steps moves, and instructions that held for the old model stop holding.

Swapping the model name is the easy part. Knowing that the new model is better for your workflow, what it costs, and what it is now able to do is the work. That knowledge comes from your own cases, scored the same way before and after.

## Baseline the current model first

- Score quality on your own cases, counting every attempt, including retries and abandoned runs.
- Measure speed where the user waits, including the time between tool calls.
- Record the cost per accepted task, including the people who review the output.
- List every prompt, tool description and setting tied to the current model.

Without the baseline, a comparison only tells you which candidate looks better in a demo. With it, every candidate is judged against the configuration your users rely on today.

## Compare all five measures

**What to compare across candidate models**

| Measure | What to compare | What a single score hides |
| --- | --- | --- |
| Quality | Task success on your cases, over every attempt | Wins on easy cases that mask losses on the hard ones |
| Speed | Time to a useful result for the person waiting | Longer outputs and extra tool calls that add seconds |
| Cost | Total cost per accepted task | A cheaper price per token spent on more tokens |
| Control | Allowed and denied actions, prompt-injection and data-access checks | A more capable model that reaches further than intended |
| Drift | Changes against the last accepted configuration | Format, refusal and instruction changes on edge cases |

## Check the checker

Automated judges and graders change with the models too. Before trusting a judge's score on a new candidate, compare its verdicts with human decisions on a sample, and plant known failures to confirm the checks still catch them. A green result is a claim about the test, not proof about the agent.

[Calibrate the judge](https://teqengine.ai/insights/llm-as-a-judge). How to check an LLM judge against human decisions before automating the release decision.

## Gate the change and write the decision down

Decide what blocks the change before the results arrive: the named cases, the thresholds for each measure, and the rule that one critical failure stops the release. Wire that gate into the pipeline so it runs on every model, prompt or setting change.

Then record the decision as adopt, hold or switch, with the evidence behind it. A hold is a useful outcome: it means the next release is judged against the same cases, and nobody has to rediscover why the last candidate was rejected.

## Roll out with a way back, then clean up

- Stage the rollout and watch the scorecard at each stage.
- Test the rollback to the last accepted configuration before the first stage.
- Plan reconciliation for actions a rollback cannot undo.
- Remove the prompts, retries and workarounds the new model no longer needs.

[The model upgrade checklist](https://teqengine.ai/resources/llm-model-upgrade-checklist). Twenty-four questions for upgrading or migrating the model behind an agent, from inventory to cleanup.

[The Upgrade Sprint](https://teqengine.ai/services/llm-model-migration). Run the comparison on your own tasks in two to three weeks and leave the regression gate in your pipeline.

## Method and scope

The decision frameworks and synthetic examples are TeqEngine’s editorial guidance.

## Continue reading

- [Model upgrade checklist](https://teqengine.ai/resources/llm-model-upgrade-checklist)
- [LLM model routing](https://teqengine.ai/insights/llm-model-routing)
- [Upgrade Sprint](https://teqengine.ai/services/llm-model-migration)
