Before you switch models.
The LLM model upgrade checklist: 24 questions behind the Upgrade Sprint, for upgrading or migrating the model behind an AI agent. Work top to bottom. Anything you can’t check is a risk to settle before the change, not after it.
It covers a single model change. For an agent’s first release, start with the launch questions.
The launch readiness checklist1. Inventory
A model change touches more than the model name. List everything that depends on it before anything moves.
- Every workflow, prompt, tool description and setting tied to the current model is listed
- Each workflow has a named owner and the decision a regression would affect
- Provider retirement and deprecation dates for the current model are on the calendar
2. Baseline
You cannot judge a new model against a number nobody measured. Score the current one first.
- Quality is scored on your own cases, counting every attempt, not just the successful ones
- Speed is measured where the user waits, including the time between tool calls
- Cost per accepted task is recorded, including retries and the people who review the output
3. Candidates
Compare the models and settings you would actually run, on the same cases, the same way.
- Two or three candidate models and their settings are chosen, with the reason for each
- Every candidate runs the same cases, with the same tools and the same scoring
- Judges and graders are checked against human decisions before their scores are trusted
4. Behavior changes
Newer models change how they work, not only how well. Look for the changes a single score hides.
- Response length and token use per task are compared, not assumed
- Tool-call patterns are checked: more calls, different arguments, new retries
- Refusals, format drift and instruction-following on your hardest cases are reviewed
5. Control
A better answer is not worth a wider blast radius. Check what the new model is able to do, not only what it says.
- Every action the agent may take is tested for allowed and denied cases on the candidate
- Prompt-injection and data-access checks run on the candidate, not only on the incumbent
- Approval steps and human takeover still trigger where they should
6. Regression gate
Decide in advance what blocks the change, so the decision does not depend on who is in the room.
- The cases that block a release are named, and one critical failure blocks it
- The gate runs in your pipeline, on every model, prompt or setting change
- Thresholds for quality, speed, cost and control are written down before the results arrive
7. Rollout
Ship the change in stages, with a tested way back.
- The rollout is staged, with the scorecard watched at each step
- Rollback to the last accepted configuration is tested before the first stage
- Actions a rollback cannot undo have a reconciliation plan
8. Decision and cleanup
Record the decision, then remove what the new model no longer needs.
- The decision is written down: adopt, hold or switch, with the evidence behind it
- Prompts, retries and workarounds the new model no longer needs are removed
- The accepted configuration is recorded, so the next change starts from a known baseline
Want this run against your agent?
The Upgrade Sprint runs these questions on your own tasks and ends in a written decision: adopt, hold or switch, with the regression gate left in your pipeline.
Explore the Upgrade Sprint