# AI vendor evaluation scorecard

Compare two engineering proposals with fixed criteria, evidence confidence and separate critical requirements. Keep the worksheet private and export it locally.

Author: TeqEngine
Published: 2026-09-11
Canonical: https://teqengine.ai/insights/ai-vendor-scorecard

## The decision

Use the same criteria and evidence standard for each candidate. Keep unassessed areas visible and critical requirements outside the weighted score. The worksheet supports a procurement discussion; it does not certify a vendor or choose a winner.

- Unknown evidence is not silently treated as a pass.
- A failed critical requirement remains visible at any score.
- Names, scores and evidence notes stay in the browser; downloads are local.

## Compare the evidence

Interactive worksheet: https://teqengine.ai/insights/ai-vendor-scorecard#worksheet

The method and complete static example appear in this article. Inputs stay in the browser and are not submitted.

## The published comparison rubric

The interactive worksheet uses the six weights in TeqEngine’s vendor-selection guide. They are an editorial starting point for a software engagement, not a validated industry ranking. If the engagement needs different priorities, establish those weights in the procurement worksheet before scoring proposals.

**Fixed weights in this worksheet**

| Criterion | Weight | Evidence to inspect |
| --- | --- | --- |
| Workflow and architecture | 20% | A task-specific design with clear failure paths. |
| Evaluation and reliability | 20% | Representative cases, failed traces and a release decision. |
| Security and data boundaries | 20% | Data flow and a demonstrated denied action. |
| Delivery and handoff | 15% | Build, deployment, acceptance and recovery artifacts. |
| Relevant engineering evidence | 15% | Permitted work or a technical validation with accurate attribution. |
| Economics and scope control | 10% | Assumptions, exclusions and operating costs. |

Rubric version 2 separates requirement fit from evidence strength. A score of 0 means the requirement is not met; 1 means major gaps; 2 means partially met; 3 means met; and 4 means exceeded with a relevant benefit. Leave unassessed criteria blank. The separate evidence field records whether support is unavailable, asserted, documented or observed. Direct inspection can establish either a poor fit or a strong one; it does not automatically earn a high score.

Record the shared evaluation scope, then the requirement, document or test reference, version/date and finding for each criterion. The worksheet requires documented or observed evidence and a reference note for completion. It cannot verify those entries. Version 1 scored evidence strength, so its totals should not be compared directly with this requirement-fit rubric.

## Keep missing information and failed requirements visible

Before a criterion is assessed, it has no score. The worksheet displays a possible total range rather than pretending the missing work is a zero or a pass. For example, a score of 3 on a 20% criterion contributes 15 points; with the other 80% unassessed, the possible total is 15–95 out of 100.

Critical requirements are separate pass, fail or unknown decisions covering data/access, ownership/handoff and release/operation. Adapt the detailed requirements to the engagement and record the basis for each decision. A failed requirement remains flagged regardless of points. An unknown requirement, missing decision basis or unsupported score leaves the assessment incomplete.

NIST’s Secure Software Development Framework provides a useful vocabulary for supplier discussions. Referencing it does not certify the vendor or replace the organization’s own contract and security review.

Sources: [NIST: Secure Software Development Framework, SP 800-218](https://csrc.nist.gov/pubs/sp/800/218/final).

## Static example: equal scores, different unresolved work

**Synthetic comparison**

| Criterion | Candidate A | Candidate B |
| --- | --- | --- |
| Architecture / 20% | 3 | 4 |
| Evaluation / 20% | 3 | 3 |
| Security / 20% | 2 | 3 |
| Handoff / 15% | 3 | 2 |
| Engineering evidence / 15% | 4 | 3 |
| Economics / 10% | 3 | 2 |
| Weighted total | 73.75 / 100 | 73.75 / 100 |

Suppose A’s evidence is recorded and all critical requirements pass, while B’s ownership and handoff requirement is still unknown. A’s worksheet can be complete; B’s remains incomplete. That does not establish that A is the better vendor. It identifies the unresolved decision that the next discussion must address.

The calculation divides each 0–4 fit score by four, multiplies by its weight and adds the contributions. Reference notes travel with the local export; supporting documents stay in the organization’s normal review system. The possible score range reflects unassessed criteria, not a statistical confidence interval or uncertainty about whether evidence is true.

## Use the scorecard with a comparable technical exercise

- Give candidates the same task boundary, constraints and completion criteria.
- Ask for one successful path and one important failure path.
- Record which evidence is observed, documented, asserted or unavailable.
- Identify the next validation that would resolve a material unknown.
- Compare scope, operating responsibilities and total economics as well as the score.

Confidentiality can limit named references. Record what cannot be verified and evaluate permitted evidence without converting a disclosure limit into either proof or a negative finding. A small reproducible technical validation can be useful when one integration or permission assumption dominates the purchase.

[Read the vendor-selection guide](https://teqengine.ai/insights/choosing-ai-agent-development-company). Use the full guide for evidence requests, proposal comparison and a bounded validation scope.

## Export and privacy

Download review record produces a readable text file; Download comparison JSON produces a structured file. Both include the rubric version, scope, weights, scores, evidence notes, critical decisions and unresolved work. Downloads are generated locally. Inputs stay only in this page’s memory; save a download before leaving or reloading. Clear both candidates resets the worksheet and offers Undo clear until another edit.

“Assessment complete” means the required entries, references and gate decisions are filled. It does not mean the evidence is independently verified, procurement approval, legal clearance or an automatic recommendation. Keep the final selection and its rationale in the organization’s decision record.

## Sources and scope

Technical references inform the cited statements. The decision frameworks and synthetic examples are TeqEngine’s editorial guidance.

- [NIST: Secure Software Development Framework, SP 800-218](https://csrc.nist.gov/pubs/sp/800/218/final)

## Continue reading

- [Technical diligence guide](https://teqengine.ai/insights/ai-technical-due-diligence)
- [Investment calculator](https://teqengine.ai/insights/ai-roi-calculator)
