Skip to content

Harness Review / two to three weeks / fixed price

Know how good your agent is, and what to fix first.

In two to three weeks: a map of where your agent's time and money go, a scorecard on quality, speed, cost, control and drift, a model and settings comparison on your own tasks, and a prioritized fix plan. The tooling stays in your repository.

Sample deliverable excerpt · synthetic example

Release decision brief

Observation
A mutating tool can act beyond the calling user’s scope.
Decision
Hold that action until authorization is checked at the tool boundary.
Evidence needed
Allowed and denied test cases, tenant isolation checks, and an auditable decision trace.
Owner & next step
Name the implementation owner, define the acceptance test, and review the result.

When to bring us in

For agents already in front of users.

Quality varies and nobody can say why

Answers change across users, documents or edge cases, and the traces don’t explain it.

Slow or expensive, with no clear cause

Latency or spend keeps rising and nobody knows which layer is responsible.

A release decision is coming

A launch, a new customer or a model change needs evidence before it goes ahead.

Choose the inspection scope

The agents you ship, or the agents that build your software.

Review the agent your customers use, the coding-agent workflow your engineers use to build it, or both.

The agents you ship

Runtime architecture, tools and permissions, evaluations, retrieval, cost and latency, and recovery, measured before the next release.

See the launch questions

The agents that build your software

Repository legibility, agent permissions, verification gates, review load and architecture drift, with a prioritized harness gap list.

See the codebase questions

The handover

Seven outputs. One prioritized plan.

Every finding comes with its evidence, an owner and the next decision, so you can choose what to fix, change or stop.

  1. Harness Map

    Every layer of one real task, timed, from the interface to the model.

  2. Scorecard

    Quality, speed, cost, control and drift, scored on your own cases.

  3. Model and settings comparison

    Your tasks run across candidate models and settings, side by side.

  4. Control findings

    What the agent may do, and where that limit is enforced.

  5. Fix plan

    Ranked by impact, each item with its evidence and owner.

  6. Test runner and trace setup

    Left in your repository, ready for the next change.

  7. Readout

    For engineering and leadership, with the decisions needed next.

The process

Agree the question. Measure the system. Decide.

01 / Agree access and scope

The decision comes first.

Confirm the workflow, the cases that decide it, the stakeholders and the access we need.

02 / Measure and test

Follow every layer of a real task.

Trace the work from the interface to the model, score it on your own cases and test the assumptions the decision depends on.

03 / Read out and hand over

A plan your team can act on.

Walk the decision-makers through the findings and hand over the evidence, the tooling and the fix plan, ready for your engineers or a forward-deployed team.

Before we start

Who is the review for?

Teams with an agent already in front of users, or a coding-agent workflow to improve, and a decision ahead: a release, a fix or a model change. For an acquisition or investment, see technical and AI diligence.

How is the review scoped?

We agree the questions, cases, access and deliverables before work begins. The scope and the price are fixed.

What do you need from us?

A decision owner, the relevant system context, and access to documentation, code, traces or test environments.

We already have evals.

Good. We start from them, add speed, cost, control and drift, and compare models and settings on the same cases.

Do we need TeqEngine to build the fixes?

Your team can work from the fix plan directly. When we do the fix work and it starts within 60 days, the review fee is credited toward it.

Can we go straight to a build?

Yes. When the outcome and ownership are clear, we scope a forward-deployed build directly.

Start a conversation

What are you building—or deciding?

Tell us what you’re building, what is getting in the way, and the outcome you need.