Quality varies and nobody can say why
Answers change across users, documents or edge cases, and the traces don’t explain it.
Harness Review / two to three weeks / fixed price
In two to three weeks: a map of where your agent's time and money go, a scorecard on quality, speed, cost, control and drift, a model and settings comparison on your own tasks, and a prioritized fix plan. The tooling stays in your repository.
Sample deliverable excerpt · synthetic example
When to bring us in
Answers change across users, documents or edge cases, and the traces don’t explain it.
Latency or spend keeps rising and nobody knows which layer is responsible.
A launch, a new customer or a model change needs evidence before it goes ahead.
Choose the inspection scope
Review the agent your customers use, the coding-agent workflow your engineers use to build it, or both.
Runtime architecture, tools and permissions, evaluations, retrieval, cost and latency, and recovery, measured before the next release.
See the launch questionsRepository legibility, agent permissions, verification gates, review load and architecture drift, with a prioritized harness gap list.
See the codebase questionsThe handover
Every finding comes with its evidence, an owner and the next decision, so you can choose what to fix, change or stop.
Every layer of one real task, timed, from the interface to the model.
Quality, speed, cost, control and drift, scored on your own cases.
Your tasks run across candidate models and settings, side by side.
What the agent may do, and where that limit is enforced.
Ranked by impact, each item with its evidence and owner.
Left in your repository, ready for the next change.
For engineering and leadership, with the decisions needed next.
The process
Confirm the workflow, the cases that decide it, the stakeholders and the access we need.
Trace the work from the interface to the model, score it on your own cases and test the assumptions the decision depends on.
Walk the decision-makers through the findings and hand over the evidence, the tooling and the fix plan, ready for your engineers or a forward-deployed team.
Teams with an agent already in front of users, or a coding-agent workflow to improve, and a decision ahead: a release, a fix or a model change. For an acquisition or investment, see technical and AI diligence.
We agree the questions, cases, access and deliverables before work begins. The scope and the price are fixed.
A decision owner, the relevant system context, and access to documentation, code, traces or test environments.
Good. We start from them, add speed, cost, control and drift, and compare models and settings on the same cases.
Your team can work from the fix plan directly. When we do the fix work and it starts within 60 days, the review fee is credited toward it.
Yes. When the outcome and ownership are clear, we scope a forward-deployed build directly.
Start a conversation
Tell us what you’re building, what is getting in the way, and the outcome you need.