Harness engineering: how we build with agents
Inspect a tenant-scoped export: its contract, observed output and verification record.
Read the guideMethods, inspection guides and selected work. How software gets built with agents, and what changes when the product itself acts.
Start here
Inspect a tenant-scoped export: its contract, observed output and verification record.
Read the guideA practical review of behavior, permissions, maintainability and handoff. Free to read.
Read the guideFor engineering leaders
Practical guides to choosing, funding and operating an AI system. Open frameworks, worked examples and no signup gate.
An evidence-based vendor scorecard for CTOs and founders: compare architecture, evaluation, security, delivery and the system you will own.
Model the scope, operating costs and cost per successful task of an AI agent. Includes a transparent calculator and a worksheet for comparing proposals.
A layer-by-layer framework for enterprise AI agent decisions, with a comparison matrix, pilot acceptance criteria and an exit test.
40 items
10 min read
An evidence-based vendor scorecard for CTOs and founders: compare architecture, evaluation, security, delivery and the system you will own.
9 min read
Model the scope, operating costs and cost per successful task of an AI agent. Includes a transparent calculator and a worksheet for comparing proposals.
10 min read
A layer-by-layer framework for enterprise AI agent decisions, with a comparison matrix, pilot acceptance criteria and an exit test.
9 min read
Define task success, test permissions and side effects, evaluate failure slices, and turn agent evaluations into an inspectable release decision.
8 min read
A practical MCP security review covering identity, authorization, tool execution, prompt injection, state handles and evidence for a release.
9 min read
A technical diligence framework for AI software investments and vendor decisions: request the right evidence, test material claims and prioritize findings.
6 min read
Use a workflow when the path is known. Introduce an agent where runtime judgment earns its complexity—and keep consequential actions under explicit control.
6 min read
A reference architecture for identity, state, tool execution, permissions, evaluation and recovery—with a concrete workflow a CTO can inspect.
5 min read
Add AI capabilities through existing identity, domain APIs and data contracts. Plan the migration, rollout and recovery before introducing autonomous writes.
6 min read
Compare MCP and direct APIs for discovery, domain behavior, permissions and execution. Decide when a shared protocol interface is useful.
5 min read
Compare single-agent and multi-agent designs through task evidence, coordination cost, context handoff and responsibility for the final action.
5 min read
Choose retrieval or fine-tuning by diagnosing missing knowledge and behavior. Compare freshness, access, evaluation evidence and operating requirements.
5 min read
Turn a useful demo into an operable product with explicit task outcomes, evaluation, rollout gates, recovery and ownership.
5 min read
Trace task outcomes, tool actions and waiting states to explain failures, control cost and investigate incidents while limiting sensitive data collection.
6 min read
Separate retrieval quality, answer support and permission failures. Use a small worked corpus to understand what each metric establishes—and what it misses.
5 min read
Use model graders for tasks they can judge reliably. Calibrate against reviewed examples, inspect missed failures and keep critical release rules explicit.
5 min read
Choose model routes by task evidence, data policy and operating limits. Treat fallback as a tested execution path—not permission to send any workload anywhere.
5 min read
Locate delays across queues, models, tools and approval. Compare latency changes against task quality, cost and reliable completion.
6 min read
Define who can do what to which resource, derive identity from trusted context and recheck authority when an agent action executes.
5 min read
Review prompt injection defenses at the data and tool boundaries, including scoped access, output validation, approval and adversarial tests.
5 min read
Bind human approval to a specific AI agent action. Review scope, expiry, changed state and the authorization needed before execution.
5 min read
Preserve tenant and source permissions through retrieval, context assembly, caching and citations. Treat revocation and index freshness as part of the design.
5 min read
Design retries for AI tool actions with operation identity, transactional boundaries and reconciliation. Inspect the actual business effect.
5 min read
Assess AI agent ROI through adoption, remaining effort, useful capacity and operating costs. Design a pilot around the assumptions that matter.
5 min read
Estimate workflow value, adoption, remaining effort and payback using your own inputs. A private browser worksheet with formulas and explicit limitations.
4 min read
Calculate a one-sided zero-failure binomial bound and the trials needed for a target. Check the statistical assumptions before using the result.
5 min read
Compare two engineering proposals with fixed criteria, evidence confidence and separate critical requirements. Keep the worksheet private and export it locally.
5 min read
Run a transaction example that commits once, loses the response and recovers the result. Inspect duplicate, conflict and tenant-scope behavior.
5 min read
Run a synthetic retrieval example with tenant and document access checks. Inspect permitted context, denied requests and the boundaries still untested.
4 min read
Inspect a runnable evaluation release gate that separates average task quality from critical failures, with fixtures you can run yourself.
2 min read
Backend AI, API and infrastructure engineering for an agent platform serving a unicorn startup and Fortune 500 companies.
2 min read
An AI-native SaaS platform that replaced manual, document-heavy logistics workflows with automated extraction, routing, and optimization.
3 min read
Vertical scaling buys time, not a future. A practical look at query optimization, rate limiting, caching, read replicas, and when sharding is actually worth it.
3 min read
Transportation Management Systems quietly decide a logistics operation's margins. Where the savings come from: routing, rate shopping, and carrier selection.
2 min read
Running distributed engineering teams without a full-time seat: the communication, time-zone, and process realities that decide whether the work ships.
2 min read
For the right workload, serverless lets a small team ship a scalable API with almost no infrastructure to run. Where it fits and where it does not.
The method, demonstrated
Free
Inspect a tenant-scoped export: its contract, observed output and verification record.
Inspection guide
Free
A practical review of behavior, permissions, maintainability and handoff. Free to read.
Newsletter
Substack
Writing on what it takes to ship agent systems: eval harnesses, tool boundaries, and the rollout discipline they need. Read on Substack.
Field manual
Free
Thirty questions across use-case fit, architecture, data and permissions, evals, observability, and ownership. Work top to bottom before you launch.
Writing on what it takes to ship agent systems: eval harnesses, tool boundaries, observability, and the operating layer around frontier models.