Skip to content

Decisions you own, not code you write

Most writing about AI reliability is aimed at the engineer doing the work. These pages are aimed at whoever has to decide whether the work happens, what it should cost, and how to tell it worked. Each one is a decision you own, with the framework I would use and the numbers to sanity-check it.

Engineers: Written to be forwarded. If you are the engineer who found this, send the relevant one up — it is framed for the person who has to approve the work.

Decision 1

What an AI agent audit should cover — and what to refuse to pay for

How to scope, price and sanity-check a reliability engagement before you sign anything.

Read the framework

Short answer

A useful audit produces four artefacts you keep: a failure taxonomy specific to your agent, a cost and latency profile, an eval suite wired to your real traffic, and a prioritised fix list with effort estimates. If a proposal does not name deliverables you own afterwards, you are buying a slide deck.

Decision 2

Fourteen questions to ask your AI team before launch

With the answer that should reassure you, and the answer that should worry you.

Read the framework

Short answer

Ask about four things, in this order: can you prove it still works after a change, can you explain a failure after the fact, is spend bounded, and does someone own it. The single most useful question is "if you changed a prompt today, what would tell you something broke?" — if the answer is "a customer would tell us", you are not ready to launch, regardless of how well the demo goes..

Decision 3

Build or buy: LLM observability

The decision costs more in engineering time than in licence fees, which is the part most comparisons omit.

Read the framework

Short answer

Buy, unless you have a hard data-residency requirement or an existing observability platform with a team who owns it — in which case extend that. Building from scratch is almost never right: the initial version takes days, and the maintenance you did not budget for is what actually costs you.

Decision 4

Five agent metrics worth reporting upward

And four that look impressive while telling you nothing.

Read the framework

Short answer

Five: cost per conversation, eval pass rate, failure rate by named category, p95 latency per operation, and containment or escalation rate. Those five answer every question an executive asks — is it working, is it getting better, what does it cost, and is it getting worse — and each one is a leading indicator rather than a post-mortem.

Decision 5

How to tell whether your AI team needs outside help

Including the cases where they clearly do not, which are more common than consultants tend to admit.

Read the framework

Short answer

Handle it internally when the gaps are clear and someone has the time — most reliability work is a few days per gap and needs no specialist. Bring in outside help when the gaps are clear but capacity is not, when you need it done by someone who has done it before, or when you need independent evidence for a customer or regulator.

Want the answer for your own system?

The scorecard scores your agent across reliability, cost and observability in four minutes and ranks the gaps by risk. It prints to a PDF — which tends to move a planning conversation faster than an opinion does.

Run the scorecard

Or just ask

Describe the decision you are weighing and I will tell you what I would do — including when the answer is to handle it internally or do nothing yet.