Skip to content

Decision support

How to tell whether your AI team needs outside help

Including the cases where they clearly do not, which are more common than consultants tend to admit.

The question

Should we bring someone in, hire, or handle this ourselves?

The short answer

Handle it internally when the gaps are clear and someone has the time — most reliability work is a few days per gap and needs no specialist. Bring in outside help when the gaps are clear but capacity is not, when you need it done by someone who has done it before, or when you need independent evidence for a customer or regulator. Hire rather than contract when agents are core to your product and the work is permanent, and be aware that a contractor is the faster way to find out what the role should even be.

Written for

  • Engineering leaders weighing outside help against internal capacity
  • Founders deciding between a contractor and a hire
  • Managers who suspect a problem but cannot name it

Why the answer is that

The default assumption should be that you do not need outside help. Most agent reliability work is unglamorous engineering — validation at boundaries, bounded resources, tracing, a test suite — and any competent senior engineer can do it once they know which gaps to close. That is why the scorecard on this site is free and gives you the ranked gap list: for a good number of teams that is the entire deliverable they needed.

What outside help actually buys is not knowledge. It is time, sequencing and pattern recognition. Someone who has closed these gaps several times will do it in a week and will not spend three days deciding whether to use an LLM-as-judge. That is worth paying for when the calendar matters, and not otherwise.

There is a third case that is less about capability than credibility: sometimes you need evidence produced by someone who does not report to you. A security reviewer, an enterprise customer or a regulator will weight an independent assessment differently, and that is a legitimate reason to buy one.

And the honest counter-case: if a consultant cannot tell you when not to hire them, their advice on everything else is worth less. I turn down work when the scorecard shows a team is already in good shape, because a referral is worth more to me than one invoice.

Which of these describes you?

Read down the first column. Most teams find themselves in the top two rows.

SituationWhat to doWhy
Gaps are clear and someone has capacityDo it internallyThe work is days per gap and needs no specialist. Use the scorecard's ranked list as your plan.
Gaps are clear, nobody has capacityContract it out, fixed scopeYou are buying calendar time, not expertise. Scope it tightly and keep the artefacts.
Something is wrong and nobody can name itStart with a diagnosticThe most common reason to buy. Diagnosis is a different skill from building, and being close to the system makes it harder.
You need evidence for a customer or regulatorBuy an independent assessmentInternal review is weighted differently by third parties, fairly or not.
Agents are core to the product, permanentlyHire — and contract meanwhileThis is a role, not a project. A contractor helps you define what the role actually is.
The team disagrees about what to fix firstAn outside read helpsSometimes what you need is a tie-break from someone with no stake in the argument.
It works fine and you are being cautiousRun the scorecard, then stopIf you score above 75 you do not need to buy anything. Recheck in six months.
You want someone to own quality long-termHire, do not contractOngoing ownership belongs inside the company. Anyone who tells you otherwise is selling a retainer.

Two of these eight rows say do not hire anyone, and one says hire an employee instead of a contractor. That distribution is roughly what I see in practice.

If you do bring someone in, judge them on this

Applies to me as much as anyone. These are the signals that separate someone who will leave you better off from someone who will leave you a document.

  1. 1They tell you what they will not do, and it is a real list rather than false modesty
  2. 2They quote a fixed price on a fixed scope, and the price does not move with your company's size or funding
  3. 3The deliverables are things you own afterwards — an eval suite, traces, a fix list with estimates — not primarily a report
  4. 4They ask what the problem costs you, not only how it works technically
  5. 5They will name a case where they would tell you not to hire them
  6. 6They do not need production credentials, and can explain what they need instead
  7. 7They can describe something that went wrong on a previous engagement, specifically
  8. 8Their fix list has effort estimates, so you can convert it into a sprint plan without them
  9. 9They leave a written handover your team can act on without further calls

The last two matter most six months later. Work that only makes sense while its author is available was not finished.

Get the answer for your own system in four minutes

The free scorecard produces a score out of 100, a breakdown across reliability, cost and observability, and your gaps ranked by production risk with a fix and effort estimate for each. It prints to a PDF you can take into a planning meeting — which is usually more persuasive than an argument.

Run the scorecard

If you are the engineer who found this page: Written to be forwarded. If you are the engineer who found this, send the relevant one up — it is framed for the person who has to approve the work.

Questions

When should we hire an AI consultant instead of doing it internally?

When the gaps are clear but nobody has capacity, when you need it done fast by someone who has done it before, or when you need independent evidence for a customer or regulator.

Not when the gaps are clear and someone has the time — that work is a few days per gap and any competent senior engineer can do it. Run the free scorecard first; for a good number of teams the ranked gap list is the whole thing they needed.

Should we hire someone full-time or use a contractor?

Hire if agents are core to your product and the work is permanent — ongoing ownership of quality belongs inside the company, and anyone telling you otherwise is selling a retainer.

Contract when the work is a bounded project, or when you are not yet sure what the role should be. A contractor is a fast, cheap way to find out what you are actually hiring for before you write the job description.

How do we know a consultant is any good before hiring them?

Ask what they will not do, and ask them to describe a case where they would tell you not to hire them. Both answers are hard to fake and highly diagnostic.

Then check the deliverables: things you own afterwards rather than a report, and a fix list with effort estimates so your team can act without them. Also ask what went wrong on a previous engagement — specificity in that answer tells you a lot.

What if we just do nothing for now?

Sometimes correct. If the agent is internal-only, low-volume, and a wrong answer is cheap, deferring is a reasonable call.

It stops being reasonable at the point where you cannot answer "what would tell us this broke?" and the agent is customer-facing. The four items on the launch checklist are the minimum I would not defer — they are days of work and they separate a bad week from an incident with a customer's name attached.

Want a second opinion on the decision?

Describe your situation and I will tell you what I would do — including when that is "nothing" or "handle it internally". No charge for that, and it is genuinely how a lot of these conversations end.