Skip to content

Independent · taking audits and one more retainer

I make AI agents
survive production.

Your agent works in the demo. Then real users arrive, costs triple, and nobody can explain why it failed at 2am. I close that gap — evals, observability, and cost control — with nine years of distributed systems behind it.

Not ready to talk? Score your agent in 4 minutes — free, no signup, nothing leaves your browser.

Durgesh Rathod
Available for new work

Based in Pune, India (+05:30). Local time .

Your time
Working-hours overlap

I hold afternoons and evenings IST open for calls, which covers European mornings and US mornings on the earlier side. Async-first either way.

The three ways agents die

A prototype has to work once. A product has to work on inputs nobody imagined.

Almost every failing agent system I have looked at is failing in one of three places. They are unglamorous and they are all fixable in days, not quarters.

01

Reliability

It works in the demo, not on real data.

No eval suite, no regression tests, no way to know whether a prompt change made things better or quietly worse.

02

Cost & latency

Spend and p95 are both a surprise.

No token budget, no model routing, no cost ceiling. Retries stack silently until the invoice arrives.

03

Observability

It is a black box when it breaks.

No tracing on tool calls, no way to reproduce a bad run, no shared vocabulary for the ways it fails.

The uncomfortable version: none of this is about model choice, and very little of it is about prompts. It is the boring engineering that conventional services got right twenty years ago and agent codebases skipped — validation at boundaries, bounded resources, traceable failures. I have catalogued 9 of these failure modes with the causes and fixes for each.

Free · 4 minutes · no signup

Is your agent production-ready, or just working today?

12 questions on the things that actually cause incidents. You get a score out of 100, a breakdown across the three pillars, and your gaps ranked by risk — each with a specific fix and an effort estimate. It runs entirely in your browser; no answers are transmitted anywhere.

Score my agentMost teams score in the 30s the first time.
  • ReliabilityCan you prove it still works after a change?
  • Cost & latencyDo you know what a conversation costs and how slow it gets?
  • ObservabilityWhen it fails at 2am, can you find out why?

How to work with me

Two ways in, both fixed-price

Prices are on the page because hiding them wastes everyone's time. Fixed scope rather than hourly, so the number does not move.

Fixed scope · 1–2 weeks

Agent Production Readiness Audit

Start here

$4,000 – $6,000

A fixed-price diagnostic on an agent you already have running. You keep everything I build.

  • Failure taxonomy for your specific agent, ranked by production risk
  • Cost and latency profile — where tokens actually go, and p50/p95 per operation
  • An eval suite wired to your real traffic that your team owns and keeps running
  • Prioritised fix list with effort estimates, sequenced by risk reduction per day of work
  • A 90-minute walkthrough with your engineers, recorded

You have an agent in or near production and a growing feeling that you cannot prove it works.

Monthly retainer · ~2 days/week, 3-month minimum

Fractional AI Reliability Lead

$5,000 – $8,000 / month

I own agent reliability as a function: evals, observability, and cost discipline, alongside your team.

  • Own and grow the eval suite as your agent surface expands
  • Tracing and alerting on the failure modes that actually cost you money
  • Model routing and token budgets — usually the fastest measurable win
  • Design review on new agent features before they ship, not after
  • Direct line to me on Slack, and a written monthly reliability report

Agents are core to your product and nobody currently owns whether they work.

Questions

The things people ask first

What does "making an AI agent production-ready" actually involve?

Three things, in this order: an eval suite so you can tell whether a change made the agent better or worse, tracing so a failure can be reproduced and diagnosed, and cost controls so spend and latency are bounded rather than discovered on an invoice.

Most teams have none of the three when they first ship. The model is rarely the problem — the missing engineering discipline around it is.

How much does it cost to work with you?

Two options, both fixed-price. The Agent Production Readiness Audit is $4,000 – $6,000 for 1–2 weeks and produces a failure taxonomy, a cost and latency profile, an eval suite you keep, and a prioritised fix list. The Fractional AI Reliability Lead is $5,000 – $8,000 / month at roughly two days a week.

I quote fixed scope rather than hourly rates, so the price is the price regardless of how long something takes me.

You are in India — does the timezone work for a US or European team?

Yes, with real overlap rather than the pretence of it. I hold 09:00–21:00 IST open, which covers the full European working day and US mornings on the earlier side. That is typically four or more shared hours with London and two to three with New York.

Beyond that I work async-first: written updates, recorded walkthroughs, and traceable work rather than requiring meetings.

Do you work with agents built on frameworks you have not used?

Yes. The failure modes are framework-independent — unvalidated tool arguments, unbounded loops, missing evals, no tracing, and cost with no ceiling look the same in LangGraph, CrewAI, the raw provider SDKs, or hand-rolled orchestration.

I have shipped production systems on CrewAI, LangChain, Model Context Protocol, AWS Bedrock, and direct provider APIs. Framework specifics are a day of reading; the reliability work is the same job.

What do you need from us to start an audit?

Read access to the agent code, whatever logs or traces already exist, and 60–90 minutes with an engineer who knows the system. Ideally a sample of real production inputs, anonymised if necessary.

If you have no traces at all, that is not a blocker — it is usually the first finding.

Can I just use the free scorecard and fix things myself?

Yes, and that is a genuinely good outcome. The 12-question scorecard names the gaps, ranks them by production risk, and gives each one a fix with an effort estimate. Plenty of teams need the map rather than a consultant.

Hire me when the gaps are clear but nobody has the time or the specific experience to close them.

Tell me what your agent does and what worries you.

A message describing your stack and volume is enough for me to say something useful back — usually the same day, and usually including at least one thing you can fix without hiring anyone.