Reliability
It works in the demo, not on real data.
No eval suite, no regression tests, no way to know whether a prompt change made things better or quietly worse.
Independent · taking audits and one more retainer
Your agent works in the demo. Then real users arrive, costs triple, and nobody can explain why it failed at 2am. I close that gap — evals, observability, and cost control — with nine years of distributed systems behind it.
Not ready to talk? Score your agent in 4 minutes — free, no signup, nothing leaves your browser.

Based in Pune, India (+05:30). Local time —.
I hold afternoons and evenings IST open for calls, which covers European mornings and US mornings on the earlier side. Async-first either way.
The three ways agents die
Almost every failing agent system I have looked at is failing in one of three places. They are unglamorous and they are all fixable in days, not quarters.
It works in the demo, not on real data.
No eval suite, no regression tests, no way to know whether a prompt change made things better or quietly worse.
Spend and p95 are both a surprise.
No token budget, no model routing, no cost ceiling. Retries stack silently until the invoice arrives.
It is a black box when it breaks.
No tracing on tool calls, no way to reproduce a bad run, no shared vocabulary for the ways it fails.
The uncomfortable version: none of this is about model choice, and very little of it is about prompts. It is the boring engineering that conventional services got right twenty years ago and agent codebases skipped — validation at boundaries, bounded resources, traceable failures. I have catalogued 9 of these failure modes with the causes and fixes for each.
Free · 4 minutes · no signup
12 questions on the things that actually cause incidents. You get a score out of 100, a breakdown across the three pillars, and your gaps ranked by risk — each with a specific fix and an effort estimate. It runs entirely in your browser; no answers are transmitted anywhere.
Evidence
Each of these includes the architecture, the number it is remembered for, and the thing that broke on the way. The failure is the part that makes the rest believable.
Reporting turnaround
A natural-language analytics agent over sensitive workforce data, serving 2,000 concurrent users across isolated tenants. Reporting went from 2–3 days to under 30 seconds — but only after we stopped trusting the model with the boundary.
Read itIn production, not a demo
Two production agent systems on MCP and CrewAI — a project-management assistant and an autonomous lead-generation pipeline. What multi-agent buys you, what it costs, and the specific cases where a single agent with good tools wins.
Read itSustained ingestion
A configuration-driven telecom data platform processing 200M configuration and 320M performance parameters per 15-minute cycle on Golang, Kafka, Kubernetes and PostgreSQL. The project that shaped how I think about bounded resources — and why I trust it more than any AI credential I have.
Read itHow to work with me
Prices are on the page because hiding them wastes everyone's time. Fixed scope rather than hourly, so the number does not move.
Fixed scope · 1–2 weeks
$4,000 – $6,000
A fixed-price diagnostic on an agent you already have running. You keep everything I build.
You have an agent in or near production and a growing feeling that you cannot prove it works.
Monthly retainer · ~2 days/week, 3-month minimum
$5,000 – $8,000 / month
I own agent reliability as a function: evals, observability, and cost discipline, alongside your team.
Agents are core to your product and nobody currently owns whether they work.
Questions
Three things, in this order: an eval suite so you can tell whether a change made the agent better or worse, tracing so a failure can be reproduced and diagnosed, and cost controls so spend and latency are bounded rather than discovered on an invoice.
Most teams have none of the three when they first ship. The model is rarely the problem — the missing engineering discipline around it is.
Two options, both fixed-price. The Agent Production Readiness Audit is $4,000 – $6,000 for 1–2 weeks and produces a failure taxonomy, a cost and latency profile, an eval suite you keep, and a prioritised fix list. The Fractional AI Reliability Lead is $5,000 – $8,000 / month at roughly two days a week.
I quote fixed scope rather than hourly rates, so the price is the price regardless of how long something takes me.
Yes, with real overlap rather than the pretence of it. I hold 09:00–21:00 IST open, which covers the full European working day and US mornings on the earlier side. That is typically four or more shared hours with London and two to three with New York.
Beyond that I work async-first: written updates, recorded walkthroughs, and traceable work rather than requiring meetings.
Yes. The failure modes are framework-independent — unvalidated tool arguments, unbounded loops, missing evals, no tracing, and cost with no ceiling look the same in LangGraph, CrewAI, the raw provider SDKs, or hand-rolled orchestration.
I have shipped production systems on CrewAI, LangChain, Model Context Protocol, AWS Bedrock, and direct provider APIs. Framework specifics are a day of reading; the reliability work is the same job.
Read access to the agent code, whatever logs or traces already exist, and 60–90 minutes with an engineer who knows the system. Ideally a sample of real production inputs, anonymised if necessary.
If you have no traces at all, that is not a blocker — it is usually the first finding.
Yes, and that is a genuinely good outcome. The 12-question scorecard names the gaps, ranks them by production risk, and gives each one a fix with an effort estimate. Plenty of teams need the map rather than a consultant.
Hire me when the gaps are clear but nobody has the time or the specific experience to close them.
A message describing your stack and volume is enough for me to say something useful back — usually the same day, and usually including at least one thing you can fix without hiring anyone.
Nothing matched. Try “cost”, “loop”, “evals”, or “rates”.