MewCP LogoAStheTech
MCPs
Use Cases

Use cases by category

Productivity & InboxInbox, calendar, and daily flowEngineering & DevOpsShip, debug, and run on-callSales & CRMPipeline, outreach, and dealsMarketing & GrowthCampaigns, SEO, and growthSupport & SuccessTriage tickets, keep customers happyFinance & OpsClose, reconcile, and expensesCreative & ContentGenerate assets and contentPeople & HiringHiring, onboarding, and HRResearch & DataSynthesize data and insights
See all use cases
Resources
BlogsProduct updates and storiesArticlesIntegration guides and code examples
PricingDocsSign in
Back to home
MewCP Logo

Infrastructure You Can Trust for Agentic Products

X

Categories

  • Productivity & Docs
  • Developer Tools
  • CRM & Sales
  • Finance & Commerce
  • Data & Analytics
  • Marketing & SEO
  • Search & Web
  • Communication
  • View All Servers →

Resources

  • Blog
  • Docs
  • Privacy Policy
  • Terms of Service

Blogs

  • View All Blogs →

Articles

  • View All Articles →
Browse Servers|Pricing|Contact

Browse by Category

Productivity & Docs

  • Gmail
  • Google Drive
  • Google Classroom
  • Google Calendar
  • Google People
  • YouTube
  • Notion
  • ClickUp
  • Figma
  • Google Tasks
  • Cal
  • Monday
  • Luma
  • Notion MCP
  • Mem MCP
  • Linear MCP
  • Calendly MCP
  • Consensus MCP
  • Craft MCP
  • Close MCP
  • Dice MCP
  • Lumin PDF MCP
  • Develop21 MCP
  • Granola MCP
  • Lucid MCP
  • Mermaid Chart MCP
  • Fireflies MCP
  • ClickUp MCP
  • Miro MCP

Developer Tools

  • Gemini
  • Veo
  • ClickUp
  • Firecrawl
  • Vercel
  • Apify
  • Github
  • HTTP
  • Chef
  • Scientific Calculator
  • Figma
  • Perplexity
  • Apify MCP
  • Hugging Face Hub MCP
  • Buildkite MCP
  • Cloudflare MCP
  • Context7 MCP
  • Ahrefs MCP
  • Sentry MCP
  • Brevo Docs MCP
  • X Docs MCP
  • Jev
  • Linear MCP
  • Calendly MCP
  • Craft MCP
  • DeepWiki MCP
  • Inspo MCP
  • Kernel MCP
  • Malwarebytes MCP
  • Mermaid Chart MCP
  • Supabase MCP
  • Microsoft Learn MCP
  • Webflow MCP

CRM & Sales

  • Google People
  • OneSignal MCP
  • Brevo Docs MCP
  • Brevo
  • Brevo MCP
  • Carbon Voice MCP
  • Clay MCP
  • Close MCP
  • Attio MCP
  • Clarify MCP

Finance & Commerce

  • Razorpay
  • Polymarket
  • Kite
  • Stripe
  • Binance
  • Upstox
  • Aiwyn MCP
  • Era-Context-MCP
  • Granted MCP
  • XDC AI MCP
  • Agentery MCP
  • Agent Embassy
  • Quick Commerce MCP
  • Longbridge MCP
  • Mercury MCP

Data & Analytics

  • Apify MCP
  • Cloudflare MCP
  • Ahrefs MCP
  • Candid MCP
  • Consensus MCP
  • Contentsquare MCP
  • Era-Context-MCP
  • Instinct MCP
  • legal Data Hunter MCP
  • Marcopolo MCP
  • Mixpanel MCP

Marketing & SEO

  • Mailchimp
  • Google Business
  • YouTube
  • Google Search Console
  • OneSignal MCP
  • Cloudflare MCP
  • Brevo Docs MCP
  • Brevo
  • Brevo MCP
  • AirOps MCP
  • Clay MCP
  • Contentsquare MCP
  • Reelsmith MCP
  • GoDaddy MCP
  • Metricool MCP
  • Webflow MCP
  • Windsor MCP

Search & Web

  • Web Scrapper
  • Firecrawl
  • Apify
  • Perplexity
  • Context.dev
  • Exa
  • Brave Search
  • Apify MCP
  • Ahrefs MCP
  • DeepWiki MCP
  • Dice MCP
  • GoDaddy MCP
  • Granted MCP
  • Microsoft Learn MCP

Communication

  • Gmail
  • Google Meet
  • Mailchimp
  • Google Calendar
  • WhatsApp
  • Slack
  • OneSignal MCP
  • Brevo Docs MCP
  • Carbon Voice MCP

© 2026 MewCP. All rights reserved.

  1. Home
  2. Blogs
  3. AI Agent Reliability: The 7 Mechanisms Between a Demo and Production

AI Agent Reliability: The 7 Mechanisms Between a Demo and Production

by Rohit Gite, Founder @MewCP·October 1, 2026·7 min read

Your agent worked in the demo because nothing went wrong. Here is the engineering layer that keeps it working when the network, the tool and the model all misbehave.

AI agent reliability is the thing nobody builds until the first incident. You wire up a model, give it four tools, run it, and it works. It works again. You ship it. Then a search API hangs for ninety seconds, your agent sits there holding an open connection, the user closes the tab, and you find out that your agent had no opinion at all about what should happen when a tool does not answer.

That is the whole gap between a demo and a production system. Not model quality. Not prompt quality. The demo ran once and nothing went wrong.

This is Day 23 of our 30 Days of AI Agents series, and it is the one where the fun stops and the engineering starts.

Why "just add a try/catch" breaks

The instinct is to wrap the agent loop in a try/catch, log the error, and return a friendly message. That handles the case where something throws. It does not handle any of the cases that actually hurt.

A hanging request never throws. It just waits.

A tool that returns 200 OK with an HTML login page never throws. Your agent reads the HTML as data and reasons about it.

A model that returns valid JSON with a hallucinated field never throws. Your downstream code reads undefined and writes it to a database.

A retry on a payment call that already succeeded never throws. It charges the customer twice.

None of those are exceptions. They are outcomes. Reliability is about having a defined response to every outcome, including the ones your language does not consider errors.

AI Agent Reliability Is a Harness, Not a Prompt

Here is the mental model worth keeping. The model call is a probabilistic operation sitting inside a network. Everything protective around it is deterministic code you write. Adding "always double check your work" to the system prompt is not a reliability mechanism, because there is no enforcement behind it.

Seven mechanisms do the real work. Four of them contain a failure while it is happening. Three of them decide how the run ends.

Diagram of a request passing through seven reliability stages from timeout through retry, error classification, tool result validation, output schema, fallback and approval gate

The four that contain the failure

1. Timeouts that actually cancel

Every outbound call gets a deadline. Not just the model call, every tool call too. And the deadline has to cancel the underlying request, not just stop waiting for it.

A Promise.race against a timer looks like a timeout but leaks. The original request stays open, still consuming a socket and still capable of resolving later. Use an abort signal and pass it all the way down.

async function withTimeout<T>(
  fn: (signal: AbortSignal) => Promise<T>,
  ms: number,
): Promise<T> {
  const controller = new AbortController();
  const timer = setTimeout(
    () => controller.abort(new Error(`timeout after ${ms}ms`)),
    ms,
  );
  try {
    return await fn(controller.signal);
  } finally {
    clearTimeout(timer);
  }
}

Two deadlines, not one. A per-attempt timeout, typically 5 to 15 seconds for a tool, and a total budget for the whole agent run. Without the total budget, three tools with three retries each can still burn six minutes before anyone notices.

2. Error classes, decided in code

Before you can retry anything, you need to know whether retrying is even sensible. Classify every failure into three buckets.

type Kind = "retryable" | "fatal" | "ambiguous";
 
function classify(error: unknown): Kind {
  const status = (error as { status?: number } | null)?.status;
 
  if (status === 429) return "retryable";
  if (status !== undefined && status >= 500) return "retryable";
  if (status !== undefined && status >= 400) return "fatal";
 
  if (error instanceof Error && error.name === "AbortError") return "ambiguous";
  if (error instanceof TypeError) return "ambiguous"; // network layer
 
  return "ambiguous";
}

fatal means stop, a bad API key or a malformed argument will fail identically forever. retryable means the server told you to come back. ambiguous is the interesting one: a timeout or a dropped connection means you genuinely do not know whether the operation ran. Most code treats ambiguous as retryable, and that is where double charges come from.

3. Retries that are safe to run

Retry with exponential backoff and full jitter, a hard attempt cap, and respect for the total budget. And retry an ambiguous failure only when the operation is idempotent.

const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
 
interface RetryOpts {
  perAttemptMs: number;
  totalBudgetMs: number;
  maxAttempts: number;
  idempotent: boolean;
}
 
export async function callTool<T>(
  run: (signal: AbortSignal) => Promise<T>,
  opts: RetryOpts,
): Promise<T> {
  const deadline = Date.now() + opts.totalBudgetMs;
  let lastError: unknown;
 
  for (let attempt = 1; attempt <= opts.maxAttempts; attempt++) {
    const remaining = deadline - Date.now();
    if (remaining <= 0) break;
 
    try {
      return await withTimeout(run, Math.min(opts.perAttemptMs, remaining));
    } catch (error) {
      lastError = error;
      const kind = classify(error);
 
      if (kind === "fatal") throw error;
      if (kind === "ambiguous" && !opts.idempotent) throw error;
      if (attempt === opts.maxAttempts) break;
 
      // full jitter: random point in [0, capped backoff)
      const cap = Math.min(8_000, 250 * 2 ** (attempt - 1));
      await sleep(Math.random() * cap);
    }
  }
 
  throw lastError;
}

Full jitter matters more than people expect. Fixed backoff means every concurrent agent run that hit the same rate limit retries at the same instant and trips it again.

To make writes idempotent, generate a stable key per logical operation, reuse it across attempts, and let the tool deduplicate on it. Read operations are free to retry. Writes are not, unless you have done this work.

const idempotencyKey = `${runId}:${stepIndex}:issue_refund`;

4. Tool result validation

A tool returning successfully is not the same as a tool returning what you asked for. Validate the payload shape before it ever reaches the model's context.

import { z } from "zod";
 
const FlightQuote = z.object({
  carrier: z.string().min(1),
  priceUsd: z.number().positive(),
  departsAt: z.string().datetime(),
});
 
const parsed = FlightQuote.safeParse(rawToolResponse);
 
if (!parsed.success) {
  return {
    ok: false as const,
    reason: "tool returned an unexpected shape",
    detail: parsed.error.issues.slice(0, 3),
  };
}

Feed the failure back to the agent as a structured tool error, not as raw garbage. An agent told "the tool returned an unexpected shape" can pick a different tool. An agent handed an HTML error page will try to summarise it.

The three that decide how it ends

Two execution traces compared on a dark panel, a demo agent hanging for 94 seconds and aborting versus a production agent timing out at 8 seconds, retrying and serving a cached fallback in 11.4 seconds

5. Output schemas

Ask for structured output, then parse it. If parsing fails, allow exactly one repair attempt where you send the validation error back and ask for a corrected object. Then stop.

let result = SchemaOut.safeParse(await generate(prompt));
 
if (!result.success) {
  const repairPrompt = `${prompt}\n\nYour previous response failed validation:\n${
    JSON.stringify(result.error.issues)
  }\nReturn only corrected JSON matching the schema.`;
  result = SchemaOut.safeParse(await generate(repairPrompt));
}
 
if (!result.success) return fallback();

Bound the repair loop. An unbounded one is a cost incident: a model that cannot satisfy a schema on attempt two will usually not satisfy it on attempt nine either, and you have paid for all nine.

6. Fallbacks

Decide in advance what a degraded answer looks like, and be honest about it. Common tiers, in order: live tool, cached result with an age stamp, a cheaper or simpler tool, a smaller model, and finally a clear statement that the agent could not complete the step.

The rule that keeps trust intact is labelling. "Prices from 40 minutes ago, live pricing is unavailable" is a good answer. Silently serving stale data as fresh is how an agent loses a user permanently.

7. Approval gates

Not every action deserves the same level of autonomy. Size the gate by blast radius: how bad is this if the agent is wrong, and how hard is it to undo?

Three tiers of agent action risk on a dark panel, auto for reversible reads, confirm for outbound actions, and review for irreversible financial operations

type Risk = "auto" | "confirm" | "review";
 
const TOOL_RISK: Record<string, Risk> = {
  search_flights: "auto",     // read only, reversible
  draft_email: "auto",        // creates nothing external
  send_email: "confirm",      // leaves your system
  issue_refund: "review",     // moves money, hard to undo
  delete_account: "review",   // irreversible
};

auto runs. confirm pauses for a one click yes. review waits for a human who can see the arguments before it executes. Getting this wrong in the cautious direction is annoying. Getting it wrong in the permissive direction ends up in a postmortem.

What still breaks after all seven

Three failure modes survive a good harness, so plan for them.

Retry storms. Every agent run retrying a struggling dependency turns a slow service into a dead one. Add a circuit breaker: after N consecutive failures against a tool, stop calling it for a cooldown window and go straight to the fallback.

Partial completion. An agent does three of five steps then fails. If those steps had side effects, you need either compensating actions or a resumable run with checkpointed state.

Silent quality decay. Everything returns 200 OK and the answers are quietly getting worse. No amount of error handling catches this. It needs evaluation, which is a separate discipline.

The pre-ship checklist

  1. Every outbound call has a per-attempt timeout that cancels the request.
  2. The whole run has a total time budget and an attempt cap.
  3. Failures are classified as retryable, fatal or ambiguous before anything reacts.
  4. Ambiguous failures are only retried on operations carrying an idempotency key.
  5. Every tool response is schema-validated before it enters model context.
  6. Model outputs are parsed against a schema with at most one repair attempt.
  7. Every tool is tagged auto, confirm or review, and the irreversible ones are never auto.

If you cannot point at the code that does each of these, that mechanism does not exist yet.

Where this lands for infrastructure

Most of these mechanisms are not agent-specific. They are the timeout, retry and validation discipline that any distributed system needs, applied at the tool boundary. That boundary is exactly where MCP servers sit, which is why we spend most of our time at MewCP on it: hosting MCP servers with sane timeout and retry behaviour, handling credentials so an expired token surfaces as a clean fatal error instead of a mystery 401 mid-run, and giving the gateway a consistent place to enforce these rules across every tool rather than reimplementing them per integration.

You can build all of this yourself. Plenty of teams do. Just build it deliberately, before the incident that forces you to.

Reliability is measured in recoveries. Nobody notices an agent that never fails, because there is no such thing. They notice the one that fails and finishes the job anyway.

ContentsOctober 1, 2026
  1. Why "just add a try/catch" breaks
  2. AI Agent Reliability Is a Harness, Not a Prompt
  3. The four that contain the failure
  4. The three that decide how it ends
  5. What still breaks after all seven
  6. The pre-ship checklist
  7. Where this lands for infrastructure
Author

Rohit Gite, Founder @MewCP

Share

Build with MewCP

Connect your AI agents to real tools in minutes.

Get started