MewCP LogoAStheTech
MCPs
Use Cases

Use cases by category

Productivity & InboxInbox, calendar, and daily flowEngineering & DevOpsShip, debug, and run on-callSales & CRMPipeline, outreach, and dealsMarketing & GrowthCampaigns, SEO, and growthSupport & SuccessTriage tickets, keep customers happyFinance & OpsClose, reconcile, and expensesCreative & ContentGenerate assets and contentPeople & HiringHiring, onboarding, and HRResearch & DataSynthesize data and insights
See all use cases
Resources
BlogsProduct updates and storiesArticlesIntegration guides and code examples
PricingDocsSign in
Back to home
MewCP Logo

Infrastructure You Can Trust for Agentic Products

X

Categories

  • Productivity & Docs
  • Developer Tools
  • CRM & Sales
  • Finance & Commerce
  • Data & Analytics
  • Marketing & SEO
  • Search & Web
  • Communication
  • View All Servers →

Resources

  • Blog
  • Docs
  • Privacy Policy
  • Terms of Service

Blogs

  • View All Blogs →

Articles

  • View All Articles →
Browse Servers|Pricing|Contact

Browse by Category

Productivity & Docs

  • Gmail
  • Google Drive
  • Google Classroom
  • Google Calendar
  • Google People
  • YouTube
  • Notion
  • ClickUp
  • Figma
  • Google Tasks
  • Cal
  • Monday
  • Luma
  • Notion MCP
  • Mem MCP
  • Linear MCP
  • Calendly MCP
  • Consensus MCP
  • Craft MCP
  • Close MCP
  • Dice MCP
  • Lumin PDF MCP
  • Develop21 MCP
  • Granola MCP
  • Lucid MCP
  • Mermaid Chart MCP
  • Fireflies MCP
  • ClickUp MCP
  • Miro MCP

Developer Tools

  • Gemini
  • Veo
  • ClickUp
  • Firecrawl
  • Vercel
  • Apify
  • Github
  • HTTP
  • Chef
  • Scientific Calculator
  • Figma
  • Perplexity
  • Apify MCP
  • Hugging Face Hub MCP
  • Buildkite MCP
  • Cloudflare MCP
  • Context7 MCP
  • Ahrefs MCP
  • Sentry MCP
  • Brevo Docs MCP
  • X Docs MCP
  • Jev
  • Linear MCP
  • Calendly MCP
  • Craft MCP
  • DeepWiki MCP
  • Inspo MCP
  • Kernel MCP
  • Malwarebytes MCP
  • Mermaid Chart MCP
  • Supabase MCP
  • Microsoft Learn MCP
  • Webflow MCP

CRM & Sales

  • Google People
  • OneSignal MCP
  • Brevo Docs MCP
  • Brevo
  • Brevo MCP
  • Carbon Voice MCP
  • Clay MCP
  • Close MCP
  • Attio MCP
  • Clarify MCP

Finance & Commerce

  • Razorpay
  • Polymarket
  • Kite
  • Stripe
  • Binance
  • Upstox
  • Aiwyn MCP
  • Era-Context-MCP
  • Granted MCP
  • XDC AI MCP
  • Agentery MCP
  • Agent Embassy
  • Quick Commerce MCP
  • Longbridge MCP
  • Mercury MCP

Data & Analytics

  • Apify MCP
  • Cloudflare MCP
  • Ahrefs MCP
  • Candid MCP
  • Consensus MCP
  • Contentsquare MCP
  • Era-Context-MCP
  • Instinct MCP
  • legal Data Hunter MCP
  • Marcopolo MCP
  • Mixpanel MCP

Marketing & SEO

  • Mailchimp
  • Google Business
  • YouTube
  • Google Search Console
  • OneSignal MCP
  • Cloudflare MCP
  • Brevo Docs MCP
  • Brevo
  • Brevo MCP
  • AirOps MCP
  • Clay MCP
  • Contentsquare MCP
  • Reelsmith MCP
  • GoDaddy MCP
  • Metricool MCP
  • Webflow MCP
  • Windsor MCP

Search & Web

  • Web Scrapper
  • Firecrawl
  • Apify
  • Perplexity
  • Context.dev
  • Exa
  • Brave Search
  • Apify MCP
  • Ahrefs MCP
  • DeepWiki MCP
  • Dice MCP
  • GoDaddy MCP
  • Granted MCP
  • Microsoft Learn MCP

Communication

  • Gmail
  • Google Meet
  • Mailchimp
  • Google Calendar
  • WhatsApp
  • Slack
  • OneSignal MCP
  • Brevo Docs MCP
  • Carbon Voice MCP

© 2026 MewCP. All rights reserved.

  1. Home
  2. Blogs
  3. Anatomy of an AI Agent: The Six Parts of Any AI Agent Architecture

Anatomy of an AI Agent: The Six Parts of Any AI Agent Architecture

by Rohit Gite, Co-Founder CTO @MewCP·August 17, 2026·10 min read

Most people draw an agent as one box, so when it breaks there is nothing to point at. Open the box and there are six parts inside. Here is what each one owns.

Ask ten engineers to sketch an AI agent and you get the same drawing ten times. A box. A prompt goes in on the left, a result comes out on the right. That drawing is fine right up until the thing misbehaves, and then the most specific sentence available is "the agent is broken." AI agent architecture exists to make that sentence better.

Because "the agent is broken" is not a bug report. Nobody can assign it, nobody can reproduce it, and nobody can tell you which line to open. You can only fix something you can point at.

Open the box and there are six parts inside. Three of them decide. Three of them act.

Why "the agent is broken" is not a bug report

Here is the practical difference between a box and a decomposition.

An agent processes a refund. The customer complains that they got refunded twice. With a box, your investigation is "read 400 lines of log and form a theory." With six named parts, it is three questions: did the plan contain two refund steps, did the model choose the refund action twice, or did one call get retried without an idempotency key?

Three different bugs, three different places, three different fixes. The first is planning. The second is reasoning. The third is execution. None of them are "the model."

The value of naming the parts is addressability. Every hour you have lost to an agent failure was an hour spent turning a symptom into a location, and the decomposition does that work in advance.

The six components of an AI agent architecture

Every agent you will build or read has these six responsibilities somewhere. They may be six modules, six functions, or six paragraphs of one prompt. The responsibility exists either way. When it is not deliberately assigned, it is being handled by accident.

The real boundary between the parts is not what they do. It is what state each one owns. Ownership is what makes a component testable, and it is the line people blur first.

ComponentOwnsFails like
PlanningThe ordered steps and the stop conditionSteps in the wrong order, work that never ends
ReasoningThe choice of the next single action and its argumentsRight plan, wrong tool, wrong arguments
MemoryWhat is already known and what already happenedRepeats work, asks for a detail it was given
ToolsThe contract with the outside world: names, schemas, permissionsA perfect plan where nothing actually changes
ExecutionTimeouts, retries, idempotency, budgetsDuplicate side effects, hangs, runaway cost
EvaluationWhether a result is good enough to move onConfidently wrong output, silent failure

Engraved anatomical detail plate showing six labelled agent organs paired with the state each one owns

A few notes that those one-line summaries flatten.

Planning is not reasoning. Planning answers "what is the shape of this job." Reasoning answers "what do I do next." A plan survives multiple steps. A reasoning decision is disposable. If you cannot print your agent's plan without running the agent, you do not have a planning component, you have reasoning that occasionally sounds strategic.

Memory is not the message array. The message array is a transport format. Memory is the decision about what a fact's lifespan is: this step, this session, this user, or your database. One growing list means one eviction rule for four different lifespans, and any rule you pick is wrong for at least three of them.

Tools are a contract, not a function call. The tool layer owns the names, the argument schemas, the descriptions the model reads, and the permission boundary. Get the contract wrong and the model produces plausible calls your system rejects.

Execution is not the tool. The tool says what issue_refund means. Execution decides what happens when issue_refund times out after 30 seconds with an unknown result. Put retry logic inside each tool and you will have six inconsistent retry policies within a month.

Evaluation is not the model saying "done." Evaluation is an independent check on the result. If the same model call that produced the answer also grades the answer, you have a component in name only.

Six seams in one small agent

Here is a complete loop with all six parts labelled. It is deliberately small enough to read in one sitting, and it is deliberately one file, because the point is that the seams are conceptual before they are structural.

from dataclasses import dataclass, field
 
 
@dataclass
class Action:
    name: str
    args: dict = field(default_factory=dict)
 
 
@dataclass
class Verdict:
    status: str          # "ok" | "retry" | "replan" | "stop"
    note: str = ""
 
 
@dataclass
class AgentState:
    goal: str
    plan: list[str] = field(default_factory=list)       # PLANNING owns this
    facts: dict = field(default_factory=dict)           # MEMORY owns this
    history: list[str] = field(default_factory=list)    # MEMORY owns this
    budget: int = 12                                    # EXECUTION owns this
 
 
def run(goal: str, tools: dict) -> AgentState:
    state = AgentState(goal=goal)
    state.plan = make_plan(goal, tools)                          # PLANNING
 
    while state.budget > 0:
        state.budget -= 1
 
        context = recall(state)                                  # MEMORY
        action = choose_action(state.plan, context, tools)        # REASONING
 
        if action.name == "finish":
            break
 
        observation = execute(action, tools)                     # TOOLS + EXECUTION
        verdict = evaluate(action, observation, state)            # EVALUATION
        remember(state, action, observation, verdict)             # MEMORY
 
        if verdict.status == "replan":
            state.plan = make_plan(goal, tools, context=recall(state))
        elif verdict.status == "stop":
            break
 
    return state

Twenty-five lines, six seams, and every one of them is a place you can put a log line, a test, or a breakpoint.

The two functions worth expanding are the two people usually skip.

import time
import uuid
 
 
def execute(action: Action, tools: dict, attempts: int = 3) -> dict:
    """EXECUTION: owns retries, timeouts, idempotency and cost. Not the tool."""
    spec = tools.get(action.name)
    if spec is None:
        return {"error": f"unknown tool {action.name}"}
 
    # One key per logical action, reused across retries, so a retried
    # write is a no-op instead of a second refund.
    key = f"{action.name}:{uuid.uuid4().hex}" if spec["writes"] else None
 
    delay = 0.5
    for attempt in range(attempts):
        try:
            return spec["fn"](**action.args, idempotency_key=key) if key \
                else spec["fn"](**action.args)
        except TimeoutError as e:
            if attempt == attempts - 1:
                return {"error": f"timeout after {attempts} attempts: {e}"}
            time.sleep(delay)
            delay *= 2
        except ValueError as e:
            # Bad arguments will fail identically every time. Do not retry.
            return {"error": f"invalid args: {e}"}
 
    return {"error": "unreachable"}

Two decisions in there carry most of the weight. Retrying only the errors that are actually transient, and generating the idempotency key once per logical action rather than once per attempt. Generate it inside the retry loop and you have rebuilt the double-refund bug with extra steps.

def evaluate(action: Action, observation: dict, state: AgentState) -> Verdict:
    """EVALUATION: an independent check, not the model's own opinion."""
    if "error" in observation:
        return Verdict("retry", observation["error"])
 
    check = POSTCONDITIONS.get(action.name)
    if check is None:
        return Verdict("ok")
 
    if not check(observation, state):
        return Verdict("replan", f"{action.name} returned 200 but the "
                                  "postcondition did not hold")
    return Verdict("ok")
 
 
POSTCONDITIONS = {
    # The call succeeded. Did the world actually change?
    "issue_refund": lambda obs, st: obs.get("status") == "refunded"
                                    and obs.get("amount", 0) > 0,
    "send_email":   lambda obs, st: obs.get("message_id") is not None,
}

That dictionary is not sophisticated and it does not need to be. It is the difference between an agent that reports success and an agent that verified success.

The three merges that cause real bugs

Six responsibilities collapse into fewer components all the time, and most collapses are harmless. Three of them are not.

Engraved plate showing three pairs of agent modules fused together with a jade fracture line through each pair

Planning merged into reasoning. No plan object exists, so the model re-derives the whole strategy on every pass with slightly different framing each time. Symptom: the agent does step 4 twice, drifts off the original goal by step 8, and you cannot answer "what is left to do" without asking the model. Fix: make the plan a stored object, even if it is only a list of strings. Something you can print before the first tool call.

Execution merged into tools. Retry, timeout and backoff logic lives inside each tool implementation. Symptom: search_orders retries five times, issue_refund retries zero times, and nobody remembers which does what. Then someone adds a retry to a non-idempotent write and you refund a customer twice. Fix: tools describe a capability and perform exactly one attempt. One executor owns the retry policy for all of them.

Evaluation merged into reasoning. The same model call decides both what to do and whether the last thing worked. Symptom: a run that reports "refund processed successfully" for a refund that returned a 500. Models are agreeable about their own work. Fix: evaluation reads the observation and the state, never the model's narration. A boolean function is a legitimate evaluator.

Evaluation is the one almost everyone skips

Of the six, evaluation is the component most likely to be missing entirely, and it is the one that decides whether anyone can trust the agent.

Without it, the loop's exit condition is effectively "the model said it was done." That is how you get a run that looks clean in the logs while the refund never processed and the email never sent.

People skip it because they think evaluation means an LLM judge, a scoring rubric and a labelled dataset. Eventually it might. On day one it means one line: after a write action, read the state back and confirm it changed.

def issue_refund_checked(order_id: str, amount: int, tools: dict) -> dict:
    tools["issue_refund"]["fn"](order_id=order_id, amount=amount)
    order = tools["get_order"]["fn"](order_id=order_id)   # read it back
    if order["refund_status"] != "refunded":
        raise RuntimeError(f"refund reported success but order {order_id} "
                           f"is still {order['refund_status']}")
    return order

One extra tool call. It catches a large share of silent failures and it costs less than a single debugging session it prevents. Start there and grow into judges and eval sets when you have the traffic to justify them.

The full symptom to component map

The carousel version of this had four rows. Here is the longer one. Print it, put it next to the trace viewer, and stop guessing.

SymptomComponent
Does the right steps in the wrong orderPlanning
Never stops, or stops before the goal is metPlanning
Picks a tool that cannot do the jobReasoning
Picks the right tool with wrong argumentsReasoning
Asks for something the user already told itMemory
Repeats an action it already completedMemory
Produces a clean plan, nothing changes in your systemsTools
Invents a tool that does not existTools
Hangs, or costs far more than the task is worthExecution
Performs a write action twiceExecution
Reports success on work that failedEvaluation
Confidently wrong final answerEvaluation

Engraved diagnostic key plate mapping agent symptoms on the left to the responsible component on the right

Notice how many of these get misdiagnosed as "the model is not smart enough." Exactly two rows in that table are genuinely reasoning problems. The other ten get fixed with code you write, not with a bigger model.

When six components is over-engineering

This is the honest part, and it is where most architecture posts stop being useful.

These are responsibilities, not folders. A 60-line agent can hold all six inside one function and still be a completely correct agent. If you scaffold six packages, six interfaces and a dependency injection container before your agent has done anything useful once, you have built a diagram, not a system.

Keep it in one function while all of these are true:

  • Five tools or fewer
  • No write actions, or writes that are trivially safe to repeat
  • One user, or one tenant with one set of credentials
  • Runs finish in under about ten steps
  • You have never had to debug the same failure twice

Start splitting the moment any one of these becomes true:

  • A write action exists that must not run twice, which makes execution its own thing
  • Two people need to change the agent in the same week, which makes the seams social as well as technical
  • You cannot say what the agent is trying to do without reading model output, which means planning needs to become data
  • The same class of bug has cost you two afternoons, which means that component has earned its own file
  • Different users have different tools available, which turns the tool layer into infrastructure rather than a dictionary

That last one surprises teams. A hardcoded tool dictionary is fine until your agent serves multiple users with different connected accounts and different permissions.

Before you build

  • You can name which component owns the plan, and print it before the first tool call
  • Reasoning chooses one action at a time, not a whole strategy per pass
  • Every stored fact has a stated lifespan: step, session, user, or system of record
  • Tools are declared in one place with names, argument schemas and descriptions
  • Every tool performs exactly one attempt, and one executor owns retry policy
  • Write actions carry an idempotency key generated once per logical action
  • There is a hard step budget that decrements on failed attempts too
  • At least one postcondition check exists after every write action
  • Your logs say which component produced each line
  • You can look at a failure and name the component before you open the code

Get those right and your agent stops being a box. It becomes a system with parts, and parts can be fixed.

Day 07 of 30 in the MewCP series on AI agents. From here the series goes inward, opening each of these six components one at a time.

ContentsAugust 17, 2026
  1. Why "the agent is broken" is not a bug report
  2. The six components of an AI agent architecture
  3. Six seams in one small agent
  4. The three merges that cause real bugs
  5. Evaluation is the one almost everyone skips
  6. The full symptom to component map
  7. When six components is over-engineering
  8. Before you build
Author

Rohit Gite, Co-Founder CTO @MewCP

Share

Build with MewCP

Connect your AI agents to real tools in minutes.

Get started