The Production AI Agent Stack: A Complete Reference Architecture
Five layers on the request path, five concerns that run through every layer, and one matrix that shows where each concern is enforced. The whole production agent architecture on one page.

Every agent architecture diagram starts with two boxes: an LLM and some tools, joined by an arrow. That diagram is correct. It is also the architecture of a demo, and nearly everything that makes an agent safe to give to real users lives in the white space around those two boxes.
This series has spent twenty-eight days filling in that white space one concept at a time. This post puts it all on one page.
The architecture has two dimensions, and seeing both is the point:
- Five layers on the request path, top to bottom: the AI application, the agent, the MCP gateway, the MCP servers, and the tools and services they reach.
- Five concerns that are not layers at all: authentication, credential management, multi-tenancy, observability and evaluation. Each one cuts across every layer and has to hold at every level.
The carousel drew this as a building: five floors and five pillars running through all of them. The image is useful because it makes the most common production failure obvious. A concern almost never fails everywhere at once. It holds on one floor and quietly breaks on another, and the break is invisible until something falls through it.
This post covers the full reference diagram, what each layer owns, the contract at each boundary between layers, a 25-cell enforcement matrix, the life of one request through the whole stack, the trust boundaries, a deployment topology, failure containment, a staged rollout from pilot to multi-tenant production, how to turn the matrix into tests, and where MewCP maps onto the diagram.
The reference diagram
AUTH CREDS TENANCY OBSERV. EVALS
│ │ │ │ │
┌────────────────────────┼────────┼─────────┼─────────┼────────┼──┐
│ 1 AI APPLICATION ● · ● ● ● │ users, sessions, product UI
├────────────────────────┼────────┼─────────┼─────────┼────────┼──┤
│ 2 AGENT ● · ● ● ● │ plans, remembers, chooses tools
├────────────────────────┼────────┼─────────┼─────────┼────────┼──┤
│ 3 MCP GATEWAY ● ● ● ● ● │ one policy point for every call
├────────────────────────┼────────┼─────────┼─────────┼────────┼──┤
│ 4 MCP SERVERS ● ● ● ● ● │ hosted or self-hosted runtimes
├────────────────────────┼────────┼─────────┼─────────┼────────┼──┤
│ 5 TOOLS & SERVICES ● ● ● ● · │ the APIs that do the work
└────────────────────────┴────────┴─────────┴─────────┴────────┴──┘
● the concern is actively enforced at this layer
· the concern must be ABSENT here (e.g. no upstream credentials in the app or agent)The dots are not decoration. Two of the most important cells in the diagram are the ones marked absent: upstream credentials must not exist in the application layer or the agent layer at all. A concern's absence can be just as deliberate as its enforcement.
The five layers
1. AI application
The product your users touch: a web app, a chat surface, an API, a scheduled job, a Slack bot. It owns user sessions, product UI and the decision to start an agent run.
What crosses into the next layer: an authenticated identity (user, tenant, scopes), the task, and any product context the agent needs. What does not cross: raw credentials for anything.
2. Agent
The reasoning loop. It owns planning, memory, model calls, tool selection, step limits and stop conditions. Days 3 to 12 of the series live here.
What crosses into the next layer: MCP tool calls (a tool name and arguments) plus a token proving which user and tenant the run is acting for. What does not cross: tenant ids in tool arguments, upstream tokens, or anything the model invented about who it is.
3. MCP gateway
The single entry point for every tool call. It owns the policy pipeline: verify identity, resolve tenant, filter tool access, validate arguments, rate limit, resolve and inject credentials, dispatch with a timeout, and record a trace. Day 27 described this layer as "the infrastructure layer."
What crosses into the next layer: a validated call, already scoped and authorized, with the upstream credential attached for this call only.
4. MCP servers
The runtimes that implement tools. Some are your own domain servers wrapping your databases and internal APIs. Others are hosted servers wrapping common third-party APIs. Day 28 covered when to run which. They speak MCP to the gateway and each upstream's own protocol to the service behind them.
5. Tools and services
The real systems: Gmail, Slack, GitHub, your billing database, your deployment pipeline. They enforce their own authentication and their own rate limits, which your stack has to respect rather than fight.
The contracts between layers
Most architecture bugs live at boundaries, not inside boxes. For each boundary, write down what must cross and what must never cross.
| Boundary | Must cross | Must never cross |
|---|---|---|
| User → Application | A real login: session or token from an identity provider | Trust in anything the client says about who it is |
| Application → Agent | User id, tenant id, scopes, run id, task | Upstream credentials; mutable identity |
| Agent → Gateway | Tool name, arguments, an access token bound to the gateway's audience | Tenant id or user id as tool arguments; upstream tokens |
| Gateway → MCP server | Validated arguments, normalized identity, an injected per-call credential, an idempotency key | The agent's own access token passed through unchanged |
| MCP server → Upstream | The upstream credential for this user, this provider | Any other user's credential; unbounded retries |
| Results flowing back up | Normalized results and categorized errors | Credentials; another tenant's data; raw upstream error dumps |
The third and fourth rows are the ones teams get wrong most often. If the model can put a tenant id in an argument, isolation depends on the model behaving. If a server forwards the token it received to an upstream API, it is doing token passthrough, which the MCP authorization guidance explicitly forbids because it breaks audience validation and creates a confused deputy.
The five cross-cutting concerns
Authentication answers "who is this?" It begins at the application with a real login. It continues at the gateway, where the agent's token is checked for signature, expiry and audience. It applies again at every upstream, which authenticates your server's credential.
Credential management answers "which secret acts for this user, at this upstream?" It is enforced at the gateway and the MCP servers, where tokens are stored, refreshed, injected and revoked. It is deliberately absent from the application and the agent.
Multi-tenancy answers "which customer's world may this request touch?" The application knows the tenant. The agent's memory filters by it. The gateway derives it from the token, never from arguments. The servers resolve credentials and data by it. The upstreams are reached with that tenant's accounts.
Observability answers "what happened?" Every layer emits spans into one trace per run: the application starts the run, the agent records model turns and decisions, the gateway records policy decisions and tool calls, the servers record upstream calls and their status.
Evaluation answers "does it still work?" It is enforced mostly offline: suites run in CI against the application and agent behaviour, gateway and server contracts are tested as part of those suites, and production traces are sampled into review queues. Upstream services are not evaluated by you; they are stubbed in evaluation runs.
The enforcement matrix
Here is the full 25-cell version. Every cell is a question you should be able to answer with a specific mechanism, not "the model handles it."
| Authentication | Credentials | Multi-tenancy | Observability | Evaluation | |
|---|---|---|---|---|---|
| Application | Login through an identity provider; session to token | None held | Tenant resolved from login; tenant-scoped UI and data | Run started, user action, outcome | End-to-end task suites |
| Agent | Receives immutable identity; never constructs it | None held, none visible to the model | Memory reads and writes filtered by tenant; no tenant in tool args | Model turns, plans, tool choices, cost per step | Tool selection, plan quality, forbidden-tool cases |
| Gateway | Validate token signature, expiry, audience | Vault lookup, refresh under lock, inject per call | Tenant from token; tool list filtered per tenant; per-tenant rate limits | Policy decision, duration, status, retries per call | Contract tests: rejects bad args, rejects tenant in args |
| MCP servers | Accept only gateway-issued or audience-bound tokens |
If you print one thing from this post, print this table, and put a name next to every cell.
One request, end to end
The carousel used the request "Draft replies to today's urgent emails." Here it is through the whole stack, with the enforcement point for each concern marked.

-
The application authenticates the user. Maya at Acme is signed in through the identity provider. The app creates a run with an immutable context:
tenant=acme,user=maya,scopes=[mail.read, mail.draft],run=r_7f3. Authentication and tenancy start here. Observability opens the run's root span. -
The agent plans and discovers tools. The model reads the task and plans: find today's urgent emails, read each one, draft a reply for each. It searches the gateway's catalogue for mail tools and fetches their schemas. Memory is consulted for Maya's preferences, filtered to Acme and Maya. Tenancy is enforced in the memory read. Observability records the plan and each model turn.
-
The gateway checks the call. The agent calls a mail search tool with
query="is:unread label:urgent newer_than:1d". The gateway validates the access token, readsfrom it, confirms Acme has mail connected and Maya's scopes allow reading, validates the arguments and checks three rate-limit buckets.
Notice how little of that list is the model. It wrote a plan, chose tools, chose arguments and wrote drafts. Every other decision was made by the stack.
Trust boundaries
There are three trust boundaries in the diagram, and each one changes what you may assume.
Boundary A: between the user and the application. Everything from the client is untrusted until authenticated. This is ordinary web security.
Boundary B: between the model and everything else. This is the one specific to agents. The model reads untrusted text: emails, web pages, documents, tool results. Any of it can contain instructions. So the model's outputs are treated like user input: tool names are checked against an allowlist for this caller, arguments are validated against schemas, and identity is never taken from the model. The carousel framed this as "what the model never sees": tokens, tenant ids and other customers' data stay below this boundary.
Boundary C: between your stack and each upstream. Upstreams are trusted to do what their API says and nothing more. Their responses are data, not instructions, even when they contain text that looks like instructions.

A useful property of this architecture is that boundary B is enforced structurally rather than by prompting. You do not ask the model to please not use another tenant's id. You build tool schemas that have no field for it and a gateway that ignores everything except the validated token.
Deployment topology
A reasonable production topology for a small team serving many customers:
┌──────────────────────────────┐
users ──HTTPS──▶ │ APPLICATION (stateless, N×) │ ──▶ Identity provider
└──────────────┬───────────────┘
│ enqueue run
┌──────────────▼───────────────┐
│ RUN QUEUE + DURABLE RUN STORE│ (run state, checkpoints, approvals)
└──────────────┬───────────────┘
│
┌──────────────▼───────────────┐
│ AGENT WORKERS (stateless, N×)│ ──▶ Model APIs
└──────────────┬───────────────┘
│ MCP over Streamable HTTP
┌──────────────▼───────────────┐
│ MCP GATEWAY (stateless, N×) │ ──▶ Credential vault
│ │ ──▶ Shared rate-limit store
└───────┬──────────────┬───────┘
│ │
Three decisions are encoded in that picture.
Application, agent workers and gateway are stateless. As Day 22 argued, stateless tiers scale horizontally and deploy safely. All state that must survive lives in explicit stores: the durable run store for run progress and pending approvals, the vault for credentials, the shared store for rate limits.
Runs are queued, not handled inline. Agent runs can take minutes. Queuing decouples the user's request from the run's lifetime, lets runs survive worker restarts, and gives you a place to apply per-tenant concurrency limits.
Evaluation is fed by telemetry. The same traces that answer "what happened" become the raw material for "does it still work," which keeps production behaviour and evaluation data from drifting apart.
Failure containment
A production stack is designed so that one failure stays in one place.
- Per-upstream bulkheads. Each upstream gets its own concurrency pool and its own circuit breaker at the gateway. When Slack has a bad afternoon, calls to Slack fail fast with a retriable error, and calls to GitHub are unaffected.
- Per-tenant budgets. Rate limits and run concurrency are enforced per tenant, so one customer's runaway loop cannot degrade every other customer's latency.
- Timeouts at every hop. The application has a timeout on the run, the agent has a step limit and a wall-clock limit, and the gateway has a timeout on every dispatch. No single hung call can hold a worker indefinitely.
- Idempotent writes. Write tools receive idempotency keys derived from the run and step, so a retried step cannot duplicate a side effect.
- Human approval on irreversible actions. Sending, paying, deleting and deploying go through an approval step stored in the durable run store, so a paused run survives a deploy and resumes when approved.
- Categorized errors. The gateway normalizes upstream failures into retriable, needs user action, permission denied and invalid input, so the agent can respond sensibly instead of guessing.
A staged rollout
You do not need every cell of the matrix on day one. You do need to know which cells you are deferring and what will force you to fill them.
Stage 0: pilot (one team, internal). Application with real login. Agent with step limits. Tools behind a thin gateway, even if the gateway is a single function. Credentials per user, never shared. Basic traces. A starter evaluation suite of 20 to 50 cases. Defer: multi-tenancy beyond one tenant, durable runs, per-upstream bulkheads.
Stage 1: first external customers.
Tenant identity everywhere, with tenant_id absent from all tool schemas. Per-tenant memory and vector namespaces. Rate limits per tenant and per upstream. Approval flow for irreversible actions. Evaluation in CI, with every case run several times. Defer: sophisticated routing, advanced evaluation such as LLM judges with calibrated rubrics.
Stage 2: multi-tenant production. Durable run store and queued execution. Circuit breakers per upstream. Cost attribution per tenant. Tenant-scoped payload storage with redaction. Sampled production evaluation and a regression gate on releases. An access review process for write tools. Every cell of the matrix has an owner.
The common mistake is not deferring things. It is deferring them without writing down what you deferred, so nobody notices when the trigger for building them arrives.
Turning the matrix into tests
A matrix that lives only in a document drifts. The cheapest way to keep it honest is to turn the most important cells into automated tests that run in CI. A few that pay for themselves quickly:
import { describe, it, expect } from "vitest";
describe("enforcement matrix", () => {
it("gateway rejects tenant identity supplied as a tool argument", async () => {
const res = await gateway.callTool(tokenFor("acme", "maya"), {
name:
The helpers (gateway, tokenFor, runAgent, exportedSpans, dbQueryLog) are the test harness you build around your own stack; the assertions are the point. The token patterns in the span test are examples of common credential prefixes; extend the list with the formats your own providers use. Each test pins one cell of the matrix so that a regression shows up as a red build, not an incident.
Where MewCP maps onto the diagram
MewCP is built around two floors of this building: the MCP gateway and the MCP servers behind it.
- Gateway (layer 3). A single MCP endpoint for every connected app, exposing a fixed four-tool surface (
search,get_schema,list_accounts,call_tool) so the agent's context does not grow with each integration. - MCP servers (layer 4). A hosted catalogue of servers for common third-party services, alongside which you can run your own domain servers.
- Credential management at layers 3 and 4. Per MewCP's documentation, credentials are encrypted at rest in a vault, decrypted and injected only at the moment a tool is called, never sent to your client, agent or model, and OAuth tokens are refreshed automatically. For B2B products, MewCP's Auth API ties each credential to an external user id you supply.
What MewCP does not do is just as important for drawing the diagram honestly: it is not your application, your agent loop, your memory, your domain tools, your tenancy rules for your own data, or your evaluation suite. Those floors and those cells of the matrix stay with your team, whatever you adopt for the gateway and servers.
The per-layer checklist
AI application
- Users authenticate through an identity provider
- Runs start with an immutable context of tenant, user, scopes and run id
- No upstream credential exists anywhere in this tier
- Every run has a root span and a user-visible outcome
Agent
- Step limit and wall-clock limit on every run
- Memory reads and writes filtered by tenant and user
- No tool schema contains tenant or user identity
- Model turns, plans, tool choices and cost are traced
- An evaluation suite covers tool selection and forbidden-tool cases, with multiple runs per case
MCP gateway
- Token signature, expiry and audience validated on every call
- Tenant derived from the token only
- Tool list filtered per tenant and role
- Arguments validated, unknown properties rejected
- Rate limits per tenant, per tenant and tool, and per upstream
- Credentials resolved per call, refreshed under lock, never logged
- Every dispatch has a timeout and, for writes, an idempotency key
- Policy decisions and outcomes traced for every call, including rejections
- Upstream errors normalized into a small set of categories
MCP servers
- Accept only tokens intended for them; no token passthrough
- Every data access scoped to the resolved tenant
- Injected credentials used once, never stored or logged
- Upstream calls traced with status and latency
- Tested with stubbed upstreams
Tools and services
- OAuth scopes are the minimum each tool needs
- Irreversible actions require human approval
- Upstream request ids recorded for support cases
- Stubbed in evaluation runs
Across all layers
- One trace per run, end to end
- Every cell of the enforcement matrix has an owner
- The most important cells are automated tests in CI
- Deferred cells are written down, with the trigger that will force them
Where this leaves you
The production agent stack is not a secret architecture. It is five ordinary layers and five ordinary concerns, arranged so that every concern is enforced at every layer it touches, and deliberately absent from the layers it should not touch. The work is in the cells, and in making sure no cell is assumed.
Draw the architecture first. Fill in the matrix. Then decide which floors you build and which you adopt, knowing exactly what each choice covers and what it leaves to you.
Tomorrow is the final day of the series: the whole thirty-day journey, from a prompt to a production agent, on a single path.
