MewCP LogoAStheTech
MCPs
Use Cases

Use cases by category

Productivity & InboxInbox, calendar, and daily flowEngineering & DevOpsShip, debug, and run on-callSales & CRMPipeline, outreach, and dealsMarketing & GrowthCampaigns, SEO, and growthSupport & SuccessTriage tickets, keep customers happyFinance & OpsClose, reconcile, and expensesCreative & ContentGenerate assets and contentPeople & HiringHiring, onboarding, and HRResearch & DataSynthesize data and insights
See all use cases
Resources
BlogsProduct updates and storiesArticlesIntegration guides and code examples
PricingDocsSign in
Back to home
MewCP Logo

Infrastructure You Can Trust for Agentic Products

X

Categories

  • Productivity & Docs
  • Developer Tools
  • CRM & Sales
  • Finance & Commerce
  • Data & Analytics
  • Marketing & SEO
  • Search & Web
  • Communication
  • View All Servers →

Resources

  • Blog
  • Docs
  • Privacy Policy
  • Terms of Service

Blogs

  • View All Blogs →

Articles

  • View All Articles →
Browse Servers|Pricing|Contact

Browse by Category

Productivity & Docs

  • Gmail
  • Google Drive
  • YouTube
  • Google Calendar
  • Google People
  • Google Classroom
  • Notion
  • ClickUp
  • Figma
  • Google Tasks
  • Cal
  • Monday
  • Luma
  • Notion MCP
  • Mem MCP
  • Linear MCP
  • Calendly MCP
  • Consensus MCP
  • Craft MCP
  • Close MCP
  • Dice MCP
  • Lumin PDF MCP
  • Develop21 MCP
  • Granola MCP
  • Lucid MCP
  • Mermaid Chart MCP
  • Fireflies MCP
  • ClickUp MCP
  • Miro MCP
  • Llamaindex MCP
  • Otter MCP
  • Mobbin MCP
  • Descript MCP

Developer Tools

  • Gemini
  • Veo
  • ClickUp
  • Firecrawl
  • Vercel
  • Apify
  • Github
  • Chef
  • Scientific Calculator
  • Figma
  • HTTP
  • Perplexity
  • Apify MCP
  • Hugging Face Hub MCP
  • Buildkite MCP
  • Cloudflare MCP
  • Context7 MCP
  • Ahrefs MCP
  • Sentry MCP
  • Brevo Docs MCP
  • X Docs MCP
  • Jev
  • Linear MCP
  • Calendly MCP
  • Craft MCP
  • DeepWiki MCP
  • Inspo MCP
  • Kernel MCP
  • Malwarebytes MCP
  • Mermaid Chart MCP
  • Supabase MCP
  • Microsoft Learn MCP
  • Webflow MCP
  • Scalar Docs MCP
  • Oneuptime MCP
  • Redocly MCP
  • Reducto Docs MCP
  • Llamaindex Docs MCP
  • B12 MCP
  • Lucid Docs MCP
  • Airwallex Docs MCP
  • Langfuse Docs MCP
  • Glen Docs MCP
  • AgentMail
  • Gogs Docs MCP
  • Netlify MCP
  • Neon MCP
  • Minlify Admin MCP
  • Mintlify Index MCP
  • Fern Docs MCP
  • Greptile MCP

CRM & Sales

  • Google People
  • OneSignal MCP
  • Brevo Docs MCP
  • Brevo
  • Brevo MCP
  • Carbon Voice MCP
  • Clay MCP
  • Close MCP
  • Attio MCP
  • Clarify MCP
  • Hunter.io
  • Plain MCP
  • Modem MCP

Finance & Commerce

  • Kite
  • Razorpay
  • Polymarket
  • Stripe
  • Binance
  • Upstox
  • Aiwyn MCP
  • Era-Context-MCP
  • Granted MCP
  • XDC AI MCP
  • Agentery MCP
  • Agent Embassy
  • Quick Commerce MCP
  • Longbridge MCP
  • Mercury MCP
  • Blockscout MCP
  • Octagon AI MCP

Data & Analytics

  • Apify MCP
  • Cloudflare MCP
  • Ahrefs MCP
  • Candid MCP
  • Consensus MCP
  • Contentsquare MCP
  • Era-Context-MCP
  • Instinct MCP
  • legal Data Hunter MCP
  • Marcopolo MCP
  • Mixpanel MCP
  • MOSPI MCP
  • Hex MCP
  • OpenRevenue MCP

Marketing & SEO

  • YouTube
  • Google Business
  • Mailchimp
  • Google Search Console
  • OneSignal MCP
  • Cloudflare MCP
  • Brevo Docs MCP
  • Brevo
  • Brevo MCP
  • AirOps MCP
  • Clay MCP
  • Contentsquare MCP
  • Reelsmith MCP
  • GoDaddy MCP
  • Metricool MCP
  • Webflow MCP
  • Windsor MCP
  • Commonroom MCP
  • B12 MCP
  • Hunter.io
  • Get MCP Ads
  • Minlify Admin MCP

Search & Web

  • Web Scrapper
  • Firecrawl
  • Apify
  • Perplexity
  • Context.dev
  • Exa
  • Brave Search
  • Apify MCP
  • Ahrefs MCP
  • DeepWiki MCP
  • Dice MCP
  • GoDaddy MCP
  • Granted MCP
  • Microsoft Learn MCP
  • Viator MCp
  • Scholargateway MCP
  • Parallel MCP
  • Mintlify Index MCP

Communication

  • Gmail
  • Google Meet
  • Google Calendar
  • Mailchimp
  • WhatsApp
  • Slack
  • OneSignal MCP
  • Brevo Docs MCP
  • Carbon Voice MCP
  • Hunter.io
  • Outlook
  • Nylas MCP

© 2026 MewCP. All rights reserved.

  1. Home
  2. Blogs
  3. The Production AI Agent Stack: A Complete Reference Architecture

The Production AI Agent Stack: A Complete Reference Architecture

by Rohit Gite, Founder @MewCP·October 9, 2026·16 min read

Five layers on the request path, five concerns that run through every layer, and one matrix that shows where each concern is enforced. The whole production agent architecture on one page.

The Production AI Agent Stack: A Complete Reference Architecture

Every agent architecture diagram starts with two boxes: an LLM and some tools, joined by an arrow. That diagram is correct. It is also the architecture of a demo, and nearly everything that makes an agent safe to give to real users lives in the white space around those two boxes.

This series has spent twenty-eight days filling in that white space one concept at a time. This post puts it all on one page.

The architecture has two dimensions, and seeing both is the point:

  • Five layers on the request path, top to bottom: the AI application, the agent, the MCP gateway, the MCP servers, and the tools and services they reach.
  • Five concerns that are not layers at all: authentication, credential management, multi-tenancy, observability and evaluation. Each one cuts across every layer and has to hold at every level.

The carousel drew this as a building: five floors and five pillars running through all of them. The image is useful because it makes the most common production failure obvious. A concern almost never fails everywhere at once. It holds on one floor and quietly breaks on another, and the break is invisible until something falls through it.

This post covers the full reference diagram, what each layer owns, the contract at each boundary between layers, a 25-cell enforcement matrix, the life of one request through the whole stack, the trust boundaries, a deployment topology, failure containment, a staged rollout from pilot to multi-tenant production, how to turn the matrix into tests, and where MewCP maps onto the diagram.

The reference diagram

                         AUTH    CREDS    TENANCY   OBSERV.   EVALS
                          │        │         │         │        │
 ┌────────────────────────┼────────┼─────────┼─────────┼────────┼──┐
 │ 1  AI APPLICATION      ●        ·         ●         ●        ●  │  users, sessions, product UI
 ├────────────────────────┼────────┼─────────┼─────────┼────────┼──┤
 │ 2  AGENT               ●        ·         ●         ●        ●  │  plans, remembers, chooses tools
 ├────────────────────────┼────────┼─────────┼─────────┼────────┼──┤
 │ 3  MCP GATEWAY         ●        ●         ●         ●        ●  │  one policy point for every call
 ├────────────────────────┼────────┼─────────┼─────────┼────────┼──┤
 │ 4  MCP SERVERS         ●        ●         ●         ●        ●  │  hosted or self-hosted runtimes
 ├────────────────────────┼────────┼─────────┼─────────┼────────┼──┤
 │ 5  TOOLS & SERVICES    ●        ●         ●         ●        ·  │  the APIs that do the work
 └────────────────────────┴────────┴─────────┴─────────┴────────┴──┘
 
   ●  the concern is actively enforced at this layer
   ·  the concern must be ABSENT here (e.g. no upstream credentials in the app or agent)

The dots are not decoration. Two of the most important cells in the diagram are the ones marked absent: upstream credentials must not exist in the application layer or the agent layer at all. A concern's absence can be just as deliberate as its enforcement.

The five layers

1. AI application

The product your users touch: a web app, a chat surface, an API, a scheduled job, a Slack bot. It owns user sessions, product UI and the decision to start an agent run.

What crosses into the next layer: an authenticated identity (user, tenant, scopes), the task, and any product context the agent needs. What does not cross: raw credentials for anything.

2. Agent

The reasoning loop. It owns planning, memory, model calls, tool selection, step limits and stop conditions. Days 3 to 12 of the series live here.

What crosses into the next layer: MCP tool calls (a tool name and arguments) plus a token proving which user and tenant the run is acting for. What does not cross: tenant ids in tool arguments, upstream tokens, or anything the model invented about who it is.

3. MCP gateway

The single entry point for every tool call. It owns the policy pipeline: verify identity, resolve tenant, filter tool access, validate arguments, rate limit, resolve and inject credentials, dispatch with a timeout, and record a trace. Day 27 described this layer as "the infrastructure layer."

What crosses into the next layer: a validated call, already scoped and authorized, with the upstream credential attached for this call only.

4. MCP servers

The runtimes that implement tools. Some are your own domain servers wrapping your databases and internal APIs. Others are hosted servers wrapping common third-party APIs. Day 28 covered when to run which. They speak MCP to the gateway and each upstream's own protocol to the service behind them.

5. Tools and services

The real systems: Gmail, Slack, GitHub, your billing database, your deployment pipeline. They enforce their own authentication and their own rate limits, which your stack has to respect rather than fight.

The contracts between layers

Most architecture bugs live at boundaries, not inside boxes. For each boundary, write down what must cross and what must never cross.

BoundaryMust crossMust never cross
User → ApplicationA real login: session or token from an identity providerTrust in anything the client says about who it is
Application → AgentUser id, tenant id, scopes, run id, taskUpstream credentials; mutable identity
Agent → GatewayTool name, arguments, an access token bound to the gateway's audienceTenant id or user id as tool arguments; upstream tokens
Gateway → MCP serverValidated arguments, normalized identity, an injected per-call credential, an idempotency keyThe agent's own access token passed through unchanged
MCP server → UpstreamThe upstream credential for this user, this providerAny other user's credential; unbounded retries
Results flowing back upNormalized results and categorized errorsCredentials; another tenant's data; raw upstream error dumps

The third and fourth rows are the ones teams get wrong most often. If the model can put a tenant id in an argument, isolation depends on the model behaving. If a server forwards the token it received to an upstream API, it is doing token passthrough, which the MCP authorization guidance explicitly forbids because it breaks audience validation and creates a confused deputy.

The five cross-cutting concerns

Authentication answers "who is this?" It begins at the application with a real login. It continues at the gateway, where the agent's token is checked for signature, expiry and audience. It applies again at every upstream, which authenticates your server's credential.

Credential management answers "which secret acts for this user, at this upstream?" It is enforced at the gateway and the MCP servers, where tokens are stored, refreshed, injected and revoked. It is deliberately absent from the application and the agent.

Multi-tenancy answers "which customer's world may this request touch?" The application knows the tenant. The agent's memory filters by it. The gateway derives it from the token, never from arguments. The servers resolve credentials and data by it. The upstreams are reached with that tenant's accounts.

Observability answers "what happened?" Every layer emits spans into one trace per run: the application starts the run, the agent records model turns and decisions, the gateway records policy decisions and tool calls, the servers record upstream calls and their status.

Evaluation answers "does it still work?" It is enforced mostly offline: suites run in CI against the application and agent behaviour, gateway and server contracts are tested as part of those suites, and production traces are sampled into review queues. Upstream services are not evaluated by you; they are stubbed in evaluation runs.

The enforcement matrix

Here is the full 25-cell version. Every cell is a question you should be able to answer with a specific mechanism, not "the model handles it."

AuthenticationCredentialsMulti-tenancyObservabilityEvaluation
ApplicationLogin through an identity provider; session to tokenNone heldTenant resolved from login; tenant-scoped UI and dataRun started, user action, outcomeEnd-to-end task suites
AgentReceives immutable identity; never constructs itNone held, none visible to the modelMemory reads and writes filtered by tenant; no tenant in tool argsModel turns, plans, tool choices, cost per stepTool selection, plan quality, forbidden-tool cases
GatewayValidate token signature, expiry, audienceVault lookup, refresh under lock, inject per callTenant from token; tool list filtered per tenant; per-tenant rate limitsPolicy decision, duration, status, retries per callContract tests: rejects bad args, rejects tenant in args
MCP serversAccept only gateway-issued or audience-bound tokens

If you print one thing from this post, print this table, and put a name next to every cell.

One request, end to end

The carousel used the request "Draft replies to today's urgent emails." Here it is through the whole stack, with the enforcement point for each concern marked.

A zigzag path of yellow light descending through five glass floors with six numbered nodes

  1. The application authenticates the user. Maya at Acme is signed in through the identity provider. The app creates a run with an immutable context: tenant=acme, user=maya, scopes=[mail.read, mail.draft], run=r_7f3. Authentication and tenancy start here. Observability opens the run's root span.

  2. The agent plans and discovers tools. The model reads the task and plans: find today's urgent emails, read each one, draft a reply for each. It searches the gateway's catalogue for mail tools and fetches their schemas. Memory is consulted for Maya's preferences, filtered to Acme and Maya. Tenancy is enforced in the memory read. Observability records the plan and each model turn.

  3. The gateway checks the call. The agent calls a mail search tool with query="is:unread label:urgent newer_than:1d". The gateway validates the access token, reads from it, confirms Acme has mail connected and Maya's scopes allow reading, validates the arguments and checks three rate-limit buckets.

Notice how little of that list is the model. It wrote a plan, chose tools, chose arguments and wrote drafts. Every other decision was made by the stack.

Trust boundaries

There are three trust boundaries in the diagram, and each one changes what you may assume.

Boundary A: between the user and the application. Everything from the client is untrusted until authenticated. This is ordinary web security.

Boundary B: between the model and everything else. This is the one specific to agents. The model reads untrusted text: emails, web pages, documents, tool results. Any of it can contain instructions. So the model's outputs are treated like user input: tool names are checked against an allowlist for this caller, arguments are validated against schemas, and identity is never taken from the model. The carousel framed this as "what the model never sees": tokens, tenant ids and other customers' data stay below this boundary.

Boundary C: between your stack and each upstream. Upstreams are trusted to do what their API says and nothing more. Their responses are data, not instructions, even when they contain text that looks like instructions.

A frosted glass plane with a glowing yellow edge separating what the model sees above from what the stack holds below

A useful property of this architecture is that boundary B is enforced structurally rather than by prompting. You do not ask the model to please not use another tenant's id. You build tool schemas that have no field for it and a gateway that ignores everything except the validated token.

Deployment topology

A reasonable production topology for a small team serving many customers:

                   ┌──────────────────────────────┐
  users ──HTTPS──▶ │ APPLICATION (stateless, N×)  │ ──▶ Identity provider
                   └──────────────┬───────────────┘
                                  │ enqueue run
                   ┌──────────────▼───────────────┐
                   │ RUN QUEUE + DURABLE RUN STORE│  (run state, checkpoints, approvals)
                   └──────────────┬───────────────┘
                                  │
                   ┌──────────────▼───────────────┐
                   │ AGENT WORKERS (stateless, N×)│ ──▶ Model APIs
                   └──────────────┬───────────────┘
                                  │ MCP over Streamable HTTP
                   ┌──────────────▼───────────────┐
                   │ MCP GATEWAY (stateless, N×)  │ ──▶ Credential vault
                   │                              │ ──▶ Shared rate-limit store
                   └───────┬──────────────┬───────┘
                           │              │







Three decisions are encoded in that picture.

Application, agent workers and gateway are stateless. As Day 22 argued, stateless tiers scale horizontally and deploy safely. All state that must survive lives in explicit stores: the durable run store for run progress and pending approvals, the vault for credentials, the shared store for rate limits.

Runs are queued, not handled inline. Agent runs can take minutes. Queuing decouples the user's request from the run's lifetime, lets runs survive worker restarts, and gives you a place to apply per-tenant concurrency limits.

Evaluation is fed by telemetry. The same traces that answer "what happened" become the raw material for "does it still work," which keeps production behaviour and evaluation data from drifting apart.

Failure containment

A production stack is designed so that one failure stays in one place.

  • Per-upstream bulkheads. Each upstream gets its own concurrency pool and its own circuit breaker at the gateway. When Slack has a bad afternoon, calls to Slack fail fast with a retriable error, and calls to GitHub are unaffected.
  • Per-tenant budgets. Rate limits and run concurrency are enforced per tenant, so one customer's runaway loop cannot degrade every other customer's latency.
  • Timeouts at every hop. The application has a timeout on the run, the agent has a step limit and a wall-clock limit, and the gateway has a timeout on every dispatch. No single hung call can hold a worker indefinitely.
  • Idempotent writes. Write tools receive idempotency keys derived from the run and step, so a retried step cannot duplicate a side effect.
  • Human approval on irreversible actions. Sending, paying, deleting and deploying go through an approval step stored in the durable run store, so a paused run survives a deploy and resumes when approved.
  • Categorized errors. The gateway normalizes upstream failures into retriable, needs user action, permission denied and invalid input, so the agent can respond sensibly instead of guessing.

A staged rollout

You do not need every cell of the matrix on day one. You do need to know which cells you are deferring and what will force you to fill them.

Stage 0: pilot (one team, internal). Application with real login. Agent with step limits. Tools behind a thin gateway, even if the gateway is a single function. Credentials per user, never shared. Basic traces. A starter evaluation suite of 20 to 50 cases. Defer: multi-tenancy beyond one tenant, durable runs, per-upstream bulkheads.

Stage 1: first external customers. Tenant identity everywhere, with tenant_id absent from all tool schemas. Per-tenant memory and vector namespaces. Rate limits per tenant and per upstream. Approval flow for irreversible actions. Evaluation in CI, with every case run several times. Defer: sophisticated routing, advanced evaluation such as LLM judges with calibrated rubrics.

Stage 2: multi-tenant production. Durable run store and queued execution. Circuit breakers per upstream. Cost attribution per tenant. Tenant-scoped payload storage with redaction. Sampled production evaluation and a regression gate on releases. An access review process for write tools. Every cell of the matrix has an owner.

The common mistake is not deferring things. It is deferring them without writing down what you deferred, so nobody notices when the trigger for building them arrives.

Turning the matrix into tests

A matrix that lives only in a document drifts. The cheapest way to keep it honest is to turn the most important cells into automated tests that run in CI. A few that pay for themselves quickly:

import { describe, it, expect } from "vitest";
 
describe("enforcement matrix", () => {
  it("gateway rejects tenant identity supplied as a tool argument", async () => {
    const res = await gateway.callTool(tokenFor("acme", "maya"), {
      name: 

































The helpers (gateway, tokenFor, runAgent, exportedSpans, dbQueryLog) are the test harness you build around your own stack; the assertions are the point. The token patterns in the span test are examples of common credential prefixes; extend the list with the formats your own providers use. Each test pins one cell of the matrix so that a regression shows up as a red build, not an incident.

Where MewCP maps onto the diagram

MewCP is built around two floors of this building: the MCP gateway and the MCP servers behind it.

  • Gateway (layer 3). A single MCP endpoint for every connected app, exposing a fixed four-tool surface (search, get_schema, list_accounts, call_tool) so the agent's context does not grow with each integration.
  • MCP servers (layer 4). A hosted catalogue of servers for common third-party services, alongside which you can run your own domain servers.
  • Credential management at layers 3 and 4. Per MewCP's documentation, credentials are encrypted at rest in a vault, decrypted and injected only at the moment a tool is called, never sent to your client, agent or model, and OAuth tokens are refreshed automatically. For B2B products, MewCP's Auth API ties each credential to an external user id you supply.

What MewCP does not do is just as important for drawing the diagram honestly: it is not your application, your agent loop, your memory, your domain tools, your tenancy rules for your own data, or your evaluation suite. Those floors and those cells of the matrix stay with your team, whatever you adopt for the gateway and servers.

The per-layer checklist

AI application

  • Users authenticate through an identity provider
  • Runs start with an immutable context of tenant, user, scopes and run id
  • No upstream credential exists anywhere in this tier
  • Every run has a root span and a user-visible outcome

Agent

  • Step limit and wall-clock limit on every run
  • Memory reads and writes filtered by tenant and user
  • No tool schema contains tenant or user identity
  • Model turns, plans, tool choices and cost are traced
  • An evaluation suite covers tool selection and forbidden-tool cases, with multiple runs per case

MCP gateway

  • Token signature, expiry and audience validated on every call
  • Tenant derived from the token only
  • Tool list filtered per tenant and role
  • Arguments validated, unknown properties rejected
  • Rate limits per tenant, per tenant and tool, and per upstream
  • Credentials resolved per call, refreshed under lock, never logged
  • Every dispatch has a timeout and, for writes, an idempotency key
  • Policy decisions and outcomes traced for every call, including rejections
  • Upstream errors normalized into a small set of categories

MCP servers

  • Accept only tokens intended for them; no token passthrough
  • Every data access scoped to the resolved tenant
  • Injected credentials used once, never stored or logged
  • Upstream calls traced with status and latency
  • Tested with stubbed upstreams

Tools and services

  • OAuth scopes are the minimum each tool needs
  • Irreversible actions require human approval
  • Upstream request ids recorded for support cases
  • Stubbed in evaluation runs

Across all layers

  • One trace per run, end to end
  • Every cell of the enforcement matrix has an owner
  • The most important cells are automated tests in CI
  • Deferred cells are written down, with the trigger that will force them

Where this leaves you

The production agent stack is not a secret architecture. It is five ordinary layers and five ordinary concerns, arranged so that every concern is enforced at every layer it touches, and deliberately absent from the layers it should not touch. The work is in the cells, and in making sure no cell is assumed.

Draw the architecture first. Fill in the matrix. Then decide which floors you build and which you adopt, knowing exactly what each choice covers and what it leaves to you.

Tomorrow is the final day of the series: the whole thirty-day journey, from a prompt to a production agent, on a single path.

ContentsOctober 9, 2026
  1. The reference diagram
  2. The five layers
  3. The contracts between layers
  4. The five cross-cutting concerns
  5. The enforcement matrix
  6. One request, end to end
  7. Trust boundaries
  8. Deployment topology
  9. Failure containment
  10. A staged rollout
  11. Turning the matrix into tests
  12. Where MewCP maps onto the diagram
  13. The per-layer checklist
  14. Where this leaves you
Author

Rohit Gite, Founder @MewCP

Share

Build with MewCP

Connect your AI agents to real tools in minutes.

Get started
Use injected credential once; never log it
Every query scoped to the resolved tenant
Upstream call, status, latency
Server tests with stubbed upstreams
Tools & servicesUpstream verifies your credentialPer-user OAuth tokens with minimal scopesReached with this tenant's account onlyUpstream request ids recorded for supportNot evaluated; stubbed in eval mode
tenant=acme
Authentication, tenancy and rate limits are all enforced here.
  • The credential is injected at call time. The gateway resolves Maya's Gmail token from the vault, refreshes it if it is near expiry, and attaches it to this single dispatch. The model never sees it. Credential management happens here and nowhere above.

  • The Gmail server executes. The server calls Gmail with Maya's token and returns a normalized list of messages. The agent reads each message, then calls a create-draft tool for each reply. Because the scopes include mail.draft but not mail.send, the agent can only draft; sending stays a human action. Upstream enforces its own auth. Least privilege keeps the irreversible action with the user.

  • The trace is recorded and sampled. Every step shares one trace id: model turns, tool calls, policy decisions, durations and cost. A sample of runs like this one is copied into an evaluation review queue, where reviewers grade draft quality and the run's cost is tracked against Acme's budget. Observability closes the run. Evaluation picks it up.

  • ┌─────────────▼───┐ ┌──────▼──────────────┐
    │ YOUR MCP SERVERS│ │ HOSTED MCP SERVERS │
    │ (your network) │ │ (provider) │
    └────────┬────────┘ └──────────┬──────────┘
    ▼ ▼
    your databases third-party APIs
    All tiers ──spans──▶ TELEMETRY PIPELINE ──sampled runs──▶ EVALUATION
    "list_invoices"
    ,
    arguments: { status: "overdue", tenant_id: "globex" },
    });
    expect(res.isError).toBe(true);
    expect(res.category).toBe("invalid_input");
    });
    it("gateway ignores the agent's token audience when it is not the gateway", async () => {
    const res = await gateway.callTool(tokenFor("acme", "maya", { audience: "some-other-api" }), {
    name: "list_invoices",
    arguments: { status: "overdue" },
    });
    expect(res.category).toBe("permission_denied");
    });
    it("tool list for a tenant excludes tools that tenant has not enabled", async () => {
    const tools = await gateway.listTools(tokenFor("globex", "sam"));
    expect(tools.map((t) => t.name)).not.toContain("issue_refund");
    });
    it("no credential ever appears in exported spans", async () => {
    await runAgent(tokenFor("acme", "maya"), "Summarise my unread email");
    for (const span of exportedSpans()) {
    expect(JSON.stringify(span.attributes)).not.toMatch(/ya29\.|xox[bp]-|ghp_/);
    }
    });
    it("two tenants with the same prompt touch only their own rows", async () => {
    await Promise.all([
    runAgent(tokenFor("acme", "maya"), "Summarise our overdue invoices"),
    runAgent(tokenFor("globex", "sam"), "Summarise our overdue invoices"),
    ]);
    expect(dbQueryLog().every((q) => q.tenant === q.requestedBy.tenant)).toBe(true);
    });
    });