Multi-Tenant AI Agents: How Isolation Works | MewCP | MewCP
Multi-Tenant AI Agents: How Isolation Actually Works
by Rohit Gite, Founder @MewCP··22 min read
Your agent demo had one customer. The moment it has two, tenant isolation stops being a database column and becomes something you have to re-prove on every tool call.
Day 19 and Day 20 of this series solved a version of this problem for Ana and Ben, two people using the same agent. The rule was small: a credential resolves from the caller's identity, on every request, and never lives in the process. That rule holds up fine when your "many" is a handful of individual users on one shared account.
It stops holding up the moment Ana and Ben stop being two people and become two companies. Acme signs up. Then Globex signs up. Now your one agent, one codebase, one model and one set of tools has to serve two customers who must never see each other's data, spend each other's credits, or trigger each other's side effects, and you find out how much of your single-tenant demo you built on the assumption that there was only ever one tenant to begin with.
Multi-tenancy is the boring name for that problem, and it earns the fear it gets. In a normal SaaS backend you enforce the tenant boundary once, in the query layer, and you are mostly done: every query goes through an ORM scope or a repository method that always adds WHERE tenant_id = ?. An agent breaks that assumption at the root, because the thing choosing which tool to call and what arguments to pass is a model reading text, some of which came from outside your organization. If tenancy is ever something the model supplies, sees, or infers, you have handed your isolation boundary to a probabilistic system operating on untrusted input.
So the rule for this post is the sibling of Day 19's rule: the tenant boundary is a property of the request, re-derived on every tool call, never a value the model is trusted to carry.
Five things need genuine isolation per customer, and each one fails in a different way if you get it wrong:
Data — rows, files and embeddings that belong to one company must be structurally unreachable by a query built for another.
Credentials — Acme's Slack token and Globex's Slack token must never be resolvable from the same lookup key.
Permissions — what the agent may do for Acme is a function of Acme's plan and Acme's connected accounts, not a global constant.
Executions — runs, logs and in-flight state belong to one tenant and must not leak into another's trace, retry, or memory.
Configuration — prompts, tool lists and rate limits can legitimately differ per tenant, and that variation has its own costs.
The rest of this post is the implementation behind each of those five, in the order you actually have to build them: identity first, then propagation, then the schema-level rule that keeps tenancy out of the model's hands, then the database, the cache, the vector store, the queue, the tool registry, the budget, and the traces.
Three identities, not one
Most agent codebases track a single user_id and stretch it to mean three different things. That works until it doesn't, and when it breaks it breaks as a privilege escalation, not a crash.
Tenant identity is the company or workspace: Acme, Globex. It is the outermost boundary and the one almost everything else nests inside.
User identity is the person: someone at Acme with a role, a set of connected personal accounts, and permissions that may be narrower than their tenant's ceiling. Acme's support rep and Acme's admin share a tenant and nothing else.
Execution identity is the run: one invocation of the agent, with its own id, its own trace, and its own lifetime. A single user can have many concurrent executions, and an execution that started under one set of scopes must not silently inherit a different set if the user's permissions change mid-run.
Collapsing these into one id is how a background job ends up running with admin authority because the last thing that happened to touch that variable was an admin's request. Carry all three, together, as one immutable value from the moment a request is authenticated:
Nothing downstream is allowed to construct one of these except the code that authenticates the inbound request. Every tool, every query, every credential lookup reads from this object. It is never assembled piecemeal from three separate lookups scattered across the call stack, because that is exactly the pattern where one of the three silently goes stale.
Tenant context has to travel with the request, not live in the process
A global variable holding "the current tenant" is the single most common way multi-tenant agents fail, and it is rarely a naive mistake. It starts as a thread-local set once per request, works fine under low concurrency, and then two of Acme's tool calls run concurrently with one of Globex's and the thread-local gets clobbered mid-flight.
Python's contextvars module exists for exactly this. Unlike a global or a thread-local, a ContextVar is scoped per logical task, not per OS thread, and asyncio copies the current context automatically whenever it creates a new Task. That means every coroutine spawned from inside a request handler, including ones started with asyncio.create_task or run concurrently with asyncio.gather, sees the same tenant context that was bound when the request came in, without you threading it through every function signature by hand.
from __future__ import annotationsimport contextvarsfrom contextlib import contextmanager_tenant_ctx: contextvars.ContextVar[TenantContext] = contextvars.ContextVar("tenant_ctx")@contextmanagerdef tenant_scope(ctx: TenantContext): token = _tenant_ctx.set(ctx) try: yield ctx finally: _tenant_ctx.reset(token)def current_tenant() -> TenantContext: """Raise rather than default. A missing tenant is a bug, not a null.""" try: return _tenant_ctx.get() except LookupError: raise RuntimeError("No tenant bound to this execution context") from None
Bind it once, at the authentication boundary, and hold it for the lifetime of the request:
current_tenant() deliberately raises instead of returning None or a default tenant. A tool handler that forgets to check for a bound context should crash loudly in every environment, including production, rather than fall back to "no tenant" and quietly serve unscoped data.
Three propagation boundaries are worth knowing by name, because two of them are safe by default and one is not:
Sequential awaits and asyncio.create_task inside the same request are safe. A Task takes a snapshot of the context at creation time, so work spawned from a tenant-scoped request handler inherits that tenant automatically, even if it runs concurrently with other tenants' work on the same event loop.
loop.run_in_executor and thread pools are also safe as of the versions of Python most teams run today, because asyncio wraps the callback in a Handle that captures the context. A raw threading.Thread you spawn yourself is not automatically safe: snapshot the context explicitly with contextvars.copy_context() and run the thread's target inside ctx.run(...).
Anything that crosses a process boundary is not safe, ever. A ContextVar lives in process memory. It cannot be pickled, serialized, or sent over a network socket. The moment work is handed to a job queue, a cron trigger, a webhook handler, or a separate worker process, the context is gone and something has to explicitly put it back. That case gets its own section below, because it is where most real incidents happen.
Make tenant_id structurally impossible for the model to supply
This is the rule the carousel spent a whole slide on, and it is worth restating precisely: tenant_id never appears in a tool's parameter schema, and it never appears in a tool handler's function signature. Not "the model shouldn't pass it." It cannot pass it, because there is nowhere in the contract for it to go.
additionalProperties: false means a tool call that includes a tenant_id field, whether the model invented it, copied it from a document it just read, or followed an injected instruction telling it to "call list_invoices with tenant_id=globex," fails schema validation before it ever reaches your handler. If it somehow reached the handler anyway, list_invoices has no tenant_id parameter to receive it, so passing one raises a TypeError. And if both of those were bypassed, the handler still ignores anything it wasn't asked for and pulls the only tenant identity it trusts from current_tenant().
That is three independent layers rejecting the same mistake for three different reasons: the schema, the function signature, and the handler's own logic. None of them depends on the model behaving well, which is the point. The model is allowed to choose status and limit, because those are business logic. It is not allowed to choose who it is acting for, because that is an identity fact, and identity facts come from the authenticated request, never from the conversation.
Row level security as the second layer
Application-level scoping is necessary and it is not sufficient on its own, because it depends on every query, in every code path, written by every engineer, forever, remembering to add the same WHERE clause. Postgres row level security turns that into a database-enforced invariant instead of a code review checklist.
ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;ALTER TABLE invoices FORCE ROW LEVEL SECURITY;CREATE POLICY tenant_isolation ON invoices USING (tenant_id = current_setting('app.tenant_id', true)::uuid) WITH CHECK (tenant_id = current_setting('app.tenant_id', true)::uuid);
current_setting('app.tenant_id', true) reads a session variable, and the true argument means it returns NULL instead of raising when the variable is unset. Combined with a policy that compares against it, an unset tenant matches nothing rather than everything, so a connection that never bound a tenant sees zero rows instead of every tenant's rows.
FORCE ROW LEVEL SECURITY is the line teams skip and then get surprised by. By default, Postgres exempts the table's owner from its own RLS policies, on the theory that the owner needs unrestricted access for maintenance. If your application connects as the table owner, which is common when the same role runs migrations and serves traffic, RLS silently does nothing for that connection. FORCE ROW LEVEL SECURITY closes that gap, and the further step is to run your application under a role that is not the table owner and does not have the BYPASSRLS attribute at all.
Set the session variable inside the same transaction as the query it protects:
SET LOCAL is scoped to the current transaction and resets automatically when it ends, which is what you want on a pooled connection you are about to hand back for another tenant's request. It also means the setting has to happen inside the transaction that runs the query, not before it. If you run PgBouncer in transaction pooling mode, this is not optional: a connection can be handed to a different client between transactions, and a SET that outlived its transaction would leak one tenant's scope onto the next request that happens to land on the same backend connection.
RLS is a second layer, not a replacement for the first. It catches the query your ORM's tenant scope forgot. It does not catch a query run by a role with BYPASSRLS, and it does nothing for data that lives outside Postgres, which is most of what the rest of this post covers.
Credentials resolve per tenant, per call, never at process start
Day 20 covered the mechanics of this in detail: envelope encryption, single-flight refresh, typed secrets that refuse to print. The multi-tenant version changes exactly one thing, which is the key you resolve by.
A single-user credential store is keyed by (user_id, provider). A multi-tenant one is keyed by (tenant_id, user_id, provider) for a personally connected account, or (tenant_id, provider) for a service-level integration the whole tenant shares, such as one Slack workspace connection used by every user at Acme. Both shapes are legitimate. What is not legitimate is a lookup that can resolve to the wrong tenant's row, which is why the resolver takes the full context object and never a bare string:
async def resolve(ctx: TenantContext, provider: str, required_scopes: frozenset[str]) -> Secret: cred = await store.get(ctx.tenant_id, ctx.user_id, provider) if cred is None: raise NotConnected(provider) if not required_scopes <= cred.scopes: raise MissingScope(provider, required_scopes - cred.scopes) return await refresh_if_needed(cred)
Same rule as Day 19 and 20: resolve at the moment of use, never at process start, never cached on a client object that outlives one request. A provider client built once at import time and reused across requests is how Acme's request ends up authenticated as whichever tenant connected first. If you cache a resolved client at all, the cache key includes the tenant id, and the entry expires no later than the credential does.
Cache keys that make cross-tenant collisions structurally impossible
Caching is where isolation quietly dies, because a cache key is usually built from whatever the function already has in scope, and "whatever is in scope" is exactly the kind of thing that is easy to get right for one tenant and wrong the day a second one shows up.
The load-bearing detail is that ctx is a required positional argument, not a keyword with a default. A cache key builder that accepts tenant_id: str | None = None will, sooner or later, get called from a code path that forgot to pass it, and the resulting key collapses onto whatever the default produces. Make the tenant identity part of the function's contract, not an optional extra, and a reviewer has to actively delete it to introduce the bug.
There is a second, less obvious cache to think about if you use a provider that offers prompt caching. Anthropic's prompt caching matches on an exact prefix, so it is not vulnerable to cross-tenant leakage in the way a naive key-value cache is. But if your system prompt includes a tenant-specific tool list or tenant-specific instructions, as the next section describes, every tenant produces a different cache prefix, which means you lose cross-tenant cache hits entirely. That is a real cost, not a security bug, and it belongs in your latency and spend budget rather than being discovered by surprise.
Vector stores: partition first, filter second
Retrieval-augmented agents add a third store to isolate, and it is the one teams most often get wrong, because a metadata filter looks identical to a namespace in a demo and behaves completely differently under a bug.
Metadata filtering keeps every tenant's vectors in one shared index and narrows results with a filter clause at query time, for example tenant_id == "acme". It works, right up until a code path constructs the query without the filter, or with the wrong tenant's value, and that query now searches the entire corpus. The vector index does not know it is supposed to protect Acme from Globex. It only knows what filter it was handed.
A namespace, or the equivalent concept under a different name (Pinecone's namespaces, Weaviate's native multi-tenancy with per-tenant partitions, a separate Qdrant collection), stores each tenant's vectors in a genuinely separate partition. A query against Acme's namespace has no path to Globex's vectors at all, filter or no filter, because they are not present in the space being searched. Pinecone's own guidance is direct about the operational difference: a query against one tenant's namespace is billed and scanned as just that namespace, deleting a tenant means deleting its namespace in one call, and if you later need to query across tenants deliberately, that is exactly the case metadata filtering is for.
If your vectors live in Postgres via pgvector, you already have the right tool: the same row level security pattern from earlier in this post applies to the vector table exactly as it does to any other, because pgvector is a column type, not a separate service with its own access model.
The rule that generalizes: use a real partition as the default isolation mechanism for tenant data, and reserve metadata filtering for narrowing within a partition you are already scoped to, such as filtering Acme's own vectors by document type. Metadata filtering as your only tenant boundary makes every query the place a leak can happen. A partition makes the leak structurally unavailable.
What happens when work crosses a queue boundary
Everything above holds inside one request, inside one process, because contextvars propagation covers exactly that span. The moment an agent hands work to a background job, a scheduled digest, a webhook retry, or any worker running in a different process, that context does not travel with it. It cannot. A ContextVar is Python process memory, and a queue message is bytes on a wire.
Whatever crosses that boundary has to carry tenant identity explicitly, as data, and the worker on the other side has to treat it as a claim to verify rather than a fact to trust.
from dataclasses import dataclass@dataclass(frozen=True)class TenantJob: tenant_id: str user_id: str run_id: str payload: dictasync def enqueue_digest(ctx: TenantContext, recipient: str) -> None: await queue.publish( "digest.send", TenantJob(ctx.tenant_id, ctx.user_id, ctx.run_id, {"recipient": recipient}), )async def handle_digest(job: TenantJob) -> None: tenant = await tenants.get(job.tenant_id) if tenant is None or not tenant.active: raise PermanentFailure(f"tenant {job.tenant_id} is gone or suspended") fresh_ctx = TenantContext(job.tenant_id, job.user_id, job.run_id, tenant.current_scopes) with tenant_scope(fresh_ctx): await send_digest(job.payload["recipient"])
Two details in handle_digest matter more than they look. The worker re-fetches the tenant's current state rather than trusting anything embedded in the job payload beyond the bare id, because a job that sat in a queue for an hour, a day, or through a retry storm was enqueued against whatever was true then, not what is true now. And scopes are re-derived from tenant.current_scopes at execution time, not carried forward from whatever was granted when the job was created, because a scope revoked between enqueue and dequeue must take effect before the job runs, not after.
This is the failure mode that actually shows up in production, more often than any deliberate attack: a retry worker picks up a job that has been sitting for a while, runs it with whatever tenant id happens to be in the payload, and nobody re-checked whether that tenant, that user, or that grant still exists. Treat every queue, cron trigger, and webhook handler as a fresh trust boundary. Re-derive, re-validate, and only then re-bind the tenant context before doing anything on that tenant's behalf.
Tenant-scoped tool registries, and what they cost you
Not every tenant should see the same tool list. Acme connected Slack and GitHub. Globex connected Salesforce and never touched GitHub. Handing every tenant every tool regardless of what they connected means the model spends context reasoning about tools it can never successfully call, and it means a tool call against an unconnected integration has to fail gracefully instead of never being offered in the first place.
Building the tool list per tenant, from that tenant's actual connected accounts and plan, is the right instinct. It has a real cost worth naming rather than assuming away: a system prompt that includes a tenant-specific tool list is a different prompt per tenant, which means you forfeit prompt-cache hits across tenants, exactly as noted above. For a platform with a handful of large enterprise tenants, that trade is usually worth it. For a platform with thousands of small tenants sharing a mostly identical tool set, it can turn into a meaningful latency and cost regression if nobody measures it.
There is a second-order problem this creates as your integration catalogue grows: if the model's available tool count scales with how many services your platform supports, rather than with what any one tenant actually connected, you end up back at the "too many tools" problem regardless of tenancy. One way this gets solved architecturally, independent of any specific vendor, is to keep the model-facing tool surface small and fixed, and push the "which of the tenant's connected integrations does this map to" decision behind that fixed surface instead of into the model's context. A hosted MCP gateway that exposes a constant handful of tools, something like search, get schema, list connected accounts, and call tool, rather than one tool per underlying integration per tenant, is one concrete way teams keep the schema flat while the number of tenants and connected services behind it grows without bound. The tenant boundary in that shape lives in how the gateway resolves an end-user or tenant identifier on the call, not in how many tools the model has to hold in its head.
Whichever shape you pick, the invariant from earlier still applies without exception: whatever narrows the tool list, tenant plan, connected accounts, feature flags, none of it is allowed to widen who the tools act for. The tool list changes per tenant. The identity resolution underneath every tool does not.
Noisy neighbour control and cost attribution
A shared agent means a shared budget, and a shared budget means one tenant's runaway loop, a tool that keeps timing out and getting retried, a query that returns far more rows than expected and drives an unusually long completion, degrades latency or spend for every other tenant on the same infrastructure unless something isolates it.
A token bucket keyed by tenant id, checked before dispatching a tool call or a model completion, is enough to stop the bleeding without building a scheduler:
Attribution is the other half of the same problem, and it is a bookkeeping gap rather than a rate-limiting one: the model provider bills your organization in aggregate, and only your own accounting layer knows which tenant caused which spend. Record input and output token counts from every completion, tagged with tenant_id, alongside every tool call's latency and outcome, in a usage ledger separate from your application logs. That ledger is what lets you answer "why did our bill triple this month" with a tenant name instead of a guess, and it is what a per-tenant plan or quota is actually enforced against.
Observability that carries tenant id without pooling every payload into one place
Tracing an agent means capturing prompts, tool arguments, and tool results, and all three of those routinely contain a tenant's actual data. A shared, fully-searchable trace index that any engineer can query across every tenant is a second, informal database that was never subjected to the row level security policy protecting the real one.
Tag every span with tenant identity at creation, from the same context object everything else in this post reads from, so cost, latency, and error-rate dashboards can be sliced per tenant without extra plumbing:
from opentelemetry import tracetracer = trace.get_tracer("agent")async def call_tool(tool_name: str, args: dict): ctx = current_tenant() with tracer.start_as_current_span(tool_name) as span: span.set_attribute("tenant.id", ctx.tenant_id) span.set_attribute("run.id", ctx.run_id) return await dispatch(tool_name, args)
The structural fields, tenant id, tool name, latency, status code, are safe in one shared index and genuinely useful there for aggregate analysis and alerting. The payload fields, the actual prompt text, tool arguments, and tool output, are the ones that need the same treatment as the credentials in Day 20: kept in storage that is scoped per tenant, or redacted before they land anywhere queryable across tenants. If a support workflow genuinely needs to inspect Acme's raw trace to debug an issue Acme reported, that access should be scoped and audited the same way access to Acme's database rows is, not implied for free by whoever has a login to the observability tool.
The multi-tenant isolation checklist
Tenant identity, user identity and execution identity are three separate fields on one immutable context object, never collapsed into a single id
tenant_id never appears in a tool's input schema or in its handler's parameter list, and the schema sets additionalProperties: false
Every tool handler resolves tenant identity from request-bound context, never from an argument, a default value, or module-level state
Tenant context is bound once at the authentication boundary using contextvars, and nothing constructs a TenantContext anywhere else
Postgres tables holding tenant data have row level security enabled and forced, and the application's runtime role does not have the BYPASSRLS attribute
The tenant-scoping session variable is set with SET LOCAL inside the same transaction as the query it protects, which is mandatory under PgBouncer transaction pooling
Credentials resolve per call, keyed by tenant and provider at minimum, and are never cached on a client object that outlives one request
Every cache key function takes tenant identity as a required argument, and there is no code path that can build a key without it
Vector data is isolated through a namespace, a per-tenant collection, or an RLS-backed table, with metadata filtering used only to narrow inside a tenant's own partition
Background jobs, webhooks and retries carry tenant identity as explicit payload data, and the worker re-validates tenant status and current scopes on execution rather than trusting what was true at enqueue time
Tenant-specific tool lists are a deliberate choice, and their prompt-cache cost has been measured, not assumed away
Tenant usage is metered against a budget that isolates one tenant's runaway loop from every other tenant's latency, with per-tenant cost attribution separate from application logs
Traces carry tenant id as a queryable structural field, while raw prompt and tool payloads live in tenant-scoped storage or are redacted before reaching a shared index
You have run two tenants concurrently against the identical prompt and confirmed, in the database's own query log, that each touched only its own rows
Where this leaves you
None of this is exotic. It is the same discipline Day 19 and Day 20 applied to one user, applied again at the next boundary up, with one addition that matters: a tenant is not just another user with a bigger id. It is a boundary that has to survive process crossings a single user's request usually never makes, because tenants are exactly where background jobs, shared caches, shared vector indexes and shared observability tooling accumulate.
The design that holds up is small enough to describe in one paragraph. One context object carrying tenant, user and run identity, bound once at the edge and re-derived by every tool rather than trusted from upstream. A tool schema that makes tenant identity ungrammatical for the model to express. A database that enforces the boundary a second time, independent of application code. And a hard rule that anything crossing a queue, a cron, or a retry carries its identity as data and re-validates it on the other side, because ambient context does not survive a process boundary and pretending it does is how a retry worker ends up running with whoever's context happened to be lying around.
Get the request-scoped part right and adding a tenant looks like adding a user. Get it wrong and the failure is invisible in every demo, because a demo only ever has one tenant to leak into.
Next in the series: what changes when the agent itself has to survive past the request, and why binding tenancy to the request instead of the conversation is also the argument for keeping agent infrastructure stateless.