MewCP LogoAStheTech
MCPs
Use Cases

Use cases by category

Productivity & InboxInbox, calendar, and daily flowEngineering & DevOpsShip, debug, and run on-callSales & CRMPipeline, outreach, and dealsMarketing & GrowthCampaigns, SEO, and growthSupport & SuccessTriage tickets, keep customers happyFinance & OpsClose, reconcile, and expensesCreative & ContentGenerate assets and contentPeople & HiringHiring, onboarding, and HRResearch & DataSynthesize data and insights
See all use cases
Resources
BlogsProduct updates and storiesArticlesIntegration guides and code examples
PricingDocsSign in
Back to home
MewCP Logo

Infrastructure You Can Trust for Agentic Products

X

Categories

  • Productivity & Docs
  • Developer Tools
  • CRM & Sales
  • Finance & Commerce
  • Data & Analytics
  • Marketing & SEO
  • Search & Web
  • Communication
  • View All Servers →

Resources

  • Blog
  • Docs
  • Privacy Policy
  • Terms of Service

Blogs

  • View All Blogs →

Articles

  • View All Articles →
Browse Servers|Pricing|Contact

Browse by Category

Productivity & Docs

  • Gmail
  • Google Drive
  • Google Classroom
  • Google Calendar
  • Google People
  • YouTube
  • Notion
  • ClickUp
  • Figma
  • Google Tasks
  • Cal
  • Monday
  • Luma
  • Notion MCP
  • Mem MCP
  • Linear MCP
  • Calendly MCP
  • Consensus MCP
  • Craft MCP
  • Close MCP
  • Dice MCP
  • Lumin PDF MCP
  • Develop21 MCP
  • Granola MCP
  • Lucid MCP
  • Mermaid Chart MCP
  • Fireflies MCP
  • ClickUp MCP
  • Miro MCP

Developer Tools

  • Gemini
  • Veo
  • ClickUp
  • Firecrawl
  • Vercel
  • Apify
  • Github
  • HTTP
  • Chef
  • Scientific Calculator
  • Figma
  • Perplexity
  • Apify MCP
  • Hugging Face Hub MCP
  • Buildkite MCP
  • Cloudflare MCP
  • Context7 MCP
  • Ahrefs MCP
  • Sentry MCP
  • Brevo Docs MCP
  • X Docs MCP
  • Jev
  • Linear MCP
  • Calendly MCP
  • Craft MCP
  • DeepWiki MCP
  • Inspo MCP
  • Kernel MCP
  • Malwarebytes MCP
  • Mermaid Chart MCP
  • Supabase MCP
  • Microsoft Learn MCP
  • Webflow MCP

CRM & Sales

  • Google People
  • OneSignal MCP
  • Brevo Docs MCP
  • Brevo
  • Brevo MCP
  • Carbon Voice MCP
  • Clay MCP
  • Close MCP
  • Attio MCP
  • Clarify MCP

Finance & Commerce

  • Razorpay
  • Polymarket
  • Kite
  • Stripe
  • Binance
  • Upstox
  • Aiwyn MCP
  • Era-Context-MCP
  • Granted MCP
  • XDC AI MCP
  • Agentery MCP
  • Agent Embassy
  • Quick Commerce MCP
  • Longbridge MCP
  • Mercury MCP

Data & Analytics

  • Apify MCP
  • Cloudflare MCP
  • Ahrefs MCP
  • Candid MCP
  • Consensus MCP
  • Contentsquare MCP
  • Era-Context-MCP
  • Instinct MCP
  • legal Data Hunter MCP
  • Marcopolo MCP
  • Mixpanel MCP

Marketing & SEO

  • Mailchimp
  • Google Business
  • YouTube
  • Google Search Console
  • OneSignal MCP
  • Cloudflare MCP
  • Brevo Docs MCP
  • Brevo
  • Brevo MCP
  • AirOps MCP
  • Clay MCP
  • Contentsquare MCP
  • Reelsmith MCP
  • GoDaddy MCP
  • Metricool MCP
  • Webflow MCP
  • Windsor MCP

Search & Web

  • Web Scrapper
  • Firecrawl
  • Apify
  • Perplexity
  • Context.dev
  • Exa
  • Brave Search
  • Apify MCP
  • Ahrefs MCP
  • DeepWiki MCP
  • Dice MCP
  • GoDaddy MCP
  • Granted MCP
  • Microsoft Learn MCP

Communication

  • Gmail
  • Google Meet
  • Mailchimp
  • Google Calendar
  • WhatsApp
  • Slack
  • OneSignal MCP
  • Brevo Docs MCP
  • Carbon Voice MCP

© 2026 MewCP. All rights reserved.

  1. Home
  2. Blogs
  3. Stateless AI Agent Infrastructure: Where Your Session State Actually Lives

Stateless AI Agent Infrastructure: Where Your Session State Actually Lives

by Rohit Gite, Founder @MewCP·September 30, 2026·23 min read

Stateless doesn't mean no state. It means the request carries it instead of the process, which is what actually buys you horizontal scale and painless restarts.

Day 21 covered tenant isolation: why the tenant id has to be re-derived from the request on every tool call instead of trusted from wherever it last showed up. This post covers the other half of the same scaling story. Once identity is a property of the request, the question that remains is where everything else about a running agent lives, and what happens to your architecture the moment that stops being "wherever the process happens to be."

Here is the failure this post exists to prevent. Your agent works. You put it behind a load balancer because traffic grew, or because you switched to a serverless runtime, or because a rolling deploy needed to happen without downtime. The very next request that lands on the new instance behaves like the user is a stranger. Nothing crashed. Nothing logged an error. The instance simply never had the context, because the context was never anywhere the instance could reach.

Stateless AI agent infrastructure is the fix, and it gets misread constantly as "the agent has no state," which is not what it means and not what it buys you. Every agent has state: a conversation, a plan, a set of credentials, a cursor into a paginated list, a pending approval. Going stateless does not delete any of that. It moves it out of the process handling this one request and into something any process can read. The state still exists. The only thing that changes is who is allowed to hold it between requests.

Why "just add another instance" breaks

The instinct when a single agent process gets overloaded is to run more of them behind a load balancer. That instinct is correct for a stateless HTTP API and wrong, by default, for an agent that has been keeping anything in memory.

The reason is mechanical, not architectural. A load balancer routes on connection health and load, not on "which process remembers this user." Unless you pin every request from one conversation to one instance with sticky routing, request two of a three-turn conversation can land anywhere. If the state lives in that process's memory, request two arrives at an instance that has never seen request one, and the agent either starts over silently or fails in a way nobody can reproduce from the logs.

Sticky routing patches this, and it introduces its own tax. Now a deploy has to drain sessions before it can recycle an instance. Now autoscaling has to account for a session's remaining lifetime before it kills a node. Now a crashed instance takes every session pinned to it down with it, and there is no other instance that can pick up where it left off, because none of them have the state either. You have not removed the coordination problem. You have hidden it inside your routing layer, where it is harder to see and harder to test.

ContentsSeptember 30, 2026
  1. Why "just add another instance" breaks
  2. Stateless AI agent infrastructure does not mean no state
  3. What actually lives in an agent session
  4. The two architectures, compared on real dimensions
  5. MCP made the same call, and it's worth reading closely
  6. Where statefulness comes back: subscriptions/listen
  7. Implementation: turning a stateful handler stateless
  8. The part nobody warns you about: idempotency
  9. The honest costs
  10. When stateful is still the right answer
  11. A migration sequence that does not take production down
  12. Shipping checklist
  13. Where this leaves you
Author

Rohit Gite, Founder @MewCP

Share

Build with MewCP

Connect your AI agents to real tools in minutes.

Get started

Stateless AI agent infrastructure does not mean no state

The reframe that actually helps here: state never disappears, it only moves. The question stateless architecture answers is not "should this data exist," it is "where is this data allowed to be found, and by whom."

A stateful design pins the answer to "one specific process, right now." A stateless design pins the answer to "anywhere reachable by an identifier the request carries." That second answer is strictly harder to build, because now you need a real store with real consistency guarantees instead of a Python dictionary. It is also the only answer that survives a second instance, a rolling deploy, or an autoscaler, because none of those events can invalidate an identifier the way they can wipe out a process's heap.

So the rule for the rest of this post: a stateless handler resolves everything it needs from the request and a backing store, and it is not allowed to assume anything survived from the last time it ran. That is a stronger constraint than "avoid global variables." It means the handler has to be correct even when a completely different instance served the previous turn, even when this is a cold start, and even when the store returns nothing because a TTL expired.

What actually lives in an agent session

Before deciding what to make stateless, inventory what "an agent session" is actually holding. In practice it is these seven things, and they do not all move the same way.

#What's hiding in the sessionWhat it actually isCan it go stateless?
1Conversation historyThe turn-by-turn transcript, including tool calls and tool resultsYes. Store it keyed by a conversation id, load a window of it per request
2Plan or scratchpad stateWhat the orchestrator has decided so far, remaining steps, intermediate resultsYes. Same treatment as history, usually a separate document so it can be read without the full transcript
3Identity and authorization contextUser id, tenant id, agent id, run id, the scopes and credential reference behind themYes, and it should already be this way. Day 19 and Day 20 covered resolving this per request rather than caching it in the process
4Protocol negotiation stateWhich MCP protocol version and capabilities this client and server agreed onYes, and as of MCP's 2026-07-28 revision it is required to be. There is no longer a handshake to remember; every request restates it
5Pending human-in-the-loop approvalsA write action that is proposed and waiting on a person to say yesMostly. The compute resolving it is stateless, but the pending decision itself is state that must live durably, with a TTL, somewhere any instance can see it and act once
6Long-running task handlesA pointer to work still running in the background, like an export or a crawlYes, as an opaque handle passed back to the client and re-supplied on the next call. The handle is stateless-referenceable even when the work behind it is pinned to a worker
7Live streaming subscriptionsAn open connection waiting on server-pushed change notificationsNo, not fully. This is the one item that resists the pattern, and it gets its own section below

Six of the seven move cleanly. The seventh is not a design mistake, it is a structural property of holding an open stream, and pretending it can be made stateless the same way as the rest is how "stateless" projects end up with one component that quietly breaks every rule the rest of the system follows.

The two architectures, compared on real dimensions

Putting the stateful and stateless versions of the same agent side by side, on dimensions that actually matter operationally rather than as an abstract preference:

DimensionStateful (session pinned to process)Stateless (session resolved from a store)
Read latencyEffectively zero, it's a memory lookupOne store round trip per request, typically low single-digit milliseconds in-region
Horizontal scalingRequires sticky routing, capacity planning per instance's session loadAny instance can serve any request, scale by adding instances
Restart / redeployKills every live session pinned to that instanceInstances are disposable, the store outlives them
Failure isolationA crashed process loses its sessions with itA crashed process loses nothing; the store is a separate failure domain (more on this below)
New failure modesSticky-routing bugs, uneven load from long sessionsStore outages, serialization bugs, stale reads, write conflicts
Best fitTight low-latency loops, large in-memory context, long-lived open streamsSpiky or bursty traffic, serverless and autoscaled compute, frequent deploys

Neither column is "correct." They are two different bets about where you want your operational complexity to live: inside your routing and process lifecycle, or inside a store you now have to run properly.

MCP made the same call, and it's worth reading closely

If you want a second opinion from people who had to solve this at protocol scale, look at what happened to the Model Context Protocol itself. The current revision, 2026-07-28, is explicit about it in the base specification:

"The Model Context Protocol (MCP) is a stateless protocol: all the information needed to process a request is contained in the request itself. A server processes each request independently; no state should be inferred from previous requests, even those on the same connection or stream."

That line did not describe MCP from the start. Earlier revisions used an initialize / notifications/initialized handshake to negotiate protocol version and capabilities once, at connection time, and a server-issued Mcp-Session-Id header to track that negotiation across subsequent requests on the Streamable HTTP transport. The 2026-07-28 revision removed both. Every request now restates its own protocol version and capabilities, and the Mcp-Session-Id header, the standalone GET stream endpoint, and Last-Event-ID resumability are all gone from the core transport.

What replaced the handshake is a set of fields carried on every single request, half in the JSON-RPC body and half mirrored into HTTP headers so a gateway or load balancer can route and log without parsing the body:

POST /mcp HTTP/1.1
Content-Type: application/json
Accept: application/json, text/event-stream
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: get_weather
 
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_weather",
    "arguments": { "location": "Seattle, WA" },
    "_meta": {
      "io.modelcontextprotocol/protocolVersion": "2026-07-28",
      "io.modelcontextprotocol/clientInfo": { "name": "ExampleClient", "version": "1.0.0" },
      "io.modelcontextprotocol/clientCapabilities": {}
    }
  }
}

Three details here are worth stealing for your own request envelope, not just admiring in the spec:

The header and the body are required to agree. MCP-Protocol-Version must match _meta["io.modelcontextprotocol/protocolVersion"], and Mcp-Method / Mcp-Name must match the body's method and params.name (or params.uri for a resource read). A server that finds a mismatch rejects the request with 400 Bad Request and a JSON-RPC HeaderMismatch error, code -32020. This exists precisely because a gateway routing on the header and a server executing on the body are two different sources of truth, and letting them silently disagree is a real security bug, not a style nit.

Missing required fields fail loudly and specifically. io.modelcontextprotocol/protocolVersion and io.modelcontextprotocol/clientCapabilities are required on every request; a request missing either is malformed and gets 400 with JSON-RPC -32602. An unsupported protocol version gets its own error, UnsupportedProtocolVersionError (-32022), listing what the server actually supports. A request relying on a capability the client never declared gets MissingRequiredClientCapabilityError (-32021), naming exactly what was missing. None of that is generic "bad request" noise, and none of it required a session lookup to produce.

Cross-call state didn't disappear, it became explicit. The spec change note for this revision says it plainly: servers that need state across calls now use "explicit, server-minted handles passed as ordinary tool arguments." That is the same move as item 6 in the inventory above, an opaque reference the client re-supplies, rather than a session the server is trusted to remember. It is also exactly the request envelope pattern this post builds out below, arrived at independently by a completely different team solving the same problem at the transport layer.

Where statefulness comes back: subscriptions/listen

The one item in the inventory that does not fold cleanly into "resolve everything from the request" is a live push channel, and MCP's own design confirms it is structurally different rather than badly designed. The 2026-07-28 revision replaced the old resources/subscribe / resources/unsubscribe methods and the standalone GET SSE endpoint with a single method, subscriptions/listen, and it is worth understanding exactly what kind of request that is.

It is still, formally, one JSON-RPC request and one JSON-RPC response. What makes it different is that the response is an SSE stream that the server deliberately keeps open, delivering notifications for as long as the client wants to keep listening:

{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "subscriptions/listen",
  "params": {
    "_meta": {
      "io.modelcontextprotocol/protocolVersion": "2026-07-28",
      "io.modelcontextprotocol/clientCapabilities": {}
    },
    "notifications": {
      "toolsListChanged": true,
      "resourceSubscriptions": ["file:///project/config.json"]
    }
  }
}

The server must send notifications/subscriptions/acknowledged as the first message, carrying the subscription's id (the JSON-RPC id of the listen request itself) in _meta["io.modelcontextprotocol/subscriptionId"], before anything else flows. Every notification after that, notifications/tools/list_changed, notifications/resources/updated, and so on, carries that same subscriptionId so a client juggling several concurrent subscriptions on one connection can tell them apart.

Here is the part that matters for your infrastructure decision: there is no resumability. Earlier Streamable HTTP revisions let a client reconnect with a Last-Event-ID header and replay what it missed. That mechanism is gone in 2026-07-28. If the underlying transport drops, whether from a load balancer idle timeout, a deploy, or a network blip, the subscription is simply over. The spec's own guidance for stdio applies just as much in spirit to HTTP: "the server holds no subscription state across reconnections," and the client's only option is to re-issue subscriptions/listen and accept that it will not get a replay of what it missed while disconnected.

Treat this as a firm design boundary rather than a gap to engineer around silently:

  • Don't put a subscriber behind a stateless-style round-robin pool the way you would a tools/call handler. The instance that accepted the subscriptions/listen request owns that TCP connection and the SSE stream on it for the whole lifetime of the subscription. That is inherently sticky, and no amount of session-store discipline changes it, because the thing being held isn't data, it's an open socket.
  • Do keep the notification source itself stateless. The instance holding the connection should not be the system of record for "did this resource change." It should subscribe to a shared fan-out (Redis pub/sub, a message bus, a change-data-capture stream from whatever store owns the underlying resource) and simply forward matching events onto whichever open connections it happens to be holding. Kill that instance and the client reconnects to a different one, opens a fresh subscription, and starts receiving from whatever the shared source is publishing now, no worse off than any reconnect.
  • Design your client for gaps, not guarantees. Since there is no redelivery, a resource that changed twice while you were disconnected will not replay both events. If your correctness depends on not missing an update, subscriptions/listen is a nice-to-have push channel on top of a resources/read you still do periodically or on reconnect, not a substitute for one.
  • Emit periodic SSE comment lines as keep-alives. The spec explicitly recommends this for long-lived listen streams so intermediaries and idle-timeout logic don't tear the connection down during quiet periods. A comment line (anything starting with :) carries no event data and clients are required to ignore it, so it costs you nothing but a few bytes.

Everything else in MCP's per-request model, tools/call, resources/read, prompts/get, stayed genuinely stateless in this revision. subscriptions/listen is the one place the protocol's own authors accepted that a live push channel cannot be, and said so directly rather than pretending otherwise. Build your own architecture with the same honesty: one component gets to be stateful because the alternative doesn't exist, and everything else does not get that excuse.

Implementation: turning a stateful handler stateless

With the inventory and the protocol precedent in hand, here is what it actually looks like to build the stateless side.

The request envelope

Dark teal blueprint diagram of a request envelope carrying identity, conversation reference, protocol fields and state handles into a resolver that reads from a shared session store

Every field in the envelope answers one question: "given nothing but this request, what do I need to look up, and where?" It is not a grab bag, it's a deliberate, minimal set:

from __future__ import annotations
 
import time
from dataclasses import dataclass, field
 
 
@dataclass(frozen=True)
class RequestEnvelope:
    """Everything a stateless handler needs, carried on the wire per call."""
 
    # Idempotency and tracing. request_id is generated by the CLIENT, not
    # the server, so a retried request carries the same value both times.
    request_id: str
    issued_at: float
 
    # Identity, resolved fresh from an authenticated source on every hop,
    # never trusted from a cache. See Day 19 and Day 20 for the full
    # credential-resolution story this plugs into.
    user_id: str
    tenant_id: str
    agent_id: str
    run_id: str
 
    # Where in an ongoing conversation this request sits. The handler
    # loads the transcript window itself; the envelope only carries the
    # pointer, never the content.
    conversation_id: str
    history_cursor: str | None
 
    # Mirrors MCP's per-request protocol fields: what this caller
    # negotiated, restated on every call instead of remembered from a
    # handshake.
    protocol_version: str
    client_capabilities: frozenset[str]
 
    # Server-minted opaque handles for any cross-call state this run has
    # already produced (a background job id, a paginated list cursor, a
    # pending-approval token). The client re-supplies these; the handler
    # never has to remember producing them.
    state_refs: dict[str, str] = field(default_factory=dict)

Two design decisions in there do most of the work. request_id is generated by the caller, never the server, which is what makes it usable as an idempotency key later. And state_refs is a flat dictionary of opaque strings rather than typed fields, because the set of things a given run needs to carry forward (a job id here, an approval token there) is open-ended and handler-specific; the envelope's job is to carry them, not to understand them.

A stateless tool handler

The handler's whole job becomes: resolve what you need from the envelope and the store, do the work, write back what changed, return. Nothing survives in the process between calls.

async def summarize_thread(envelope: RequestEnvelope, thread_id: str) -> ToolResult:
    """Stateless handler: every dependency is resolved per call."""
 
    # Identity resolves fresh, per Day 19/20's credential resolver.
    token = await credentials.resolve(
        envelope.user_id, "google", {GMAIL_READONLY}
    )
 
    # Conversation state resolves from the store, keyed by conversation
    # id, starting from wherever this caller last left off.
    history = await session_store.get_window(
        key=f"conv:{envelope.tenant_id}:{envelope.conversation_id}",
        cursor=envelope.history_cursor,
        max_turns=20,
    )
 
    async with httpx.AsyncClient(timeout=15.0) as http:
        response = await http.get(
            "https://gmail.googleapis.com/gmail/v1/users/me/threads/"
            f"{thread_id}",
            headers={"Authorization": f"Bearer {token.reveal()}"},
        )
    response.raise_for_status()
    summary = summarize(response.json(), prior_turns=history.turns)
 
    # Write back the one thing this call changed. Conditional on the
    # version this handler read, so a concurrent writer loses instead of
    # silently overwriting (see the store interface below).
    new_cursor = await session_store.append_turn(
        key=f"conv:{envelope.tenant_id}:{envelope.conversation_id}",
        expected_version=history.version,
        turn=Turn(role="tool", content=summary),
        ttl_seconds=86_400,
    )
 
    return ToolResult(
        content=summary,
        state_refs={**envelope.state_refs, "history_cursor": new_cursor},
    )

Nothing in this function references a module-level variable, a class attribute set by a previous call, or an in-memory cache with no expiry. If this instance gets killed the moment after it returns, the next call for this same conversation, on any instance, produces the identical result, because everything it depended on is sitting in the store, not in this process's heap.

The session store interface

Dark teal blueprint diagram of a session store record with a version number and TTL, showing a compare-and-swap write succeeding for one caller and being rejected for a concurrent second caller

The store is where the actual engineering happens. Three properties are non-negotiable: every record has a TTL so abandoned sessions don't accumulate forever, every write is versioned so two concurrent handlers can't silently clobber each other, and every write is conditional on the version the caller last read.

from typing import Protocol
 
 
class VersionConflict(Exception):
    """Raised when a write's expected_version no longer matches the store."""
 
 
class SessionRecord(Protocol):
    version: int
    turns: list["Turn"]
 
 
class SessionStore(Protocol):
    async def get_window(
        self, key: str, cursor: str | None, max_turns: int
    ) -> SessionRecord:
        """Read the current record. Returns an empty record, version 0,
        if the key doesn't exist or its TTL has already expired."""
        ...
 
    async def append_turn(
        self, key: str, expected_version: int, turn: "Turn", ttl_seconds: int
    ) -> str:
        """Conditional write. Succeeds only if the stored version still
        equals expected_version, then resets the TTL and returns a new
        cursor. Raises VersionConflict otherwise."""
        ...
 
    async def put_if_absent(
        self, key: str, data: bytes, ttl_seconds: int
    ) -> bool:
        """Idempotency primitive: True if this call created the key,
        False if it already existed. See idempotency below."""
        ...

A Redis-backed implementation gets TTL for free and needs one Lua script to make the read-check-write atomic across a network hop, which a plain GET then SET cannot guarantee under concurrency:

import json
import redis.asyncio as redis
 
_CAS_APPEND = """
local raw = redis.call('GET', KEYS[1])
local current_version = 0
local turns = {}
if raw then
    local record = cjson.decode(raw)
    current_version = record.version
    turns = record.turns
end
if current_version ~= tonumber(ARGV[1]) then
    return {err = "VERSION_CONFLICT"}
end
table.insert(turns, cjson.decode(ARGV[2]))
local new_record = { version = current_version + 1, turns = turns }
redis.call('SET', KEYS[1], cjson.encode(new_record), 'EX', ARGV[3])
return current_version + 1
"""
 
 
class RedisSessionStore:
    def __init__(self, client: redis.Redis) -> None:
        self.client = client
        self._cas_append = client.register_script(_CAS_APPEND)
 
    async def append_turn(self, key, expected_version, turn, ttl_seconds) -> str:
        try:
            new_version = await self._cas_append(
                keys=[key],
                args=[expected_version, json.dumps(turn.__dict__), ttl_seconds],
            )
        except redis.ResponseError as exc:
            if "VERSION_CONFLICT" in str(exc):
                raise VersionConflict(key)
            raise
        return str(new_version)
 
    async def put_if_absent(self, key: str, data: bytes, ttl_seconds: int) -> bool:
        # NX: only set if the key doesn't already exist. Atomic on its own,
        # no script needed.
        return bool(await self.client.set(key, data, nx=True, ex=ttl_seconds))

The Lua script runs entirely inside Redis, so the read, the version check, and the write are one atomic operation from the caller's point of view. No lock, no separate transaction, no window where a second writer can sneak in between your read and your write.

The part nobody warns you about: idempotency

Making a handler stateless does not automatically make it safe to retry, and this is where teams get hurt during migration, not before it.

The moment any instance can serve any request, retries stop being rare. A load balancer health check times out and reissues. A rolling deploy interrupts an in-flight call and the client retries against the new instance. A client-side timeout fires just as the server was about to respond. Every one of these produces the exact same tools/call, with the exact same arguments, arriving a second time, possibly at a different instance than the first attempt.

Two things have to be true together, not just one, or the second execution silently repeats the side effect instead of being recognized as a duplicate:

An idempotency key derived from the request, not generated by the server. This is exactly why request_id in the envelope above is client-generated. If the server minted it, the two attempts of the same logical request would get two different keys, and the whole mechanism is void before it starts.

A conditional write on the state store, so the second execution loses instead of overwriting. put_if_absent above is the primitive:

async def with_idempotency(envelope: RequestEnvelope, handler, *args) -> ToolResult:
    key = f"idem:{envelope.tenant_id}:{envelope.request_id}"
 
    claimed = await session_store.put_if_absent(key, b"in-flight", ttl_seconds=120)
    if not claimed:
        cached = await session_store.get_window(key + ":result", None, 1)
        if cached.turns:
            return cached.turns[0].content  # already ran, return the same result
        raise RetryTooSoon(envelope.request_id)  # first attempt is still in flight
 
    result = await handler(envelope, *args)
    await session_store.append_turn(
        key=key + ":result", expected_version=0, turn=Turn("result", result), ttl_seconds=120
    )
    return result

Stateless without idempotency is not stateless architecture, it is distributed double-charging with extra steps. If the side effect is sending an email, archiving a thread, or placing an order, the second execution of the "same" request is a bug with a customer-visible consequence, not a harmless retry.

The honest costs

None of the above is free, and pretending otherwise is how a migration gets rolled back six months in because nobody budgeted for what changed.

Load latency on every call. A memory read is nanoseconds. A store round trip, even to an in-region Redis or Postgres instance, adds real, measurable latency to every single request that used to be instant. For a chatty agent making several tool calls per turn, that adds up per turn, not just per session. Budget for it explicitly rather than discovering it in a P99 dashboard after launch.

A new failure domain. Your agent's uptime is now the lower of your compute's uptime and your store's uptime, not the same number. A healthy fleet of stateless instances in front of a store that is down is a fully outaged product, and it fails in a way your old stateful setup structurally could not: partial failure that looks like the compute layer's fault when it isn't. The store needs its own SLO, its own on-call runbook, and its own failover story, or you have just relocated your single point of failure rather than removed it.

Serialization overhead. Every read and write now pays the cost of encoding and decoding a conversation window, a plan document, or a set of state refs, on every hop, in both directions. A session that has quietly grown to hundreds of turns costs real CPU and real bytes to move on every call that touches it. This is exactly why the handler above windows the history (max_turns=20) rather than round-tripping the entire transcript every time; unbounded state size is a cost curve you do not want discovering itself in production traffic.

When stateful is still the right answer

None of this is a blanket argument for statelessness everywhere, and treating it as one is the mistake this whole post is trying to prevent. Decide per component, using the actual shape of that component's workload, not a company-wide policy.

ComponentLean statelessLean stateful
Request routing / tool call handlersSpiky or bursty traffic, frequent deploys, autoscaled or serverless computeRarely; this is the default stateless case
Conversation history and plan stateAlmost always; it's exactly what a session store exists forOnly if turn-to-turn latency is the dominant product constraint and the fleet is small and stable enough that sticky routing's costs stay low
Credentials and authorizationAlways; per-request resolution is a security requirement, not just a scaling oneNever; caching a credential in a process is the vulnerability this post's Day 19/20 companions exist to close
Pending human approvalsThe compute checking for a decision, yesThe decision record itself needs a durable, TTL'd store regardless, this isn't really an either/or
Live streaming subscriptions (subscriptions/listen and equivalents)The event source behind it, via pub/sub fan-outThe connection itself, structurally, for its lifetime
Long-running background jobsThe handle referencing the jobThe worker actually executing it, which legitimately owns process-local state (open file handles, an in-memory index build) for its duration
Tight, low-latency inner loops (a hot retrieval cache, an in-memory vector index serving one workload)Only if you can tolerate the added hopYes, when the whole point of the design is avoiding exactly that hop

The pattern across every stateful row: either the workload has a genuine low-latency requirement a store round trip would violate, or the thing being held is a live connection or in-process resource rather than data that can be serialized and handed to another process. Everything else belongs in a store.

A migration sequence that does not take production down

Doing this in one cutover is how "we migrated to stateless infrastructure" turns into an incident report. A sequence that keeps you shippable at every step:

  1. Externalize identity and authorization resolution first. This is the lowest-risk, highest-value move and it should already be true independent of this migration; see Day 19 and Day 20. Nothing else here works safely if a credential is still cached in the process.

  2. Introduce the request envelope additively. Start carrying request_id, the identity fields, and state_refs on every request, even while handlers still also read from in-memory state. This costs you nothing operationally and lets you validate the envelope's shape against real traffic before anything depends on it.

  3. Add the session store and dual-write. Every write goes to both the in-memory session and the new store. Reads still come from memory. Log every case where a dual-read of the same key would have produced different results between the two sources; those divergences are your test suite for the store's correctness before you trust it.

  4. Flip reads to the store, gated by a rollout flag. Move a fraction of traffic, or one request type, to reading exclusively from the store while the memory write path stays in place as a fallback you can flip back to instantly. Watch version-conflict rates and store latency percentiles before widening the rollout, not after.

  5. Remove in-memory state and sticky routing, except for subscriptions/listen. Once reads and writes are fully on the store and stable, delete the in-memory path and any routing rules that existed only to keep sessions pinned. Keep, deliberately, whatever sticky handling your streaming subscriptions need, since that one component was never a candidate for removal in the first place.

Each step is independently shippable and independently revertible. If step 4 shows conflict rates you don't like, you stop there, fix the store's write pattern, and try again, rather than having already deleted the fallback that would let you.

Shipping checklist

  • Every tool handler resolves identity, conversation state, and credentials from the request and a store, never from module-level state, a class attribute, or an in-memory cache with no expiry
  • request_id is generated by the client and used as the idempotency key; the server never mints its own
  • Every session store record carries a TTL and a version, and every write is conditional on the version the caller read
  • A version conflict on write is handled explicitly (retry-with-reread, or surface a conflict to the caller), never silently swallowed or overwritten
  • Cross-call state travels as opaque, server-minted handles in the request (state_refs), matching the pattern MCP itself adopted in its 2026-07-28 revision
  • Conversation history is windowed, not fully round-tripped on every call, and the window size is a deliberate, monitored choice
  • The session store has its own SLO, its own on-call path, and a tested failover, separate from your compute layer's
  • subscriptions/listen (or any equivalent live push channel) runs on infrastructure that expects sticky connections, backed by a stateless pub/sub fan-out underneath it, not on the same round-robin pool as your request handlers
  • Clients that use streaming subscriptions are built to resubscribe and tolerate gaps, since there is no Last-Event-ID-style redelivery to lean on
  • The migration ran dual-write before dual-read, and dual-read before removing the fallback, with real divergence data reviewed at each stage
  • Load latency, store failure scenarios, and serialization cost were measured against real traffic before this shipped as the only path, not assumed acceptable

Where this leaves you

The mental model is small enough to hold in your head: state never disappears, it only moves, and the entire discipline is deciding where it is allowed to be found. Six of the seven things hiding in an agent session move cleanly into a request envelope and a versioned, TTL'd store. The seventh, a live push channel, doesn't move, because an open connection isn't data you can hand to another process, and MCP's own maintainers reached the identical conclusion when they built subscriptions/listen as the one deliberately stateful exception in an otherwise stateless 2026-07-28 revision.

Get the store's TTL, versioning, and conditional writes right, pair every stateless handler with a real idempotency key, and be honest about the store as a new failure domain rather than a free win, and stateless infrastructure earns the horizontal scale and painless restarts it promises. Skip any one of those and you've just moved the bug from your process's memory into a database, where it's harder to see and much easier to trust.

Next in the series: what changes when an agent's work outlives a single request entirely, and a run has to survive across minutes or hours instead of milliseconds.