Skip to content

MCP session stale after server scale-to-zero - @retry_on_errors reuses dead cached session #7060

Description

@sachiantany

Bug Description

When an MCP server (e.g. on Cloud Run) scales to zero and back, the server's in-memory sessions are lost. The agent's @retry_on_errors decorator retries the failed call, but MCPSessionManager.create_session() returns the same cached dead session because _is_session_disconnected() only checks transport-level stream closure - the transport is still alive, only the server-side session is gone.

This means every retry reuses the stale session and fails again, until all retries are exhausted.

Steps to Reproduce

  1. Deploy an MCP server on Cloud Run (or any auto-scaling platform) with scale-to-zero enabled
  2. Connect an ADK agent using McpToolset with a Streamable HTTP connection
  3. Make a successful tool call (session is created and cached)
  4. Wait for the MCP server to scale to zero (~15 minutes idle)
  5. Make another tool call - the server scales back up with fresh state

Expected Behavior

The agent should detect the server-side session loss and create a fresh session on retry.

Actual Behavior

  • @retry_on_errors catches the error and retries
  • create_session() finds the cached session
  • _is_session_disconnected() returns False (transport streams are still open)
  • The same dead session is reused → retry fails with the same error
  • All retries exhausted → permanent failure

Root Cause

_is_session_disconnected() only checks _read_stream._closed and _write_stream._closed. In a scale-to-zero scenario, the HTTP transport is fine - the server responds normally with
a 404/error for the unknown session ID. The check needs a way to detect **servtion, not just transport failure.

The existing _MCP_GRACEFUL_ERROR_HANDLING flag handles the case where the asyncio task dies, but not this scenario where the task is alive and the server responds normally.

Environment

  • ADK version: 2.7.1
  • MCP server: Cloud Run with Streamable HTTP transport
  • Python: 3.12+

Suggested Fix

Add an invalidate_session() method to MCPSessionManager that marks a session key as stale. Have _execute_with_session() call it on any exception before re-raising, so the next @retry_on_errors attempt builds a fresh session. PR incoming.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

mcp[Component] This issues is related to MCP supportrequest clarification[Status] The maintainer need clarification or more information from the author

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions