Skip to content

Python: Decide recovery policy for ended MCP subscription streams #9109

Description

Summary

Track an optional resilience policy for Python MCP subscriptions/listen streams after the base MCP v2 migration in #8245 is complete.

The base migration opens one modern subscription for enabled tool/prompt list-change notifications, preserves the legacy message_handler path, and recreates the subscription when Agent Framework reconnects the whole MCP connection. If the individual listen stream ends while the client connection remains usable, refresh currently stops until a later connection reset.

This is not required for MCP 2026 protocol compliance and must not block #8245. The specification says an abrupt loss may trigger reconnect; it does not require automatic same-connection recovery. The Python SDK's application guide separately recommends that watchers which still care should back off, re-listen, and refetch because events are not replayed.

Design questions

  • Should Agent Framework promise continuous catalog watching after either graceful stream completion or SubscriptionLost?
  • If so, which failures are transient versus permanent (SubscriptionLost, timeout, MCPError, method-not-found, pre-2026 ListenNotSupportedError)?
  • What retry/backoff policy should be used, and should it be bounded?
  • Should one combined tools/prompts subscription be re-opened, with both honored catalogs refetched after the new acknowledgment?
  • How should recovery coordinate with the existing lifecycle owner without reconnecting from inside—and cancelling—the subscription consumer task?
  • What diagnostics should distinguish graceful completion, abrupt loss, and permanent rejection?

Possible acceptance criteria

  • Recovery behavior is explicitly decided and documented as Agent Framework policy.
  • If enabled, graceful completion and SubscriptionLost back off before re-listening.
  • Catalog refetch occurs only after the replacement subscription is acknowledged; no event replay is assumed.
  • Tool and prompt catalogs share one subscription/recovery lifecycle.
  • Close, reset, cancellation, and retry delay leave no background tasks or listen contexts behind.
  • Permanent unsupported/rejected cases do not spin in a retry loop.
  • Existing full-connection reconnect behavior continues to establish exactly one new subscription.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    mcpUsage: [Issues, PRs], Target: MCPpythonUsage: [Issues, PRs], Target: Python

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions