Skip to content

Socket Mode: reconnects behind NAT leak server-side connection registrations → too_many_websockets cap → silent loss of interactive payloads; mitigation proposals #1940

Description

@lubosxyz

Slack SDK version: slack-sdk 3.43.0, slack-bolt 1.29.0
Python: 3.13 (aiohttp Socket Mode client)
OS/platform: Linux container (Docker) on macOS host, behind NAT

Summary

This is part bug report, part mitigation proposal, backed by wire-level data.
When SocketModeClient reconnects while the network path is degraded (half-open
TCP: the close frame never reaches Slack), the old connection remains registered
server-side for an extended period. Repeated reconnects therefore accumulate
"ghost" registrations up to Slack's 10-connection cap (disconnect: too_many_websockets). Slack then delivers envelopes across all registered
connections, so most interactive (block_actions) payloads — which unlike
events are not retried on non-ack — are silently lost. From the app's
perspective the client looks perfectly healthy: ping/pong fine, events flowing.

Wire evidence

(from a run_message_listeners wrapper logging hello and disconnect frames)

  • fresh start: hello reports num_connections=1
  • ~1.5 h later, on a reconnect: 3x disconnect reason=too_many_websockets,
    then hello num_connections=10 — while the process verifiably held ONE
    established TCP connection to Slack the whole time
  • ghost registrations age out at roughly one per 30-45 minutes
  • while num_connections is high, most button clicks never arrive on any
    connection we hold; with a clean pool, every click arrives (tested across
    message sizes 0.5-5 KB — size is irrelevant)
  • reproduced on a SECOND app in the same workspace: first hello after a
    process restart reported num_connections=7 for an app that also runs as a
    single instance
  • observed approximate_connection_time (inside hello.debug_info) is
    consistently 18060 (~5 h), which sets the ghost age-out horizon

Why this is hard to see with the current SDK

  1. The hello envelope (carrying num_connections) never reaches
    message_listeners — SocketModeRequest.from_dict requires
    type+envelope_id+payload, so apps cannot observe the most important signal
    without wrapping internals.
  2. disconnect frames (including too_many_websockets) are handled by
    run_message_listeners before the listener loop and only visible at debug
    logging.
  3. A degraded connection still passes is_connected() / ping-pong checks, so
    client-side health monitoring cannot detect the server-side pool state.

Proposals (any subset would help)

  1. Surface hello metadata (num_connections, approximate_connection_time,
    host) and disconnect reasons via a public callback or at INFO logging.
  2. Emit a loud warning when num_connections in hello exceeds a threshold
    (e.g., 4) while the client manages fewer connections — this is direct
    evidence of ghost registrations and imminent interactive-payload loss.
  3. Consider make-before-break reconnects with close-confirmation, or documenting
    that reconnect-heavy operation behind NAT can poison the server-side pool.
  4. Still-open PRs Fix aiohttp Socket Mode close lifecycle #1914 / fix(socket-mode): only release connect_operation_lock when this task acquired it #1926 address orphaned client-side sessions in the
    same failure family — this issue is their server-side counterpart, and since
    neither is merged/released (latest release is 3.43.0), apps currently have no
    upstream remedy at all; that raises the priority of surfacing the diagnostics
    from proposals 1-2.

Happy to share full logs and reproduction notes. We have also filed a parallel
report with Slack developer support regarding the server-side routing/eviction
behavior; will cross-link.

Activity

  1. lubosxyz commented on Aug 13, 2026

    @lubosxyz
    Author

    Follow-up with measured data after deploying client-side mitigations (the wire
    instrumentation from the report, plus an exponential-backoff idle watchdog:
    900 -> 1800 -> 3600 s cap, +-15% jitter, reset only on a delivered envelope).

    One night before vs. one night after, same app, same host:

    metric (per night) before after
    idle-watchdog reconnects 73 4
    hello num_connections climbed to 10 (cap) 1 on every hello
    too_many_websockets bursts 2 (11 frames) 0

    Two observations that may help triage:

    1. Reducing the reconnect rate keeps the primary app's pool clean, which
      supports the accumulation model (ghosts are minted by reconnects whose close
      never reaches the server).
    2. However, other single-instance apps on the same host still sit at a steady
      num_connections of 4-5 overnight (network-level flaps still force
      occasional SDK reconnects, and each ghost lingers for the ~5 h
      approximate_connection_time horizon). Client-side tuning cannot get below
      that floor — which is why surfacing num_connections (proposal 1/2) and a
      server-side eviction or force-disconnect API would still be valuable even
      for well-behaved clients.
  2. WilliamBergamin commented on Sep 3, 2026

    @WilliamBergamin
    Contributor

    Hi @lubosxyz thanks for the exceptionally detailed report 💯

    the wire-level instrumentation and the before/after num_connections data make this failure mode very clear.

    Status on referenced issues: both are now merged and released in 3.44.1:

    Both reduce the spurious/concurrent reconnects that are one of the mechanisms minting the ghost registrations you describe, so they should help at the margins.

    That said, I don't think these resolve this issue, so I'd like to keep it open. As you framed it, this is the server-side counterpart: the root cause is registrations aging out on Slack's side over ~5 h, which the SDK might not fix directly. Your follow-up data makes the decisive point, after cutting reconnects 73 → 4, other well-behaved single-instance apps still sat at num_connections 4–5 overnight, so reconnect-correctness fixes alone can't get below that floor.

    I'm narrowing the scope of this issue to the SDK-actionable proposals that remain unaddressed:

    1. Surface hello metadata (num_connections, approximate_connection_time) and disconnect reasons. Maybe via a public callback and/or INFO logging.
    2. Warn when hello's num_connections exceeds a threshold while the client holds fewer connections.
    3. make-before-break reconnect with close-confirmation, or at minimum documenting the NAT hazard.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    auto-triage-skipdiscussionM-T: An issue where more input is needed to reach a decision

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions