You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Socket Mode: reconnects behind NAT leak server-side connection registrations → too_many_websockets cap → silent loss of interactive payloads; mitigation proposals #1940
This is part bug report, part mitigation proposal, backed by wire-level data.
When SocketModeClient reconnects while the network path is degraded (half-open
TCP: the close frame never reaches Slack), the old connection remains registered
server-side for an extended period. Repeated reconnects therefore accumulate
"ghost" registrations up to Slack's 10-connection cap (disconnect: too_many_websockets). Slack then delivers envelopes across all registered
connections, so most interactive (block_actions) payloads — which unlike
events are not retried on non-ack — are silently lost. From the app's
perspective the client looks perfectly healthy: ping/pong fine, events flowing.
Wire evidence
(from a run_message_listeners wrapper logging hello and disconnect frames)
fresh start: hello reports num_connections=1
~1.5 h later, on a reconnect: 3x disconnect reason=too_many_websockets,
then hello num_connections=10 — while the process verifiably held ONE
established TCP connection to Slack the whole time
ghost registrations age out at roughly one per 30-45 minutes
while num_connections is high, most button clicks never arrive on any
connection we hold; with a clean pool, every click arrives (tested across
message sizes 0.5-5 KB — size is irrelevant)
reproduced on a SECOND app in the same workspace: first hello after a
process restart reported num_connections=7 for an app that also runs as a
single instance
observed approximate_connection_time (inside hello.debug_info) is
consistently 18060 (~5 h), which sets the ghost age-out horizon
Why this is hard to see with the current SDK
The hello envelope (carrying num_connections) never reaches message_listeners — SocketModeRequest.from_dict requires
type+envelope_id+payload, so apps cannot observe the most important signal
without wrapping internals.
disconnect frames (including too_many_websockets) are handled by run_message_listeners before the listener loop and only visible at debug
logging.
A degraded connection still passes is_connected() / ping-pong checks, so
client-side health monitoring cannot detect the server-side pool state.
Proposals (any subset would help)
Surface hello metadata (num_connections, approximate_connection_time,
host) and disconnect reasons via a public callback or at INFO logging.
Emit a loud warning when num_connections in hello exceeds a threshold
(e.g., 4) while the client manages fewer connections — this is direct
evidence of ghost registrations and imminent interactive-payload loss.
Consider make-before-break reconnects with close-confirmation, or documenting
that reconnect-heavy operation behind NAT can poison the server-side pool.
Happy to share full logs and reproduction notes. We have also filed a parallel
report with Slack developer support regarding the server-side routing/eviction
behavior; will cross-link.
Follow-up with measured data after deploying client-side mitigations (the wire
instrumentation from the report, plus an exponential-backoff idle watchdog:
900 -> 1800 -> 3600 s cap, +-15% jitter, reset only on a delivered envelope).
One night before vs. one night after, same app, same host:
metric (per night)
before
after
idle-watchdog reconnects
73
4
hello num_connections
climbed to 10 (cap)
1 on every hello
too_many_websockets bursts
2 (11 frames)
0
Two observations that may help triage:
Reducing the reconnect rate keeps the primary app's pool clean, which
supports the accumulation model (ghosts are minted by reconnects whose close
never reaches the server).
However, other single-instance apps on the same host still sit at a steady num_connections of 4-5 overnight (network-level flaps still force
occasional SDK reconnects, and each ghost lingers for the ~5 h approximate_connection_time horizon). Client-side tuning cannot get below
that floor — which is why surfacing num_connections (proposal 1/2) and a
server-side eviction or force-disconnect API would still be valuable even
for well-behaved clients.
Both reduce the spurious/concurrent reconnects that are one of the mechanisms minting the ghost registrations you describe, so they should help at the margins.
That said, I don't think these resolve this issue, so I'd like to keep it open. As you framed it, this is the server-side counterpart: the root cause is registrations aging out on Slack's side over ~5 h, which the SDK might not fix directly. Your follow-up data makes the decisive point, after cutting reconnects 73 → 4, other well-behaved single-instance apps still sat at num_connections 4–5 overnight, so reconnect-correctness fixes alone can't get below that floor.
I'm narrowing the scope of this issue to the SDK-actionable proposals that remain unaddressed:
Surface hello metadata (num_connections, approximate_connection_time) and disconnect reasons. Maybe via a public callback and/or INFO logging.
Warn when hello's num_connections exceeds a threshold while the client holds fewer connections.
make-before-break reconnect with close-confirmation, or at minimum documenting the NAT hazard.
Slack SDK version: slack-sdk 3.43.0, slack-bolt 1.29.0
Python: 3.13 (aiohttp Socket Mode client)
OS/platform: Linux container (Docker) on macOS host, behind NAT
Summary
This is part bug report, part mitigation proposal, backed by wire-level data.
When
SocketModeClientreconnects while the network path is degraded (half-openTCP: the close frame never reaches Slack), the old connection remains registered
server-side for an extended period. Repeated reconnects therefore accumulate
"ghost" registrations up to Slack's 10-connection cap (
disconnect: too_many_websockets). Slack then delivers envelopes across all registeredconnections, so most interactive (
block_actions) payloads — which unlikeevents are not retried on non-ack — are silently lost. From the app's
perspective the client looks perfectly healthy: ping/pong fine, events flowing.
Wire evidence
(from a
run_message_listenerswrapper logginghelloanddisconnectframes)helloreportsnum_connections=1disconnect reason=too_many_websockets,then
hello num_connections=10— while the process verifiably held ONEestablished TCP connection to Slack the whole time
num_connectionsis high, most button clicks never arrive on anyconnection we hold; with a clean pool, every click arrives (tested across
message sizes 0.5-5 KB — size is irrelevant)
helloafter aprocess restart reported
num_connections=7for an app that also runs as asingle instance
approximate_connection_time(insidehello.debug_info) isconsistently
18060(~5 h), which sets the ghost age-out horizonWhy this is hard to see with the current SDK
helloenvelope (carryingnum_connections) never reachesmessage_listeners—SocketModeRequest.from_dictrequirestype+envelope_id+payload, so apps cannot observe the most important signal
without wrapping internals.
disconnectframes (includingtoo_many_websockets) are handled byrun_message_listenersbefore the listener loop and only visible at debuglogging.
is_connected()/ ping-pong checks, soclient-side health monitoring cannot detect the server-side pool state.
Proposals (any subset would help)
hellometadata (num_connections,approximate_connection_time,host) and
disconnectreasons via a public callback or at INFO logging.num_connectionsinhelloexceeds a threshold(e.g., 4) while the client manages fewer connections — this is direct
evidence of ghost registrations and imminent interactive-payload loss.
that reconnect-heavy operation behind NAT can poison the server-side pool.
same failure family — this issue is their server-side counterpart, and since
neither is merged/released (latest release is 3.43.0), apps currently have no
upstream remedy at all; that raises the priority of surfacing the diagnostics
from proposals 1-2.
Happy to share full logs and reproduction notes. We have also filed a parallel
report with Slack developer support regarding the server-side routing/eviction
behavior; will cross-link.