Skip to content

Paired hosts echo channel listings forever, starving the gateway #200

Description

@iamnbutler

Observed while dogfooding (Ace Canary 0.0.16/0.0.17, two Macs on the tailnet): opening a channel shows "Loading chat" for 20–60 s, on the peer host and locally.

Evidence on natembp's Ace Helper (pid 64359):

  • Gateway request durations rose steadily from ~250 ms (14:39Z) to ~6 s (14:50Z), starting right after the peer nate16-1 connected (14:38:57Z), including local-only requests.
  • The helper sat at 83–97% CPU; sample shows the main thread busy in synchronous openat/read/close/access/kill/getdirentries — the catalog listing (local()).
  • One authenticated loopback WebSocket, one channels request: handshake 9.4 s, then 45 unsolicited identical channels frames (1 unique) in 0.6 s, ~1.97 MB (~75 frames/s). Loopback /health timed out at 3 s/8 s.

Cause: a peer's channels frame makes GatewayClient emit, peers.watch calls broadcast(), and broadcast() sends to every socket, including the inbound socket of that same peer's client. Both hosts echo unchanged listings to each other forever, re-reading the catalog for each frame. Every new trigger adds another circulating frame, so latency keeps growing until the helper restarts.

Fix: peer and directory changes go to the owner's app sockets only; only this host's own catalog changes go to peers. Local listings are computed once per broadcast.

Activity

  1. iamnbutler commented on Oct 7, 2026

    @iamnbutler
    ContributorAuthor

    Signed Canary 0.0.19 (7b55b4a, #203; build-only run 37641790495) is installed and reopened on both natembp and Nate16.

    Install

    • Version confirmed on both Macs.
    • On Nate16, the transferred DMG matched SHA-256 9fa18ad1a628de14f4e97f121cbac01cb9054fb79eb1b17f2f01c4eab9f033df, and the staged app passed Gatekeeper as a notarized Developer ID build.
    • On both Macs, the old helpers exited normally and the new app started on the existing data. Afterwards, the installer mounts were detached and the temporary transfer server was stopped.

    natembp (source host), before → after

    • Ace Helper CPU: 83–97% → 0.0%.
    • Loopback WebSocket connect: 9,403 ms → 11–14 ms.
    • /health: timed out at 3 s and 8 s → 0.5–5.2 ms.
    • Channel watch until live and transcript complete: 262 ms for modal-top-offset, 25 ms for canary-github-releases.
    • After the update, the resume receipt cleared and the workers resumed.
    • In the real app, switching to a channel showed its content immediately (within one click and a 0.5 s snapshot).

    Nate16 (viewing natembp)

    • The updated app reopened, and canary-github-releases showed the full transcript from natembp.
    • We don't have a precise remote timing: the user went back to using Screen Sharing, so that measurement wasn't run.
  2. iamnbutler commented on Oct 7, 2026

    @iamnbutler
    ContributorAuthor

    Additional read-only check of Nate16's gateway (ws://Nate16:4142/ws), Canary 0.0.19:

    • The connection opened in 137 ms.
    • hello identified nate16-1 and replied in 12 ms.
    • channels replied in 12 ms with one local channel.
    • Over 10,001 ms the connection received exactly two frames, both replies to these requests: no unsolicited channel listings, errors, or failed replies. The connection was then closed.

    The listing loop is not running on the remote host either. How long a remote transcript takes to load is still not measured.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions