Skip to content

fix(channel): keep hosted runs resident while busy - #186

Merged
iamnbutler merged 1 commit into
mainfrom
cloud-pilot-ops/cloud-pilot
Oct 7, 2026
Merged

iamnbutler merged 1 commit into
mainfrom
cloud-pilot-ops/cloud-pilot

Conversation

@iamnbutler

Copy link
Copy Markdown
Contributor

A hosted channel's Durable Object hibernated while a chat was busy and a workspace tool call was pending. The workspace socket is hibernatable and a silent shell sends nothing, so once no client sockets were open the object went idle. The keepalive alarm then woke a new instance. pi resumed and told the model the shell was interrupted, but the process kept running on the workspace. A later Kill could not reach it, because the new instance had no record of the old call.

Fix: onBusy already reports busy transitions, including the initial snapshot taken before harness.resume(). HostedChannel now uses those transitions to hold one pending timer while any chat is busy. Each timer re-arms itself after KEEPALIVE, so every hold is a new, bounded operation, and the busy-to-idle transition clears it synchronously. A pending timer stops Cloudflare from hibernating the object. Idle and dormant channels have no timer and hibernate as before. The durable alarm path is unchanged and still brings back an object that the runtime resets. There are no changes to storage, packages/channel, or the workspace protocol. The architecture doc's hosted-channels section explains active-run residency vs idle hibernation.

Fixes #183. Refs #176, #13.

Real Cloudflare evidence (dedicated ace-channel-pilot / ace-directory-pilot workers, one hosted channel on an isolated host, model anthropic/claude-sonnet-5-5, tools on a local Mac workspace). The existing deployed services were not touched.

  • Before the fix (ad3f731): a detached 75s silent shell with no cloud clients was reported to the model as "interrupted and may have partially run" about 38s in. Its PID kept running. The workspace logs show new cloud call generations (2b08a256, then 387df656) and no workspace.drop. A 120s repeat was interrupted about 30s in. Kill afterwards returned ok, but the original PID kept running and wrote its end marker.
  • After the fix, with the same channel and history kept across deploys:
    • Detached 75s silent shells finished in a single call generation with error=false and no interruption. This passed on both the interval draft and the final timer version.
    • Kill after 56s of a detached silent 120s shell, with no clients: the cancel came from the original generation. The PID and its child were gone within 3s, and no end marker was written after the deadline.
    • Stop and Kill at about 13s also removed both processes.
    • After the run finished, the idle object hibernated. The next tool call showed a new generation and the workspace socket survived.
  • Same channel across 5 deploys: full protocol replays (entry, call, and result identities and values) matched before and after each deploy, and earlier snapshots are exact prefixes of the final replay.
  • Real runtime reset: ctx.abort() was injected through a temporary pilot-only wrapper that was never committed and has since been removed.
    • Mid-shell reset: workspace dropped with code 1006 and reconnected within about 1s. The workspace killed the old process, and pi reported the call as interrupted with no replay.
    • Reset after the first streamed delta: pi kept the partial reply marked stopped and then produced a complete regenerated reply.
  • Workspace offline/online: with the workspace host stopped, human messages persisted and tools returned "The workspace is offline". After restarting the host, the workspace reconnected and the shell worked.
  • Checks: bun types and bun run ci pass.

Not verified here: a second physical machine, and sleep/network loss on a physical host. A plain redeploy during an active shell did not measurably reset the object, so reset coverage comes from ctx.abort() only.

Built in Ace

The following ran in the Ace operator channel cloud-pilot-ops, using Ace's lane, shell, and file tools:

  • Code and doc edits in Ace lane cloud-pilot
  • Pilot Worker deploys and secret installation
  • Real model and tool validation against the hosted channel
  • bun types / bun run ci
  • Git commit, push, and this PR

Codex (outside Ace) coordinated the investigation and review, inspected local evidence, and organized the GitHub issues (#176, #183), because Ace cannot yet host the Codex harness (#9). The first repository and credential-readiness inspection and the tracker commands also ran in Codex's shell. That was an external escape, not an Ace shell limitation.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Hosted idle hibernation interrupts active tools while local shell keeps running

1 participant