Skip to content

Verify a deployed hosted channel with a local workspace #176

Description

@iamnbutler

Run a fresh, team-deployed hosted channel on Cloudflare while its workspace stays on a real team machine. The channel, model loop and durable pi state live in the cell; repository files, shell processes, lanes, browser and desktop tools remain on the workspace. This is the first milestone of #13, not a cloud execution environment.

Status — October 7

The initial deployed qualification passed on one physical workspace machine after reproducing and fixing #183. PR #186, reviewed head 4d534550b0a36d9c8272113018c71c0b9b199830, merged as 80c67ca on October 7 at 13:37:52 UTC after GitHub CI passed; #183 is closed. The deployed evidence remains tied to the pilot versions below. This pilot stays open for the actual second-machine, teammate directory-discovery and physical sleep/network-partition checks below.

Detailed scenario results and repeat procedure cover T1–T16, including the pre-fix failures. Real model: anthropic/claude-sonnet-5-5; files, lanes and tools stayed on the local Mac.

  • Dedicated channel service, initial version 360a7bca-a1fc-4bbb-8fa5-e64b599a8273 from source ad3f7313ea18f2fd9c5f9e3fb2ca9c5b1df67861.
  • Dedicated directory service, version 0df59e7b-4f07-41a7-9722-33b232398e07.
  • Final channel-service version b58efd2d-579e-4c70-a9aa-c326bca1e086 includes the PR fix(channel): keep hosted runs resident while busy #186 fix and uses the original service entry point. The temporary pilot-only ctx.abort() wrapper was removed, history was compared again, and shell work succeeded afterward.
  • The same hosted channel and history were preserved throughout. Existing non-pilot services were unchanged; the isolated pilot remains available for the next acceptance pass.

Existing foundation

The channel service, workspace connection and ace new --hosted <url> landed in 3ff0b97; 1a2e1f3 fixed the first connection. PR #85 records real model/tool validation in a local Workers/Durable Objects runtime. That evidence does not establish deployed Cloudflare behavior.

Acceptance

Checked items below are supported by the recorded deployed evidence. PR #186 has merged; the whole pilot remains incomplete until the remaining physical-machine and discovery checks pass.

  • Deploy dedicated pilot channel/directory services from recorded source revisions, use an isolated project/profile and a new channel, and preserve existing services/data/credentials. Service URLs and versions are recorded above without secrets.
  • Create the hosted channel with the supported CLI and complete real model/local shell and lane file read/write work. Verify actual local execution and durable transcript, tool results, metadata and usage (T1/T3).
  • Redeploy the service against the same channel, compare full protocol replay and channel info including usage, reopen and continue work without resetting the store or replacing the channel (T6). A plain deployment during an active shell also completed without duplicate execution (T10), but did not measurably reset the runtime.
  • Cleanly stop the workspace host with SIGTERM, retain cloud human chat, return an explicit offline tool outcome, then restart/reconnect and complete shell work (T15).
  • Verify physical workspace sleep/wake and network loss/partition recovery. Clean SIGTERM is not evidence for these cases.
  • Open the hosted channel through Ace on another actual team machine: verify attributed human messages, live updates/reconnect, permitted model invocation and execution remaining on its workspace. Another context or self-tailnet loopback is not a second machine.
  • Verify a teammate host discovers and reaches the hosted channel through the pilot directory; exercise its offline/reconnect presentation. This was not simulated with a second host process on the same machine.
  • Cause a real runtime reset during an active shell and separately after the first model-stream delta using temporary pilot-only ctx.abort() (T13/T14). Confirm shell process cleanup, unsafe-tool interruption without replay, preserved historical prefix and model continuation. The stopped partial reply plus regenerated complete reply is recorded as observed behavior, not an established defect. Reset coverage comes from this fault injection, not plain deployment.
  • Verify a detached silent 75-second tool completes once after the residency fix; exercise Stop, ordinary Kill and Kill after 56 seconds with no client watching. Confirm parent/child cleanup, no completion marker after cancellation/deadline, and a usable channel afterward (T4 final/T8/T9/T11).
  • Observe idle hibernation after work ends, then reopen and execute successfully while preserving the workspace socket (T7). This establishes observed residency/hibernation behavior, not long-term resource or billing guarantees.
  • Publish source/deployment and scenario evidence; retain the pre-fix Hosted idle hibernation interrupts active tools while local shell keeps running #183 reproduction, distinguish passed checks from unavailable cases, and link the defect from [Meta] Run hosted channels in the cloud with local workspaces #13 and [Meta] Dogfood Ace end to end #5.
  • Review and merge fix(channel): keep hosted runs resident while busy #186 so the validated Hosted idle hibernation interrupts active tools while local shell keeps running #183 fix is part of the supported source. Merged as 80c67ca; Hosted idle hibernation interrupts active tools while local shell keeps running #183 closed automatically.

Use real runtimes, models and provider credentials. Keep channel state in pi-durable and the channel package runtime-neutral. Existing channels are permanent data; any format change requires a versioned migration checked against an existing-channel backup.

App setup, general hosted backup/restore, host-independent lifecycle/discovery and outbound messaging are later focused work in #13. This pilot can expose their limits without silently broadening its scope. Execution-host handoff #49, directory push #14 and cross-gateway terminals #11 are related work, not prerequisites.

Tracked by #13; deployed-path dogfooding evidence for #5. No acceptance checkbox above is complete merely because the implementation exists or local Workers checks passed.

Activity

  1. iamnbutler commented on Oct 7, 2026

    @iamnbutler
    ContributorAuthor

    Historical pre-fix report. The completed initial qualification and candidate regression results supersede the in-progress status below.

    Initial deployed pilot evidence on October 7, using source ad3f7313ea18f2fd9c5f9e3fb2ca9c5b1df67861 and the dedicated real Cloudflare ace-channel-pilot service:

    Scenario Status Evidence
    T1: real model and local shell Passed A real provider/model run used the actual hosted-channel/workspace route and executed its shell work locally.
    T3: actual lane file read/write Passed The hosted run read and changed a file in the real local lane; work remained on the workspace.
    T4: detached, quiet in-flight shell Failed — #183 A 75-second silent shell kept running locally after idle hibernation reconstructed the hosted runtime and pi recorded the tool as interrupted. The workspace socket remained connected; there was no deploy or deliberate disconnect.
    T5: Kill after the call was orphaned by hibernation Failed — #183 A 120-second shell outlived pi's interrupted result and the agent's final reply. Kill returned successfully while the old shell and its sleep child remained alive.

    The #183 reproduction started the shell at 11:35:03 UTC, reconstructed the runtime at 11:35:38 UTC and observed the original shell's end marker at 11:36:18 UTC. A later wait was interrupted again after another recovery. The model avoided replaying the original command; this is not evidence that arbitrary side effects are safe to continue across that loss of pending-call state.

    Bounded replay, process and runtime-generation artifacts are retained by the pilot operator. #183 is a child issue of this pilot and the first blocker linked from #13; it is also recorded in #5.

    T5 started PID 20265 (sleep child 20268) at 11:37:09 UTC in generation fdc6b613:6. Pi reported interruption roughly 30 seconds later; the model finished without retrying at 11:37:38 UTC. Hosted Kill at 11:38:12 UTC returned successfully with pi already idle, but both original processes remained alive afterward. The operator retained the T5 replay, process checks, Kill timestamp and shell marker evidence.

    This is a demonstrated failure to cancel the call already orphaned by hibernation, not proof that ordinary Kill fails while a current active workspace call remains attached. The service busy/residency fix is being prepared. The complete evidence matrix and remaining upgrade, interruption, cancellation, idle and second-machine results will follow; no acceptance checkboxes are changed by this partial report.

  2. iamnbutler commented on Oct 7, 2026

    @iamnbutler
    ContributorAuthor

    October 7 deployed qualification update: the initial one-machine pilot and #183 regression checks are complete. PR #186, head 4d534550b0a36d9c8272113018c71c0b9b199830, is ready for review and unmerged. GitHub CI passed at 11:55:09 UTC; bun types and bun run ci also passed in the Ace operator channel.

    The dedicated services are ace-channel-pilot and ace-directory-pilot. Initial source was ad3f7313ea18f2fd9c5f9e3fb2ca9c5b1df67861. The service fix is the one in PR #186; the final deployment is b58efd2d-579e-4c70-a9aa-c326bca1e086, using the original service entry point with the temporary reset wrapper removed. The same hosted channel was retained throughout; existing non-pilot services were unchanged.

    Scenario Observed result
    T1/T3: real model and local lane tools anthropic/claude-sonnet-5-5 executed shell work on the local Mac and read/wrote an actual lane file.
    T6: deployment persistence Full protocol replays and channel info, including usage, matched before/after service deployments. Earlier histories are exact prefixes of later histories. No replacement channel or store reset.
    T4/T5 before fix Quiet work was interrupted by hibernation while its shell continued; subsequent Kill could not reach the orphaned call. Tracked in #183.
    T4 after final fix A detached silent 75-second shell completed once in the same generation, with no client watching, interrupted result, retry or workspace drop.
    T7: idle afterward After the run ended, the cell hibernated; a later successful tool used a new generation while the workspace socket survived. This observes residency/hibernation, not a billing benchmark.
    T8/T9: ordinary Stop/Kill Both shell and child exited within three seconds; no completion marker appeared, pi recorded abort and the channel stayed usable.
    T11: delayed Kill regression A 120-second silent shell had no clients for 56 seconds. Kill came from the original generation, removed shell and child within three seconds, and no completion marker appeared after the original deadline.
    T10: plain deploy during shell The 60-second shell completed once across deployment. No actual runtime reset was observed, so this is continuity evidence rather than reset-recovery evidence.
    T13: actual reset during shell Temporary pilot-only ctx.abort() caused a real socket drop and reconnect in about one second. The old processes were gone within five seconds, no end marker appeared, and pi recorded interruption without replaying the unsafe command. Prior history stayed intact.
    T14: actual reset during streamed model reply Reset after the first streamed delta retained the stopped partial reply and resumed with a fresh complete reply. No tool side effects occurred; prior history stayed intact. This observed recovery behavior is not being filed as a defect without an established contrary expectation.
    T15: clean host shutdown/reconnect With the workspace host stopped by SIGTERM, human chat persisted and tools returned “The workspace is offline.” Restart restored the connection and successful shell work. This is not a physical sleep/network-partition test.
    T16: unwrapped final deployment The reset wrapper was removed; unchanged history and a successful shell run were verified on the final original service entry point.

    To repeat #183's regression: dispatch a silent shell longer than the former hibernation window, close all ordinary channel clients, confirm one start/result in the same generation, then repeat with Kill after at least 56 seconds and inspect both parent/child exit and absence of an end marker after the original deadline. After the run ends, verify an idle cell can hibernate and reopen. Genuine-reset checks require the separately identified pilot fault injection; a plain deployment must not be reported as an observed reset.

    Still open: review/merge #186; another actual physical client/teammate machine, including attributed collaboration and pilot-directory discovery; physical workspace sleep/wake and network partition/recovery. No second host was simulated with another context on this machine. Keep this pilot open for those checks. The isolated pilot deployment, channel and workspace connection remain available for the next acceptance pass.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions