Repository navigation
feat(e2e): preserve Kubernetes clusters for debugging #3675
Copy link
Copy link
Open
Labels
state:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Sep 24, 2026 /assign
🏗️ build-plan
Implementation Plan
Issue type:
feat
Complexity: Low
Confidence: High — clear pathSummary
Add an opt-in preservation mode for wrapper-created Kubernetes E2E clusters while retaining default and caller-owned-context cleanup behavior. Validate the option before creating resources and print safe inspection and cleanup guidance for preserved credentials and state.
Scope
e2e/with-kube-gateway.sh: validate the option, preserve only wrapper-owned clusters and workdirs, retain failure diagnostics, and print inspection and cleanup commands.tasks/scripts/test-kubernetes-e2e-preserve.sh: cover validation, preservation, default cleanup, and caller-owned contexts with deterministic command stubs.tasks/test.toml: include the focused test in the standard test suite..agents/skills/helm-dev-environment/SKILL.md: document ownership semantics, retained credentials, inspection, and cleanup.
Implementation Steps
- Validate and normalize
OPENSHELL_E2E_KUBE_PRESERVE_CLUSTERbefore creating the work directory or cluster. - Preserve wrapper-owned ephemeral state after failure diagnostics but before destructive cleanup.
- Print cluster, context, kubeconfig, workdir, inspection command, both cleanup commands, and a temporary-credential warning.
- Add deterministic regression tests for preserved, default, invalid, and external-context behavior.
Test Plan
- Unit/focused shell tests: Exercise wrapper lifecycle behavior through stubbed cluster commands.
- Integration:
bash -n, focused task, andmise run pre-commit. - E2E: Run the focused Kubernetes wrapper lifecycle path without requiring a live cluster; the change controls the harness lifecycle rather than OpenShell runtime behavior.
Risks & Open Questions
- Preserve the wrapped command's exit status and safely quote paths that may contain spaces.
- Preserve only wrapper-created clusters; never reinterpret caller-owned contexts.
Documentation Impact
Contributor workflow documentation only; no published product or gateway configuration documentation changes.
Revision 1 — initial plan
🏗️ build-from-issue-agent
Implementation Complete
PR: #3680
What was built
Added opt-in preservation for wrapper-created Kubernetes E2E clusters and work directories, with ownership-safe cleanup behavior and explicit credential warnings.
Tests
- Focused shell coverage: invalid input, success/failure preservation, default cleanup, and caller-owned contexts
mise run pre-commit- Live ephemeral k3d failure-path validation
Docs updated
.agents/skills/helm-dev-environment/SKILL.md
The issue will auto-close when the PR is merged.
- added 2 commits that reference this issue
on Sep 24, 2026
Metadata
Metadata
Assignees
Labels
state:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
User Story
As an OpenShell contributor debugging a Kubernetes E2E failure, I want to preserve the ephemeral cluster and test artifacts after the run so that I can inspect the failed state directly.
Problem Statement
The Kubernetes E2E wrapper always deletes the k3d cluster and temporary work directory from its exit trap. Its failure diagnostics capture common resources, but they cannot anticipate every investigation and remove the live state needed for follow-up commands.
Impact / Why This Matters
Contributors must reproduce a potentially slow or intermittent failure, modify the wrapper locally, or run the setup manually. This increases investigation time and can hide evidence that exists only immediately after a failure.
Proposed Design
Add an opt-in
OPENSHELL_E2E_KUBE_PRESERVE_CLUSTER=1mode for ephemeral clusters created by the E2E wrapper. When enabled, the wrapper preserves its cluster and work directory and prints their locations plus explicit cleanup commands. Existing cleanup remains the default, and externally supplied clusters are not reclassified as wrapper-owned resources.Preserved work directories can contain short-lived test credentials, so the output and documentation must tell contributors to remove them after investigation.
Acceptance Criteria
OPENSHELL_E2E_KUBE_PRESERVE_CLUSTER=1preserves an ephemeral cluster created by the wrapper and its work directory after success or failure.Alternatives Considered
Editing the exit trap locally is error-prone and not reproducible. Reusing a manually managed cluster helps some investigations but does not preserve the exact state produced by the standard ephemeral E2E workflow.
Agent Investigation
The cleanup trap in
e2e/with-kube-gateway.shcollects a fixed diagnostic bundle, uninstalls fixtures, deletes a wrapper-created k3d cluster, and removes its temporary work directory. An early, opt-in preservation branch can retain the live environment without changing default behavior.Checklist