Skip to content

[Feature]: capture mode for command and prompt workflow steps #4613

Description

@Quratulain-bilal

Problem

Workflow command and prompt steps always stream output to the terminal. The stdout and stderr fields in the step output dict are always empty strings, making them inaccessible to downstream steps via {{ steps.<id>.output.stdout }}.

This limits workflow authors who need to:

  • Parse AI agent output and branch on its content
  • Pipe command output into a later step as structured data
  • Conditionally act on specific output patterns (e.g. exit messages, generated file paths)

The command step docstring (line 26) notes this as a planned enhancement:

Full stdout/stderr capture is a planned enhancement.

Proposed Solution

Add an opt-in capture: true field to command and prompt step configs. When set, dispatch uses capture_output=True so stdout and stderr are returned in the step output dict and accessible to downstream steps.

# Before (streaming, stdout/stderr always empty):
- id: plan
  type: command
  config:
    command: speckit.plan
    integration: claude

# After (capture mode, stdout/stderr available):
- id: plan
  type: command
  config:
    command: speckit.plan
    integration: claude
    capture: true
    timeout: 300

Downstream steps can then reference:

- id: parse
  type: shell
  config:
    command: "echo '{{ steps.plan.output.stdout | from_json }}'"

Design Notes

  • Backward compatible: capture defaults to false; existing workflows are unaffected
  • Follows existing pattern: ShellStep already captures output with capture_output=True and supports output_format: json
  • Timeout support: A timeout field (seconds, default 600) prevents hung commands from blocking the workflow, consistent with dispatch_command() default
  • Implementation: Requires forwarding stream=not capture to IntegrationBase.dispatch_command() which already supports both modes

Affected Files

  • src/specify_cli/workflows/steps/command/__init__.py - config schema, dispatch call, output assembly
  • src/specify_cli/workflows/steps/prompt/__init__.py - same pattern for prompt steps
  • src/specify_cli/workflows/validators.py - capture (bool) and timeout (int) validation
  • tests/test_workflows.py - capture mode regression tests

Questions

  1. Does this direction align with your roadmap for the workflow engine?
  2. Should capture mode also apply to the prompt step, or only command steps?
  3. Any preferences on the field name (capture vs capture_output vs something else)?

Happy to implement if this is something you'd like to move forward with.

Activity

  1. added
    triage-nice-to-haveVerdict: evidence-backed fix or greenlit feature — land after review
    feature-assessRun the Spec Kit idea-assessment pipeline on this feature request
    on Sep 17, 2026
  2. github-actions commented on Sep 22, 2026

    @github-actions
    Contributor

    Feature assessment — capture-workflow-output · Stage 1/5: Intake

    Idea Intake: Capture mode for command and prompt workflow steps

    Idea (as captured)

    Workflow command and prompt steps always stream output to the terminal. The stdout and stderr fields in the step output dict are always empty strings, making them inaccessible to downstream steps via {{ steps..output.stdout }}.

    Proposed capability: add an opt-in capture: true field to command and prompt step configs. In capture mode, stdout and stderr are returned in the step output dict and can be consumed by downstream steps. The proposal also calls for a timeout field, defaulting to 600 seconds, and forwarding stream=not capture to the existing dispatch mechanism. Existing workflows should remain unchanged because capture defaults to false.

    The issue asks whether this direction fits the workflow-engine roadmap, whether capture should apply to prompt steps as well as command steps, and whether capture or another field name is preferred.

    Restated

    Workflow authors need an opt-in way to retain command and prompt output as step data so later workflow steps can parse it or branch on it. The proposed improvement preserves current streaming behavior by default while exposing captured stdout/stderr and timeout control when explicitly requested.

    Origin & Context

    First-Glance Unknowns

    • [NEEDS CLARIFICATION: Should capture mode apply identically to both command and prompt steps, or only command steps?]
    • [NEEDS CLARIFICATION: Is capture the preferred public configuration name, or should it be capture_output?]
    • [NEEDS CLARIFICATION: What output size, memory, and log-redaction limits are required when retaining output?]
    • [NEEDS CLARIFICATION: Should timeout be introduced or changed for both step types, and what default/range is appropriate?]
    • [NEEDS CLARIFICATION: What exact stdout/stderr behavior and output shape should downstream expressions rely on for failures and timeouts?]

    Generated by 💡 Assess a Feature Request by Installing and Running Spec Kit for #4613 · copilot · gpt52codex · 4.59 AIC · ⌖ 6.17 AIC · ⊞ 24.2K · ◷

  3. github-actions commented on Sep 22, 2026

    @github-actions
    Contributor

    Feature assessment — capture-workflow-output · Stage 2/5: Research

    Idea Research: Capture mode for command and prompt workflow steps

    • Slug: capture-workflow-output
    • Created: 2026-09-22T18:08:43Z
    • Evidence confidence (overall): medium

    Users & Demand

    • The requesting workflow author needs to parse agent/command output and branch or pass structured data to later steps; this is stated directly in issue [Feature]: capture mode for command and prompt workflow steps #4613, but no usage volume or additional user reports are provided — [source: GitHub issue [Feature]: capture mode for command and prompt workflow steps #4613] (confidence: medium, cited)
    • The existing workflow expression language supports references to prior step output and a from_json filter, indicating a documented downstream-consumption use case — [source: docs/reference/workflows.md, Expressions section] (confidence: medium, cited)
    • Observed demand beyond the single issue author is unknown — [NEEDS CLARIFICATION: collect issue/repository usage or additional user requests] (confidence: low, assumption)

    Prior Art

    • shell steps already capture stdout/stderr with subprocess.run(capture_output=True) and optionally parse stdout as JSON under output.data — [source: src/specify_cli/workflows/step/shell/__init__.py, lines 48-87] (confidence: high, cited)
    • command and prompt steps currently document live streaming and always expose an exit code, while describing full stdout/stderr capture as a planned enhancement — [source: src/specify_cli/workflows/step/command/__init__.py, lines 17-29; src/specify_cli/workflows/step/prompt/__init__.py, lines 17-31] (confidence: high, cited)
    • Both command and prompt execution paths already assemble stdout and stderr fields from dispatch results, so the present behavior and the requested behavior may differ at the integration-dispatch boundary rather than in the result schema alone — [source: src/specify_cli/workflows/step/command/__init__.py, lines 170-194; src/specify_cli/workflows/step/prompt/__init__.py, lines 107-137] (confidence: high, cited)
    • Existing prompt validation already checks a positive finite timeout and prompt dispatch accepts a timeout parameter; the proposal's timeout scope may overlap current behavior — [source: src/specify_cli/workflows/step/prompt/__init__.py, lines 92-104 and 145-180] (confidence: high, cited)

    Market & Context

    • The repository's workflow documentation presents prior-step output references and JSON parsing as supported expression capabilities, while warning that prompt output can be influenced by files, tickets, or web content — [source: docs/reference/workflows.md, Expressions and Interpolation sections] (confidence: high, cited)
    • No external market, competitor, adoption, or cost-of-inaction data was supplied or fetched; the issue is the sole demand signal — [NEEDS CLARIFICATION: establish whether comparable workflow engines or integrations set output-size/redaction conventions] (confidence: low, assumption)

    Data & Constraints

    • Capturing output changes process I/O behavior from streaming to retained text and may increase memory use; the issue does not specify maximum output size, truncation, encoding, or secret-redaction rules — [source: GitHub issue [Feature]: capture mode for command and prompt workflow steps #4613] (confidence: medium, cited)
    • Downstream interpolation into a shell step is explicitly documented as potentially unsafe because values are inserted without shell escaping; captured agent output must not be treated as safe command text — [source: docs/reference/workflows.md, Interpolation and shell safety] (confidence: high, cited)
    • Backward compatibility depends on preserving streaming as the default; this is a proposal requirement, not an independently verified implementation constraint — [source: GitHub issue [Feature]: capture mode for command and prompt workflow steps #4613] (confidence: medium, cited)
    • The affected implementation surfaces named by the issue are present, but the repository path for a separate src/specify_cli/workflows/validators.py module is not present in this checkout; validation is distributed in step classes — [source: repository inspection; src/specify_cli/workflows/step/command/__init__.py, prompt/__init__.py] (confidence: high, cited)

    Evidence Against the Idea

    • The strongest counterpoint is scope and safety uncertainty: retaining arbitrary AI/command output can expose secrets, consume significant memory, or be interpolated into shell commands without escaping — [source: docs/reference/workflows.md, Interpolation and shell safety] (confidence: high, cited)
    • The proposal may duplicate existing timeout behavior for prompt steps and may overstate that stdout/stderr are absent from all result paths; the exact dispatch semantics need confirmation — [source: src/specify_cli/workflows/step/prompt/__init__.py, lines 92-104; command/prompt execute paths] (confidence: medium, cited)
    • Demand is currently supported by one issue and no measured workflow adoption data — [source: GitHub issue [Feature]: capture mode for command and prompt workflow steps #4613] (confidence: medium, cited)

    Gaps & Open Questions

    • [NEEDS CLARIFICATION: What are acceptable output-size, truncation, encoding, and memory limits?]
    • [NEEDS CLARIFICATION: How should captured stdout/stderr be redacted or handled when they contain secrets or untrusted agent text?]
    • [NEEDS CLARIFICATION: What does IntegrationBase.dispatch_command() return in streaming versus capture mode, and which step types need a changed call contract?]
    • [NEEDS CLARIFICATION: Is timeout a new requirement for command, prompt, or both, given existing prompt timeout handling?]
    • [NEEDS CLARIFICATION: What measurable adoption or reliability improvement would justify the change beyond the requesting issue?]

    Sources

    • [Feature]: capture mode for command and prompt workflow steps #4613 (host: github.com, policy: allowlisted)
    • src/specify_cli/workflows/step/command/__init__.py (repository source)
    • src/specify_cli/workflows/step/prompt/__init__.py (repository source)
    • src/specify_cli/workflows/step/shell/__init__.py (repository source)
    • docs/reference/workflows.md (repository documentation)

    Generated by 💡 Assess a Feature Request by Installing and Running Spec Kit for #4613 · copilot · gpt52codex · 4.59 AIC · ⌖ 6.17 AIC · ⊞ 24.2K · ◷

  4. github-actions commented on Sep 22, 2026

    @github-actions
    Contributor

    Feature assessment — capture-workflow-output · Stage 3/5: Problem

    Problem Definition: Capture mode for command and prompt workflow steps

    • Slug: capture-workflow-output
    • Created: 2026-09-22T18:08:43Z
    • Inputs used: intake.md | research.md | user input only

    Problem Statement

    Workflow authors cannot reliably use the textual output produced by command and prompt steps as data for later workflow decisions or processing, even though the workflow language supports prior-step output references. This forces authors to lose potentially useful output or use indirect workarounds, while any direct reuse of untrusted text can create shell-safety and data-retention risks.

    Affected Users & Stakeholders

    • Users: Workflow authors — cannot parse or branch on command/agent text from earlier steps.
    • Users: Workflow operators — may need to inspect or process generated paths, status messages, or structured output across steps.
    • Stakeholders: Spec Kit workflow maintainers — own the step execution contract, validation, and compatibility behavior.
    • Stakeholders: Workflow users and repository maintainers — are affected by memory use, log visibility, secret exposure, and shell-injection risks from retained output.

    Goals

    • Make an explicit, documented output-data contract possible for command and prompt steps.
    • Preserve existing live-progress behavior for workflows that do not opt into retained output.
    • Enable downstream steps to consume output predictably, including success, failure, and timeout cases.
    • Make output-retention limits and unsafe downstream usage explicit enough for authors to operate safely.

    Non-Goals

    • Redesigning the workflow expression language or adding arbitrary code execution capabilities.
    • Making captured output safe to interpolate into shell commands; shell safety remains the workflow author's responsibility.
    • Adding broad workflow persistence, external storage, or cross-run output retention.
    • Selecting a final configuration field name, output limit, or timeout default before the open questions are answered.

    Success Metrics

    • A workflow with an opted-in capture setting can reference the intended command/prompt stdout and stderr fields in a later step without changing the default behavior of existing workflows (baseline: unavailable/unknown).
    • Existing workflows without the setting continue to stream output and pass the current workflow test suite (baseline: current behavior; exact pass count unknown).
    • Invalid capture and timeout configurations fail validation or the step cleanly with actionable errors (baseline: unknown for capture; timeout validation exists for prompt).
    • The documented output contract covers both successful and failed dispatches and does not silently exceed an agreed retention limit (baseline: limits unknown).

    Cost of Inaction

    Authors who need output-driven branching or parsing must continue without a supported command/prompt capture contract, limiting workflow composition or encouraging ad hoc workarounds. The current absence also avoids new retention and injection risks, so the cost of inaction should be weighed against the lack of measured demand and the unresolved safety constraints.

    Open Questions

    • [NEEDS CLARIFICATION: Should command and prompt steps share one capture contract, or should prompt response capture be scoped separately?]
    • [NEEDS CLARIFICATION: What maximum output size, truncation behavior, and redaction policy are acceptable?]
    • [NEEDS CLARIFICATION: What exact dispatch return values and output shape should be guaranteed in stream and capture modes?]
    • [NEEDS CLARIFICATION: Is a new timeout requirement needed for command steps, prompt steps, or both?]
    • [NEEDS CLARIFICATION: What adoption or reliability threshold would demonstrate sufficient value?]

    Generated by 💡 Assess a Feature Request by Installing and Running Spec Kit for #4613 · copilot · gpt52codex · 4.59 AIC · ⌖ 6.17 AIC · ⊞ 24.2K · ◷

  5. github-actions commented on Sep 22, 2026

    @github-actions
    Contributor

    Feature assessment — capture-workflow-output · Stage 4/5: Concept

    Concept: Capture mode for command and prompt workflow steps

    • Slug: capture-workflow-output
    • Created: 2026-09-22T18:08:43Z
    • Recommended option: Option A — Bounded opt-in capture

    Options

    Option A — Bounded opt-in capture

    • Sketch: Let workflow authors explicitly opt selected command and/or prompt steps into a documented retained-output mode, while leaving current streaming behavior unchanged for all existing configurations. The result would be available to later workflow expressions under a clearly defined success/failure contract, with explicit limits and warnings for untrusted output.
    • Appetite: medium
    • Trade-offs: Directly addresses the stated need and follows the repository's existing shell-step capture pattern, but requires agreement on output size, redaction, timeout overlap, and whether command and prompt response capture are truly equivalent.
    • Rabbit holes: Unbounded output retention, inconsistent integration behavior, secret leakage in state/logs, and unsafe interpolation into shell steps could expand the work substantially.

    Option B — Capture only command-step output

    • Sketch: Provide a narrowly scoped retained-output capability for command steps first, treating prompt response capture as a separate future decision. Keep prompt steps on their current exit-code/live-progress contract until demand and safe response handling are better understood.
    • Appetite: small
    • Trade-offs: Reduces initial scope and risk while addressing structured command workflows, but leaves the issue author's prompt use case unresolved and creates two related contracts to explain and maintain.
    • Rabbit holes: Authors may expect parity between command and prompt steps; integration-specific command behavior and output limits still need definition.

    Option C — Do nothing for now

    • Sketch: Keep command and prompt output streaming and expose only the current exit-code-oriented results, while gathering more demand and clarifying safety and retention requirements before changing the contract.
    • Appetite: small
    • Trade-offs: Avoids new memory, privacy, and shell-safety exposure and requires no compatibility change, but preserves the reported inability to compose workflows around textual output.
    • Rabbit holes: Workflows may continue using unsupported workarounds, and the documented prior-step output capabilities remain incomplete for these step types.

    Recommendation

    Prefer Option A at concept level because it addresses the stated composition problem while preserving backward compatibility and aligning with the existing shell-step capture precedent. It should not advance to specification until output limits/redaction, exact dispatch semantics, prompt-versus-command scope, and timeout overlap are answered; if those questions cannot be resolved, Option B or C is safer.

    Out of Scope (for the recommended option)

    • Redesigning expressions, adding shell escaping that would make arbitrary interpolated text safe, or sandboxing shell steps.
    • Persisting captured output beyond the workflow run or introducing external output storage.
    • Guaranteeing identical content across integrations whose CLIs expose different stream behavior.
    • Adding unrelated workflow step types, routing, or agent features.

    Assumptions to Validate

    • The existing dispatch layer can distinguish streaming from retained output without breaking current integration behavior.
    • A bounded output contract is sufficient for the primary parsing and branching use cases.
    • Workflow authors understand that retained agent output is untrusted and must not be inserted into unconstrained shell commands.
    • Prompt response capture has enough value to justify inclusion in the same initial contract as command output.
    • The current issue represents a meaningful workflow need beyond one request.

    Generated by 💡 Assess a Feature Request by Installing and Running Spec Kit for #4613 · copilot · gpt52codex · 4.59 AIC · ⌖ 6.17 AIC · ⊞ 24.2K · ◷

  6. github-actions commented on Sep 22, 2026

    @github-actions
    Contributor

    Feature assessment — capture-workflow-output · Stage 5/5: Decision — verdict needs-clarification

    Decision: Capture mode for command and prompt workflow steps

    • Slug: capture-workflow-output
    • Decided: 2026-09-22T18:08:43Z
    • Verdict: needs-clarification
    • Artifacts reviewed: intake.md | research.md | problem.md | concept.md

    Scorecard

    Criterion Rating Justification
    Problem validity adequate The issue identifies a concrete gap between documented prior-step output expressions and command/prompt text availability; the repository documents exit-code access and planned full capture.
    Evidence strength adequate The request and repository implementation/documentation provide cited evidence, but demand is represented by one issue and lacks usage data or external validation.
    Value vs. inaction adequate Output-driven workflow composition would address the reported limitation, while inaction avoids new retention and safety costs; the relative value is not measured.
    Feasibility / appetite adequate A bounded opt-in concept follows the existing shell-step capture precedent and is plausibly medium appetite, but dispatch semantics and limits are unconfirmed.
    Strategic fit unknown No explicit constitution or roadmap decision about workflow output capture was identified in the assessment evidence.
    Risk posture weak The artifacts identify unresolved output-size, redaction, memory, timeout, and unsafe shell-interpolation risks; no mitigation contract has been agreed.

    Verdict & Rationale

    Needs clarification. The problem is credible and a bounded opt-in concept is plausible, but the assessment cannot responsibly hand off to specification while core behavior and safety boundaries remain undefined. Resolve the blocking questions below, then revisit research/shape before deciding whether the full command-and-prompt scope is justified.

    If needs-clarification

    • Blocking questions:
      • [NEEDS CLARIFICATION: What maximum output size, truncation behavior, encoding, and memory policy apply to retained stdout/stderr and prompt responses?]
      • [NEEDS CLARIFICATION: What redaction or handling rules apply to secrets and untrusted agent text in captured output and persisted run state?]
      • [NEEDS CLARIFICATION: What exact stream-versus-capture contract does the dispatch layer provide for each supported integration, including failures and timeouts?]
      • [NEEDS CLARIFICATION: Should the initial scope include prompt response capture, command output capture, or both?]
      • [NEEDS CLARIFICATION: Is timeout a new command-step requirement, an existing prompt-step behavior, or a separate concern?]
      • [NEEDS CLARIFICATION: What usage/adoption or reliability threshold demonstrates enough value beyond the single issue request?]
    • Revisit stage: research and shape

    Generated by 💡 Assess a Feature Request by Installing and Running Spec Kit for #4613 · copilot · gpt52codex · 4.59 AIC · ⌖ 6.17 AIC · ⊞ 24.2K · ◷

  7. mnriem commented on Sep 23, 2026

    @mnriem
    Collaborator

    @Quratulain-bilal Please look at the assessment and add clarifications where appropriate

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    feature-assessRun the Spec Kit idea-assessment pipeline on this feature requesttriage-nice-to-haveVerdict: evidence-backed fix or greenlit feature — land after review

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions