Skip to content

[Bug]: preset-provided scripts are never resolved — type: script and $CORE_SCRIPT have no effect #4551

Description

@pavel-neumann

Bug Description

The preset reference documents scripts as a first-class layer of the resolution stack:

"Presets can provide command files, template files (like plan-template.md), and script files.
Templates and scripts are looked up from the stack when Spec Kit needs them.
Scripts support replace and wrap; script wrappers use $CORE_SCRIPT as the placeholder."

and lists the project-local override location as .specify/templates/overrides/scripts/.

None of this takes effect. A preset may declare type: script, and specify preset add accepts and
installs it without complaint, but nothing ever asks the resolver for a script, so the file that
actually runs is always the core one.

Three independent places in the shipped code show why:

  1. No caller requests the script type. The only production call into preset composition
    determines the type as
    template_type = "command" if is_command else "template"
    (specify_cli/presets/_commands.py:495). The value "script" is never passed, so
    PresetResolver.resolve() / collect_all_layers() are never invoked for scripts — even though
    both handle template_type == "script" (specify_cli/presets/__init__.py:5524, 5531, 5613,
    and the $CORE_SCRIPT substitution at :6168).
  2. The Bash runtime resolves only .md. resolve_template() in
    core_pack/scripts/bash/common.sh:509 looks for "$base/overrides/${template_name}.md", and
    every subsequent tier appends .md as well. There is no .sh branch and no
    overrides/scripts/ lookup anywhere in the file.
  3. The Python twin mirrors that. core_pack/scripts/python/common.py:308 documents its order as
    "mirrors resolve_template in scripts/bash/common.sh" and resolves
    overrides/f"{template_name}.md" only.
    So the composition engine for scripts exists and is reachable by unit test, but is not wired to
    anything that runs.

This looks like the script half of #2132 / #2133 never landed on the runtime side, while #3143
documented it as shipped. I could not find an existing report: no open issue or PR with "script" in
the title covers it (#1964 is adjacent but is about extensions and has been open since March 2026),
and #4443 / #4445 concern PyYAML availability, not resolution.
runs the command.
Suggestion. Either wire a caller for template_type="script" and add .sh resolution to
common.sh and its Python twin, or — if scripts are not meant to be resolvable yet — adjust
docs/reference/presets.md and have preset.yml validation reject type: script so the
declaration fails loudly instead of silently.

Steps to Reproduce

# 1. Scaffold a project
mkdir demo && cd demo && git init
specify init . --integration claude --script sh --force
 
# 2. Record the core script's hash
sha256sum .specify/scripts/bash/setup-plan.sh
 
# 3a. Try the documented project-local override
mkdir -p .specify/templates/overrides/scripts
printf '#!/usr/bin/env bash\necho "OVERRIDE RAN"\n' \
  > .specify/templates/overrides/scripts/setup-plan.sh
chmod +x .specify/templates/overrides/scripts/setup-plan.sh
 
# 3b. Or try a preset that provides the script
mkdir -p /tmp/p/scripts
cat > /tmp/p/preset.yml <<'YAML'
schema_version: "1.0"
preset:
  id: "gate"
  name: "Gate"
  version: "1.0.0"
  description: "Wraps setup-plan."
requires:
  speckit_version: ">=1.0.0"
provides:
  templates:
    - type: "script"
      name: "setup-plan"
      file: "scripts/setup-plan.sh"
      strategy: "wrap"
YAML
printf '#!/usr/bin/env bash\necho "BEFORE"\n$CORE_SCRIPT "$@"\n' > /tmp/p/scripts/setup-plan.sh
specify preset add --dev /tmp/p
 
# 4. Re-scaffold and compare
specify init . --integration claude --script sh --force
sha256sum .specify/scripts/bash/setup-plan.sh

Expected Behavior

.specify/scripts/bash/setup-plan.sh is composed from the stack, so the wrapper runs first and
delegates to the core script through $CORE_SCRIPT — matching the documented behaviour that
"templates and scripts are looked up from the stack when Spec Kit needs them".

Actual Behavior

The hash in step 4 is identical to step 2 and still matches
.specify/integrations/speckit.manifest.json. The override file and the preset-provided script are
never read. specify preset add reports success and specify preset list shows the preset as
enabled, so there is no signal that the declared script is inert.

Tried additionally, all without effect: strategy: replace as well as wrap,
specify integration upgrade, specify init --here --force. Same result on 1.0.1.

Specify CLI Version

1.0.6

AI Agent

Claude Code

Operating System

macOS 26.6.2

Python Version

Python 3.11.15

Error Logs

# There is no error output — that is part of the problem.
# A preset declaring a script installs cleanly and is reported as active:
 
$ specify preset add --dev /tmp/p
Installing preset from /tmp/p...
✓ Preset 'Gate' v1.0.0 installed (priority 10)
 
$ specify preset list
Installed Presets (in resolution order — highest precedence first)
 
  Gate (gate) v1.0.0 — enabled — priority 10
    Wraps setup-plan.
    Templates: 1
 
# …while the script it provides is never resolved or executed.

Additional Context

Why this matters. The documented mechanism is the natural way to add a step in front of a core
script: the preset wraps setup-plan.sh, the core script stays untouched underneath, and Spec Kit
maintains the relationship. Any preset that needs to run a check, a fetch, or a validation before
the command body reads its inputs wants exactly this.

Current workaround, for anyone hitting the same wall. Override the command instead of the
script. A preset providing speckit.plan with strategy: "wrap" and a five-line body:

---
scripts:
  sh: scripts/bash/setup-plan-wrapper.sh --json
---
{CORE_TEMPLATE}

points the materialised command at a project-owned script, which ends with
exec "$ROOT/.specify/scripts/bash/setup-plan.sh" "$@". This works and survives upgrades, because
{CORE_TEMPLATE} keeps pulling the command body from core.

It has two costs that the documented mechanism would remove:

  • two artefacts instead of one. The script cannot travel inside the preset, so it lives beside
    it as a separate committed file, and it needs a distinct filename because it cannot replace
    setup-plan.sh.
  • nothing validates the link. A preset whose frontmatter names a script that does not exist
    installs without any warning and is listed as healthy; the dead path only surfaces when a user
    runs the command.
    Suggestion. Either wire a caller for template_type="script" and add .sh resolution to
    common.sh and its Python twin, or — if scripts are not meant to be resolvable yet — adjust
    docs/reference/presets.md and have preset.yml validation reject type: script so the
    declaration fails loudly instead of silently.

Activity

  1. github-actions commented on Sep 12, 2026

    @github-actions
    Contributor

    Bug assessment — preset-scripts-inert: Valid · severity low


    Bug Assessment: preset-provided scripts are never resolved

    Report (summarized)

    Issue author @pavel-neumann reports that the preset reference documents script layers, including project overrides under .specify/templates/overrides/scripts/, type: script preset entries, replace/wrap strategies, and $CORE_SCRIPT. In a 1.0.6 macOS/Python 3.11.15 setup, a documented override and a preset wrapping setup-plan.sh were accepted but did not affect the materialized script; no error was emitted. The report provides a workaround using a wrapped command that points to a separate project-owned script.

    No additional issue comments were present, and no external URL needed to be fetched.

    Symptom

    Script preset declarations and project-local script overrides are accepted and displayed as installed, but the runtime script resolution path continues to use the core .sh file. The expected behavior is for script layers to be resolved and composed, including $CORE_SCRIPT delegation for wrappers.

    Reproduction

    1. Create and initialize a project with specify init . --integration claude --script sh --force.
    2. Add .specify/templates/overrides/scripts/setup-plan.sh and make it executable.
    3. Alternatively, install a preset whose manifest declares type: script, name: setup-plan, a script file, and strategy: wrap using $CORE_SCRIPT.
    4. Re-run initialization or the relevant workflow and compare the generated .specify/scripts/bash/setup-plan.sh with the core script.
    5. The reported result is that the override/wrapper is not used, while specify preset add and specify preset list report success.

    The report does not establish whether the same behavior was exercised on all supported operating systems or with the Python/PowerShell runtime variants: [NEEDS CLARIFICATION: cross-platform scope].

    Suspected Code Paths

    • src/specify_cli/presets/_commands.py:521-523 — the production preset resolve command selects only command or template based on the name; no runtime caller in the searched code requests script composition.
    • src/specify_cli/presets/__init__.py:271, 490-505 — preset manifests explicitly accept type: script and validate script strategies, so installation/validation does not reject the feature.
    • src/specify_cli/presets/__init__.py:5536-5572 and 5817-5861 — PresetResolver.resolve() and collect_all_layers() contain script-aware .sh lookup for overrides, presets, extensions, and core content.
    • src/specify_cli/presets/__init__.py:6045-6204 — the Python resolver implements script composition and substitutes $CORE_SCRIPT for wrap layers, but this API is not connected to the shipped runtime script resolver.
    • scripts/bash/common.sh:501-586 — resolve_template() only accepts a template name and searches .md files in template locations; it has no script-type argument or .sh/overrides/scripts/ branch.
    • scripts/bash/common.sh:610 onward — resolve_template_content() is likewise a template-content resolver and is used by workflow scripts for .md templates, not executable script layers.
    • scripts/python/common.py:308-344 — the Python twin documents and implements only .md template lookup in the template directories; it has no script lookup path.

    Root Cause Hypothesis

    The preset composition model was implemented in PresetResolver for scripts, but the runtime workflow layer still exposes only the older markdown-template resolver. Core workflow commands invoke scripts by their installed paths, while no materialization or dispatch step asks PresetResolver for template_type="script"; consequently, accepted script declarations are inert. Confidence: high for the missing runtime wiring; medium for the exact intended materialization point because the report was not independently executed here.

    Proposed Remediation

    Preferred: Add a single script-layer materialization/resolution path that is invoked when the selected script variant is installed or refreshed. It should resolve the script name with template_type="script", compose replace and wrap layers using $CORE_SCRIPT, and write the resulting executable to the selected .specify/scripts/{bash,powershell,python}/ location while preserving the existing manifest/hash and upgrade behavior. Keep the resolver’s path validation and manifest-authoritative lookup rules, and make the Bash, PowerShell, and Python runtime paths behaviorally consistent.

    The runtime helper functions should either call a shared generated result or gain an explicit script resolver; simply adding another .md convention would not address the documented .sh and $CORE_SCRIPT contract.

    Alternatives:

    • If script presets are not ready for support, reject type: script during preset validation and revise the preset reference to mark the capability unsupported. This avoids silent success but removes the documented feature.
    • Extend the existing shell/Python template helpers with a script mode. This is smaller locally, but risks duplicating the already-implemented Python preset composition logic and creating Bash/Python/PowerShell drift.

    Files likely to change:

    • src/specify_cli/presets/__init__.py or the initialization/materialization caller that installs workflow scripts
    • scripts/bash/common.sh
    • scripts/python/common.py
    • the corresponding PowerShell common/runtime helper, if script resolution is supported there
    • targeted preset and integration tests, likely under tests/test_presets.py and tests/integrations/

    Tests to add or update:

    • Project-local overrides/scripts/<name>.sh replaces the core script.
    • A preset type: script with strategy: replace is materialized and executed.
    • A preset strategy: wrap receives a usable $CORE_SCRIPT target and delegates to the lower-priority core script.
    • Priority ordering, missing declared files, and manifest-authoritative behavior remain consistent with command/template resolution.
    • Equivalent behavior is covered for each supported script type, or explicitly documented if only shell scripts are supported.
    • Re-initialization and integration upgrade preserve the composed script and manifest bookkeeping.

    Risks & Considerations

    • Script composition changes executable behavior for projects that install presets; it should be opt-in through declared script layers and covered by regression tests.
    • Generated wrapper paths must remain inside the project and must not permit manifest path traversal or execution of undeclared files.
    • Upgrade/uninstall must not clobber user-modified generated scripts; existing manifest/hash semantics should be reused.
    • The current issue is feature-scoped and has a documented workaround, with no evidence of data loss or security impact; this supports low severity.

    Open Questions

    • [NEEDS CLARIFICATION: Should script presets be materialized during specify init, integration upgrade, or only resolved dynamically at workflow execution time?]
    • [NEEDS CLARIFICATION: Is script-layer support intended for Bash only, or for the PowerShell and Python variants as well?]
    • [NEEDS CLARIFICATION: Should extension-provided script entries follow the same runtime contract as preset-provided scripts?]

    Generated by 🐛 Assess Bug from Labeled Issue for #4551 · copilot · gpt52codex · 2.68 AIC · ⌖ 6.29 AIC · ⊞ 25.6K · ◷

  2. Ashfaqbs commented on Sep 15, 2026

    @Ashfaqbs

    I'd like to take this on. Verified all three points against current main before commenting:

    • specify_cli/presets/_commands.py's preset_resolve (the specify preset resolve debug command) is indeed the only call site in src/ that passes a literal template_type into PresetResolver.collect_all_layers()/.resolve(), and it only ever passes "command" or "template" — confirmed via a full-tree search, no call site anywhere passes "script", even though "script" is a fully valid member of VALID_PRESET_TEMPLATE_TYPES and every branch in presets/__init__.py (resolve(), collect_all_layers(), the $CORE_SCRIPT wrap/replace handling) already implements it.
    • scripts/bash/common.sh's resolve_template() (line 509 on current main) only ever builds overrides/${template_name}.md — no .sh branch, no overrides/scripts/ path.
    • scripts/python/common.py mirrors that with the same .md-only resolution.

    So this isn't a partial implementation with a rough edge — it's a fully-built, fully-tested composition engine with zero production callers, on both the CLI-inspection side and the runtime side.

    I'd lean toward wiring it up rather than rejecting type: script at validation time: the composition logic already exists and is exercised by unit tests, so ripping it out (or blocking it at validation) would be discarding working code to paper over a wiring gap rather than closing it. My proposed scope:

    1. Identify and fix the actual installation path that should request template_type="script" when a preset declares a script layer (this is a different call site than preset_resolve, which is read-only/debug-only — still tracing exactly where preset-provided command/template files get written to disk during preset add/apply, to wire the script case in alongside it).
    2. Add .sh resolution (and an overrides/scripts/ lookup, per the docs) to resolve_template()/resolve_template_content() in common.sh, mirrored in common.py.
    3. Regression tests: a preset with a type: script layer that actually gets picked up by resolve_template/resolve_template_content at runtime (currently nothing exercises this end-to-end, only the composition engine in isolation).

    Before I put time into a PR: does this direction match what you'd want to land, or is there a reason type: script support was scoped out at runtime after #3143 documented it as shipped? Happy to narrow scope further if there's a preferred first slice.

    Disclosure per CONTRIBUTING.md: I used Claude Code (agentic coding assistant) to search the codebase and verify the three claims above against current main before writing this comment; the investigation and this proposal reflect my own read of the code, not an unverified AI suggestion.

  3. mnriem commented on Sep 16, 2026

    @mnriem
    Collaborator

    Thanks for investigating this. To clarify the intended contract: $CORE_SCRIPT should provide a callable reference to the wrapped script, resolved through the stack at runtime—not inline that script’s source text.

    With multiple wrapping presets, that means the next lower-priority script layer, ultimately reaching the built-in script. It must not bypass intervening wrappers by always calling core directly. The wrapper then controls its before/after work and delegates with the appropriate arguments.

    Please scope the implementation around that execution path, including runtime-appropriate invocation for Bash, PowerShell, and Python, and preservation of existing ownership and upgrade behavior. The current text-substitution implementation should not determine the intended contract.

    Before opening the implementation PR, outline how the wrapped target will be resolved and invoked. Include end-to-end evidence that an installed wrapper actually runs, stacked wrappers execute in the intended order, and projects without script overrides retain their existing behavior. Isolated composition tests alone are insufficient for this change.

    Drafted for @mnriem with assistance from GitHub Copilot (model: GPT-6 Astra; interactive comment drafting).

  4. Ashfaqbs commented on Sep 19, 2026

    @Ashfaqbs

    Design outline, per your request — not opening the PR until this is confirmed.

    The mechanism

    Today resolve_content's wrap strategy textually splices content, which is exactly what you flagged as the wrong contract for scripts. But the repo already has the right pattern for a runtime-resolved contract — just applied to templates, not scripts:

    • resolve_template_content in scripts/bash/common.sh / scripts/python/common.py, and Resolve-TemplateContent in scripts/powershell/common.ps1, all walk the live priority stack (project override → .specify/presets/.registry-sorted presets → extensions → core) at runtime, not at install/compose time.
    • Preset layers already persist independently on disk after install — shutil.copytree(source_dir, dest_dir) copies each preset's own files into .specify/presets/<pack_id>/ wholesale, and PresetResolver.collect_all_layers already reads each layer straight from that per-preset path. Nothing about a preset's own script file gets deleted or flattened away today.

    So the substrate for "resolve through the stack, reach the next lower layer, don't bypass intervening wrappers" already exists — it's just never been exposed as a path lookup instead of a content lookup.

    Proposal: add a sibling function alongside resolve_template_content — resolve_next_script_layer(template_name, repo_root, below_priority) (bash/python) / Resolve-NextScriptLayer (PowerShell) — that runs the same stack walk but:

    1. starts strictly below the calling layer's own priority (so a 3-preset stack can't skip a middle wrapper — this is the specific case you called out), and
    2. returns a path, not composed content, terminating at the core template's own script (.specify/templates/scripts/<name>{.sh,.py,.ps1}) as the guaranteed final layer, same terminus collect_all_layers's Priority 4 already uses for templates.

    At composition time, when strategy is wrap and template_type == "script", $CORE_SCRIPT stops being replaced with inlined text. Instead it's replaced with a call to that resolver, so the installed wrapper script itself does the runtime resolution and invocation — the composed/installed file is the top-of-stack wrapper's own source, minimally rewritten, not a merged blob.

    Runtime-appropriate invocation

    • Bash: $CORE_SCRIPT → "$(resolve_next_script_layer ...)" "$@", invoked via exec so exit code and stdio pass through unchanged (matches how these scripts are already invoked as standalone processes, not sourced).
    • PowerShell: & (Resolve-NextScriptLayer ...) @args, preserving $LASTEXITCODE.
    • Python: since these are standalone .py files run via python script.py (not imported), same-name stacked layers would collide on sys.modules if imported directly — so this needs subprocess.run([sys.executable, next_path, *sys.argv[1:]]) with the child's exit code propagated, rather than an in-process import. Flagging this now since it's the one place the three runtimes can't share an identical strategy.

    Ownership / upgrade behavior

    This only changes what a wrap-strategy script's $CORE_SCRIPT expands to — which single file is "the installed, owned" file for a given script name doesn't change (still whatever collect_all_layers's Priority 1 override or highest-priority preset resolves to). Ownership/ provenance tracking is untouched by this.

    Evidence I'll gather before opening the PR

    Per your point that isolated composition tests aren't sufficient:

    • A real 3-layer stack (core → preset A wrap → preset B wrap) installed into a scratch project via the actual specify CLI, then the installed script run end-to-end with each layer writing an ordered marker (not just asserting text presence) — verifying B's before/after runs, then A's (not skipped), then core, in that order.
    • A project with zero script overrides/presets installed and run unmodified, confirming today's non-wrapping behavior is bit-for-bit unaffected.
    • The above for both Bash and Python runtimes in this sandbox; I don't have a Windows/PowerShell environment available here, so I'll say so plainly rather than claim PowerShell coverage I can't produce, and ask for a maintainer or CI check on that leg specifically.

    Let me know if this direction (runtime path-resolution reusing/extending the existing resolve_template_content pattern, exec/subprocess invocation per runtime) matches what you had in mind before I start building it.

    Comment drafted with the assistance of Claude Code (Sonnet 5) — design research (reading collect_all_layers, resolve_content, resolve_template_content/Resolve-TemplateContent across all three runtime script directories) and comment drafting; the proposed design and scoping decisions are mine.

  5. mnriem commented on Sep 21, 2026

    @mnriem
    Collaborator

    Thanks for outlining this. The runtime path-resolution direction is closer to the intended contract, but there is one important correction before implementation: script layers should be resolved dynamically when the script is invoked, not selected, composed, or materialized during scaffolding. The references to “composition time” and an “installed top-of-stack wrapper” suggest that the winning wrapper would still be determined ahead of execution.

    The script resolver should follow the same live stack and validation rules as template resolution:

    1. Project override
    2. Enabled presets ordered by (priority, preset id)
    3. Enabled extensions
    4. Built-in core script

    Instead of returning composed content, it should provide the invoked wrapper with a callable continuation. That continuation resolves and executes the exact next layer in the live stack, supplies itself as that layer’s continuation, and eventually reaches the built-in script. A replace layer is terminal; a wrap layer receives the continuation. $CORE_SCRIPT therefore means “continue with the next resolved layer,” not “jump directly to core.”

    The continuation must identify the exact current layer or stack position rather than accepting only below_priority, because multiple presets can have the same numeric priority and are then ordered by ID. It should also be a dispatcher handle rather than merely the next wrapper’s path; otherwise that lower wrapper has no continuation of its own.

    The language-specific invocation can adapt the same continuation contract:

    • Bash invokes the continuation with "$@".
    • PowerShell invokes it with @args.
    • Python invokes it as a subprocess with sys.argv[1:].
    • Each wrapper captures the continuation’s exit status, performs any after-work, and exits with that status.

    The wrapper itself must not exec the continuation because that would prevent its after-work from running. The continuation dispatcher, running as the wrapper’s child process, may exec the resolved lower layer because control still returns to the parent wrapper when that child terminates.

    With this model, changing preset enablement or priority takes effect on the next invocation without re-scaffolding, preset files remain independent, and the existing core-script ownership and upgrade behavior can remain intact.

    Please revise the outline around that dynamic continuation model. The end-to-end evidence should include:

    • B-before → A-before → core → A-after → B-after
    • Two wrappers with the same numeric priority
    • Argument, working-directory, stdin/stdout/stderr, and exit-status propagation
    • Preset enablement and priority changes taking effect without re-scaffolding
    • Unmodified behavior for projects without script layers
    • Bash, PowerShell, and Python coverage, with PowerShell exercised through CI if unavailable locally

    Posted on behalf of @mnriem by GitHub Copilot (model: GPT-5.6 Sol, interactive); comment fully AI-drafted from @mnriem’s design direction.

  6. Ashfaqbs commented on Sep 22, 2026

    @Ashfaqbs

    Understood — dynamic resolution at invocation, not scaffold-time selection. Revising around the continuation model:

    Continuation identity, not index or priority

    The one place I'd refine your outline: a wrapper's own $CORE_SCRIPT substitution should not bake in a numeric stack position (index or below_priority) at all, static or otherwise — because the live stack can reorder between scaffolding and invocation (preset disabled, priority changed), and a wrapper installed when it was position 2 of 4 has no way to know it's now position 1 of 3 unless something re-derives its position fresh every call.

    So $CORE_SCRIPT expands to a call carrying the wrapper's own layer identity — (layer_kind, layer_id), e.g. (preset, "my-preset"), (extension, "my-ext"), or (project-override,) — not a position. The dispatcher, invoked with that identity:

    1. Re-resolves the full live stack fresh (project override → enabled presets ordered by (priority, id) → enabled extensions → core) — same resolver collect_all_layers/resolve_template_content already use for templates.
    2. Finds where (layer_kind, layer_id) currently sits in that freshly-resolved list.
    3. Continues from the entry immediately below it — whatever that is now, regardless of where it was when the wrapper was installed.

    This is what makes "priority changes take effect without re-scaffolding" actually true rather than true-until-the-stack-shifts, and it's also the natural answer to the same-priority-tie case you flagged: ties are already broken by id in the resolver, so identity-based lookup gets tie-breaking for free instead of needing special-casing.

    Process shape, confirming my read of your model

    • Wrapper (top of stack) reaches $CORE_SCRIPT → spawns the continuation dispatcher as a child (not exec — the wrapper still has after-work to run), passing its own (layer_kind, layer_id).
    • Dispatcher re-resolves the stack, locates the next layer down, and execs into it directly (dispatcher has no after-work of its own, so exec is safe here — control returns to the wrapper, not the dispatcher, once that exec'd process exits).
    • If the next layer is itself wrap, its own $CORE_SCRIPT repeats the same spawn-child-dispatcher step with its identity — so the chain is self-similar all the way to core, and I don't need the top-level caller to precompute or pass down remaining depth.
    • replace and core are both terminal — no continuation, no $CORE_SCRIPT substitution to make.

    Per-runtime, only the last hop differs as you described: Bash exec "$(resolve)" "$@" inside the dispatcher, PowerShell & (Resolve-...) @args with $LASTEXITCODE propagated, Python subprocess.run([sys.executable, next_path, *sys.argv[1:]]) from the wrapper (not exec, since CPython has no process-image exec equivalent that preserves the interpreter cleanly) with the child's exit code returned as the wrapper's own.

    One open question before I start

    Where does (layer_kind, layer_id) get embedded in the installed wrapper file — as a literal token substituted into the script at preset add/apply time (simple, but means the file's bytes encode identity, which upgrade/hash bookkeeping needs to tolerate as an expected diff), or does the dispatcher accept the wrapper's filename/path and derive identity from where the installed file physically lives under .specify/presets/<id>/... (avoids touching content, but only works if a wrapper's installed path deterministically implies its identity — true today per collect_all_layers, but worth confirming that's a contract you want load-bearing here too)? I'd lean toward the second — less to keep in sync on upgrade — but want your read before committing to it.

    Evidence plan otherwise matches what you listed: B-before → A-before → core → A-after → B-after, same-priority tie ordering, arg/cwd/stdio/exit-status propagation, live enablement/priority changes without re-scaffolding, unmodified no-override behavior, Bash+Python run end-to-end locally with PowerShell flagged for CI verification since I don't have a Windows environment here.

    Disclosure per CONTRIBUTING.md: I used Claude Code to re-read collect_all_layers, resolve_template_content, and the preset install path while drafting this reply, and to help write it up; the design decisions (identity-over-index, dispatcher-as-child-then-exec, where to embed identity) are my own analysis of your stated contract, not an unverified AI suggestion.

  7. mnriem commented on Sep 22, 2026

    @mnriem
    Collaborator

    You are close, but the wrapper should not identify itself or reconstruct its position in the stack. The runtime resolver should resolve the complete effective script chain when the top-level script is invoked, using the same ordering, validation, and strategy rules as the preset resolver. It then invokes the first layer with an opaque continuation bound to the remaining resolved layers.

    When a wrapper calls $CORE_SCRIPT, it invokes that supplied continuation. The continuation advances to the next layer and supplies that layer’s continuation. Therefore there is no identity to embed in the wrapper file and no need to derive identity from its installed path.

    Resolution is refreshed for each new top-level invocation, so priority and enablement changes take effect without rescaffolding. It should not normally be refreshed between nested continuation calls within the same invocation; that invocation should execute against one consistent resolved stack.

    The new work is the runtime-specific execution adapter for that continuation. The stack resolution itself should reuse or extract the existing preset resolver rules rather than establish a second identity-based resolution protocol.

  8. Ashfaqbs commented on Sep 23, 2026

    @Ashfaqbs

    Agreed — dropping identity entirely. Re-deriving position was solving a problem the continuation model doesn't have: if the continuation itself carries "what's left," a wrapper never needs to know or reconstruct where it sits.

    One thing your correction lets me ground concretely: I checked how a script actually gets invoked in this codebase, since that determines what "when the top-level script is invoked" means mechanically. A command template (e.g. templates/commands/plan.md) declares scripts.sh: scripts/bash/setup-plan.sh --json in its frontmatter — the agent shells out to that fixed, materialized path directly. There's no CLI-side dispatch call site in between; whatever file preset add/apply installed at .specify/scripts/bash/setup-plan.sh is the thing that gets run. Only that one file — whichever layer currently wins priority 1 — physically lives at the canonical path; every lower layer stays wherever it was installed (.specify/presets/<id>/scripts/...) and is never copied to the canonical path itself.

    That maps directly onto your model with no extra machinery:

    • The resolved chain and the continuation are the same object. The top-level resolver builds the ordered list once — project override → enabled presets by (priority, id) → enabled extensions → core, reusing collect_all_layers's/resolve_template_content's existing ordering and validation, not a second implementation of it. It hands the first layer's $CORE_SCRIPT an opaque handle over the rest of that same list. There's no separate "identity" concept to design out — I'd already have needed the full list either way, and the continuation is just "the part of the list I haven't consumed yet," so removing identity actually removes a whole parallel piece of state, not just a lookup key.
    • Only the canonical-path script does a fresh resolution. It's the one thing actually invoked by shelling out, so it's the only place a "top-level invocation" boundary exists to hook into. Every nested $CORE_SCRIPT call is a lower layer consuming the continuation it was handed, not re-entering the resolver — which is what makes "not normally refreshed between nested continuation calls" true by construction rather than a rule I'd have to enforce separately.
    • The continuation needs to be an actual runtime value, not a token I invent conventions for. Concretely: the resolver serializes the remaining ordered layers (path + strategy, nothing else) once per top-level invocation, and a small shared continuation-runner (one per runtime, called by every wrap layer's $CORE_SCRIPT) consumes the head, execs/subprocess's it, and — only if there's more remaining and that head was itself a wrap layer — passes the tail along to the next hop. This is the one runtime-specific execution adapter piece; the ordering/validation/strategy rules it consumes are entirely reused, not reimplemented.

    Per-runtime adapter, otherwise unchanged from what I described last round:

    • Bash: continuation-runner execs the resolved next layer with "$@"; the calling wrapper never execs the continuation itself, since it still has after-work.
    • PowerShell: & (next-layer) @args, $LASTEXITCODE propagated back to the wrapper.
    • Python: subprocess.run([sys.executable, next_path, *sys.argv[1:]]) from the wrapper, exit code returned — no in-process import, since these run as standalone processes and same-named stacked layers would otherwise collide on sys.modules.

    Evidence plan is unchanged from your last message and I think it fully covers this model:

    • B-before → A-before → core → A-after → B-after
    • two wrappers with the same numeric priority (tie-broken by id, for free, since it's the same ordering the resolver already does for templates)
    • arg / cwd / stdin / stdout / stderr / exit-status propagation through the chain
    • preset enablement and priority changes taking effect on the next invocation without re-scaffolding
    • unmodified behavior for a project with zero script layers
    • Bash, PowerShell, and Python, PowerShell exercised via CI if not available locally

    Does this match the intended design closely enough to start on, or is there another correction before I put time into the PR?

    Disclosure per CONTRIBUTING.md: I used Claude Code (agentic coding assistant) to search the codebase (confirming how templates/commands/plan.md invokes scripts) and to help draft this response; the design judgment and the final proposal above are my own.

  9. mnriem commented on Sep 23, 2026

    @mnriem
    Collaborator

    Lets start and we will address any issues along the way

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions