Skip to content

Latest commit

 

History

History
437 lines (350 loc) · 19.4 KB

File metadata and controls

437 lines (350 loc) · 19.4 KB

ARC + DinD Configuration

AWF supports ARC runners where the runner filesystem and Docker daemon filesystem are split (DinD sidecar patterns).

Runner topology selector

The simplest way to configure AWF for ARC/DinD is through the runner.topology config key:

{
  "runner": {
    "topology": "arc-dind"
  }
}

When runner.topology is set to "arc-dind", AWF enables ARC/DinD-specific sysroot staging behavior:

Behavior Default Override
Sysroot image for /host base build-tools:<tag> runner.sysrootImage
Tool cache warning if under /opt Emitted Set RUNNER_TOOL_CACHE to shared path

Other ARC/DinD settings (for example network.isolation and dind.preStageDirs) are configured explicitly through their own fields.

Build-tools sysroot image

On ARC/DinD, the standard system mounts (/usr:/host/usr:ro, etc.) resolve to the runner container's filesystem, which is invisible to the Docker daemon's split filesystem. The build-tools sysroot image solves this by providing a pre-built Ubuntu 22.04 image containing system-level build infrastructure:

  • Compilers & linkers: gcc, g++, make, cmake, autoconf, binutils
  • Dev libraries: libssl-dev, libc6-dev, libicu-dev, zlib1g-dev
  • System utilities: bash, coreutils, git, curl, wget, jq
  • Agent dependencies: libcap2-bin (capsh), gosu, gnupg, gh

How it works

  1. AWF emits a sysroot-stage init service in the compose file
  2. The init container copies the build-tools image FS into a named sysroot volume
  3. The agent mounts the sysroot volume read-only at /host
  4. entrypoint.sh finds /host/bin/sh and capsh, chroots successfully
# Generated docker-compose.yml (simplified)
services:
  sysroot-stage:
    image: ghcr.io/github/gh-aw-firewall/build-tools:0.28.0
    volumes: ["sysroot:/sysroot"]
    entrypoint: ["/bin/sh", "-c"]
    command: ["cp -a /usr /lib /bin /sbin /etc /sysroot/ ..."]

  agent:
    depends_on:
      sysroot-stage: { condition: service_completed_successfully }
    volumes:
      - sysroot:/host:rw
      - /tmp/gh-aw/tool-cache:/host/tmp/gh-aw/tool-cache:ro

volumes:
  sysroot: {}

Custom sysroot image

Override the default build-tools image:

{
  "runner": {
    "topology": "arc-dind",
    "sysrootImage": "ghcr.io/my-org/custom-sysroot:latest"
  }
}

Tool cache for language SDKs

Language SDKs (Go, Node, Java, .NET) are NOT baked into the sysroot image. They are installed on-demand by setup-* actions into a shared tool cache volume.

Important: On ARC, RUNNER_TOOL_CACHE must point to a shared path visible to both the runner container and the DinD daemon (e.g., /tmp/gh-aw/tool-cache). The default /opt/hostedtoolcache is invisible to the DinD daemon.

# Early in workflow, before setup-* actions:
- run: echo "RUNNER_TOOL_CACHE=/tmp/gh-aw/tool-cache" >> "$GITHUB_ENV"

Staging additional CLI tools (not just the invoking engine binary)

AWF never discovers or stages arbitrary tools that a sandboxed command invokes (for example a threat-detection binary, a linter, or any helper CLI installed by a separate workflow step). The only automatic staging is narrow:

  • First-command staging copies the first token of the AWF command (the invoking engine binary — copilot, claude, codex, …) into a daemon-visible staging root and mounts it at /tmp/awf-runner-bin/<name>. It runs only when --docker-host-path-prefix is a /tmp-rooted path shared by the runner and the DinD daemon (for example /tmp/gh-aw). A /host-style prefix does not trigger it, so with that prefix even the engine binary must be staged explicitly.
  • dind.stageEngineBinary is an explicitly configured fallback (you give it path/targetPath), not automatic discovery.

On runner.topology: arc-dind, a tool installed only onto the runner's filesystem (via actions/setup-*, a manual curl install, or an action that drops a binary under $RUNNER_TOOL_CACHE//opt//usr/local/bin) is invisible to the sysroot the agent chroots into: that sysroot comes from the build-tools image plus daemon-visible bind mounts, not from the runner pod's filesystem. Calling such a tool inside AWF without staging it fails with exit code 127 ("command not found").

Staging recipe

Copy the tool into a shared, daemon-visible /tmp-rooted directory and mount that directory, because staging alone does not expose a path — --docker-host-path-prefix only rewrites bind-mount sources, it does not publish arbitrary runner paths into the sandbox.

- name: Stage my-tool and its output directory
  run: |
    mkdir -p /tmp/gh-aw/bin /tmp/gh-aw/my-tool
    cp "$(command -v my-tool)" /tmp/gh-aw/bin/my-tool
    chmod +x /tmp/gh-aw/bin/my-tool

- name: Run my-tool under AWF
  run: |
    sudo awf --docker-host-path-prefix /tmp/gh-aw \
      --mount /tmp/gh-aw/bin:/tmp/gh-aw/bin:ro \
      --mount /tmp/gh-aw/my-tool:/tmp/gh-aw/my-tool:rw \
      --allow-domains api.example.com \
      -- /tmp/gh-aw/bin/my-tool --output /tmp/gh-aw/my-tool/result.json

Notes on this recipe:

  • Every --mount host path must already exist — AWF validates it before launch and aborts with Host path does not exist otherwise. Create both the staged bin directory and the output directory in the preceding step.
  • Copying the file returned by command -v is only sufficient for a self-contained executable (a static binary, or one whose interpreter, shared libraries, Node/Python modules and data files already exist in the sysroot). Wrapper scripts, setup-*-installed SDK tools, and package-managed linters are typically shims over a runtime and support tree, so copying the entry point alone still fails — with a missing-interpreter, missing-.so, or missing-module error rather than exit 127. For those, stage/mount the whole install tree (and its runtime, e.g. the tool cache at /tmp/gh-aw/tool-cache) instead of a single file.
  • chroot.binariesSourcePath is the alternative when you want staged tools on PATH: AWF mounts that runner-side directory at /tmp/awf-runner-bin inside the chroot and prepends it to PATH, so the command can be invoked by bare name.

Output paths must be writable — and consistent

Staging the binary only fixes exit code 127. A tool that writes output also needs its --output/destination path to resolve to a mount that is:

  1. writable — a :ro mount used to expose staged inputs does not become writable just because the tool's output path lives underneath it. Add a more specific :rw mount for the output subdirectory, as in the recipe above:

    --mount /tmp/gh-aw/bin:/tmp/gh-aw/bin:ro \
    --mount /tmp/gh-aw/my-tool:/tmp/gh-aw/my-tool:rw

    Docker orders bind mounts by destination depth, so the more specific rw mount of the subdirectory is applied after (and therefore takes precedence over) a broader ro mount of its parent.

  2. the same path for the producer and the consumer — pick one canonical sandbox path (e.g. always /tmp/gh-aw/my-tool/result.json) and use it for both the AWF-side write and the later step or artifact-upload read. AWF mounts custom targets under the chroot, so the in-sandbox path matches the --mount container path; the runner-side source is what --docker-host-path-prefix rewrites. A mismatch (writing under ${RUNNER_TEMP}/… but reading /tmp/gh-aw/…, which --docker-host-path-prefix does not alias) produces no error — the tool "succeeds" but the output never appears where the consumer looks.

For gh-aw engine streaming logs (including Pi), do not write directly to ${RUNNER_TEMP}/gh-aw/pi-streaming.jsonl when ${RUNNER_TEMP}/gh-aw is mounted :ro. Mount an existing writable child such as ${RUNNER_TEMP}/gh-aw/sandbox/agent with :rw, and write the log at ${RUNNER_TEMP}/gh-aw/sandbox/agent/pi-streaming.jsonl. The engine command, log parser, and artifact upload must all use that same path. Mounting the child :rw does not make the rest of the read-only parent writable; the workflow/compiler that chooses the log location must also update its consumers.

Locating API proxy token-usage logs

The api-proxy writes token-usage.jsonl into <proxy-logs-dir>/api-proxy-logs/ (or <workDir>/api-proxy-logs/, moved to /tmp/api-proxy-logs-<ts>/ after cleanup when --proxy-logs-dir is not set). Under runner.topology: arc-dind the gh-aw compiler passes ${RUNNER_TEMP}/gh-aw/... paths, so the file is not at the literal /tmp/gh-aw/sandbox/firewall/logs/... path a post-run step may hardcode.

After cleanup AWF logs the final runner-visible location (Token usage log available at: <path>) and, when $GITHUB_ENV is available, exports it for later steps of the same job:

- name: Parse token usage
  run: |
    if [ -n "${AWF_TOKEN_USAGE_LOG:-}" ]; then
      jq -s 'length' "$AWF_TOKEN_USAGE_LOG"
    fi

AWF_TOKEN_USAGE_LOG is only exported when the file exists. Post-run consumers should prefer it over hardcoded paths.

--docker-host-path-prefix translation is accounted for: with a daemon-only prefix (/host, /tmp/gh-aw, ...) the daemon writes back to the runner path unchanged. With a shared /tmp prefix, a --proxy-logs-dir outside /tmp (for example under ${RUNNER_TEMP}) is rewritten to /tmp<dir>; AWF warns at startup, pre-creates that directory so the non-root api-proxy can write to it, and reports/exports the /tmp<dir>/api-proxy-logs/token-usage.jsonl path. Keep --proxy-logs-dir under the shared prefix to avoid the rewrite.

Writable home under sysroot staging

Sysroot staging drops agent bind mounts whose sources the DinD daemon cannot resolve, including AWF's own ${workDir}-chroot-home volume for /host$HOME. An explicitly supplied mount is exempt from that filter: if the caller passes --mount <daemon-visible-home>:$HOME:rw (the gh-aw compiler does this for ${RUNNER_TEMP}/gh-aw/home), the resulting /host$HOME mount is kept, because the caller vouches for the source being visible to the daemon. The exemption matches on both source and target, so AWF's own mounts to the same target stay subject to the filter.

A writable /host$HOME matters for two reasons:

  • the /dev/null credential-hiding overlays are mounted under /host$HOME, and runc cannot create those mountpoints under a read-only parent;
  • entrypoint.sh pre-seeds JVM build tool proxy config (~/.m2, ~/.gradle) under the chroot home.

If no writable /host$HOME survives the filter, AWF logs a warning and skips the /host$HOME credential overlays instead of failing container creation — the overlays at the un-prefixed $HOME path (on the agent's own rootfs) are still applied, but credential files under the chroot home are not masked for that run. The entrypoint likewise warns and skips JVM proxy pre-seeding rather than aborting.

What AWF handles automatically

  • Split-filesystem probing for --docker-host-path-prefix
  • Chroot staging for:
    • invoking CLI binary (copilot, claude, codex, etc.)
    • /etc/passwd
    • /etc/group
    • generated chroot /etc/hosts
  • DinD DOCKER_HOST propagation into agent/MCP environments when DinD is detected

Explicit ARC/DinD config surface

For fine-grained control (or when not using runner.topology):

{
  "container": {
    "enableDind": true,
    "dockerHostPathPrefix": "/tmp/gh-aw"
  },
  "chroot": {
    "binariesSourcePath": "/tmp/gh-aw/runner-bin",
    "identity": {
      "home": "/tmp/gh-aw/home",
      "user": "runner",
      "uid": 1001,
      "gid": 1001
    }
  },
  "dind": {
    "preStageDirs": true,
    "workDir": "/tmp/gh-aw",
    "stagingImage": "ghcr.io/github/gh-aw-firewall/agent:latest",
    "stageEngineBinary": {
      "path": "/usr/local/bin/copilot",
      "targetPath": "/usr/local/bin/copilot"
    }
  },
  "runner": {
    "topology": "arc-dind",
    "sysrootImage": "ghcr.io/github/gh-aw-firewall/build-tools:latest"
  }
}

Field behavior

  • chroot.identity.*: applied inside entrypoint after chroot /host to override HOME/USER/LOGNAME and identity mapping hints.
  • chroot.binariesSourcePath: mounts a runner-side binaries directory at /host/tmp/awf-runner-bin (inside chroot: /tmp/awf-runner-bin) and prepends it to PATH, so runner-installed CLIs are visible even when /usr comes from the DinD daemon filesystem.
  • dind.preStageDirs: runs a short-lived staging container in DinD mode to create required workdir tree with open permissions.
  • dind.stageEngineBinary: copies an engine binary from the runner path into daemon-visible filesystem before compose startup.
  • dind.stagingImage: image used for short-lived staging containers.
  • dind.workDir: target root for DinD pre-staged directory tree (/tmp/gh-aw default).
  • runner.topology: "arc-dind": enables sysroot staging (sysroot-stage init service + sysroot volume mounted on agent at /host:rw).
  • runner.sysrootImage: optional override for the sysroot image used by runner.topology=arc-dind.

Sysroot staging lifecycle

When runner.topology is arc-dind, AWF starts a one-shot sysroot-stage service that copies the filesystem from a build-tools image derived from the same --image-registry and --image-tag settings as the other AWF containers (unless runner.sysrootImage overrides it) into a named sysroot volume. The agent mounts that volume at /host:rw.

This image pre-installs root-required system build dependencies (for example gcc/make/cmake, libssl-dev/libc6-dev/libicu-dev, capsh/gosu/gh) so ARC workflow steps can stay non-root.

Tool cache path guidance for ARC

If RUNNER_TOOL_CACHE points under /opt (for example /opt/hostedtoolcache) AWF logs a warning in runner.topology=arc-dind mode because /opt is commonly not visible from the DinD daemon filesystem. Prefer a shared runner/daemon path under /tmp/gh-aw when possible.

Auto-detection of split filesystem setups

AWF detects likely ARC/DinD environments at startup and warns when --docker-host-path-prefix is missing:

  • non-default unix DOCKER_HOST socket paths (outside /var/run/docker.sock and /run/docker.sock)
  • loopback TCP DOCKER_HOST endpoints (tcp://localhost:* or tcp://127.0.0.1:*) — the standard ARC RunnerScaleSet DinD sidecar configuration
  • AWF_DIND=1

Recommended DinD base image

For ARC DinD chroot workloads, prefer the glibc companion image:

  • ghcr.io/github/gh-aw-firewall/dind-ubuntu:latest

It includes docker-ce, libcap2-bin (capsh), and Node.js preinstalled.

Runtime prerequisite

Copilot CLI still requires node to be available inside the chrooted runtime PATH.

Joining services: containers to awf-net for direct protocol access

On runner.topology: arc-dind, GitHub Actions services: containers are started by the runner on the runner's own bridge network (github_network_<hash>), while the AWF agent runs on AWF's isolated awf-net. These two bridges are not routed to each other, so workloads inside the sandbox that need to speak a service's native wire protocol (database drivers, migration tools, client libraries under test, etc.) cannot reach services: containers — only HTTP(S) egress through Squid is available from inside the agent.

The existing host-iptables-based service-port routing (hostServicePorts) does not help here: on ARC/DinD, network.isolation: true uses the Docker-network topology and never programs host iptables NAT rules, since network isolation is enforced entirely at the Docker network level, not at the host firewall level.

Verified pattern: join the service to awf-net

The supported workaround is to join the service container — never the agent — onto awf-net after AWF creates it, using a pre-step "waiter" that polls for the network's existence and then attaches the service container to it with an alias the agent can resolve by name:

services:
  postgres:
    image: postgres:16
    env:
      POSTGRES_PASSWORD: postgres

steps:
  - name: Attach postgres service to awf-net
    run: |
      # Remove a stale AWF network so the waiter cannot attach to a network
      # that AWF will reclaim. Do this only when no other AWF run is active.
      if docker network inspect awf-net >/dev/null 2>&1; then
        docker network rm awf-net >/dev/null 2>&1 || {
          echo "cannot replace stale awf-net; another AWF run may be active" >&2
          exit 1
        }
      fi

      # Keep the waiter running while the later generated AWF step creates the
      # network. job.services exposes the exact container ID; no image lookup
      # or label is required.
      service_container='${{ job.services.postgres.id }}'
      nohup bash -c '
        for i in $(seq 1 120); do
          if docker network inspect awf-net >/dev/null 2>&1; then
            docker network connect --alias postgres awf-net "$1" && exit 0
          fi
          sleep 1
        done
        echo "awf-net was not created" >&2
        exit 1
      ' -- "$service_container" >/tmp/awf-net-waiter.log 2>&1 &

The nohup background process must remain alive until the later generated AWF step starts. Check /tmp/awf-net-waiter.log after that step if the service is not reachable. Once joined, code running inside the AWF sandbox can connect directly to postgres:5432 (or whatever alias/port the service exposes) using the service's native protocol, bypassing the Squid HTTP(S)-only egress path entirely for that one container-to-container link.

This pattern has been end-to-end validated with 669/669 dotnet/Npgsql integration tests passing against a Postgres services: container joined to awf-net this way.

Security note

Only the service container may join awf-net — never the agent. The whole point of AWF is to restrict what the agent (the untrusted/AI-driven process) can reach on the network. Joining the agent to the runner's own bridge network (or otherwise bypassing Squid) would defeat the egress firewall entirely, since traffic on that bridge is not subject to domain allowlisting. Joining a service container is safer than attaching the agent to the runner bridge, but it brings that service into the security boundary. A service that supports proxying, outbound callbacks, or command execution can become an application-layer pivot; for example, do not give the example database a privileged role unless it is required. Use least-privileged credentials and constrain the service's outbound access where possible. This attachment does not itself give the agent a direct route to the runner bridge, but it is not a guarantee that the service cannot provide an egress path.

Future direction

A longer-term improvement under consideration is compiler sugar such as services.<name>.attach: true, which would emit the waiter/join steps shown above automatically instead of requiring hand-written shell in workflow frontmatter. No such field exists yet — the manual pattern above is the only currently supported mechanism.

See also