diff --git a/.github/extension/README.md b/.github/extension/README.md index 62006d903a5..a51b4c7470e 100644 --- a/.github/extension/README.md +++ b/.github/extension/README.md @@ -63,10 +63,21 @@ To keep the two provider paths from duplicating the ~80% of steps they share, it - **`run-rad-commands.yml`** — the unified **dispatcher** and the only file that is dispatched. It owns the dispatch contract (`workflow_dispatch` inputs and the `Radius - Verify Credentials` auto-trigger). A `detect` job binds the GitHub Environment, reads which provider variable is set (`AZURE_CLIENT_ID` / `AWS_ROLE_ARN`), and calls the matching provider workflow via `workflow_call` with `secrets: inherit`. - **`run-rad-commands-azure.yml`** — a reusable (`workflow_call`) workflow with only the Azure-specific steps: Azure OIDC login, AKS connection (`az aks get-credentials`), workload-identity credential registration, and the `azure-avm` recipe pack (Azure Verified Modules) downloaded from [resource-types-contrib](https://github.com/radius-project/resource-types-contrib). - **`run-rad-commands-aws.yml`** — a reusable (`workflow_call`) workflow with only the AWS-specific steps: AWS OIDC login, EKS connection (access entry + static token kubeconfig), IRSA credential registration, and the `aws-terraform` recipe pack. -- **`actions/*`** — composite actions holding the provider-agnostic phases both provider workflows share: [`setup-control-plane`](actions/setup-control-plane/action.yml), [`restore-state`](actions/restore-state/action.yml), [`run-rad-commands`](actions/run-rad-commands/action.yml), and [`teardown`](actions/teardown/action.yml). The provider workflows reference them from `radius-project/radius` at a pinned ref (the `{{RADIUS_REF}}` placeholder the generator fills in), so the shared logic has a single reviewed home and is not copied into user repos. Third-party actions in these workflows are pinned to full commit SHAs (with a `# vX` comment); only the first-party Radius composite actions are referenced by ref. +- **`actions/*`** — composite actions holding the provider-agnostic phases both provider workflows share: [`setup-control-plane`](actions/setup-control-plane/action.yml), [`restore-state`](actions/restore-state/action.yml), [`apply-custom-recipe-packs`](actions/apply-custom-recipe-packs/action.yml), [`run-rad-commands`](actions/run-rad-commands/action.yml), [`delete-resource`](actions/delete-resource/action.yml), and [`teardown`](actions/teardown/action.yml). The provider workflows reference them from `radius-project/radius` at a pinned ref (the `{{RADIUS_REF}}` placeholder the generator fills in), so the shared logic has a single reviewed home and is not copied into user repos. Third-party actions in these workflows are pinned to full commit SHAs (with a `# vX` comment); only the first-party Radius composite actions are referenced by ref. The deploy flow generates the dispatcher and both provider workflows, commits them to the target repo under `.github/workflows/`, and dispatches `run-rad-commands.yml`. +## `delete-application.yml` / `delete-environment.yml` (delete dispatchers and provider workflows) + +Radius deletes a deployed application or an environment with the same ephemeral-control-plane model as the deploy flow. Because deleting recipe-backed resources runs the recipes' delete path (e.g. `terraform destroy`) against the target cluster and cloud, the delete workflows restore the persisted Radius state first, run the delete, then persist the updated state again — so subsequent runs plan against the post-delete state. + +- **`delete-application.yml`** — dispatcher to delete one application. `workflow_dispatch` inputs: `environment` (GitHub Environment name) and `application` (application name). A `detect` job binds the environment, reads the provider variable, and calls the matching provider delete workflow with `resource_type: application`. +- **`delete-environment.yml`** — dispatcher to delete one environment. `workflow_dispatch` inputs: `environment` (GitHub Environment name) and optional `environment_name` (the Radius environment name, defaulting to the GitHub Environment name, since the deploy flow names the Radius environment after it). Calls the provider delete workflow with `resource_type: environment`. +- **`delete-azure.yml`** / **`delete-aws.yml`** — reusable (`workflow_call`) workflows with the provider-specific steps (OIDC login, cluster connection, cloud OIDC token projection, and credential registration) shared with the deploy provider workflows. They reuse the `setup-control-plane`, `restore-state`, [`delete-resource`](actions/delete-resource/action.yml), and `teardown` composite actions. Like the deploy provider workflows they log in to GHCR and set the `RADIUS_STATE_*` variables so `rad startup`/`rad shutdown` can open the OCI-backed state archive. Unlike the deploy provider workflows they do **not** create the environment, recipe pack, or the in-pod image-push registry credentials — the environment and its recipes are restored from state, and deleting builds no images. + +The `delete-resource` composite action runs `rad app delete --yes --preview` or `rad env delete --yes --preview` (`--preview` selects the Radius.Core surface the deploy flow provisions) and writes a `rad-delete-result` artifact — a JSON document with `outcome`, `exitCode`, `resourceType`, `name`, and the command `output`. + + ### What it does The dispatcher routes to the matching provider workflow, which runs on `ubuntu-latest`. It stands up an ephemeral [k3d](https://k3d.io) cluster to host the Radius control plane on the runner, points that control plane at the user's existing EKS/AKS cluster, and deploys the application there. The control-plane setup, state restore, and run/teardown phases below run from the shared composite actions; the OIDC login, cluster connection, token projection, credential registration, and recipe-pack creation are the provider-specific steps. When a provider's identifying variable is empty, its steps are skipped and resources deploy to the ephemeral control-plane cluster instead of an external target. @@ -82,9 +93,10 @@ The dispatcher routes to the matching provider workflow, which runs on `ubuntu-l 9. **Restore persisted state (`rad startup`).** Restores the control-plane databases and the Terraform recipe-state Secrets saved by the previous run, so `rad deploy` plans against prior state rather than an empty backend. A no-op on the first run. 10. **Register cloud credentials.** Registers the cloud identity with `rad credential register azure wi` / `aws irsa` so Radius holds the identity selector and reads the projected token at runtime. 11. **Create the Radius environment and recipe pack.** `rad deploy`s a `radius-env.bicep` that defines a `Radius.Core/recipePacks` resource and the `Radius.Core/environments` resource that references it. Azure downloads the `azure-avm` pack (Azure Verified Modules) from [resource-types-contrib](https://github.com/radius-project/resource-types-contrib); AWS generates an inline `aws-terraform` pack. `radius-env.bicep` is written to the app file's directory (e.g. `.radius/`) and deployed from there, so `rad deploy` resolves the repo's own `bicepconfig.json` (which declares the `radius` extension) — bicep resolves the config nearest the `.bicep` file. The `Radius.Compute/containerImages` type ships with the Radius extension, so no separate resource-type registration is needed. -12. **Run the requested rad commands.** Validates each command in `rad_commands` against the allowed-command set, then runs them in order (stopping on the first failure) and writes a combined `rad-commands-result` artifact. When `rad_commands` is empty it runs the default `rad deploy --environment `, passing the `image` parameter (the `image` input, defaulting to `github.sha`), any application parameters from the `RADIUS_DEPLOY_PARAMS` secret, and the registry push/pull credentials as `registryUsername` (`github.actor`) and `registryPassword` (the built-in `GITHUB_TOKEN`). Those feed the app's `Radius.Security/secrets` resource (`radius-ghcr-registry-creds`), which materializes the registry Secret on the target cluster so the containerImages recipe's in-pod BuildKit can push the application image. The secret value is passed via an argv array and never written into the recorded command string. -13. **Persist state (`rad shutdown`).** Backs the control-plane databases and Terraform recipe-state Secrets up to the `radius-state` git orphan branch. This runs even when the deploy fails (`if: always()`), so a partially-applied Terraform run is not lost. -14. **Tear down.** Runs `rad app list`, and always deletes the ephemeral `radius-cp` cluster. On failure, Radius and application logs are collected and uploaded as the `radius-logs` artifact (three-day retention). +12. **Register custom types and apply custom recipe pack.** When the app's `.radius/` folder carries a `custom-types.yaml` file, the shared `apply-custom-recipe-packs` action registers those resource types with `rad resource-type create --from-file` (skipped when absent). When it carries a `custom-recipe-pack.bicep` file, the action snapshots the recipe-pack IDs before and after `rad deploy`ing that pack to identify the newly-created pack(s), reads the environment's existing `recipePacks` with `rad env show --preview`, and runs `rad env update --recipe-packs --preview` so the environment keeps the default provider pack and gains the custom pack — without pulling in unrelated packs the control plane may know about (skipped when absent). When neither file exists this step is a no-op and the default pack stays in place. +13. **Run the requested rad commands.** Validates each command in `rad_commands` against the allowed-command set, then runs them in order (stopping on the first failure) and writes a combined `rad-commands-result` artifact. When `rad_commands` is empty it runs the default `rad deploy --environment `, passing the `image` parameter (the `image` input, defaulting to `github.sha`), any application parameters from the `RADIUS_DEPLOY_PARAMS` secret, and the registry push/pull credentials as `registryUsername` (`github.actor`) and `registryPassword` (the built-in `GITHUB_TOKEN`). Those feed the app's `Radius.Security/secrets` resource (`radius-ghcr-registry-creds`), which materializes the registry Secret on the target cluster so the containerImages recipe's in-pod BuildKit can push the application image. The secret value is passed via an argv array and never written into the recorded command string. +14. **Persist state (`rad shutdown`).** Backs the control-plane databases and Terraform recipe-state Secrets up to the state archive — the OCI-backed archive by default (pushed to GHCR, selected by the `RADIUS_STATE_*` variables), or the `radius-state` git orphan branch when `RADIUS_STATE_BACKEND=git`. This runs even when the deploy fails (`if: always()`), so a partially-applied Terraform run is not lost. +15. **Tear down.** Runs `rad app list`, and always deletes the ephemeral `radius-cp` cluster. On failure, Radius and application logs are collected and uploaded as the `radius-logs` artifact (three-day retention). ### Triggers and permissions @@ -102,7 +114,7 @@ Triggers and permissions live on the **dispatcher** (`run-rad-commands.yml`); th | `rad_commands` | No | A single `rad` command string, or a JSON array of command strings run in order (the `rad` prefix omitted, e.g. `["deploy .radius/app.bicep --environment dev", "app graph my-app -o json"]`). Each command is validated against the allowed-command set. Falls back to the `RADIUS_RAD_COMMANDS` variable. When empty, the workflow runs its default `rad deploy` of the app bicep. | - **Outputs:** a combined `rad-commands-result` artifact — a JSON document with a top-level `outcome`/`exitCode` and a `commands` array (one entry per command, in input order, with each command's exit code and output). -- **Permissions:** `id-token: write` (required for OIDC), `contents: write` (so `rad shutdown` can push the `radius-state` branch), and `packages: write` (to push the application image built by the containerImages recipe). +- **Permissions:** `id-token: write` (required for OIDC), `contents: write` (so `rad shutdown` can push the `radius-state` branch when the git state backend is selected), and `packages: write` (to push the OCI-backed state archive to GHCR and the application image built by the containerImages recipe). ### Required environment variables @@ -125,7 +137,7 @@ This workflow also reads GitHub Actions **secrets** for image push and applicati ### State persistence (`rad startup` / `rad shutdown`) -`rad startup` and `rad shutdown` are kind-agnostic CLI commands that restore and back up all durable Radius state (control-plane PostgreSQL + Terraform recipe-state Secrets) to a `radius-state` git orphan branch. They do not manage cluster lifecycle — the workflow owns creating and destroying the ephemeral control plane around them. `rad startup` runs after the install (so `rad deploy` plans against prior state) and `rad shutdown` runs after the commands with `if: always()` (so state survives a failed deploy). +`rad startup` and `rad shutdown` are kind-agnostic CLI commands that restore and back up all durable Radius state (control-plane PostgreSQL + Terraform recipe-state Secrets). These workflows use the OCI-backed state archive by default — the `RADIUS_STATE_*` variables select an OCI repository and the workflow logs in to GHCR before `rad startup`/`rad shutdown` — and fall back to the `radius-state` git orphan branch only when `RADIUS_STATE_BACKEND=git`. They do not manage cluster lifecycle — the workflow owns creating and destroying the ephemeral control plane around them. `rad startup` runs after the install (so `rad deploy` plans against prior state) and `rad shutdown` runs after the commands with `if: always()` (so state survives a failed deploy). ### Prerequisites diff --git a/.github/extension/actions/apply-custom-recipe-packs/action.yml b/.github/extension/actions/apply-custom-recipe-packs/action.yml new file mode 100644 index 00000000000..9015649ef0b --- /dev/null +++ b/.github/extension/actions/apply-custom-recipe-packs/action.yml @@ -0,0 +1,112 @@ +# Provider-agnostic custom resource-type / recipe-pack apply shared by +# run-rad-commands-aws.yml and run-rad-commands-azure.yml. Runs after the +# provider-specific "Create Radius environment and recipe pack" step. When the +# app's `.radius/` folder (alongside `.radius/app.bicep`) carries custom resource +# types, this: +# 1. registers those types from `.radius/custom-types.yaml` via +# `rad resource-type create --from-file` (skipped when the file is absent), and +# 2. deploys `.radius/custom-recipe-pack.bicep`, then attaches only the pack(s) +# newly created by that deploy to the environment -- preserving the +# environment's existing recipe packs (e.g. the default provider pack) and +# adding the custom pack, without pulling in unrelated packs the control +# plane may know about (skipped when the file is absent). +# When neither file exists the action is a no-op, leaving the default pack in place. +name: Radius - Apply custom recipe packs +description: Register the repo's custom resource types and recipe pack (if present) and attach the new pack to the environment alongside its existing packs. + +inputs: + environment: + description: Radius environment name to update with the full recipe pack list. + required: true + app-file: + description: Application bicep file. Its directory (e.g. .radius/) is searched for custom-types.yaml and custom-recipe-pack.bicep. + required: true + +runs: + using: composite + steps: + - name: Register custom types and apply custom recipe pack + shell: bash + env: + ENVIRONMENT: ${{ inputs.environment }} + APP_FILE: ${{ inputs.app-file }} + run: | + set -euo pipefail + + # Custom type / recipe-pack files live next to the app file (e.g. .radius/). + # Deploying from that directory lets `rad deploy` resolve the repo's own + # .radius/bicepconfig.json (which declares the `radius` extension) -- bicep + # resolves the config nearest the .bicep file. + APP_DIR=$(dirname "$APP_FILE") + CUSTOM_TYPES_YAML="$APP_DIR/custom-types.yaml" + RECIPE_PACK_BICEP="$APP_DIR/custom-recipe-pack.bicep" + + # 1. Register custom resource types with Radius. --from-file registers every + # type defined in the manifest. Must run before deploying the recipe pack, + # which references these types. Absent file => nothing to register. + if [ -f "$CUSTOM_TYPES_YAML" ]; then + echo "Registering custom resource types from $CUSTOM_TYPES_YAML..." + rad resource-type create --from-file "$CUSTOM_TYPES_YAML" + echo "✅ Custom resource types registered." + else + echo "No custom resource types at $CUSTOM_TYPES_YAML; skipping type registration." + fi + + # 2. Deploy the custom recipe pack and attach it to the environment. Absent + # file => nothing to deploy, keep the default pack. + if [ ! -f "$RECIPE_PACK_BICEP" ]; then + echo "No custom recipe pack at $RECIPE_PACK_BICEP; keeping the default recipe pack." + exit 0 + fi + + # Capture the set of recipe-pack IDs before deploying so we can identify + # exactly which pack(s) this step creates. `rad recipe-pack list` returns + # every pack across scopes (including ones restored from prior state), so we + # must not blindly attach all of them -- that could pull unrelated packs into + # the environment and trip server-side recipe-pack conflict validation. + # ids_json emits a compact JSON array of the non-empty string IDs. + ids_json() { + rad recipe-pack list -o json \ + | jq -c '(if type == "array" then . else [.] end) | map(.id) | map(select(type == "string" and . != ""))' + } + + echo "Recording existing recipe packs..." + PACKS_BEFORE=$(ids_json) + + # Pass --environment explicitly: the recipe pack bicep only creates a + # Radius.Core/recipePacks resource (no environment), so rad deploy would + # otherwise require a default environment to be set. The environment was + # created by the preceding step. + echo "Deploying custom recipe pack from $RECIPE_PACK_BICEP..." + rad deploy "$RECIPE_PACK_BICEP" --environment "$ENVIRONMENT" + + PACKS_AFTER=$(ids_json) + + # New pack(s) = those present after the deploy but not before. + NEW_PACKS=$(jq -nc --argjson before "$PACKS_BEFORE" --argjson after "$PACKS_AFTER" '$after - $before') + + # Preserve the environment's current recipe packs (e.g. the default provider + # pack attached when the environment was created) and add only the new pack(s). + # `rad env update --recipe-packs` REPLACES the list, so we compute the full + # desired set here rather than passing every pack the control plane knows. + # --preview reads the Radius.Core/environments resource (which carries + # properties.recipePacks); without it env show hits the legacy surface and + # can't see recipe packs. That command prints the resource first followed by + # optional provider/recipe JSON docs, so jq -s slurps the stream and takes the + # environment resource ([0]) rather than choking on the trailing documents. + EXISTING_PACKS=$(rad env show "$ENVIRONMENT" --preview -o json | jq -sc '((.[0].properties.recipePacks) // []) | map(select(type == "string" and . != ""))') + + PACK_IDS=$(jq -nr --argjson existing "$EXISTING_PACKS" --argjson new "$NEW_PACKS" '($existing + $new) | unique | join(",")') + + if [ -z "$PACK_IDS" ]; then + echo "No recipe packs found to attach to environment '$ENVIRONMENT'." >&2 + exit 1 + fi + + # Replace the environment's recipe pack list with the preserved + new set so + # the custom types resolve alongside the default provider recipes. --preview + # selects the Radius.Core implementation; without it `rad env update` + # dispatches to the legacy command, which does not understand --recipe-packs. + echo "Attaching recipe packs to environment '$ENVIRONMENT': $PACK_IDS" + rad env update "$ENVIRONMENT" --recipe-packs "$PACK_IDS" --preview + echo "✅ Environment '$ENVIRONMENT' updated with recipe packs." diff --git a/.github/extension/actions/delete-resource/action.yml b/.github/extension/actions/delete-resource/action.yml new file mode 100644 index 00000000000..3d9c08bf4b0 --- /dev/null +++ b/.github/extension/actions/delete-resource/action.yml @@ -0,0 +1,117 @@ +# Provider-agnostic resource delete shared by delete-azure.yml and delete-aws.yml. +# Runs after restore-state (rad startup) has brought back the control-plane state, +# so the environment, its recipe packs, the application, and the Terraform recipe +# state all exist and recipe deletes can run. Deletes a single application or +# environment with the Radius.Core preview surface, then writes a rad-delete-result +# artifact. State is persisted again by the separate `teardown` action (rad shutdown), +# which runs unconditionally after this step. +name: Radius - Delete resource +description: Delete a Radius application or environment (Radius.Core preview) and write a result artifact. + +inputs: + resource-type: + description: What to delete -- "application" or "environment". + required: true + name: + description: Name of the application or environment to delete. + required: true + +runs: + using: composite + steps: + - name: Delete Radius resource + shell: bash + env: + RESOURCE_TYPE: ${{ inputs.resource-type }} + RESOURCE_NAME: ${{ inputs.name }} + run: | + # pipefail so a failed `rad` whose output is piped through `tee` is detected + # via PIPESTATUS rather than masked by tee's exit code. -e/-u make setup and + # result-writing failures (mkdir, mktemp, jq, redirection) fail fast instead + # of silently producing a missing or partial rad-delete-result artifact; -e + # is relaxed only around the `rad | tee` pipeline below so the delete's own + # exit code is captured and reported rather than aborting the step. + set -euo pipefail + mkdir -p /tmp/radius-output + RESULT_FILE=/tmp/radius-output/rad-delete-result.json + + # Write the combined result on exit so the rad-delete-result artifact is + # complete even when the delete fails and the step exits early. + # Capture rad's combined output to a file (not a shell variable) so the + # result artifact can embed it via `jq --rawfile` regardless of size -- + # passing large output through `jq --arg` risks exceeding OS argv limits and + # failing the artifact write even on a successful delete. Created up front and + # left empty so write_result can always read it, including on an early exit. + OUTCOME="succeeded" + EXIT=0 + OUTFILE=$(mktemp) + write_result() { + jq -n \ + --arg outcome "$OUTCOME" \ + --argjson exitCode "$EXIT" \ + --arg resourceType "$RESOURCE_TYPE" \ + --arg name "$RESOURCE_NAME" \ + --rawfile output "$OUTFILE" \ + '{schemaVersion:"1.0", outcome:$outcome, exitCode:$exitCode, resourceType:$resourceType, name:$name, output:$output}' \ + > "$RESULT_FILE" + } + trap write_result EXIT + + # Sanitize untrusted inputs for log output only: strip control characters + # (newlines/CR in particular) so a crafted resource type/name cannot start a + # new log line or inject GitHub Actions workflow commands (e.g. a name + # containing `\n::warning::`). The raw values are still passed verbatim to + # `rad` and to write_result (jq --arg JSON-encodes them safely). + SAFE_TYPE=$(printf '%s' "$RESOURCE_TYPE" | tr -d '[:cntrl:]') + SAFE_NAME=$(printf '%s' "$RESOURCE_NAME" | tr -d '[:cntrl:]') + + if [ -z "${RESOURCE_NAME//[[:space:]]/}" ]; then + echo "No resource name supplied to delete." >&2 + OUTCOME="invalid_input" + EXIT=2 + exit 2 + fi + + # Map the resource type to its rad command. Only application and environment + # are supported; anything else fails fast without touching the control plane. + case "$RESOURCE_TYPE" in + application) VERB=(app delete) ;; + environment) VERB=(env delete) ;; + *) + echo "Unsupported resource type: '$SAFE_TYPE' (expected 'application' or 'environment')." >&2 + OUTCOME="invalid_input" + EXIT=2 + exit 2 + ;; + esac + + # --yes bypasses the interactive confirmation prompt; --preview selects the + # Radius.Core implementation the deploy workflow provisions (without it the + # base command dispatches to the legacy, non-Radius.Core surface). + # Keep the group label constant: RESOURCE_NAME is untrusted and a newline + # or workflow-command-like text in it could corrupt log rendering or inject + # workflow commands. The actual `rad` invocation below still carries the name. + echo "::group::rad ${VERB[*]} --yes --preview" + # Relax -e only for the pipeline so a non-zero `rad` exit is captured via + # PIPESTATUS and reported below, rather than aborting the step before we + # can write the result artifact. -e is restored immediately after. + set +e + rad "${VERB[@]}" "$RESOURCE_NAME" --yes --preview 2>&1 | tee "$OUTFILE" + EXIT=${PIPESTATUS[0]} + set -e + echo "::endgroup::" + + if [ "$EXIT" -ne 0 ]; then + OUTCOME="failed" + echo "Failed to delete $SAFE_TYPE '$SAFE_NAME'." >&2 + exit "$EXIT" + fi + echo "✅ Deleted $SAFE_TYPE '$SAFE_NAME'." + + - name: Upload delete result + if: always() + uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4 + with: + name: rad-delete-result + path: /tmp/radius-output/ + retention-days: 1 diff --git a/.github/extension/delete-application.yml b/.github/extension/delete-application.yml new file mode 100644 index 00000000000..86b0b06976d --- /dev/null +++ b/.github/extension/delete-application.yml @@ -0,0 +1,70 @@ +# This workflow is auto-generated by Radius. It deletes a Radius application: it +# detects whether the selected GitHub Environment is wired for Azure or AWS and +# calls the matching reusable delete workflow (delete-azure.yml / delete-aws.yml) +# with resource_type=application. Deleting runs on an ephemeral k3d control plane +# that restores the persisted Radius state, runs `rad app delete` (including any +# recipe deletes against the target cluster), and persists the updated state again. +name: Radius - Delete Application + +on: + workflow_dispatch: + inputs: + environment: + description: 'GitHub Environment name' + required: true + default: '{{ENV}}' + application: + description: 'Name of the application to delete' + required: true + +permissions: + id-token: write + contents: write + packages: write + +jobs: + # Bind to the GitHub Environment so environment-scoped variables are visible, then + # pick the provider from whichever identifying variable is set. A job that calls a + # reusable workflow (`uses:`) can't bind an environment itself, so this routing + # decision has to happen in a regular job first. + detect: + name: Detect provider + runs-on: ubuntu-latest + environment: ${{ inputs.environment || '{{ENV}}' }} + outputs: + provider: ${{ steps.detect.outputs.provider }} + steps: + - name: Determine provider from environment variables + id: detect + run: | + if [ -n "${{ vars.AZURE_CLIENT_ID }}" ]; then + echo "provider=azure" >> "$GITHUB_OUTPUT" + elif [ -n "${{ vars.AWS_ROLE_ARN }}" ]; then + echo "provider=aws" >> "$GITHUB_OUTPUT" + else + echo "No AZURE_CLIENT_ID or AWS_ROLE_ARN set on this environment." >&2 + echo "provider=none" >> "$GITHUB_OUTPUT" + exit 1 + fi + + azure: + name: Azure + needs: detect + if: ${{ needs.detect.outputs.provider == 'azure' }} + uses: ./.github/workflows/delete-azure.yml + with: + environment: ${{ inputs.environment || '{{ENV}}' }} + resource_type: application + name: ${{ inputs.application }} + secrets: inherit + + aws: + name: AWS + needs: detect + if: ${{ needs.detect.outputs.provider == 'aws' }} + uses: ./.github/workflows/delete-aws.yml + with: + environment: ${{ inputs.environment || '{{ENV}}' }} + resource_type: application + name: ${{ inputs.application }} + secrets: inherit diff --git a/.github/extension/delete-aws.yml b/.github/extension/delete-aws.yml new file mode 100644 index 00000000000..009f0395084 --- /dev/null +++ b/.github/extension/delete-aws.yml @@ -0,0 +1,222 @@ +# This workflow is auto-generated by Radius to delete a Radius application or +# environment from an AWS environment. It is a reusable (workflow_call) workflow +# invoked by the delete-application.yml / delete-environment.yml dispatchers; it is +# not dispatched directly. It creates an ephemeral k3d cluster for the Radius control +# plane, connects to the user's EKS cluster, restores persisted state, deletes the +# requested resource (running any recipe deletes against the target cluster), then +# persists the updated state again and tears the control plane down. The provider- +# agnostic phases are shared composite actions in radius-project/radius; only the +# AWS-specific steps live here. +name: Radius - Delete (AWS) + +on: + workflow_call: + inputs: + environment: + description: 'GitHub Environment name' + type: string + required: true + resource_type: + description: 'What to delete: application or environment' + type: string + required: true + name: + description: 'Name of the application or environment to delete' + type: string + required: true + +permissions: + id-token: write + contents: write + packages: write + +env: + ENVIRONMENT: ${{ inputs.environment }} + +jobs: + delete: + name: Delete with Radius + runs-on: ubuntu-latest + environment: ${{ inputs.environment }} + # OCI-backed Radius control-plane state, provisioned per-environment by the + # deploy tooling (canvas extension) as environment variables. Set at job + # level (not workflow level) so environment-scoped vars.* resolve, and so the + # shared restore-state (rad startup) and teardown (rad shutdown) actions — + # plus the delete commands — inherit them and can open the state archive. + # Without these, rad startup fails with "OCI archive repository is not + # configured; set RADIUS_STATE_REGISTRY or RADIUS_GRAPH_REGISTRY". + env: + RADIUS_STATE_BACKEND: ${{ vars.RADIUS_STATE_BACKEND }} + RADIUS_STATE_REGISTRY: ${{ vars.RADIUS_STATE_REGISTRY }} + RADIUS_STATE_ARCHIVE: ${{ vars.RADIUS_STATE_ARCHIVE }} + steps: + - name: Checkout + uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4 + + - name: Configure AWS Credentials (OIDC) + if: ${{ vars.AWS_ROLE_ARN != '' }} + uses: aws-actions/configure-aws-credentials@7474bc4690e29a8392af63c5b98e7449536d5c3a # v4 + with: + role-to-assume: ${{ vars.AWS_ROLE_ARN }} + aws-region: ${{ vars.AWS_REGION }} + + - name: Log in to GHCR for the state archive + # The OCI-backed Radius state archive authenticates to GHCR using the + # runner's Docker credential store (~/.docker/config.json). rad startup + # (restore-state) and rad shutdown (teardown) open that archive, so the + # runner must be logged in BEFORE those steps run -- otherwise the pull + # is anonymous and GHCR rejects the private radius-state package with a + # 401. docker login persists for the whole job, covering teardown too. + # GITHUB_TOKEN has packages: write (set at job level) for repo-linked packages. + # If the state archive lives under a different owner/org, replace the password with a PAT stored as a secret that can read/write that GHCR package. + uses: docker/login-action@af1e73f918a031802d376d3c8bbc3fe56130a9b0 # v4.4.0 + with: + registry: ghcr.io + username: ${{ github.actor }} + password: ${{ secrets.GITHUB_TOKEN }} + + - name: Get target cluster kubeconfig + run: | + mkdir -p "$HOME/.kube" + echo "RADIUS_TARGET_KUBECONFIG=$HOME/.kube/target-cluster" >> "$GITHUB_ENV" + + - name: Connect to EKS cluster + if: ${{ vars.AWS_EKS_CLUSTER_NAME != '' }} + run: | + CLUSTER="${{ vars.AWS_EKS_CLUSTER_NAME }}" + REGION="${{ vars.AWS_REGION }}" + ROLE_ARN="${{ vars.AWS_ROLE_ARN }}" + TARGET="$RADIUS_TARGET_KUBECONFIG" + + # Ensure the IAM role has access to the EKS cluster. These calls are + # idempotent by intent: a ResourceInUseException means the entry/policy is + # already present, which is fine to ignore. Any other failure (auth, + # permissions, throttling) is surfaced with its stderr and fails the step + # rather than being masked behind an "already exists" message. + echo "Ensuring EKS access entry for $ROLE_ARN..." + if ! err=$(aws eks create-access-entry \ + --cluster-name "$CLUSTER" \ + --principal-arn "$ROLE_ARN" \ + --type STANDARD \ + --region "$REGION" 2>&1); then + if printf '%s' "$err" | grep -q 'ResourceInUseException'; then + echo "Access entry already exists" + else + printf '%s\n' "$err" >&2 + exit 1 + fi + fi + if ! err=$(aws eks associate-access-policy \ + --cluster-name "$CLUSTER" \ + --principal-arn "$ROLE_ARN" \ + --policy-arn arn:aws:eks::aws:cluster-access-policy/AmazonEKSClusterAdminPolicy \ + --access-scope type=cluster \ + --region "$REGION" 2>&1); then + if printf '%s' "$err" | grep -q 'ResourceInUseException'; then + echo "Access policy already associated" + else + printf '%s\n' "$err" >&2 + exit 1 + fi + fi + + # Build a static kubeconfig with a bearer token instead of exec-based auth. + ENDPOINT=$(aws eks describe-cluster --name "$CLUSTER" --region "$REGION" --query 'cluster.endpoint' --output text) + CA_DATA=$(aws eks describe-cluster --name "$CLUSTER" --region "$REGION" --query 'cluster.certificateAuthority.data' --output text) + TOKEN=$(aws eks get-token --cluster-name "$CLUSTER" --region "$REGION" --output json | jq -r '.status.token') + printf 'apiVersion: v1\nclusters:\n- cluster:\n certificate-authority-data: %s\n server: %s\n name: eks\ncontexts:\n- context:\n cluster: eks\n user: eks-user\n name: eks\ncurrent-context: eks\nkind: Config\nusers:\n- name: eks-user\n user:\n token: %s\n' "$CA_DATA" "$ENDPOINT" "$TOKEN" > "$TARGET" + echo "EKS kubeconfig saved with static token" + kubectl --kubeconfig "$TARGET" cluster-info || echo "WARNING: Could not connect to EKS cluster" + + - name: Set up control plane + uses: radius-project/radius/.github/extension/actions/setup-control-plane@{{RADIUS_REF}} + + - name: Project cloud OIDC tokens into Radius pods + run: | + # Mint GitHub OIDC tokens and mount them where Radius expects them, so the + # recipe deletes triggered by the delete can authenticate to AWS exactly as + # the deploy workflow does. GitHub Actions is already a trusted OIDC issuer + # for AWS via the IAM role trust policy. + if [ -n "${{ vars.AWS_ROLE_ARN }}" ]; then + echo "Projecting AWS OIDC token..." + AWS_TOKEN=$(curl -sS -H "Authorization: bearer $ACTIONS_ID_TOKEN_REQUEST_TOKEN" \ + "$ACTIONS_ID_TOKEN_REQUEST_URL&audience=sts.amazonaws.com" | jq -r '.value') + kubectl create secret generic aws-oidc-token -n radius-system \ + --from-literal=token="$AWS_TOKEN" --dry-run=client -o yaml | kubectl apply -f - + + # AWS IRSA: the UCP AWS proxy (ucp) and the Terraform AWS provider + # (dynamic-rp) both read the token from the hard-coded path + # /var/run/secrets/eks.amazonaws.com/serviceaccount/token. + AWS_PATCH='[ + {"op":"add","path":"/spec/template/spec/volumes/-","value":{"name":"aws-oidc-token","secret":{"secretName":"aws-oidc-token"}}}, + {"op":"add","path":"/spec/template/spec/containers/0/volumeMounts/-","value":{"name":"aws-oidc-token","mountPath":"/var/run/secrets/eks.amazonaws.com/serviceaccount","readOnly":true}} + ]' + for deploy in applications-rp dynamic-rp ucp; do + kubectl patch deployment $deploy -n radius-system --type=json -p="$AWS_PATCH" 2>/dev/null || true + done + fi + + for deploy in applications-rp dynamic-rp bicep-de ucp; do + kubectl rollout status deployment/$deploy -n radius-system --timeout=300s || true + done + echo "✅ Cloud OIDC tokens projected into Radius pods." + + - name: Refresh external deployment target credentials + run: | + TARGET_KUBECONFIG="$RADIUS_TARGET_KUBECONFIG" + + if [ ! -f "$TARGET_KUBECONFIG" ]; then + echo "No target kubeconfig found, resources are on the k3d control plane" + exit 0 + fi + + # Refresh the EKS token right before delete (EKS tokens are short-lived and + # the one minted earlier may have expired during install). + if [ -n "${{ vars.AWS_ROLE_ARN }}" ] && [ -n "${{ vars.AWS_EKS_CLUSTER_NAME }}" ]; then + echo "Generating fresh EKS token..." + CLUSTER="${{ vars.AWS_EKS_CLUSTER_NAME }}" + REGION="${{ vars.AWS_REGION }}" + ENDPOINT=$(aws eks describe-cluster --name "$CLUSTER" --region "$REGION" --query 'cluster.endpoint' --output text) + CA_DATA=$(aws eks describe-cluster --name "$CLUSTER" --region "$REGION" --query 'cluster.certificateAuthority.data' --output text) + TOKEN=$(aws eks get-token --cluster-name "$CLUSTER" --region "$REGION" --output json | jq -r '.status.token') + printf 'apiVersion: v1\nclusters:\n- cluster:\n certificate-authority-data: %s\n server: %s\n name: eks\ncontexts:\n- context:\n cluster: eks\n user: eks-user\n name: eks\ncurrent-context: eks\nkind: Config\nusers:\n- name: eks-user\n user:\n token: %s\n' "$CA_DATA" "$ENDPOINT" "$TOKEN" > "$TARGET_KUBECONFIG" + fi + + # Update the secret the chart mounted at install with the refreshed + # kubeconfig, then restart the recipe-executing pods so they re-read it. + kubectl create secret generic target-kubeconfig --namespace radius-system \ + --from-file=kubeconfig="$TARGET_KUBECONFIG" --dry-run=client -o yaml | kubectl apply -f - + for deploy in applications-rp dynamic-rp bicep-de; do + kubectl rollout restart deployment/$deploy -n radius-system + done + + echo "Waiting for rollouts..." + kubectl rollout status deployment/applications-rp -n radius-system --timeout=300s + kubectl rollout status deployment/dynamic-rp -n radius-system --timeout=300s + kubectl rollout status deployment/bicep-de -n radius-system --timeout=300s + echo "External deployment target configured." + + - name: Restore Radius state + uses: radius-project/radius/.github/extension/actions/restore-state@{{RADIUS_REF}} + with: + namespace: ${{ vars.KUBERNETES_NAMESPACE || 'default' }} + + - name: Register cloud credentials with Radius + run: | + # Register the AWS identity so the recipe deletes can authenticate, + # reading the projected OIDC token at runtime. + if [ -n "${{ vars.AWS_ROLE_ARN }}" ]; then + echo "Registering AWS IRSA credential..." + rad credential register aws irsa --iam-role "${{ vars.AWS_ROLE_ARN }}" + fi + echo "✅ Cloud credentials registered with Radius." + + - name: Delete Radius resource + uses: radius-project/radius/.github/extension/actions/delete-resource@{{RADIUS_REF}} + with: + resource-type: ${{ inputs.resource_type }} + name: ${{ inputs.name }} + + - name: Teardown + if: always() + uses: radius-project/radius/.github/extension/actions/teardown@{{RADIUS_REF}} diff --git a/.github/extension/delete-azure.yml b/.github/extension/delete-azure.yml new file mode 100644 index 00000000000..9cd34b327ec --- /dev/null +++ b/.github/extension/delete-azure.yml @@ -0,0 +1,175 @@ +# This workflow is auto-generated by Radius to delete a Radius application or +# environment from an Azure environment. It is a reusable (workflow_call) workflow +# invoked by the delete-application.yml / delete-environment.yml dispatchers; it is +# not dispatched directly. It creates an ephemeral k3d cluster for the Radius control +# plane, connects to the user's AKS cluster, restores persisted state, deletes the +# requested resource (running any recipe deletes against the target cluster), then +# persists the updated state again and tears the control plane down. The provider- +# agnostic phases are shared composite actions in radius-project/radius; only the +# Azure-specific steps live here. +name: Radius - Delete (Azure) + +on: + workflow_call: + inputs: + environment: + description: 'GitHub Environment name' + type: string + required: true + resource_type: + description: 'What to delete: application or environment' + type: string + required: true + name: + description: 'Name of the application or environment to delete' + type: string + required: true + +permissions: + id-token: write + contents: write + packages: write + +env: + ENVIRONMENT: ${{ inputs.environment }} + +jobs: + delete: + name: Delete with Radius + runs-on: ubuntu-latest + environment: ${{ inputs.environment }} + # OCI-backed Radius control-plane state, provisioned per-environment by the + # deploy tooling (canvas extension) as environment variables. Set at job + # level (not workflow level) so environment-scoped vars.* resolve, and so the + # shared restore-state (rad startup) and teardown (rad shutdown) actions — + # plus the delete commands — inherit them and can open the state archive. + # Without these, rad startup fails with "OCI archive repository is not + # configured; set RADIUS_STATE_REGISTRY or RADIUS_GRAPH_REGISTRY". + env: + RADIUS_STATE_BACKEND: ${{ vars.RADIUS_STATE_BACKEND }} + RADIUS_STATE_REGISTRY: ${{ vars.RADIUS_STATE_REGISTRY }} + RADIUS_STATE_ARCHIVE: ${{ vars.RADIUS_STATE_ARCHIVE }} + steps: + - name: Checkout + uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4 + + - name: Azure Login (OIDC) + if: ${{ vars.AZURE_CLIENT_ID != '' }} + uses: azure/login@532459ea530d8321f2fb9bb10d1e0bcf23869a43 # v3.0.0 + with: + client-id: ${{ vars.AZURE_CLIENT_ID }} + tenant-id: ${{ vars.AZURE_TENANT_ID }} + subscription-id: ${{ vars.AZURE_SUBSCRIPTION_ID }} + + - name: Log in to GHCR for the state archive + # The OCI-backed Radius state archive authenticates to GHCR using the + # runner's Docker credential store (~/.docker/config.json). rad startup + # (restore-state) and rad shutdown (teardown) open that archive, so the + # runner must be logged in BEFORE those steps run -- otherwise the pull + # is anonymous and GHCR rejects the private radius-state package with a + # 401. docker login persists for the whole job, covering teardown too. + # GITHUB_TOKEN has packages: write (set at job level) for repo-linked packages. + # If the state archive lives under a different owner/org, replace the password with a PAT stored as a secret that can read/write that GHCR package. + uses: docker/login-action@af1e73f918a031802d376d3c8bbc3fe56130a9b0 # v4.4.0 + with: + registry: ghcr.io + username: ${{ github.actor }} + password: ${{ secrets.GITHUB_TOKEN }} + + - name: Get target cluster kubeconfig + run: | + mkdir -p "$HOME/.kube" + echo "RADIUS_TARGET_KUBECONFIG=$HOME/.kube/target-cluster" >> "$GITHUB_ENV" + + - name: Connect to AKS cluster + if: ${{ vars.AZURE_AKS_CLUSTER_NAME != '' }} + run: | + az aks get-credentials \ + --resource-group "${{ vars.AZURE_RESOURCE_GROUP }}" \ + --name "${{ vars.AZURE_AKS_CLUSTER_NAME }}" \ + --subscription "${{ vars.AZURE_SUBSCRIPTION_ID }}" \ + --file "$RADIUS_TARGET_KUBECONFIG" + + - name: Set up control plane + uses: radius-project/radius/.github/extension/actions/setup-control-plane@{{RADIUS_REF}} + + - name: Project cloud OIDC tokens into Radius pods + run: | + # Mint GitHub OIDC tokens and mount them where Radius expects them, so the + # recipe deletes triggered by the delete can authenticate to Azure exactly + # as the deploy workflow does. GitHub Actions is already a trusted OIDC + # issuer for Azure via the AAD federated credential. + if [ -n "${{ vars.AZURE_CLIENT_ID }}" ]; then + echo "Projecting Azure OIDC token..." + AZ_TOKEN=$(curl -sS -H "Authorization: bearer $ACTIONS_ID_TOKEN_REQUEST_TOKEN" \ + "$ACTIONS_ID_TOKEN_REQUEST_URL&audience=api://AzureADTokenExchange" | jq -r '.value') + # The secret data key becomes the mounted file name, so it must be + # 'azure-identity-token' -- the file Radius and the Terraform azurerm + # provider read from /var/run/secrets/azure/tokens/. + kubectl create secret generic azure-oidc-token -n radius-system \ + --from-literal=azure-identity-token="$AZ_TOKEN" --dry-run=client -o yaml | kubectl apply -f - + + AZ_PATCH='[ + {"op":"add","path":"/spec/template/spec/volumes/-","value":{"name":"azure-oidc-token","secret":{"secretName":"azure-oidc-token"}}}, + {"op":"add","path":"/spec/template/spec/containers/0/volumeMounts/-","value":{"name":"azure-oidc-token","mountPath":"/var/run/secrets/azure/tokens","readOnly":true}}, + {"op":"add","path":"/spec/template/spec/containers/0/env/-","value":{"name":"AZURE_FEDERATED_TOKEN_FILE","value":"/var/run/secrets/azure/tokens/azure-identity-token"}} + ]' + for deploy in applications-rp dynamic-rp bicep-de; do + kubectl patch deployment $deploy -n radius-system --type=json -p="$AZ_PATCH" 2>/dev/null || true + done + fi + + for deploy in applications-rp dynamic-rp bicep-de ucp; do + kubectl rollout status deployment/$deploy -n radius-system --timeout=300s || true + done + echo "✅ Cloud OIDC tokens projected into Radius pods." + + - name: Refresh external deployment target credentials + run: | + TARGET_KUBECONFIG="$RADIUS_TARGET_KUBECONFIG" + + if [ ! -f "$TARGET_KUBECONFIG" ]; then + echo "No target kubeconfig found, resources are on the k3d control plane" + exit 0 + fi + + # Update the secret the chart mounted at install with the refreshed + # kubeconfig, then restart the recipe-executing pods so they re-read it. + kubectl create secret generic target-kubeconfig --namespace radius-system \ + --from-file=kubeconfig="$TARGET_KUBECONFIG" --dry-run=client -o yaml | kubectl apply -f - + for deploy in applications-rp dynamic-rp bicep-de; do + kubectl rollout restart deployment/$deploy -n radius-system + done + + echo "Waiting for rollouts..." + kubectl rollout status deployment/applications-rp -n radius-system --timeout=300s + kubectl rollout status deployment/dynamic-rp -n radius-system --timeout=300s + kubectl rollout status deployment/bicep-de -n radius-system --timeout=300s + echo "External deployment target configured." + + - name: Restore Radius state + uses: radius-project/radius/.github/extension/actions/restore-state@{{RADIUS_REF}} + with: + namespace: ${{ vars.KUBERNETES_NAMESPACE || 'default' }} + + - name: Register cloud credentials with Radius + run: | + # Register the Azure identity so the recipe deletes can authenticate, + # reading the projected OIDC token at runtime. + if [ -n "${{ vars.AZURE_CLIENT_ID }}" ]; then + echo "Registering Azure workload identity credential..." + rad credential register azure wi \ + --client-id "${{ vars.AZURE_CLIENT_ID }}" \ + --tenant-id "${{ vars.AZURE_TENANT_ID }}" + fi + echo "✅ Cloud credentials registered with Radius." + + - name: Delete Radius resource + uses: radius-project/radius/.github/extension/actions/delete-resource@{{RADIUS_REF}} + with: + resource-type: ${{ inputs.resource_type }} + name: ${{ inputs.name }} + + - name: Teardown + if: always() + uses: radius-project/radius/.github/extension/actions/teardown@{{RADIUS_REF}} diff --git a/.github/extension/delete-environment.yml b/.github/extension/delete-environment.yml new file mode 100644 index 00000000000..81a8160a62e --- /dev/null +++ b/.github/extension/delete-environment.yml @@ -0,0 +1,73 @@ +# This workflow is auto-generated by Radius. It deletes a Radius environment: it +# detects whether the selected GitHub Environment is wired for Azure or AWS and +# calls the matching reusable delete workflow (delete-azure.yml / delete-aws.yml) +# with resource_type=environment. Deleting runs on an ephemeral k3d control plane +# that restores the persisted Radius state, runs `rad env delete`, and persists the +# updated state again. The Radius environment name defaults to the GitHub Environment +# name (the deploy workflow names the Radius environment after it), but can be +# overridden with the environment_name input. +name: Radius - Delete Environment + +on: + workflow_dispatch: + inputs: + environment: + description: 'GitHub Environment name' + required: true + default: '{{ENV}}' + environment_name: + description: 'Radius environment name to delete (defaults to the GitHub Environment name)' + required: false + default: '' + +permissions: + id-token: write + contents: write + packages: write + +jobs: + # Bind to the GitHub Environment so environment-scoped variables are visible, then + # pick the provider from whichever identifying variable is set. A job that calls a + # reusable workflow (`uses:`) can't bind an environment itself, so this routing + # decision has to happen in a regular job first. + detect: + name: Detect provider + runs-on: ubuntu-latest + environment: ${{ inputs.environment || '{{ENV}}' }} + outputs: + provider: ${{ steps.detect.outputs.provider }} + steps: + - name: Determine provider from environment variables + id: detect + run: | + if [ -n "${{ vars.AZURE_CLIENT_ID }}" ]; then + echo "provider=azure" >> "$GITHUB_OUTPUT" + elif [ -n "${{ vars.AWS_ROLE_ARN }}" ]; then + echo "provider=aws" >> "$GITHUB_OUTPUT" + else + echo "No AZURE_CLIENT_ID or AWS_ROLE_ARN set on this environment." >&2 + echo "provider=none" >> "$GITHUB_OUTPUT" + exit 1 + fi + + azure: + name: Azure + needs: detect + if: ${{ needs.detect.outputs.provider == 'azure' }} + uses: ./.github/workflows/delete-azure.yml + with: + environment: ${{ inputs.environment || '{{ENV}}' }} + resource_type: environment + name: ${{ inputs.environment_name || inputs.environment || '{{ENV}}' }} + secrets: inherit + + aws: + name: AWS + needs: detect + if: ${{ needs.detect.outputs.provider == 'aws' }} + uses: ./.github/workflows/delete-aws.yml + with: + environment: ${{ inputs.environment || '{{ENV}}' }} + resource_type: environment + name: ${{ inputs.environment_name || inputs.environment || '{{ENV}}' }} + secrets: inherit diff --git a/.github/extension/run-rad-commands-aws.yml b/.github/extension/run-rad-commands-aws.yml index 475f54d48c6..d5e3ff63d60 100644 --- a/.github/extension/run-rad-commands-aws.yml +++ b/.github/extension/run-rad-commands-aws.yml @@ -94,20 +94,37 @@ jobs: ROLE_ARN="${{ vars.AWS_ROLE_ARN }}" TARGET="$RADIUS_TARGET_KUBECONFIG" - # Ensure the IAM role has access to the EKS cluster. - # Creates an access entry if one doesn't already exist. + # Ensure the IAM role has access to the EKS cluster. These calls are + # idempotent by intent: a ResourceInUseException means the entry/policy is + # already present, which is fine to ignore. Any other failure (auth, + # permissions, throttling) is surfaced with its stderr and fails the step + # rather than being masked behind an "already exists" message. echo "Ensuring EKS access entry for $ROLE_ARN..." - aws eks create-access-entry \ + if ! err=$(aws eks create-access-entry \ --cluster-name "$CLUSTER" \ --principal-arn "$ROLE_ARN" \ --type STANDARD \ - --region "$REGION" 2>/dev/null || echo "Access entry already exists" - aws eks associate-access-policy \ + --region "$REGION" 2>&1); then + if printf '%s' "$err" | grep -q 'ResourceInUseException'; then + echo "Access entry already exists" + else + printf '%s\n' "$err" >&2 + exit 1 + fi + fi + if ! err=$(aws eks associate-access-policy \ --cluster-name "$CLUSTER" \ --principal-arn "$ROLE_ARN" \ --policy-arn arn:aws:eks::aws:cluster-access-policy/AmazonEKSClusterAdminPolicy \ --access-scope type=cluster \ - --region "$REGION" 2>/dev/null || echo "Access policy already associated" + --region "$REGION" 2>&1); then + if printf '%s' "$err" | grep -q 'ResourceInUseException'; then + echo "Access policy already associated" + else + printf '%s\n' "$err" >&2 + exit 1 + fi + fi # Build a static kubeconfig with a bearer token instead of exec-based auth. # The exec-based config requires aws CLI inside the container, which Radius @@ -349,6 +366,12 @@ jobs: rad deploy "$ENV_BICEP" echo "✅ Environment '$ENVIRONMENT' created with recipe pack." + - name: Apply custom recipe packs + uses: radius-project/radius/.github/extension/actions/apply-custom-recipe-packs@{{RADIUS_REF}} + with: + environment: ${{ inputs.environment }} + app-file: ${{ env.APP_FILE }} + - name: Run rad commands uses: radius-project/radius/.github/extension/actions/run-rad-commands@{{RADIUS_REF}} with: diff --git a/.github/extension/run-rad-commands-azure.yml b/.github/extension/run-rad-commands-azure.yml index 01b5b4010b6..df075096f1d 100644 --- a/.github/extension/run-rad-commands-azure.yml +++ b/.github/extension/run-rad-commands-azure.yml @@ -255,6 +255,12 @@ jobs: --parameters containerImagesRegistrySecretName="$REGISTRY_SECRET_NAME" echo "✅ Environment '$ENVIRONMENT' created with recipe pack." + - name: Apply custom recipe packs + uses: radius-project/radius/.github/extension/actions/apply-custom-recipe-packs@{{RADIUS_REF}} + with: + environment: ${{ inputs.environment }} + app-file: ${{ env.APP_FILE }} + - name: Run rad commands uses: radius-project/radius/.github/extension/actions/run-rad-commands@{{RADIUS_REF}} with: diff --git a/eng/design-notes/environments/2026-06-repo-radius-deploy-workflow.md b/eng/design-notes/environments/2026-06-repo-radius-deploy-workflow.md index b857ae8018f..582c978bba3 100644 --- a/eng/design-notes/environments/2026-06-repo-radius-deploy-workflow.md +++ b/eng/design-notes/environments/2026-06-repo-radius-deploy-workflow.md @@ -23,7 +23,7 @@ The workflow was originally a single file that branched on which provider variab Explicitly **out of scope**: - **Cloud-side OIDC / permission provisioning** — creating the AWS IAM role + trust policy or the Entra app registration + federated credential. The workflow *consumes* an environment that is already federated; standing that up is tracked separately. -- **The state-storage mechanism** (`rad startup` / `rad shutdown`, the `radius-state` git orphan branch) — owned by the [state-storage design](../2026-06-repo-radius-state-storage.md). +- **The state-storage mechanism** (`rad startup` / `rad shutdown`, the OCI-backed state archive or the `radius-state` git orphan branch) — owned by the [state-storage design](../2026-06-repo-radius-state-storage.md). - **The multi-cluster seam internals** (`global.targetCluster`, the cluster access resolver) — owned by the [multi-cluster design](2026-06-multi-cluster.md). This document only describes how the workflow *drives* that seam. - **Mid-run cloud-token refresh** beyond the single pre-deploy EKS refresh — a long Azure run may outlive the one-time token exchange; refreshing it mid-run is a deferred fast follow. @@ -91,7 +91,7 @@ All **third-party actions** used anywhere in the generated workflows and shared ### Permissions -`id-token: write` (OIDC), `contents: write` (so `rad shutdown` can push the `radius-state` branch), and `packages: write` (so container-image recipes can push to GHCR). Declared on the dispatcher and, because reusable workflows run with the caller's grants, inherited by the provider workflows. +`id-token: write` (OIDC), `contents: write` (so `rad shutdown` can push the `radius-state` branch when the git state backend is selected), and `packages: write` (so container-image recipes and the OCI-backed state archive can push to GHCR). Declared on the dispatcher and, because reusable workflows run with the caller's grants, inherited by the provider workflows. ## Workflow stages @@ -107,9 +107,10 @@ flowchart TD H --> I[rad startup
restore PostgreSQL + Terraform state] I --> J[rad credential register
aws irsa / azure wi] J --> K[Create environment + recipe pack] - K --> L[Provision registry creds on control plane] + K --> KA[Apply custom recipe packs
register custom types + attach new pack
optional] + KA --> L[Provision registry creds on control plane] L --> M[Run rad_commands or default deploy
upload rad-commands-result artifact] - M --> N[rad shutdown
back up + push radius-state] + M --> N[rad shutdown
back up + push state archive] N --> O[Delete k3d cluster] ``` @@ -150,7 +151,7 @@ The action contract hides which model is in use: the dispatch inputs and the res ### State persistence — `rad startup` / `rad shutdown` -`rad startup` and `rad shutdown` are kind-agnostic CLI commands that back up and restore all durable Radius state (control-plane PostgreSQL + Terraform recipe-state Secrets) to a `radius-state` git orphan branch pushed to the repo's `origin`. They do not manage cluster lifecycle — the workflow owns creating and destroying the ephemeral control plane around them. The mechanism is the plan of record; see the [state-storage design](../2026-06-repo-radius-state-storage.md). +`rad startup` and `rad shutdown` are kind-agnostic CLI commands that back up and restore all durable Radius state (control-plane PostgreSQL + Terraform recipe-state Secrets). These workflows use the OCI-backed state archive by default (the `RADIUS_STATE_*` variables select an OCI repository, pushed to GHCR after a docker login), falling back to a `radius-state` git orphan branch pushed to the repo's `origin` only when `RADIUS_STATE_BACKEND=git`. They do not manage cluster lifecycle — the workflow owns creating and destroying the ephemeral control plane around them. The mechanism is the plan of record; see the [state-storage design](../2026-06-repo-radius-state-storage.md). ### Recipe pack and environment @@ -158,6 +159,17 @@ The workflow provisions a `Radius.Core/recipePacks` resource and a `Radius.Core/ The `radius-env.bicep` that carries the pack is written to the app file's directory (e.g. `.radius/`) and deployed from there. bicep resolves `bicepconfig.json` nearest the `.bicep` file, so deploying from that directory picks up the repo's own `.radius/bicepconfig.json` — which declares the `radius` extension — rather than a (non-existent) config at the workspace root. The `Radius.Compute/containerImages` type ships with the published `radius` Bicep extension, so the workflow no longer registers resource types or wires a local Bicep extension at deploy time. +**Custom resource types and recipe packs.** A repo that defines its own custom resource types places two files in the app's `.radius/` folder (next to `.radius/app.bicep`): a `custom-types.yaml` resource-type manifest and a `custom-recipe-pack.bicep` recipe pack. After the default provider pack and environment are created, the shared `apply-custom-recipe-packs` composite action registers the types with `rad resource-type create --from-file custom-types.yaml` (so the recipe pack's referenced types exist), then snapshots the recipe-pack IDs before and after `rad deploy`ing `custom-recipe-pack.bicep` to identify the newly-created pack(s), reads the environment's current `recipePacks` with `rad env show --preview`, and runs `rad env update --recipe-packs --preview`. Because `--recipe-packs` replaces the list, the action computes the union of the environment's existing packs and only the newly-created pack(s) — it deliberately does **not** attach every pack from `rad recipe-pack list`, which spans all scopes and could pull unrelated packs (e.g. restored from prior state) into the environment and trip recipe-pack conflict validation. Each file is optional and independent: an absent `custom-types.yaml` skips registration, an absent `custom-recipe-pack.bicep` keeps the default pack. The recipe pack is deployed from `.radius/` for the same `bicepconfig.json` resolution reason as `radius-env.bicep`. + +### Delete workflows + +Deleting an application or an environment reuses the deploy composition, minus the create stages. `delete-application.yml` and `delete-environment.yml` are thin dispatchers that mirror `run-rad-commands.yml`: a `detect` job binds the GitHub Environment, picks the provider from `AZURE_CLIENT_ID` / `AWS_ROLE_ARN`, and calls a reusable provider workflow (`delete-azure.yml` / `delete-aws.yml`) with `resource_type` (`application` or `environment`) and the target `name`. One provider workflow pair, parameterized by `resource_type`, keeps provider setup defined once per provider rather than once per resource type. + +The provider delete workflows run the same provider setup as deploy — OIDC login, cluster connection, cloud OIDC token projection, `rad credential register` — because `rad app delete` runs the resources' recipe delete path (e.g. `terraform destroy`) against the target cluster and cloud. They skip the deploy-only stages (create environment, deploy recipe pack, register registry credentials): the environment and its recipes come back from restored state, and deleting builds no images. The order is therefore restore state → delete → persist state. Restoring first (`rad startup`) is what makes the delete see the environment, recipe packs, resources, and Terraform state; persisting after (`rad shutdown`, run with `if: always()` in `teardown`) is what satisfies the requirement that the post-delete state is stored again, so the next operation plans against it. Both use the state archive — OCI-backed by default (`RADIUS_STATE_*` + GHCR login), or the `radius-state` git orphan branch when `RADIUS_STATE_BACKEND=git`. + +The shared `delete-resource` composite action runs `rad app delete --yes --preview` or `rad env delete --yes --preview`. `--preview` is required: without it these commands fall through to the legacy implementation instead of the Radius.Core surface the deploy flow provisions. It writes a `rad-delete-result` artifact (JSON with `outcome`, `exitCode`, `resourceType`, `name`, `output`) so a frontend can report the result the same way it reads the deploy result. Deleting an application also deletes that application's resources; deleting an environment removes the environment and its recipe-pack associations. + + ### Control plane startup (Investment 5) Because the control plane is created and torn down on every operation, startup time is on the critical path for every user-facing action and is the primary determinant of perceived responsiveness. Two backend decisions follow from that, both owned by this technical design rather than the feature spec: