open-swe/.github/scripts/roll-box.sh
Adam Moussa 3224d7abf4
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
Agent CI / Agent lint (push) Waiting to run
Agent CI / Agent format check (push) Waiting to run
Agent CI / Agent unit tests (push) Waiting to run
Agent CI / Playwright E2E (push) Waiting to run
ci: gate dev→prod promotion on green checks + add rollback safety net (#28)
T20 CD safety nets. Two gaps closed before the first real prod deploy:

1. Promotion gate. promote_dev_to_prod.yml previously fast-forwarded
   dev→main unconditionally. It now hard-gates on check-dev-green.sh:
   every check-run on the dev HEAD must be completed+passing AND the
   Agent CI suite (lint/format/unit/E2E) must be present+success, or the
   promotion blocks (fails safe on a missing/renamed check). The promote
   run excludes its OWN check-run by run-id (unforgeable), never by the
   mutable name "promote", so a colliding red check cannot hide. Fields
   are read with a 0x1F separator so an empty conclusion (every
   in_progress check) cannot shift columns. ci.yml now also runs on
   push:dev so dev HEAD actually carries that signal (a PR check alone
   can be admin-merged past).

2. Rollback + last-good. publish-and-deploy.sh advances
   releases/last-good/ only after a successful roll (deploy.sh gates on
   `systemctl is-active`), and makes releases/latest/ transactional —
   reverting to the prior release if the roll fails so a replaced box
   never self-deploys a broken release. New rollback.yml + rollback.sh
   re-point latest at last-good (or an explicit sha) and re-fire the
   deploy; prod is gated by the `prod` Environment approval, same as a
   deploy. The shared fire/wait/aggregate-gate logic is factored into
   roll-box.sh (used by both forward and backward rolls).

Least-privilege: drop the unused s3:DeleteObject from the app deploy
role — publish/rollback/deploy only Get+Put (S3-to-S3 copy), and the
rollback fallback now depends on immutable release history staying
intact. Lifecycle expiry (not CI) handles old-version cleanup.

Gate logic unit-tested (7 cases + jq round-trip). IAM change +
release-safety control cross-reviewed by GPT-4.1: APPROVE, no blocks.

Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi
2026-06-27 20:21:59 -04:00

59 lines
2.8 KiB
Bash
Executable file

#!/usr/bin/env bash
# Fire the env's SSM deploy document (tag-targeted) and wait for it to finish,
# gating on the AGGREGATE command status. Shared by publish-and-deploy.sh (forward
# roll) and rollback.sh (backward roll) so the fire/wait/gate logic lives in ONE
# place. Requires ENV + DEPLOY_DOC in the environment and aws creds already set.
#
# Tag-targeting (project=open-swe,env=<env>) is exactly what the app deploy role's
# tag-scoped ssm:SendCommand allows — no ec2:DescribeInstances, no instance id.
set -euo pipefail
: "${ENV:?}" "${DEPLOY_DOC:?}"
COMMENT="${ROLL_COMMENT:-roll ${ENV}}"
echo "==> fire ${DEPLOY_DOC} via SSM (tag-targeted: project=open-swe, env=${ENV})"
CMD_ID="$(aws ssm send-command \
--document-name "${DEPLOY_DOC}" \
--targets "Key=tag:project,Values=open-swe" "Key=tag:env,Values=${ENV}" \
--comment "${COMMENT}" \
--query 'Command.CommandId' --output text)"
echo "command: ${CMD_ID}"
echo "==> wait for the deploy to finish"
IID=""
for _ in $(seq 1 60); do
sleep 10
IID="$(aws ssm list-command-invocations --command-id "${CMD_ID}" \
--query 'CommandInvocations[0].InstanceId' --output text 2>/dev/null || echo None)"
[ -z "${IID}" ] || [ "${IID}" = "None" ] && continue
STATUS="$(aws ssm list-command-invocations --command-id "${CMD_ID}" \
--query 'CommandInvocations[0].Status' --output text 2>/dev/null || echo Pending)"
case "${STATUS}" in
Success | Failed | Cancelled | TimedOut) break ;;
esac
done
if [ -z "${IID}" ] || [ "${IID}" = "None" ]; then
echo "ERROR: no box picked up the deploy command (is a running open-swe ${ENV} box registered with SSM?)" >&2
exit 1
fi
echo "==> deploy.sh output from ${IID}:"
echo "----- stdout -----"
aws ssm get-command-invocation --command-id "${CMD_ID}" --instance-id "${IID}" \
--query 'StandardOutputContent' --output text || true
echo "----- stderr -----"
aws ssm get-command-invocation --command-id "${CMD_ID}" --instance-id "${IID}" \
--query 'StandardErrorContent' --output text || true
# Gate on the AGGREGATE command status (Success only if EVERY targeted invocation
# succeeded), not CommandInvocations[0] — during a userDataCausesReplacement window
# two instances can briefly share the project/env tags, and a partial failure on the
# other instance must not be reported as success.
TARGETS="$(aws ssm list-commands --command-id "${CMD_ID}" \
--query 'Commands[0].TargetCount' --output text 2>/dev/null || echo 1)"
[ "${TARGETS}" = "1" ] || echo "WARNING: deploy fanned out to ${TARGETS} instances (expected 1)"
AGG="$(aws ssm list-commands --command-id "${CMD_ID}" \
--query 'Commands[0].Status' --output text 2>/dev/null || echo Failed)"
echo "==> aggregate deploy status: ${AGG} (across ${TARGETS} target(s))"
[ "${AGG}" = "Success" ] || { echo "ERROR: deploy did not succeed (${AGG})" >&2; exit 1; }