mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 17:23:15 +00:00
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
Agent CI / Agent lint (push) Waiting to run
Agent CI / Agent format check (push) Waiting to run
Agent CI / Agent unit tests (push) Waiting to run
Agent CI / Playwright E2E (push) Waiting to run
T20 CD safety nets. Two gaps closed before the first real prod deploy: 1. Promotion gate. promote_dev_to_prod.yml previously fast-forwarded dev→main unconditionally. It now hard-gates on check-dev-green.sh: every check-run on the dev HEAD must be completed+passing AND the Agent CI suite (lint/format/unit/E2E) must be present+success, or the promotion blocks (fails safe on a missing/renamed check). The promote run excludes its OWN check-run by run-id (unforgeable), never by the mutable name "promote", so a colliding red check cannot hide. Fields are read with a 0x1F separator so an empty conclusion (every in_progress check) cannot shift columns. ci.yml now also runs on push:dev so dev HEAD actually carries that signal (a PR check alone can be admin-merged past). 2. Rollback + last-good. publish-and-deploy.sh advances releases/last-good/ only after a successful roll (deploy.sh gates on `systemctl is-active`), and makes releases/latest/ transactional — reverting to the prior release if the roll fails so a replaced box never self-deploys a broken release. New rollback.yml + rollback.sh re-point latest at last-good (or an explicit sha) and re-fire the deploy; prod is gated by the `prod` Environment approval, same as a deploy. The shared fire/wait/aggregate-gate logic is factored into roll-box.sh (used by both forward and backward rolls). Least-privilege: drop the unused s3:DeleteObject from the app deploy role — publish/rollback/deploy only Get+Put (S3-to-S3 copy), and the rollback fallback now depends on immutable release history staying intact. Lifecycle expiry (not CI) handles old-version cleanup. Gate logic unit-tested (7 cases + jq round-trip). IAM change + release-safety control cross-reviewed by GPT-4.1: APPROVE, no blocks. Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi
64 lines
3 KiB
Bash
Executable file
64 lines
3 KiB
Bash
Executable file
#!/usr/bin/env bash
|
|
# Publish the packaged artifacts to the env's S3 bucket and roll the box to them.
|
|
# Run by build-artifacts.yml AFTER aws creds are configured (env: ENV, BUCKET,
|
|
# DEPLOY_DOC). Each release is stored immutably under releases/<sha>/.
|
|
#
|
|
# releases/latest/ (what the box's deploy.sh pulls) is advanced TRANSACTIONALLY:
|
|
# it is pointed at the new release, the box is rolled, and ONLY on a successful
|
|
# roll is it kept — a failed roll reverts releases/latest/ to the prior release so
|
|
# a later box boot / replacement never self-deploys a release that failed to come
|
|
# up. releases/last-good/ (rollback fallback) is advanced only after success and
|
|
# means "last release whose deploy.sh brought the service up active" (deploy.sh
|
|
# gates on `systemctl is-active`), not merely "last uploaded".
|
|
#
|
|
# The fire/wait/gate against the box lives in roll-box.sh (shared with rollback.sh);
|
|
# the deploy is fired by TAG (project=open-swe,env=<env>), exactly what the app
|
|
# deploy role's tag-scoped ssm:SendCommand allows.
|
|
set -euo pipefail
|
|
|
|
: "${ENV:?}" "${BUCKET:?}" "${DEPLOY_DOC:?}"
|
|
SHA="${GITHUB_SHA:?}"
|
|
[ -f app.tar.gz ] && [ -f spa.tar.gz ] || { echo "ERROR: artifacts not built" >&2; exit 1; }
|
|
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
|
|
# Re-point releases/latest/ at the release stored under releases/<sha>/, recording
|
|
# the sha so the pointer is self-describing (used to revert on failure).
|
|
point_latest() {
|
|
local s="$1" f
|
|
for f in app.tar.gz spa.tar.gz; do
|
|
aws s3 cp "s3://${BUCKET}/releases/${s}/${f}" "s3://${BUCKET}/releases/latest/${f}"
|
|
done
|
|
printf '%s\n' "${s}" | aws s3 cp - "s3://${BUCKET}/releases/latest/sha.txt"
|
|
}
|
|
|
|
# Which release does latest point at right now? (empty on the very first deploy.)
|
|
PREV_SHA="$(aws s3 cp "s3://${BUCKET}/releases/latest/sha.txt" - 2>/dev/null | tr -d '[:space:]' || true)"
|
|
|
|
echo "==> upload release ${SHA} to s3://${BUCKET}/releases/${SHA}/"
|
|
for f in app.tar.gz spa.tar.gz; do
|
|
aws s3 cp "${f}" "s3://${BUCKET}/releases/${SHA}/${f}"
|
|
done
|
|
|
|
echo "==> point releases/latest/ -> ${SHA} (was ${PREV_SHA:-<none>})"
|
|
point_latest "${SHA}"
|
|
|
|
# Roll the box to releases/latest/. On a non-Success aggregate, roll-box.sh exits
|
|
# non-zero; revert latest to the prior release so no later boot pulls the bad one.
|
|
if ! ROLL_COMMENT="release ${SHA}" bash "${HERE}/roll-box.sh"; then
|
|
if [ -n "${PREV_SHA}" ]; then
|
|
echo "!! deploy failed — reverting releases/latest/ -> ${PREV_SHA}" >&2
|
|
point_latest "${PREV_SHA}"
|
|
else
|
|
echo "!! deploy failed on the FIRST release — leaving releases/latest/ = ${SHA} (no prior release to revert to)" >&2
|
|
fi
|
|
exit 1
|
|
fi
|
|
|
|
# Roll succeeded (service came up active) -> this release is now the known-good one.
|
|
echo "==> mark releases/last-good/ = ${SHA} (rollback fallback target)"
|
|
for f in app.tar.gz spa.tar.gz; do
|
|
aws s3 cp "s3://${BUCKET}/releases/${SHA}/${f}" "s3://${BUCKET}/releases/last-good/${f}"
|
|
done
|
|
printf '%s\n' "${SHA}" | aws s3 cp - "s3://${BUCKET}/releases/last-good/sha.txt"
|
|
|
|
echo "==> ${ENV} rolled to release ${SHA} (now last-good)"
|