shoc-pr-review-runner/scripts/generate-evidence.sh
Adam Moussa c3cd8f7765
feat: SHOC PR review runner, phase 1
Manually-dispatched GitHub Actions workflow that reviews SHOC pull requests in
a clean environment: exact-head checkout of shoc-frontend-new and shoc-backend,
clean build/test gates, a truthful evidence report, a single-shot Fireworks
review, deterministic output validation, and published artifacts. The runner
never writes to the product repositories or their pull requests.

The review checklists move here from the reviewers' local Cursor commands so
the instructions live outside both product repos.

Phase 1 does not provision a database, start either application, or run live
browser flows; the evidence report records those as NOT_RUN so a review cannot
claim them.

Security architecture: building a PR executes its author's code, so the
workflow is split. The gates job runs that code holding no Fireworks key and
revokes its App token first; the review job holds the key, executes no product
code, and re-checks out this repo fresh. Product checkouts live outside the
workspace, the App token is downscoped at mint time, gate results fail closed
on any duplicate key, changed files are read from git objects rather than the
filesystem, and the validator re-checks every claim against the gate table.
2026-07-29 12:05:38 -04:00

138 lines
5.7 KiB
Bash
Executable file

#!/usr/bin/env bash
# Render /workspace/artifacts/review-evidence.md (spec §18) from the recorded
# gate statuses and PR metadata. Every check appears with an explicit status:
# PASS / FAIL / NOT_APPLICABLE / NOT_RUN / BLOCKED. Phase-2/3 checks the Phase-1
# runner cannot execute are stated NOT_RUN with the reason — never omitted,
# never converted into a pass.
#
# Reads: REVIEW_TYPE, TICKET, REVIEW_NOTES, RUN_ID/GITHUB_RUN_ID, artifacts from
# earlier steps ($ARTIFACTS_DIR/{frontend,backend}-pr.json, checkout.json,
# gate-status.tsv)
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=lib.sh
source "$SCRIPT_DIR/lib.sh"
require_env REVIEW_TYPE
EVIDENCE_FILE="$ARTIFACTS_DIR/review-evidence.md"
run_id="${GITHUB_RUN_ID:-local}"
# The evidence report is only as trustworthy as the gate table it renders.
assert_gate_table_intact
# Untrusted strings (PR titles, branch names, dispatcher-supplied ticket and
# notes) must not be able to forge lines inside the evidence report — the agent
# is told the evidence report is the only source of truth for what executed.
# Control characters and newlines are stripped and the length is capped, so a
# crafted value stays on the single line it was rendered into.
sanitize() {
printf '%s' "$1" | tr -d '\000-\037' | cut -c1-200
}
pr_field() { # pr_field <side> <jq-expr> [fallback]
local f="$ARTIFACTS_DIR/$1-pr.json"
if [ -f "$f" ]; then sanitize "$(jq -r "$2" "$f")"; else echo "${3:-not in scope}"; fi
}
g() { gate_status "$1"; }
# A side that has no PR under review has all its gates NOT_APPLICABLE.
side_in_scope() { # side_in_scope <frontend|backend>
case "$REVIEW_TYPE" in
paired) return 0 ;;
"$1") return 0 ;;
*) return 1 ;;
esac
}
fe_gate() { if side_in_scope frontend; then g "$1"; else echo "NOT_APPLICABLE (backend-only review)"; fi; }
be_gate() { if side_in_scope backend; then g "$1"; else echo "NOT_APPLICABLE (frontend-only review)"; fi; }
companion_note=""
if [ "$REVIEW_TYPE" != "paired" ] && [ -f "$ARTIFACTS_DIR/checkout.json" ]; then
cb="$(jq -r '.companion_branch' "$ARTIFACTS_DIR/checkout.json")"
companion_note=" (companion checked out at \`$cb\` head for contract context, not under review)"
fi
cat >"$EVIDENCE_FILE" <<EOF
# Review Evidence
## Review Request
- Review type: $REVIEW_TYPE
- Frontend PR: $(pr_field frontend '"#\(.pr)"') (title and branch names appear in the untrusted section, not here)
- Backend PR: $(pr_field backend '"#\(.pr)"')
- Ticket: $(sanitize "${TICKET:-not provided}")
- Reviewer notes: $(sanitize "${REVIEW_NOTES:-none}")
- Workflow run: $run_id
## Exact Heads
- Frontend SHA: $(pr_field frontend '.short_sha')$( side_in_scope frontend || printf '%s' "$companion_note")
- Backend SHA: $(pr_field backend '.short_sha')$( side_in_scope backend || printf '%s' "$companion_note")
- Frontend base: $(pr_field frontend '.base_ref')
- Backend base: $(pr_field backend '.base_ref')
## Governance Signals
- Frontend PR CI status: $(pr_field frontend '.ci_status')
- Backend PR CI status: $(pr_field backend '.ci_status')
- Frontend PR mergeable: $(pr_field frontend '.mergeable // "unknown"')
- Backend PR mergeable: $(pr_field backend '.mergeable // "unknown"')
- CI statuses come from the PR head's check-runs; red/pending/missing is a governance signal, not silently omitted.
## Stack Status
- Frontend parent: NOT_RUN (stacked-PR resolution is Phase 4)
- Backend parent: NOT_RUN (stacked-PR resolution is Phase 4)
- Base integrity: NOT_RUN (base comparison is Phase 4)
## Backend Gates
- Restore: $(be_gate backend.restore)
- Release build: $(be_gate backend.build)
- Tests: $(be_gate backend.test)
- Migration list: NOT_RUN (database provisioning is Phase 2)
- Migration script: NOT_RUN (database provisioning is Phase 2)
- Migration apply: NOT_RUN (database provisioning is Phase 2)
- Startup: NOT_RUN (runtime environment is Phase 2)
- Health endpoint: NOT_RUN (runtime environment is Phase 2)
- API runtime scenarios: NOT_RUN (runtime environment is Phase 2)
## Frontend Gates
- Clean install: $(fe_gate frontend.install)
- Lint: $(fe_gate frontend.lint)
- TypeScript + production build: $(fe_gate frontend.build) (tsc -b runs inside the build script)
- Unit/component tests: $(fe_gate frontend.unit_tests)
- Development startup: NOT_RUN (manual route exercise is Phase 2)
- Production preview: NOT_RUN (runtime environment is Phase 2)
## Browser Validation
- Mocked Playwright: $(fe_gate frontend.e2e_mocked)
- Live Playwright: NOT_RUN (live backend integration is Phase 3; MOCKED COVERAGE IS NOT LIVE COVERAGE)
- Affected routes: NOT_RUN (live browser validation is Phase 3)
- Console errors: NOT_RUN (live browser validation is Phase 3)
- Failed requests: NOT_RUN (live browser validation is Phase 3)
## Runtime Limitations
- The Phase 1 runner does not provision a database, start either application, or run live browser flows. Any conclusion about runtime behavior must come from code inspection and is not runtime-verified.
- Disabled integrations: all (no runtime environment in Phase 1)
- Mocked external systems: the Playwright suite mocks ALL backend API calls via page.route
- Unexecuted checks: listed NOT_RUN above with reasons
## Logs and Artifacts
$(if [ -d "$LOG_DIR" ] && [ -n "$(ls -A "$LOG_DIR" 2>/dev/null)" ]; then
for f in "$LOG_DIR"/*; do
printf -- '- logs/%s\n' "$(basename "$f")"
done
else
printf -- '- none\n'
fi)
## Gate Detail
$(if [ -f "$GATE_STATUS_FILE" ]; then
while IFS=$'\t' read -r key status det; do
# shellcheck disable=SC2016 # backticks are literal markdown
printf -- '- `%s`: %s%s\n' "$key" "$status" "${det:+ — $det}"
done <"$GATE_STATUS_FILE"
else
printf -- '- no gates recorded\n'
fi)
EOF
log "evidence written to $EVIDENCE_FILE"