shoc-pr-review-runner/docs/phase-2-todo.md
Adam Moussa c3cd8f7765
feat: SHOC PR review runner, phase 1
Manually-dispatched GitHub Actions workflow that reviews SHOC pull requests in
a clean environment: exact-head checkout of shoc-frontend-new and shoc-backend,
clean build/test gates, a truthful evidence report, a single-shot Fireworks
review, deterministic output validation, and published artifacts. The runner
never writes to the product repositories or their pull requests.

The review checklists move here from the reviewers' local Cursor commands so
the instructions live outside both product repos.

Phase 1 does not provision a database, start either application, or run live
browser flows; the evidence report records those as NOT_RUN so a review cannot
claim them.

Security architecture: building a PR executes its author's code, so the
workflow is split. The gates job runs that code holding no Fireworks key and
revokes its App token first; the review job holds the key, executes no product
code, and re-checks out this repo fresh. Product checkouts live outside the
workspace, the App token is downscoped at mint time, gate results fail closed
on any duplicate key, changed files are read from git objects rather than the
filesystem, and the validator re-checks every claim against the gate table.
2026-07-29 12:05:38 -04:00

8.7 KiB

Phase 2 TODO — Runtime Environment (disposable DB, startup, API checks)

Scope (spec Phase 2): disposable database, migration validation, backend startup + health gate, frontend startup, shared environment variables, API runtime checks, process/log management. Everything below turns an existing NOT_RUN line in review-evidence.md into a real gate recorded through scripts/lib.sh (record_gate/run_gate). Facts about the product repos were verified 2026-07-29; trust them over the original spec draft (which wrongly said PostgreSQL — the backend is EF Core 8.0.8 + SQL Server).

Ordered work plan

1. Workflow: SQL Server service container

  • Add a services: mssql block to the review job in .github/workflows/review-pr.yml: mcr.microsoft.com/mssql/server:2022-latest, port 1433:1433, ACCEPT_EULA=Y, MSSQL_SA_PASSWORD (see open questions), container health-cmd so the job waits for readiness. No compose file exists in shoc-backend; the service container is the whole database story.
  • Install sqlcmd on the runner (mssql-tools18 apt package) in a step gated on backend/paired, before provisioning.
  • Generate per-run runtime secrets in an early step: JWT_SECRET (openssl rand -hex 32 — must be ≥32 chars or login 500s) and the DB password if per-run. Add both to redact-check.sh's scan set.

2. New scripts (all source lib.sh, all gated on backend/paired unless noted)

  • scripts/proc.sh — shared process-management helpers sourced next to lib.sh: start_bg <name> <logfile> <cmd...> (nohup, PID to $ARTIFACTS_DIR/pids/<name>.pid, log to $LOG_DIR), stop_bg, stop_all_bg (kill + wait, idempotent, never fails the caller).
  • scripts/provision-database.sh — wait for SQL Server, CREATE DATABASE ShocReview via sqlcmd; gate backend.db_provision.
  • scripts/run-migration-gates.sh — migration validation (see §3); gates backend.migration_list, backend.migration_script, backend.migration_apply.
  • scripts/seed-admin-user.sh — direct-SQL identity seed (see §4); gate backend.seed_admin.
  • scripts/start-backend.sh — start the API via proc.sh, poll health, run the DB-touching assertion (see §5); gates backend.startup, backend.health.
  • scripts/run-api-checks.sh — authenticated runtime scenarios (see §6); gate backend.api_runtime (plus per-scenario detail in the gate table).
  • scripts/start-frontend.sh — frontend runtime startup (see §7); gates frontend.dev_startup, frontend.preview_build, frontend.preview_startup. Gated on frontend/paired; runtime API checks require paired (else NOT_RUN with reason "no live backend in frontend-only review").
  • scripts/stop-runtime.sh — calls stop_all_bg; wired as an if: always() workflow step so backend/frontend processes are stopped even after failed stages (spec requirement), before evidence generation and artifact upload.

3. Migration validation detail

  • 43 migrations live in Data.SeaHavenIndustries/Migrations/; there is NO migrate-on-startup and NO IDesignTimeDbContextFactory.
  • dotnet tool restore must run from Api.SeaHavenIndustries/ (the tool manifest — dotnet-ef 8.0.8, rollForward false — is at Api.SeaHavenIndustries/.config/dotnet-tools.json).
  • Every dotnet ef call needs --project Data.SeaHavenIndustries/... --startup-project Api.SeaHavenIndustries/....
  • Gates: migrations list (enumerates cleanly), migrations script --idempotent -o $ARTIFACTS_DIR/migrations.sql (SQL generation as artifact), then apply. Prefer database update; the repo's own deploy uses ef migrations bundle --self-contained -r linux-x64 (see its scripts/package-elastic-beanstalk.sh) — bundle is the fidelity option if database update misbehaves.

4. Seeding strategy (first-admin bootstrap gap)

  • No register endpoint; UserController.AddUser is [Authorize]; API user creation emails passwords via SendGrid; the Program.cs role-seeding block is fully commented out (Program.cs:199-253). So: seed by direct SQL.
  • seed-admin-user.sh runs a templates/seed-admin.sql via sqlcmd inserting AspNetRoles ("Admin" — the only role enforced in [Authorize] attributes), AspNetUsers (fixed reviewer account, precomputed ASP.NET Identity v3 PBKDF2 password hash), AspNetUserRoles.
  • The known password + hash pair is committed (disposable localhost-only DB; document as non-sensitive). Run after migration apply, before startup.

5. Backend startup + health

  • Start dotnet run --project Api.SeaHavenIndustries/... --no-build -c Release (or run the built DLL) via proc.sh with the shared env (§8).
  • NO anonymous /health endpoint exists (the only health-ish route is Admin-authorized). Health probe = GET /swagger/v1/swagger.json with retry budget; requires ASPNETCORE_ENVIRONMENT=Development (anything else disables Swagger AND enables HTTPS redirect).
  • Swagger 200 ≠ DB configured: the committed appsettings placeholder "${CONNECTION_STRING}" lets the app start and fail per-request. So the health gate also asserts DB wiring: POST /api/Authentication/login with bad creds must return 401 — a 500 means the connection string didn't take.
  • Expect log noise: 4 hosted services start at boot; the vendor-document scan worker polls the DB every 30s. Capture full stdout/stderr to logs/backend-runtime.log; evidence links it.

6. API runtime checks

  • With the seeded admin: login → 200 + JWT; call one Admin-authorized endpoint with the token → 200; call it unauthenticated → 401.
  • Record known limitation: ClamAV__Host is empty, so vendor uploads return 423 — assert-and-record as limitation, not failure.

7. Frontend runtime

  • Dev startup gate: npm run dev with VITE_API_TARGET=http://127.0.0.1:5141 (the dev proxy target env var), probe http://127.0.0.1:3000/ for 200.
  • Production preview: vite preview (4173) has NO /api proxy, so a preview against the local backend needs a second build with absolute VITE_API_URL=http://127.0.0.1:5141/api baked at BUILD time (must end in /api or the build-time contract guard throws; the Phase-1 gate build uses relative /api and cannot be reused). Then probe 4173.
  • Playwright's own dev server on 4173 is untouched (mocked suite unchanged).

8. Shared environment variables (one place: workflow env + runner-config)

  • ConnectionStrings__DefaultConnection = Server=127.0.0.1,1433;Database=ShocReview;User Id=sa;Password=...;TrustServerCertificate=True
  • JWT__Secret = per-run ≥32 chars (the committed "${JWT_SECRET}" placeholder passes the null check but breaks login — never rely on it)
  • ASPNETCORE_ENVIRONMENT=Development, ASPNETCORE_URLS=http://127.0.0.1:5141
  • WorkOrderIngest__Enabled=false, Sync__Enabled=false, WorkOrderReconciliation__Enabled=false, ClamAV__Host= (empty)

9. review/runner-config.yml additions

  • database: extend with db name, sa-password sourcing, sqlcmd tooling.
  • backend: add runtime_env block (§8 values), health retry budget, seed account name, migration project paths.
  • frontend: add preview_port: 4173, dev_api_target, preview build env.
  • Keep the workflow-env mirror rule (CI cross-checks the pair).

10. Evidence report (generate-evidence.sh)

  • Backend Gates: Migration list/script/apply, Startup, Health endpoint, API runtime scenarios switch from hardcoded NOT_RUN to $(be_gate ...); add DB provision + admin seed lines.
  • Frontend Gates: Development startup and Production preview switch to $(fe_gate ...).
  • Runtime Limitations rewritten: enumerate disabled integrations (ingest, sync, reconciliation, ClamAV → uploads 423), note JWT secret and DB are runner-provided, keep "mocked Playwright ≠ live coverage".
  • Tests: new fixtures in tests/fixtures/artifacts/ covering runtime-gate PASS/FAIL rows; extend CI bash tests for proc.sh start/stop semantics.

Out of scope (Phase 3+)

  • Live Playwright against the running stack, affected-route walking, console/ failed-request capture (Phase 3). Stacked/paired PR intelligence (Phase 4).
  • Fixing shoc-backend itself (health endpoint, seeding block) — record gaps.

Open questions

  1. MSSQL_SA_PASSWORD: fixed throwaway (service env can't consume step outputs) vs repo secret? Leaning fixed + documented non-sensitive.
  2. sqlcmd via apt mssql-tools18 vs docker exec into the service container?
  3. Migration apply: dotnet ef database update vs the repo's own bundle path?
  4. Is the dev-server startup gate worth its runtime once preview startup exists, or is preview + dev-proxy config check enough?
  5. Which Admin endpoint is the canonical authenticated smoke check (stable, read-only, no side effects/emails)?
  6. Should backend.health failing hard-block the frontend preview gates in paired runs (BLOCKED) or let them probe independently?