shoc-pr-review-runner/docs/phase-2-todo.md
Adam Moussa c3cd8f7765
feat: SHOC PR review runner, phase 1
Manually-dispatched GitHub Actions workflow that reviews SHOC pull requests in
a clean environment: exact-head checkout of shoc-frontend-new and shoc-backend,
clean build/test gates, a truthful evidence report, a single-shot Fireworks
review, deterministic output validation, and published artifacts. The runner
never writes to the product repositories or their pull requests.

The review checklists move here from the reviewers' local Cursor commands so
the instructions live outside both product repos.

Phase 1 does not provision a database, start either application, or run live
browser flows; the evidence report records those as NOT_RUN so a review cannot
claim them.

Security architecture: building a PR executes its author's code, so the
workflow is split. The gates job runs that code holding no Fireworks key and
revokes its App token first; the review job holds the key, executes no product
code, and re-checks out this repo fresh. Product checkouts live outside the
workspace, the App token is downscoped at mint time, gate results fail closed
on any duplicate key, changed files are read from git objects rather than the
filesystem, and the validator re-checks every claim against the gate table.
2026-07-29 12:05:38 -04:00

147 lines
8.7 KiB
Markdown

# Phase 2 TODO — Runtime Environment (disposable DB, startup, API checks)
Scope (spec Phase 2): disposable database, migration validation, backend
startup + health gate, frontend startup, shared environment variables, API
runtime checks, process/log management. Everything below turns an existing
`NOT_RUN` line in `review-evidence.md` into a real gate recorded through
`scripts/lib.sh` (`record_gate`/`run_gate`). Facts about the product repos were
verified 2026-07-29; trust them over the original spec draft (which wrongly
said PostgreSQL — the backend is EF Core 8.0.8 + SQL Server).
## Ordered work plan
### 1. Workflow: SQL Server service container
- Add a `services: mssql` block to the `review` job in
`.github/workflows/review-pr.yml`: `mcr.microsoft.com/mssql/server:2022-latest`,
port `1433:1433`, `ACCEPT_EULA=Y`, `MSSQL_SA_PASSWORD` (see open questions),
container health-cmd so the job waits for readiness. No compose file exists in
shoc-backend; the service container is the whole database story.
- Install `sqlcmd` on the runner (`mssql-tools18` apt package) in a step gated
on backend/paired, before provisioning.
- Generate per-run runtime secrets in an early step: `JWT_SECRET`
(`openssl rand -hex 32` — must be ≥32 chars or login 500s) and the DB
password if per-run. Add both to `redact-check.sh`'s scan set.
### 2. New scripts (all source `lib.sh`, all gated on backend/paired unless noted)
- `scripts/proc.sh` — shared process-management helpers sourced next to
`lib.sh`: `start_bg <name> <logfile> <cmd...>` (nohup, PID to
`$ARTIFACTS_DIR/pids/<name>.pid`, log to `$LOG_DIR`), `stop_bg`,
`stop_all_bg` (kill + wait, idempotent, never fails the caller).
- `scripts/provision-database.sh` — wait for SQL Server, `CREATE DATABASE
ShocReview` via sqlcmd; gate `backend.db_provision`.
- `scripts/run-migration-gates.sh` — migration validation (see §3); gates
`backend.migration_list`, `backend.migration_script`, `backend.migration_apply`.
- `scripts/seed-admin-user.sh` — direct-SQL identity seed (see §4); gate
`backend.seed_admin`.
- `scripts/start-backend.sh` — start the API via `proc.sh`, poll health, run
the DB-touching assertion (see §5); gates `backend.startup`, `backend.health`.
- `scripts/run-api-checks.sh` — authenticated runtime scenarios (see §6);
gate `backend.api_runtime` (plus per-scenario detail in the gate table).
- `scripts/start-frontend.sh` — frontend runtime startup (see §7); gates
`frontend.dev_startup`, `frontend.preview_build`, `frontend.preview_startup`.
Gated on frontend/paired; runtime API checks require paired (else NOT_RUN
with reason "no live backend in frontend-only review").
- `scripts/stop-runtime.sh` — calls `stop_all_bg`; wired as an `if: always()`
workflow step so backend/frontend processes are stopped even after failed
stages (spec requirement), before evidence generation and artifact upload.
### 3. Migration validation detail
- 43 migrations live in `Data.SeaHavenIndustries/Migrations/`; there is NO
migrate-on-startup and NO `IDesignTimeDbContextFactory`.
- `dotnet tool restore` must run from `Api.SeaHavenIndustries/` (the tool
manifest — dotnet-ef 8.0.8, rollForward false — is at
`Api.SeaHavenIndustries/.config/dotnet-tools.json`).
- Every `dotnet ef` call needs `--project Data.SeaHavenIndustries/... --startup-project
Api.SeaHavenIndustries/...`.
- Gates: `migrations list` (enumerates cleanly), `migrations script
--idempotent -o $ARTIFACTS_DIR/migrations.sql` (SQL generation as artifact),
then apply. Prefer `database update`; the repo's own deploy uses
`ef migrations bundle --self-contained -r linux-x64` (see its
`scripts/package-elastic-beanstalk.sh`) — bundle is the fidelity option if
`database update` misbehaves.
### 4. Seeding strategy (first-admin bootstrap gap)
- No register endpoint; `UserController.AddUser` is `[Authorize]`; API user
creation emails passwords via SendGrid; the Program.cs role-seeding block is
fully commented out (Program.cs:199-253). So: seed by direct SQL.
- `seed-admin-user.sh` runs a `templates/seed-admin.sql` via sqlcmd inserting
`AspNetRoles` ("Admin" — the only role enforced in `[Authorize]` attributes),
`AspNetUsers` (fixed reviewer account, precomputed ASP.NET Identity v3
PBKDF2 password hash), `AspNetUserRoles`.
- The known password + hash pair is committed (disposable localhost-only DB;
document as non-sensitive). Run after migration apply, before startup.
### 5. Backend startup + health
- Start `dotnet run --project Api.SeaHavenIndustries/... --no-build -c Release`
(or run the built DLL) via `proc.sh` with the shared env (§8).
- NO anonymous /health endpoint exists (the only health-ish route is
Admin-authorized). Health probe = `GET /swagger/v1/swagger.json` with retry
budget; requires `ASPNETCORE_ENVIRONMENT=Development` (anything else disables
Swagger AND enables HTTPS redirect).
- Swagger 200 ≠ DB configured: the committed appsettings placeholder
`"${CONNECTION_STRING}"` lets the app start and fail per-request. So the
health gate also asserts DB wiring: `POST /api/Authentication/login` with bad
creds must return 401 — a 500 means the connection string didn't take.
- Expect log noise: 4 hosted services start at boot; the vendor-document scan
worker polls the DB every 30s. Capture full stdout/stderr to
`logs/backend-runtime.log`; evidence links it.
### 6. API runtime checks
- With the seeded admin: login → 200 + JWT; call one Admin-authorized
endpoint with the token → 200; call it unauthenticated → 401.
- Record known limitation: ClamAV__Host is empty, so vendor uploads return
423 — assert-and-record as limitation, not failure.
### 7. Frontend runtime
- Dev startup gate: `npm run dev` with `VITE_API_TARGET=http://127.0.0.1:5141`
(the dev proxy target env var), probe `http://127.0.0.1:3000/` for 200.
- Production preview: `vite preview` (4173) has NO /api proxy, so a preview
against the local backend needs a second build with absolute
`VITE_API_URL=http://127.0.0.1:5141/api` baked at BUILD time (must end in
`/api` or the build-time contract guard throws; the Phase-1 gate build uses
relative `/api` and cannot be reused). Then probe 4173.
- Playwright's own dev server on 4173 is untouched (mocked suite unchanged).
### 8. Shared environment variables (one place: workflow env + runner-config)
- `ConnectionStrings__DefaultConnection` = `Server=127.0.0.1,1433;Database=ShocReview;User Id=sa;Password=...;TrustServerCertificate=True`
- `JWT__Secret` = per-run ≥32 chars (the committed `"${JWT_SECRET}"`
placeholder passes the null check but breaks login — never rely on it)
- `ASPNETCORE_ENVIRONMENT=Development`, `ASPNETCORE_URLS=http://127.0.0.1:5141`
- `WorkOrderIngest__Enabled=false`, `Sync__Enabled=false`,
`WorkOrderReconciliation__Enabled=false`, `ClamAV__Host=` (empty)
### 9. `review/runner-config.yml` additions
- `database:` extend with db name, sa-password sourcing, sqlcmd tooling.
- `backend:` add `runtime_env` block (§8 values), health retry budget,
seed account name, migration project paths.
- `frontend:` add `preview_port: 4173`, `dev_api_target`, preview build env.
- Keep the workflow-env mirror rule (CI cross-checks the pair).
### 10. Evidence report (`generate-evidence.sh`)
- Backend Gates: Migration list/script/apply, Startup, Health endpoint, API
runtime scenarios switch from hardcoded NOT_RUN to `$(be_gate ...)`; add DB
provision + admin seed lines.
- Frontend Gates: Development startup and Production preview switch to
`$(fe_gate ...)`.
- Runtime Limitations rewritten: enumerate disabled integrations (ingest,
sync, reconciliation, ClamAV → uploads 423), note JWT secret and DB are
runner-provided, keep "mocked Playwright ≠ live coverage".
- Tests: new fixtures in `tests/fixtures/artifacts/` covering runtime-gate
PASS/FAIL rows; extend CI bash tests for `proc.sh` start/stop semantics.
## Out of scope (Phase 3+)
- Live Playwright against the running stack, affected-route walking, console/
failed-request capture (Phase 3). Stacked/paired PR intelligence (Phase 4).
- Fixing shoc-backend itself (health endpoint, seeding block) — record gaps.
## Open questions
1. `MSSQL_SA_PASSWORD`: fixed throwaway (service env can't consume step
outputs) vs repo secret? Leaning fixed + documented non-sensitive.
2. sqlcmd via apt `mssql-tools18` vs `docker exec` into the service container?
3. Migration apply: `dotnet ef database update` vs the repo's own bundle path?
4. Is the dev-server startup gate worth its runtime once preview startup
exists, or is preview + dev-proxy config check enough?
5. Which Admin endpoint is the canonical authenticated smoke check (stable,
read-only, no side effects/emails)?
6. Should `backend.health` failing hard-block the frontend preview gates in
paired runs (BLOCKED) or let them probe independently?