Wire the box-side build->dispatch->verify run identity so the verifier gate can bind to the CI run the dispatcher triggered: - task_model: add run_id / ci_correlation_tag / dispatched_at to TaskRecord + PipelineState (+ dict round-trip). - dispatcher: RunLocator seam + DispatchResult; dispatch_apply_verify stamps a dispatched-at watermark, fires, then resolves the run via the workflow run-name (gh run list; the per-task_id concurrency group makes it unambiguous). Fails closed to run_id=None. - dispatch_invoker: persist run_id/dispatched_at/ci_correlation_tag into state. - workflow: additive run-name surfacing inputs.task_id as the correlation key (flagged for the C1 /sh-security-review + GPT-4.1 cross-review re-run). - docs: P3-PHASE0-DESIGN.md records the async-resume design decision. Part of Phase 0 (feat/agent-team-p3-box-integration). No behavior change on the default path: P3 wiring is still opt-in/inert.
4.8 KiB
P3 Phase 0 — box-side build→dispatch→verify integration (design decision)
Branch: feat/agent-team-p3-box-integration. This note records the load-bearing
design decision for Phase 0 so the security/cross-review gates and the parallel
WebUI session have the rationale in-tree. It is the result of the /sh-plan-review
loop (3 rounds, GPT-4.1) + a code-level investigation of the as-built dispatcher,
ci_fetcher, and coordinator execution model.
Context (verified on-disk, 2026-06-23)
The CI apply/verify workflow is already live + provisioned (the agent-apply
environment, the AGENT_APPLY_APP_* secrets, and dispatched runs all exist as of
2026-06-22). What is not built is the box-side integration that makes a task
flow through BUILD → (trigger CI) → VERIFY automatically:
dispatcher.pyfiresgh workflow runbut never captures the resulting run id.dispatch_invokerreturns{}— nothing writesstate["run_id"].ci_fetcherreadsstate["run_id"](so it always fails closed → gate BLOCKs).- The graph orders BUILD → VERIFY → DISPATCH, but VERIFY needs a CI conclusion that only exists after DISPATCH triggers CI. (semantic inversion)
Decision 1 — node order: BUILD → DISPATCH → VERIFY (reorder)
DISPATCH triggers the CI run and must run before VERIFY reads its conclusion. The P3 subgraph is reordered accordingly (graph.py + build_verify_subgraph.py). The reorder adds no new graph nodes (BUILD/DISPATCH/VERIFY already exist), so the parallel WebUI branch's graph introspection + NODE_META coverage are unaffected; only edge wiring changes.
Decision 2 — CI wait: async resume-on-CI-complete, NOT a blocking poll
A CI run takes ~7 min. The coordinator is a single durable daemon (LangGraph
interrupt()/resume + a tick() maintenance sweep). A multi-minute blocking
VERIFY node would stall the tick loop and every other task. The durable
interrupt/resume machinery already exists for exactly the "external event resumes
a suspended task" shape (the Slack responder; the deadline timer). So:
BUILD → DISPATCH (push branch, trigger CI, capture run_id, suspend)
→ [CI-watcher resumes on terminal conclusion] → VERIFY (read result, gate)
DISPATCH captures run_id + dispatched_at into state and the task suspends. A
new CI-watcher (a tick()-driven sweep, mirroring deadline_timer) polls the
in-flight run_ids read-only and resumes each task once its run reaches a
terminal conclusion (or its dispatched_at + timeout elapses → park). VERIFY then
reads the authenticated conclusion via the existing read-only fetcher and the
pure-code gate decides pass/fail. The LLM remains a fix-proposer only.
Decision 3 — run_id capture is poll-based + correlation-tagged (anti-race)
gh workflow run does not return a run id. The dispatcher polls
gh run list --workflow … --json databaseId,headBranch,createdAt,event filtered
to this task's unique per-dispatch head branch + a per-dispatch
correlation tag (a nonce carried as a workflow input and echoed in the run),
bounded to runs created after the dispatch timestamp. This unambiguously matches
the dispatched run even with multiple tasks or rapid re-dispatch. A None/unfound
run_id fails closed (the gate BLOCKs / the task parks) — never a vacuous pass.
Decision 4 — per-task expected_run_id
VerifierConfig.expected_run_id was a static wiring-time constant. It is now
resolved per-task from state["run_id"] (the id the dispatcher captured), so the
gate binds each task's verdict to its own dispatched run and rejects a substituted
run id.
Decision 5 — fail-safe serve default
Per Adam's decision the bound P3 wiring becomes the new serve default. Because
the wiring factories are called eagerly at graph-build, the binding is wrapped so
a missing AGENT_TEAM_REPO_OWNER/_NAME / CI-read token degrades to the INERT P3
path (task parks, one WARNING + a #agent-team inert-mode notice) — never a
RuntimeError at serve-start that would crash-loop the daemon.
Out of scope (tracked follow-ups)
- The §4.3 box-native diff transport (signed artifact / branch-only token) that
would let the always-on box trigger CI without operator credentials. Until then,
the branch push +
gh workflow runuse operator-host credentials (the box holds no standing write token). - Tier-3 fixer off
--dry-run; the cross-plane checker→draft-PR loop.
Coordination with the WebUI branch (feature/agent-team-webui-makeover)
Shared files: graph.py (they wrap nodes via _instrument; we reorder edges),
coordinator.py (both edit the build_graph(...) call block). WebUI merges
first; this branch rebases onto the new main before deploy. Wrapper composition:
instrument(failsafe(node)) so a fail-safe park is still logged. The
project_r720_agent_team memory fix is owned by the WebUI session.