agent-team Plane-2: bind P1+P2 to real models, live transport, coordinator #12

Merged
amoussa1229 merged 6 commits from feature/agent-team-plane2-p1-p2 into main 2026-06-18 17:29:03 +00:00
amoussa1229 commented 2026-06-18 16:57:31 +00:00 (Migrated from github.com)

What

Takes the Plane-2 scaffold from deterministic stubs to real model bindings + a live runtime, through P1 (human gate) and P2 (planner + adversarial review loop) of the depth-first track. Pre-deployment: nothing is provisioned or wired to live Slack/CI.

INTAKE -> CLARIFY (Claude, human gate) -> PLAN (Claude) -> REVIEW (GPT-4.1)
                                              ^                 |
                                              +-- loop-back ----+
                                                  approve/escalate -> END

Commits (logical)

  1. feat: P1 clarifier -> real Claude (subscription-OAuth invoker + Claude clarifier callables, fail-safe gate)
  2. feat: planner/review/builder/verifier node bindings (GPT-4.1 review, DeepSeek builders inert, ci_gate sole PASS authority) + review_loop hardening
  3. feat: live Slack transport + Socket Mode listener with owner allowlist
  4. feat: P1/P2 graph wiring + coordinator daemon + run-team start/serve (+ guarded re-delivery CAS)
  5. docs: R720 deploy runbook + coordinator systemd unit

Security review (/sh-security-review)

4 fresh-context detectors + proof-or-kill verifier. Gate: PASS after fixes.

  • AUTHZ-01 (HIGH, fixed): Slack listener had no sender-identity check; any channel member could steer the pipeline. Fixed with a fail-closed owner allowlist (AGENT_TEAM_SLACK_OWNER_IDS) + the open-status CAS as anti-replay.
  • RACE-REDELIVER (MEDIUM, fixed): unconditional DELETE in re-delivery could clobber a concurrently-accepted answer. Fixed with a guarded status='open' AND channel_ref IS NULL CAS under BEGIN IMMEDIATE + rowcount-skip.
  • injection / secrets-crypto detectors: clean. Subprocess calls to run.py confirmed list-argv (no shell).

Residual (non-blocking, follow-ups)

  • AUTHZ-02 (med): question_id is derivable; mitigated by the allowlist. Optional: mint random ids as defense-in-depth.
  • RACE-SUBMIT (low, unverified): optional hardening of snapshot_interrupt_turns to distinguish an unparseable turn.
  • LOGIC-VERIFIER-RUNID (low, unverified): the CI-as-verifier trust boundary is the P3 GPT-4.1 cross-review gate item (verifier/CI path is not wired live in this PR).

Tests

733 passing, ruff clean. New coverage spans the invoker, clarifier/planner/review/builder/verifier bindings, live transport + listener authz, the P1/P2 graph routing (approve-terminate, loop-to-cap-then-escalate), the coordinator, and the two security regressions.

Not in this PR (gated / later)

  • P3 builders+verifier live: HARD-GATED behind /sh-security-review + the mandatory GPT-4.1 cross-review of the CI/OIDC trust boundary.
  • Provisioning (VM snapshot, rsync, tokens, systemd enable) + the live 4-criteria demo.
  • Plane-1 checkers + Phase 0 substrate.
## What Takes the Plane-2 scaffold from deterministic stubs to **real model bindings + a live runtime**, through **P1 (human gate) and P2 (planner + adversarial review loop)** of the depth-first track. Pre-deployment: nothing is provisioned or wired to live Slack/CI. ``` INTAKE -> CLARIFY (Claude, human gate) -> PLAN (Claude) -> REVIEW (GPT-4.1) ^ | +-- loop-back ----+ approve/escalate -> END ``` ## Commits (logical) 1. `feat`: P1 clarifier -> real Claude (subscription-OAuth invoker + Claude clarifier callables, fail-safe gate) 2. `feat`: planner/review/builder/verifier node bindings (GPT-4.1 review, DeepSeek builders inert, ci_gate sole PASS authority) + review_loop hardening 3. `feat`: live Slack transport + Socket Mode listener **with owner allowlist** 4. `feat`: P1/P2 graph wiring + coordinator daemon + run-team start/serve (+ guarded re-delivery CAS) 5. `docs`: R720 deploy runbook + coordinator systemd unit ## Security review (`/sh-security-review`) 4 fresh-context detectors + proof-or-kill verifier. **Gate: PASS** after fixes. - **AUTHZ-01 (HIGH, fixed):** Slack listener had no sender-identity check; any channel member could steer the pipeline. Fixed with a fail-closed owner allowlist (`AGENT_TEAM_SLACK_OWNER_IDS`) + the open-status CAS as anti-replay. - **RACE-REDELIVER (MEDIUM, fixed):** unconditional DELETE in re-delivery could clobber a concurrently-accepted answer. Fixed with a guarded `status='open' AND channel_ref IS NULL` CAS under `BEGIN IMMEDIATE` + rowcount-skip. - injection / secrets-crypto detectors: clean. Subprocess calls to `run.py` confirmed list-argv (no shell). ### Residual (non-blocking, follow-ups) - AUTHZ-02 (med): `question_id` is derivable; mitigated by the allowlist. Optional: mint random ids as defense-in-depth. - RACE-SUBMIT (low, unverified): optional hardening of `snapshot_interrupt_turns` to distinguish an unparseable turn. - LOGIC-VERIFIER-RUNID (low, unverified): the CI-as-verifier trust boundary is **the P3 GPT-4.1 cross-review gate item** (verifier/CI path is not wired live in this PR). ## Tests 733 passing, ruff clean. New coverage spans the invoker, clarifier/planner/review/builder/verifier bindings, live transport + listener authz, the P1/P2 graph routing (approve-terminate, loop-to-cap-then-escalate), the coordinator, and the two security regressions. ## Not in this PR (gated / later) - **P3 builders+verifier live**: HARD-GATED behind `/sh-security-review` + the mandatory GPT-4.1 cross-review of the CI/OIDC trust boundary. - Provisioning (VM snapshot, rsync, tokens, systemd enable) + the live 4-criteria demo. - Plane-1 checkers + Phase 0 substrate.
This repo is archived. You cannot comment on pull requests.
No description provided.