Address security-review F1/F3: the workflow re-check used a POSIX
[[:cntrl:]] grep (missed the C1 range the box-side sanitizer strips) and
wc -m (byte count, not chars). Replace both with a single python3 check
whose accept condition is byte-for-byte identical to sanitize_pr_title —
rejects [\x00-\x1f\x7f-\x9f] and caps at 70 characters — so the
defense-in-depth re-validation genuinely matches the box path.
Address cross-review FIX: the title re-validation used GNU-only
`grep -P`. Switch to POSIX `[[:cntrl:]]` under LC_ALL=C so the check
is portable; C1 chars are already stripped box-side by sanitize_pr_title.
The apply pipeline's draft PRs were titled `agent-apply: <task_id>
(diff <hash>)` with a flat body — neither within Sea Haven PR
conventions. That format was deliberate injection-hardening (only
sanitized tokens, never model free-text; §4.6).
Thread the approved plan's title through dispatch as an optional
`pr_title` input, sanitized box-side (single line, no control chars,
70-char cap, capitalized) and RE-VALIDATED in the workflow as defense-
in-depth, with the hardened `agent-apply: <task_id>` title as the
fallback when empty/unsafe. Provenance (task id, diff hash, head) moves
to the PR body. gh consumes both as argv data, never shell-interpolated.
Adds .github/workflows/ci-web.yaml, a thin caller over the org
Sea-Haven-Industries/.github ci-typescript-frontend.yaml reusable workflow,
scoped to agent-team/web changes. Runs build (tsc -b + vite build = typecheck)
and vitest on PRs to main. Distinct job id ci-web (status 'ci-web / ci') so it
does not collide with the Python 'ci / ci' context. format:check/lint/test:e2e
steps toggled off until Prettier/ESLint/Playwright are wired.
Verified locally: npm ci + npm run build + npm test (36/36) all green.
Wire the box-side build->dispatch->verify run identity so the verifier gate
can bind to the CI run the dispatcher triggered:
- task_model: add run_id / ci_correlation_tag / dispatched_at to TaskRecord +
PipelineState (+ dict round-trip).
- dispatcher: RunLocator seam + DispatchResult; dispatch_apply_verify stamps a
dispatched-at watermark, fires, then resolves the run via the workflow
run-name (gh run list; the per-task_id concurrency group makes it
unambiguous). Fails closed to run_id=None.
- dispatch_invoker: persist run_id/dispatched_at/ci_correlation_tag into state.
- workflow: additive run-name surfacing inputs.task_id as the correlation key
(flagged for the C1 /sh-security-review + GPT-4.1 cross-review re-run).
- docs: P3-PHASE0-DESIGN.md records the async-resume design decision.
Part of Phase 0 (feat/agent-team-p3-box-integration). No behavior change on the
default path: P3 wiring is still opt-in/inert.
Reworked per GPT-4.1 cross-family review BLOCK. The original PR removed the
agent-apply GitHub Environment (the live required-reviewer human gate) and
replaced it with a Slack notice that FAILS OPEN when its webhook secret is
absent (which it is) plus an audit-log line. The cross-review correctly
flagged this as trading a preventive control for detective controls, one of
which silently no-ops.
This commit:
- Restores environment: agent-apply on gate-and-pr (the human approval pause).
- Drops the fail-open Slack notify step.
- Keeps the unconditional audit-log step as an additive detective control.
- Restores the MANDATORY-INVARIANT assertion (env must be present) and adds
an assertion that the audit step is retained.
WS3's auto-dispatch (dispatch_invoker.py + graph/coordinator wiring) is
unchanged: it fires workflow_dispatch, which now pauses at the restored gate
for human approval — auto-dispatch up to the approval, then one click.
Live smoke exposed a self-pollution bug (pre-existing from PR #17): the post-build
denied-path check wrote its own _build_diff_z.bin / _build_status_z.bin into the
working tree, then its own `git status --untracked-files=all` flagged them as
out-of-scope writes — failing any run with a narrow declared_scope (the smoke's
docs/**). Write them to $RUNNER_TEMP instead (read via $_DIFF_Z/$_STATUS_Z), so
the check no longer sees its own temp files. The agent-team pytest artifacts were
already correctly gitignored; only the check's own files tripped it. Test harness
updated to pass the env paths. 1044 tests, ruff clean.
GitHub Actions only runs workflows under .github/workflows/, so the apply/verify
workflow at agent-team/ci/ was never registered (workflow_dispatch 404'd). Move it
to .github/workflows/agent-team-apply-verify.yml so it is a real, dispatchable
workflow. Its only trigger is workflow_dispatch + it is gated by the agent-apply
required-reviewer environment, so it never auto-runs and nothing privileged runs
unapproved. Updated the two workflow test files' path refs (parents[2]/.github/
workflows) and the ci/README pointer. Dispatcher push gains --no-verify: the apply
path is scanned CI-side (guard + the PR's checks), so it must not be blocked by the
operator's LOCAL human-commit pre-push dev hook (which flags pre-existing whole-repo
FPs like .env.example). 1044 tests, ruff clean.
Repo-root 'pytest --collect-only' failed with ImportPathMismatchError because
agent-team/tests/ and the root tests/ are both the 'tests' package. Add a root
conftest.py that excludes agent-team from the root collection, and a dedicated
agent-team-tests CI job that runs the (API-key-free) agent-team suite in its own
working dir.
Closes the no-CI gap surfaced in the post-merge retrospective. The org
reusable workflows live under Sea-Haven-Industries and assume SAM/CDK
projects — orchestrator is a personal CLI tool with neither, so this is
a standalone workflow on the same conventions (actions/checkout@v6,
Python 3.12, ruff).
- lint job: ruff check + ruff format --check.
- test-collect job: installs requirements + runs `pytest --collect-only`.
Catches import errors and golden-set test discovery regressions without
needing live ANTHROPIC/COMPOSIO secrets — full pytest stays a local
pre-push responsibility.
Also adds .DS_Store to the local .gitignore (also covered by the user
global gitignore, but belt-and-suspenders).