Park notifications were opaque ('Task 6c3c3202 parked — needs your attention'):
no idea what the task is, where it got to, or what's blocking it. Now each
message names the task DESCRIPTION (not just the short id), the PHASE it reached,
and the actual BLOCKER — _summarize_blocker() pulls the last review_verdict's
findings (the GPT-4.1 REQUEST_CHANGES text, collapsed + truncated), falling back
to 'no plan built' / 'revision cap hit'. Applies to parked + needs-more-input +
plan-ready emits. +1 test (description + phase + blocker present). 1148 passed.
The bot only ever posted clarifier questions; parks/completions were silent and
multi-turn follow-up questions were never posted during normal operation (only
the startup recover sweep posted them). So an answered task was a black box.
- Coordinator gains a 'notify' sink + _emit() (guarded, never breaks the loop).
- tick() now runs _post_resume_followups(results) after the drain: for each
resumed thread it (a) POSTS a newly-pending clarifier question (fixes silent
multi-turn — the drain path left it unposted) and emits 'needs more input';
else emits 'parked — needs attention' or 'plan ready for review' from the
settled state.
- run-team serve wires notify -> Slack channel (build_slack_poster) and an
alarm_hook that logs the deadline-park WARNING AND posts a parked notice.
Live-Slack only; dry-run/non-Slack/no-channel = silent (None), no token needed.
Tests: +4 (needs-input/parked/plan-ready emits + notify-failure swallow);
_FakeCoordinator gains notify/alarm_hook. 1147 passed, ruff clean.
/new-task (and every intake: GitHub issue, /sh-assign-task) reached the clarifier
with NO description -> the clarifier asked 'no task description provided'. Root
cause: coordinator.start_task only LOGGED task_text (a P1-era decision when the
deterministic clarifier didn't consume a description), graph.start_task took no
task arg, and PipelineState/TaskRecord had no 'task' channel at all.
Fix: add a first-class 'task' field to PipelineState + TaskRecord (+ round-trip
in task_from_dict); graph.start_task seeds task into the initial invoke (persists
through intake_node's partial-state return into CLARIFY); coordinator.start_task
passes task=task_text. The clarifier already reads state['task'] via
_task_description, so it now sees the real description.
Test: start_task(task='build a login form') -> suspended CLARIFY state carries
task. 1143 passed, ruff clean.
'the app did not respond': serve() registered @app.action/@app.event but NO
@app.command handler, so Bolt never acked the /new-task slash command within
Slack's ~3s deadline. Add an @app.command(/new-task) handler that ack()s first,
re-stamps type:slash_commands onto Bolt's inner command body (Bolt strips the
Socket Mode envelope type that _is_new_task_command/_discriminating_type expect),
then forwards to handle_event (AUTHZ-01 + new-task dispatch). The resulting
payload shape is the one already covered by test_new_task_calls_callback_*.
The agent-team venv was missing langchain-anthropic/-openai/-google-genai/
-community, so models.py failed to import and the in-process GPT-4.1 review /
Gemini scan / DeepSeek build silently fail-closed to REQUEST_CHANGES (the
non-Claude models never ran on the box). Step 3 now installs the full pinned
requirements.txt into the venv instead of just fastapi/uvicorn. Installed +
verified live on the box: all three model factories construct.
Gap-audit findings:
- G1 (blocks-feature): make_cross_reviewer_invoker / make_fast_coder_invoker did
'from models import' without putting the orchestrator root on sys.path. The
run-team serve daemon only bootstraps agent-team/, so on the live box every
GPT-4.1 plan review hit ModuleNotFoundError -> review_plan's blanket except
silently fail-closed to REQUEST_CHANGES (GPT-4.1 never actually ran). Both
in-process invokers now call invoker_multi._ensure_orchestrator_on_path()
before the deferred import. WS1 introduced this when it swapped the review
default from the subprocess invoker to in-process.
- G4 (degrades): the handbook context_provider was wired into the planner only;
default_clarify_node_factory now accepts + forwards it, and run-team wires it
into build_clarify_node too, so clarifying questions are handbook-aware.
Tests: +2 regression tests (path-bootstrap, clarifier threading); _FakeCoordinator
gains build_clarify_node. 1142 passed, ruff clean.
The slack_listener handles {type:slash_commands, command:/new-task} but the app
manifest declared no slash commands and no 'commands' scope, so Slack never
offered /new-task (the command can't be invoked). Add the slash command +
commands bot scope. Socket Mode delivers it over the socket (no request URL).
APPLY: update the app A0BCC7TTU66 from this manifest + reinstall to pick up the
new scope.
Idempotent attended update of the live coordinator to WS0-WS5: rsync repo +
handbook, install fastapi/uvicorn, append AGENT_TEAM_API_TOKEN/SEA_HAVEN_HANDBOOK_DIR
to secrev.env if absent, restart the daemon, smoke tests. HTTP API is an opt-in
separate step; P3 dispatch stays inert. Snapshot-first + confirm before restart.
Integration branch combining WS0-WS5 (PRs #43-#46) + the activation wiring that
flips the safe seams ON in the run-team serve path:
- WS5 (D10): inject the Sea Haven handbook conventions into the planner prompt
via context_provider (zero-arg handbook loader; fail-safe to '' when absent).
- WS2: an allowlisted Slack /new-task starts a task on this coordinator
(set_new_task_callback adapter -> start_task; AUTHZ-01 gates it upstream).
- WS1 bind_multi_invoker() is already wired in _cmd_serve.
Coordinator gains a new_task_callback param + set_new_task_callback() (resolves
the constructor chicken-and-egg of referencing the coordinator's own start_task);
default_slack_listener_factory forwards it to the SlackListener.
NOT wired (deliberately): the P3 dispatch_node / build_verify path. Activating
it correctly needs a per-task expected_run_id bound into gated_build_verify_wiring
(plumbing that does not exist yet) AND the CI trust-boundary security re-review.
It stays inert pending that work.
Tests: +5 activation-wiring tests; _FakeCoordinator stub gains
set_new_task_callback. Full agent-team suite: 1140 passed, ruff clean.
Security follow-up from the per-PR review (non-blocking, defense-in-depth):
- Replace the save_memory name blocklist with an allowlist regex
(^[A-Za-z0-9][A-Za-z0-9._-]*$, max 128) so dot-only/hidden/backslash/NUL/
over-long names are rejected outright, not written as malformed-but-contained
files.
- Write via os.open(..., O_NOFOLLOW): the open fails (ELOOP) if the final
path component is a pre-planted symlink, closing the TOCTOU where a symlink
in _box-drafts/ could redirect the write outside the dir. O_CREAT|O_TRUNC
keeps overwrite-on-resave for regular files.
Tests: adds allowlist-rejection + symlink-refusal cases (21 pass).
Security follow-ups from the per-PR review (non-blocking MEDIUMs):
- Eager _get_token() at make_app build time so a missing AGENT_TEAM_API_TOKEN
fails fast instead of serving requests first (matches the docstring contract).
- Disable /docs, /redoc, /openapi.json (no auth dependency in FastAPI) — the
API is VPN-only/127.0.0.1 and should not expose its schema unauthenticated.
- Scrub raw exception text and subprocess stderr from 500 response bodies;
log server-side instead (avoid internal-path/state disclosure).
- Bound /orchestrator/invoke concurrency with a semaphore (429 over the cap)
so an authenticated caller cannot exhaust the box via many 600s subprocesses.
Also pin fastapi/uvicorn in requirements.txt (WS1 dep). With fastapi now
installed in CI, the previously skip-guarded TestClient tests run for real;
the importorskip guard stays as a no-op safety net.
Tests: 23 pass (adds docs-disabled + concurrency-429 cases).
Reworked per GPT-4.1 cross-family review BLOCK. The original PR removed the
agent-apply GitHub Environment (the live required-reviewer human gate) and
replaced it with a Slack notice that FAILS OPEN when its webhook secret is
absent (which it is) plus an audit-log line. The cross-review correctly
flagged this as trading a preventive control for detective controls, one of
which silently no-ops.
This commit:
- Restores environment: agent-apply on gate-and-pr (the human approval pause).
- Drops the fail-open Slack notify step.
- Keeps the unconditional audit-log step as an additive detective control.
- Restores the MANDATORY-INVARIANT assertion (env must be present) and adds
an assertion that the audit step is retained.
WS3's auto-dispatch (dispatch_invoker.py + graph/coordinator wiring) is
unchanged: it fires workflow_dispatch, which now pauses at the restored gate
for human approval — auto-dispatch up to the approval, then one click.
CI has no fastapi (box-only dependency); api.py imports it lazily. Guard the
7 TestClient smoke tests with skipif(find_spec('fastapi') is None) so they
skip in CI instead of failing collection, leaving the 14 invoker_multi tests
running. Also apply ruff format to the 5 WS1 files CI flagged.
- invoker_multi.py: multi_invoke(prompt, *, model) dispatches to GPT-4.1
(cross_reviewer), DeepSeek (fast_coder), or Gemini (scanner) in-process;
bind_multi_invoker() wires the review seam; lazy models import
- review_loop_llm: swap default_plan_reviewer from make_run_py_invoker()
to make_cross_reviewer_invoker() (in-process GPT-4.1); subprocess path
retained as make_run_py_invoker() for opt-in use
- builders_llm: add make_fast_coder_invoker() in-process DeepSeek path;
rename subprocess path to subprocess_build (opt-in fallback); default_build
now delegates to make_fast_coder_invoker()
- api.py: FastAPI app with bearer-token auth (AGENT_TEAM_API_TOKEN env),
POST /tasks, GET /tasks/{id}, POST /orchestrator/invoke; binds 127.0.0.1;
code only — not started here
- run-team.py: bind_multi_invoker() called in serve() alongside
bind_subscription_invoker()
- tests: 21 new WS1 tests + updated builders_llm tests (1065 total, all passing)
- retriever: add save_memory() writing to _box-drafts/ review queue, add
memory_dir param to retrieve() for isolated test routing
- handbook: new load_handbook_conventions() with safe no-op contract (returns
"" when dir absent/empty/unreadable, never raises)
- clarifier_llm: add context_provider=None seam to ClaudeClarifier and
build_claude_clarifier_callables(); failure in provider is silent
- planner: add context_provider=None seam to build_plan_prompt() and plan_node()
- coordinator: thread context_provider through default_plan_node_factory(),
only pass kwarg when non-None to preserve stub-monkeypatching in tests
- tests: 19 new WS5 tests covering all seams (1063 total, all passing)
confluence-doc is provisioned (OAuth 2.0 service-account creds in ~/secrev.env +
PAGE_MAP_FILE from the live IT space). Drop it from COORDINATOR_SKIP_ROLES (now
just aws-posture) and add Environment=PAGE_MAP_FILE. Verified live on the box:
'Confluence API: oauth auth ready', 21 recommend-only doc gaps reported (D7, no
writes). aws-posture stays skipped (IAM/step-ca not stood up).
The public /_edge/tenant_info endpoint returns an HTML 'Page Unavailable' on the
seahavenind site, so cloudId auto-resolution failed. Fetch the token first, then
resolve cloudId from the OAuth-native https://api.atlassian.com/oauth/token/
accessible-resources (Bearer), preferring the resource whose url matches the
configured site, else the first. CONFLUENCE_CLOUD_ID override still honored.
Canary 3/3 (offline), shellcheck clean.
Atlassian org service accounts have no classic API token — they authenticate via
OAuth 2.0 client-credentials (2LO). Add a dual-mode auth seam to confluence-doc:
- OAuth (preferred when CONFLUENCE_OAUTH_CLIENT_ID/_SECRET set): POST
auth.atlassian.com/oauth/token (client_id+client_secret+grant_type=client_credentials)
→ 60-min Bearer; calls go to api.atlassian.com/ex/confluence/<cloudId>/wiki/api/v2/...
cloudId auto-resolves from the site's public /_edge/tenant_info (no input needed).
- Basic (email+API token) retained as a fallback.
conf_api_init() picks the mode once; conf_get() does the authenticated GET. Any
failure (no cloudId / token request fails) → RUN_API=0, live checks SKIPPED, NO
false alarm (matches the existing no-data discipline). Secret passed in the request
body (--data-urlencode), never logged. Still read + recommend-only (D7); canary
unaffected (offline) — 3/3. shellcheck clean (accepted SC1091).
Live smoke exposed a self-pollution bug (pre-existing from PR #17): the post-build
denied-path check wrote its own _build_diff_z.bin / _build_status_z.bin into the
working tree, then its own `git status --untracked-files=all` flagged them as
out-of-scope writes — failing any run with a narrow declared_scope (the smoke's
docs/**). Write them to $RUNNER_TEMP instead (read via $_DIFF_Z/$_STATUS_Z), so
the check no longer sees its own temp files. The agent-team pytest artifacts were
already correctly gitignored; only the check's own files tripped it. Test harness
updated to pass the env paths. 1044 tests, ruff clean.
GitHub Actions only runs workflows under .github/workflows/, so the apply/verify
workflow at agent-team/ci/ was never registered (workflow_dispatch 404'd). Move it
to .github/workflows/agent-team-apply-verify.yml so it is a real, dispatchable
workflow. Its only trigger is workflow_dispatch + it is gated by the agent-apply
required-reviewer environment, so it never auto-runs and nothing privileged runs
unapproved. Updated the two workflow test files' path refs (parents[2]/.github/
workflows) and the ci/README pointer. Dispatcher push gains --no-verify: the apply
path is scanned CI-side (guard + the PR's checks), so it must not be blocked by the
operator's LOCAL human-commit pre-push dev hook (which flags pre-existing whole-repo
FPs like .env.example). 1044 tests, ruff clean.
- checker_coordinator.sh: add COORDINATOR_SKIP_ROLES env (comma-separated) to drop
roles whose credentials are not provisioned from the registry entirely (never
canaried/run/ALARMed). Fail-safe: empty/unset = run all.
- systemd: sea-haven-checkers.{service,timer} run the coordinator nightly at ~03:30
UTC (90 min after the secrev sweep so they don't contend on $MIRROR_DIR / the
Claude pool). The unit sets COORDINATOR_SKIP_ROLES=aws-posture,confluence-doc
(aws-posture needs IAM Roles Anywhere; confluence-doc needs the confluence-bot
token — both intentionally unprovisioned).
Deployed + enabled on the box (deploy-before-merge): canary 4/4 with the skip,
timer scheduled for 2026-06-23 03:35 UTC.
The build-test step ran the generic root `pytest -q`, but this repo's root
suite needs live ANTHROPIC/COMPOSIO keys (collect-only in CI) — so on the
orchestrator target it would fail in the credential-less build sandbox and block
every apply. Retarget the authoritative test to the self-contained agent-team/
subproject (exactly what the repo's green `ci / subproject-tests` runs). ruff
stays repo-wide. Command is AUTHOR-FIXED (not a dispatch input) so no injection
surface is added; multi-target parameterization remains a provisioning item.
This unblocks the first live smoke test.
* feat(agent-team): P3-flip Phase 1 — expand denylist vectors (§4.2) + runner-trust assertion (§4.1)
First controls of the P3-live-flip Phase-1 CI hardening (workflow stays INERT;
this only tightens the trust boundary). Whole Phase-1 surface is gated by
/sh-security-review + GPT-4.1 cross-review before any flip.
§4.2 — expand the trust-control denylist with direct code-execution / supply-chain
vectors, kept byte-identical across all three copies (ci_gate.DENYLIST_GLOBS + the
guard + post-build inline DENY_GLOBS), drift-guarded:
.gitmodules, .husky/**, .githooks/**, .gitattributes, .npmrc, and generated/build
artifacts (__generated__, *.generated.*, dist/**, build/**, *.min.js).
Deliberate: lockfiles are NOT wholesale denied — lockfile-postinstall RCE is already
contained by the credential-less egress-blocked build sandbox, and the Tier-3 dep-CVE
fixer rewrites lockfiles to produce its draft PRs; a blanket deny would make it
un-shippable. Flagged in-code for the security gate. Direct code-execution config
(hooks/filters/npmrc/submodules) is the actual §4.2 RCE surface.
§4.1 — runner-trust: assert no job (esp. the privileged gate-and-pr) can run on a
self-hosted/user-provided runner; all must be GitHub-hosted.
998 tests pass, ruff clean.
* feat(agent-team): P3-flip Phase 1 — gate-weakening detector (§4.5)
A diff that ADDS a lint/type/coverage/security suppression (noqa, type: ignore,
pragma: no cover, nosec, nosemgrep), a test skip/xfail, or a hook bypass
(--no-verify) could make CI pass falsely. The pure-code gate now flags these via
gate_weakening_violations() and BLOCKs in evaluate_ci_gate as a top-priority trust
violation (step 1b, alongside the denylist) — regardless of the authenticated CI
conclusion. A build cannot pass itself by disabling its own checks; flagged diffs
escalate to a human. Only ADDED lines are inspected (removing a suppression is fine).
1015 tests pass, ruff clean.
* feat(agent-team): P3-flip — diff transport (§4.3) + flip privileged apply path live
Completes the box->CI diff handoff and flips the apply/verify privileged job
live (gated behind the agent-apply environment's required reviewer).
Transport (§4.3): the read-only box (D2) emits a diff but holds no write token.
- New credential-less `materialize` job decodes the untrusted `diff_b64`
dispatch input via env (CWE-94), fail-closed re-hashes it against
`expected_diff_hash`, and uploads it as the named artifact so guard/build-test
download it same-run. guard now `needs: materialize`.
- New `dispatcher.py` (the trusted apply path, operator/Mac-side — never the
box): pushes the diff as a head branch then `gh workflow run`s the workflow.
Pure input-assembly (sha256 == sha256sum, b64 round-trip, head ref) is
unit-tested; git/gh are injected seams. Push-before-dispatch; fail-closed on
empty diff/scope, unsafe task_id/owner/repo.
Flip: gate-and-pr binds `environment: agent-apply` (required reviewer
amoussa1229) + grants exactly `pull-requests: write`; the App-token + draft-PR
steps run only on `steps.gate.outputs.gate == 'pass'` (no more if:false); the
draft PR opens with an explicit `--head`; task_id/head_branch charset-validated
(§4.6). Updated the hardening tests from inert-state to live-state assertions +
added transport tests. 1039 tests, ruff clean, workflow YAML valid.
NOTE: workflow only runs on manual workflow_dispatch and the privileged job is
held at the required-reviewer gate, so nothing privileged runs unapproved.
* fix(agent-team): P3-flip — address GPT-4.1 cross-review (size bound, ref-traversal guard)
- BLOCK: cap candidate diff at 40 KB in the dispatcher (the diff rides a base64
workflow_dispatch input; GitHub caps inputs at ~64 KB so an oversized diff
cannot dispatch at all) + a defense-in-depth decoded-size bound in materialize.
- FIX: harden the draft-PR HEAD_BRANCH guard to reject leading/trailing slash,
'..' segments, and '//' (CWE-88 git ref-traversal), not just bad charset.
- NIT: document the mandatory invariants on gate-and-pr (required-reviewer
environment must stay; runs-on must stay GitHub-hosted).
- QUESTION (lockfiles): answered in-code — the build-test sandbox is
credential-less + egress-blocked, so lockfile-postinstall RCE is contained.
Tests added for all guards. 1042 tests, ruff clean, YAML valid.
* fix(agent-team): P3-flip — resolve /sh-security-review findings (LOGIC-1/2/3)
High-recall fan-out (injection/logic/iac+secrets) + proof-or-kill on the LIVE
apply path found 3 real issues the cross-review missed; all fixed:
- LOGIC-2 (HIGH, was a live hole): build-test ran `ruff check . || echo` /
`pytest -q || echo`, swallowing failures so the job was always 'success' and
the gate would open draft PRs on RED builds. ruff/pytest now run
authoritatively under set -e (pytest exit 5 'no tests' is the only non-fatal
case); the exit code IS the build-test conclusion the gate keys on.
- LOGIC-1 (verified!=shipped): the dispatcher used `git apply` + `git add -A`,
staging stray untracked content into the pushed PR head. Now `git apply
--index` stages exactly the diff, so the head tree is precisely base+diff —
bound to the bytes CI hash-verified.
- LOGIC-3 (§4.5 on the live path): gate-weakening was enforced only box-side;
added a gate-weakening check to the guard job so the live PR-opening path
rejects a diff that adds suppressions/skips, even on a green build.
Injection / secrets / least-privilege / flip-correctness / no-untrusted-checkout
all came back clean. 1044 tests, ruff clean, YAML valid.
* docs(agent-team): fold round-2 GPT-4.1 plan-review findings into P3-live-flip plan
Round-2 cross-review (REQUEST CHANGES) folded:
- B1 deploy-before-merge made a concrete CI-enforced gate (required status check
fed by a box-exercised dispatch dry-run), not prose.
- B2 added rollback for a prematurely-flipped privileged job that actually RAN
(token rotate, revert opened PR/branch, audit the window) — distinct from an
accidental merge.
- B3 rollback is a tested, re-runnable script over ALL privileged surfaces
(workflow, environment, App perms, branch protection), not a one-time manual run.
- B4 Phase 3 explicitly gated on Phase 2 being fully provisioned + verified.
- B5 /sh-security-review + GPT-4.1 cross-review RE-RUN on the actual enabled
workflow before the flip, not only the inert version.
- B6 docs/memory updated incrementally at each privileged step; Phase 6 is the
final reconciliation pass.
Plus FIX/NIT/QUESTION: gate-weakening pattern review, check-name discovery test,
memory update on denylist change, snapshot retention in Phase 0, draft-PR
notification (Phase 4) + conservative-rollout controls (Phase 5) made concrete.
GH_TOKEN->GITHUB_TOKEN gate marked done (box alias added).
* docs(agent-team): close B1 — deploy-before-merge is a committed hard gate, not a manual fallback
Round-3 re-review resolved B2-B6 but flagged B1 still-open: the prior wording left
a 'enforced by hand until the check exists' escape hatch. Reframe B1 as a REQUIRED
Phase-1 build deliverable that blocks the flip (no manual fallback), with the
required-status-check + admin-bypass-disabled branch protection, and a dry-run
faithfulness note (same workflow file/jobs as live, only the privileged if: differs).
* docs(agent-team): close B1 ordering — admin-bypass disabled before any flip via the B4 precondition
Round-4 confirmatory re-review noted (c) admin-bypass-disable sits in Phase 2 while
the gate is framed Phase-1. Clarify there is no ordering window: the flip (Phase 3)
is gated on Phase 2 completion (B4), so branch protection incl. admin-bypass-disable
is necessarily in place before any flip. The check's implementation being a Phase-1
build task (not yet physically built) is expected for a pre-build plan; it is
non-optional and flip-blocking, enforced at the Phase-1 hard stop. Stopping the
plan-review cycle here per the project's '3 cycles, residual is build-time' rule.
* feat(agent-team): durable GitHub-issue intake de-dup (schema v2)
The intake poller de-duped ingested issues in an in-memory set that does
not survive a process restart. A scheduled/cron intake (each run a fresh
process) would therefore re-ingest every still-open labeled issue on every
run and spawn duplicate pipeline tasks. Since the box is read-only (no write
token to remove the intake label), durable de-dup is the only correct guard.
- schema v2: new ingested_issues(source, issue_id, ingested_at) table +
issue_already_ingested / record_issue_ingested helpers; migrate() adds the
table to a legacy v1 DB and restamps; init_db creates it.
- github_intake: pluggable IngestStore seam (in-memory default preserved for
tests/one-off; durable build_ledger_ingest_store for production). Record is
after start_task succeeds, so a failed intake stays retryable.
- run-team.py intake-github wires the ledger store keyed by github:owner/repo,
making a scheduled timer idempotent across runs.
978 tests pass (ruff clean).
* fix(agent-team): harden github intake per /sh-security-review (CWE-918, idempotency)
Fixes from the high-recall detector fan-out on the durable-dedup change:
- INTAKE-LOGIC-01 (idempotency): switch the IngestStore seam from
check-then-record (seen/mark) to claim-then-do (claim/release). The id is
now reserved BEFORE the non-idempotent start_task side effect, so a crash in
that window cannot re-spawn a duplicate task on the next run; a raising
start_task releases the claim so transient failures stay retryable. Adds
delete_issue_ingested to the schema layer for the release path.
- INTAKE-SSRF-001 / INTAKE-PATHSPLICE-002 (CWE-918) in build_default_issue_client:
drop the caller-overridable api_root (hardcode GITHUB_API_ROOT) and validate
owner/repo against an anchored charset before splicing them into the
token-bearing API URL — mirrors the sibling ci_fetcher BLOCK-3/FIX-3 fixes.
Tests cover cross-process duplicate prevention, release-on-failure retry, and
the owner/repo + api_root rejection. 982 tests pass, ruff clean.
Follow-up (pre-existing, not introduced here): the label-only intake has no
author allowlist (cf. AGENT_TEAM_SLACK_OWNER_IDS on the Slack listener); the
Slack answer gate bounds the blast radius. Track as separate hardening.
Three robustness/hardening fixes surfaced by /sh-security-review on the
agent-team listener/coordinator surface. None alter AUTHZ-01 allowlist
behavior or the first-answer-wins compare-and-set semantics.
1. Respawn close() leak (CWE-772). The watchdog _supervise_slack_listener
respawned the inbound Slack listener without tearing down the dead one,
leaking a Socket Mode WebSocket / SDK thread set per flap. Now the dead
listener is closed before respawn (new _close_dead_listener, idempotent),
AND _run_listener has a finally that always closes the listener so a
crashed serve() releases its socket. SlackListener.close() is idempotent,
so the belt-and-braces close stays a safe no-op.
2. Broadened exception guard in handle_event (CWE-248). The submit block
only caught ValueError; the accept path (submit_answer ->
_question_turn/_question_thread) can raise KeyError on a concurrently
mutated row, and the CAS can raise sqlite3.Error. An uncaught exception
would escape into the Bolt dispatch. Added a separate `except Exception`
that logs at WARNING (not silent, not debug) and returns None. The
existing ValueError-as-debug behavior is unchanged; authorization still
runs first, so the trust boundary is not widened.
3. channel_ref partial-unique index (defense-in-depth). Added
uq_pending_questions_open_channel_ref — a PARTIAL UNIQUE index on
(channel_ref) WHERE channel_ref IS NOT NULL AND status='open' — so two
OPEN rows can never share a non-null channel_ref (a thread_ts can never
map to two open questions). Installed in init_db AND unconditionally in
migrate (idempotent IF NOT EXISTS) so existing v1 DBs gain it. NULLs and
closed rows are excluded; mirrored verbatim into schema.sql.
Tests: +8 (was 960, now 968). New: schema partial-unique reject/null/closed/
migrate cases; handle_event KeyError + sqlite3.Error swallow cases;
coordinator close-before-respawn + run_listener-closes-on-crash. Fixed the
operator-cli test fixture to use a per-question channel_ref (it previously
inserted multiple open rows sharing one ref, which the new index correctly
rejects).
Phased plan to take Plane-2 from clarify+plan to producing reviewable draft PRs:
locked decisions (GitHub App pull-requests:write, zero-AWS, read-only box), the
mandatory gates (/sh-plan-review + /sh-security-review + GPT-4.1 cross-review on
the CI surface), the B4 CI trust boundary (split CI, denylist, diff-hash, pure-code
green gate), phases 0-6 with owners + exercised rollbacks, what changes vs what
does not, risks, and definition of done. Input to /sh-plan-review before any build.
The org-owned app A0BC7AT8NUD (created via the org config token) connected its
Socket Mode socket but received ZERO workspace Events API events, so the
clarifier never heard answers. Root cause: Socket Mode event delivery only works
for WORKSPACE-LEVEL apps. Recreated from the dashboard scoped to the Sea Haven
Industries workspace -> A0BCC7TTU66 (bot @agentteam2); events now deliver and the
live human gate works end to end. Old org app deleted.
Updates the slack/ README: new app id/bot, a hard-lesson callout (must be
workspace-level, never org-owned), the real-envelope answer-matching note, and a
corrected dashboard setup/update flow.
* fix(agent-team): handle real slack_bolt event envelope + map thread replies to open questions
The Socket Mode inbound listener was unit-tested against a SYNTHETIC payload
shape that does not match what slack_bolt actually delivers, so the suite was
green while a real Slack thread reply was silently dropped (the clarifier
question stayed `open`). Real slack_bolt delivers an Events API message /
app_mention as `{"type":"event_callback","event":{"type":"message",...}}` and
a free-text thread reply carries NO callback_id/question_id/metadata.
Three breaks fixed (all on the free-text reply path):
1. Type gate — handle_event gated on the OUTER `type`, which is
"event_callback" for a real message/app_mention, so the event fell outside
_ANSWER_BEARING_TYPES and was dropped. Now collapsed to the discriminating
INNER `event.type` via _discriminating_type / _inner_event.
2. question_id recovery — a real reply has no callback_id/question_id/metadata
(the bot's metadata is on the QUESTION message, not the reply). When explicit
id recovery fails, the listener now resolves the question by the inner
event's `thread_ts` against the OPEN ledger row whose `channel_ref` equals it
(new schema helper find_open_question_by_channel_ref, constrained to
status='open' as anti-replay). Explicit id recovery still takes precedence.
3. answer extraction — a real message event carries its text at `event.text`,
not a top-level `answer`/`text`. The thread-reply path now takes the inner
`event.text` (stripped) as the answer value.
AUTHZ-01 is unchanged and still runs FIRST: authorization gates on the sender's
Slack user id (`event.user` for the Events API shape) and fails closed on an
empty/unknown allowlist or unrecoverable sender. The new mapping only resolves
WHICH question is answered, never WHO may answer. Answers stay opaque DATA
(parameterized SQL + json.dumps; never eval/exec/interpolate).
Tests: replaced the synthetic events-API fixtures with REAL Bolt envelopes and
added regression coverage — real thread reply maps via channel_ref and is
accepted, text is stripped, non-owner reply rejected (row stays open), thread_ts
matching no open row is a no-op, reply to an already-answered row is a no-op
(anti-replay), and app_mention is normalized identically. block_actions /
slash_command paths retained.
* fix(agent-team): bind subscription invoker in the start CLI
`run-team.py start` runs the clarifier graph to the first human gate IN the CLI
process, and the clarifier calls Claude (assess_confidence). The invoker is a
process-local binding that only `serve` set, so `start` failed with
"claude_invoke has no invoker bound". Bind the real subscription invoker here,
mirroring Coordinator.serve(). Found during the live R720 P1 bring-up.
The Socket Mode inbound listener crashed at registration time on first live
run: `@app.action({})` raised `BoltError: action ({}) must be any of str,
Pattern, and dict` under slack_bolt 1.28.0, killing the listener thread (the
whole inbound answer path — message/app_mention/block_actions — went down,
caught only by the coordinator's respawn watchdog). serve() is marked
`# pragma: no cover - live socket`, so this was never exercised until the R720
bring-up. Replace the unsupported empty-dict matcher with a catch-all
`re.compile(r".*")` action_id regex; handle_event still does the real filtering
+ AUTHZ-01 owner-allowlist gate, so over-matching is safe.
Verified on sh-secrev: listener connects (live Socket Mode WebSocket), outbound
chat.postMessage works, 0 errors. /sh-security-review PASS (no confirmed
critical/high; matcher change introduces no new findings).
Also adds the dedicated Slack app (manifest + README) backing the clarifier
gate — "Sea Haven agent-team" (A0BC7AT8NUD), workspace-scoped install to avoid
the Enterprise-Grid `scope_not_allowed_on_enterprise` org-install trap — and
patches the provisioning runbook's stale langgraph pin (1.1.10 -> 1.2.5).
* fix(agent-team): serve() starts the inbound Slack listener (D-1)
Coordinator.serve() now constructs and starts the SlackListener concurrently
with the tick/drain loop on a background daemon thread, but ONLY when the live
transport is a SlackTransport AND SLACK_APP_TOKEN is configured. When Slack is
not the transport or the app token is absent, serve() behaves exactly as before
(tick/recover only) — Slack is never made mandatory.
- New injectable build_listener seam + default_slack_listener_factory sharing
the coordinator's own transport, ledger db_path, and resume_queue put.
- AUTHZ-01 owner-allowlist + open-status CAS untouched: serve() sources
AGENT_TEAM_SLACK_OWNER_IDS in SlackListener.serve, which still fails closed.
- SlackListener.close() added for clean Socket Mode teardown on shutdown;
serve() stops the listener + joins the thread in a finally.
- Tests: start-when-Slack+app-token, no-start otherwise, clean shutdown,
idempotent start, serve start/stop around the loop, listener close().
* fix(agent-team): systemd unit loads ~/orchestrator/.env + uses venv python (D-2/D-7)
D-2: add EnvironmentFile=-/home/adam/orchestrator/.env (optional '-') so the P2
GPT-4.1 review loop's cross_reviewer sub-process can read the non-Claude provider
key once a task reaches REVIEW. Mirrors the sea-haven-secrev unit.
D-7: point ExecStart at the agent-team venv interpreter
(/home/adam/orchestrator/agent-team/.venv/bin/python) instead of
/usr/bin/env python3, which resolved the system interpreter without the
installed deps under systemd's PATH.
All hardening (NoNewPrivileges / ProtectSystem=full / ProtectHome=read-only /
ReadWritePaths) is retained unchanged (locked decision).
* docs(agent-team): land provisioning + operator runbooks under docs/provisioning
- PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5
intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub
App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked
FIXED so the demo can use the live Slack answer path.
- P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the
Slack and operator-CLI answer paths documented for all four exit criteria.
- DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the
operator-CLI divergence kept as provisioning notes.
- OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks,
failed HITL resumes, budget exhaustion, transport outages, and
COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the
re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6).
* fix(agent-team): supervise the Slack listener thread — recurring ALARM + respawn
sh-security-review (logic) MEDIUM: a crashed listener thread was logged once,
then the daemon ran on 'deaf' — posting clarifier questions but receiving no
answers, every gate silently parking, process never exiting so systemd
Restart=on-failure never fired. serve() now calls _supervise_slack_listener()
each pass: when the listener is enabled but its thread is dead, it emits a
recurring ERROR ALARM and respawns via the idempotent starter (self-heal).
No-op when alive or disabled. +3 tests. (authz detector: wiring clean — AUTHZ-01
fail-closed allowlist + open-status CAS intact, dead listener fails SAFE.)
* feat(secrev): wire full Plane-1 roster into the checker coordinator registry
Register doc-drift, aws-posture, plan-groomer, confluence-doc (weekly cadence)
alongside compliance-drift + dependency-cve (nightly). The coordinator canary
suite now runs all 6 roles' canaries (all PASS) under the one shared budget +
versioned rotation/coverage state; --squeeze-dry-run still proves defer-not-drop
+ COVERAGE alarm. Central integration after the parallel Phase-3/4 PRs landed.
* fix(secrev): valid SAM in doc-drift fixture templates (cfn-lint E0001)
The doc-drift sample-stack fixtures declared AWS::Serverless::Function with no
Properties; cfn-lint's SAM transform errored (HIGH). Added minimal valid
Properties (Handler/Runtime/InlineCode). Pre-existing on main — #19 pushed with
--no-verify (xargs overflow) and CI runs no cfn-lint, so it slipped through.
doc-drift still detects the stack (keys on template presence).
* feat(secrev): doc-drift Plane-1 Tier-1 checker (UNGATED)
Third Plane-1 checker on the Phase-0 shared substrate, mirroring
compliance-drift.sh / dependency-cve.sh conventions verbatim (set -euo pipefail,
sourced substrate, --canary/--dry-run/--no-api/--refresh/--targets, mode-600
reports under $REPORT_ROOT/doc-drift/<UTC-date>/, ALARM-only, finding.schema
spirit JSON, exit 0/2/3, dotgit->.git fixture trick).
Detects documentation drift deterministically (design §4 doc-drift row):
- readme-omits-component: README omits an existing major component in the tree
(top-level service dir, SAM/CDK stack, Lambda handler dir, openapi/docs spec)
- readme-stale-vs-code: README last-touch far older than newest code commit
(two-factor: >=DOC_DRIFT_STALE_DAYS AND >=DOC_DRIFT_STALE_COMMITS)
A repo with NO README is SKIPPED (compliance-drift owns readme-present; no
double-flag). Future Gemini large-context judge (§4) is an inert stub (maybe_judge),
off in canary/dry-run/offline.
Planted-drift fixture corpus + EXPECTED_DRIFT_COUNT=4, canary-asserted (exit 3 on
miss). shellcheck -x clean (only accepted SC1091 source-line info).
Does NOT touch checker_coordinator.sh, requirements.txt, or aws-posture.
Wiring/systemd is gated (PROVISIONING footer). Design refs §4, §7 Phase 3.
* feat(secrev): Phase-3 IAM artifacts for cross-review (aws-posture gated)
Authored FILES (not applied to AWS — provisioning gated behind the mandatory
GPT-4.1 IAM cross-review + Adam, design §7 B3) for the aws-posture checker's
read-only AWS identity. Decision D5: box stays read-only, auths via IAM Roles
Anywhere short-lived leaf certs from a new internal step-ca; NO long-lived AWS key.
- aws-posture-readonly-policy.json least-privilege read-only (ce:Get*,
cloudwatch:GetMetric*/DescribeAlarms, ec2/elb/rds:Describe*, lambda list +
GetFunctionConfiguration, s3:ListAllMyBuckets/GetBucketLocation). No write,
no iam:* mutation, no s3:GetObject/secrets/kms/logs data reads, no wildcard
actions. Resource:* only where AWS has no resource-level support.
- aws-posture-readonly-policy.rationale.md per-statement least-privilege rationale.
- aws-posture-trust-policy.json pins Roles Anywhere principal + leaf subject CN +
issuer CN + trust-anchor SourceArn (three conditions, all required).
- roles-anywhere-config.json trust anchor (pins step-ca root) + profile (1h session).
- step-ca-config-sketch.md internal CA config + systemd-timer leaf auto-renewal.
- CROSS-REVIEW-PACKET.md end-to-end trust model, blast radius, EXERCISED rollback,
reviewer scrutiny list.
Does NOT build aws-posture.sh, touch checker_coordinator.sh, or requirements.txt.
* fix(secrev): apply IAM cross-review FIXes
GPT-4.1 IAM cross-review 2026-06-18: APPROVE, no BLOCKs. Applied FIXes:
- trust policy: add aws:SourceAccount=328440206208 (confused-deputy guard)
alongside the existing aws:SourceArn trust-anchor pin
- readonly policy: remove ec2:DescribeImages (data minimization — AMIs are
not an idle-spend signal)
- aws:RequestedRegion NIT: deliberately SKIPPED — ce:* and s3:ListAllMyBuckets
are global-endpoint services a blanket region condition could DENY; rationale
recorded in aws-posture-readonly-policy.rationale.md
- rationale.md + CROSS-REVIEW-PACKET.md: record APPROVE + FIXes + NIT answers
(snapshots=account-owned idle signal; s3 list=names-only; no logs:* needed)
* feat(secrev): aws-posture checker (Tier-2, provisioning-gated)
Read-only Tier-2 idle/anomalous-spend + idle-resource posture checker for the
R720 agent-team (design D5 / §4 / §6.3 / §7 Phase 3). Mirrors the Tier-1 checker
conventions verbatim (flags --canary/--dry-run/--no-api/--targets, mode-600
report under $REPORT_ROOT/aws-posture/<date>/, ALARM-only, finding.schema.json
spirit, exit 0/2/3, shared substrate redact/post_slack_alarm).
Detectors (complement GuardDuty/SecurityHub/Config, do not replace):
- anomalous Cost Explorer deltas (ce get-anomalies, $-impact threshold)
- stopped EC2 still paying for attached EBS
- unattached EBS volumes
- unassociated Elastic IPs
- idle NAT gateways (≈0 bytes out)
- idle load balancers (0 healthy targets)
- idle RDS (0 connections over window)
Live AWS calls are PROVISIONING-GATED: they run ONLY when Roles Anywhere creds
are available (STS identity probe) AND not --no-api/--canary. With no creds or
--no-api/--canary the checker SKIPS live calls and notes them — NEVER alarms on
missing data (memory feedback_cloudwatch_alarms). Roles Anywhere/step-ca are not
stood up (IAM cross-review PASSED 2026-06-18; see security-review/iam/).
Offline canary: fixtures of mocked AWS responses (cost/describe-* JSON) under
fixtures/aws-posture/ + EXPECTED_FINDING_COUNT=7, asserted fully offline (no aws,
no network). Identical detector code runs online and offline. shellcheck-clean
(only accepted SC1091), chmod +x.
* fix(secrev): doc-drift fixture py ruff-clean (root CI runs check + format --check)
The repo-root CI lint runs both 'ruff check .' and 'ruff format --check .' over
all fixtures. Fixed E701 one-liners and ruff-formatted the sample-service .py
files (handlers/*, feature_*.py). Fixture content is irrelevant to doc-drift
(keys on file/dir presence + git staleness).
* feat(agent-team): Plane-1 Tier-3 fixer — dependency-cve finding -> patch + CI dispatch (opt-in/inert)
The fixer (design §4 fixer row, §7 Phase 5, §3.3.2) takes a CONFIRMED,
low-risk dependency-cve finding (the narrowest fix class) and produces:
* a fix SPEC (Claude, via the §3.1 billing seam), and
* a minimal bump PATCH (DeepSeek fast_coder, via the orchestrator run.py
path that builders_llm uses),
records the candidate diff + its content-hash, and emits the org-CI
workflow_dispatch inputs (task_id / diff_artifact_name / expected_diff_hash /
declared_scope) for the gate-passed P3-live apply/verify surface.
INERT / opt-in / fail-safe, mirroring build_verify_wiring:
* plan_fix dispatches NOTHING; dispatch_fix has NO default dispatcher
(the box holds no write token, D2) so an un-wired call can never fire a
workflow.
* no git/patch/subprocess/fs-write in executable code — the patch is emitted
as diff TEXT only; CI applies it and opens a DRAFT PR, the box never
applies/pushes/merges.
* untrusted-patch hygiene: the generated diff is confined box-side to the
single dependency manifest (declared_scope) and rejected via
ci_gate.denylist_violations if it escapes scope or touches the
trust-control surface — defense-in-depth with the CI guard.
* bad/ambiguous findings (wrong check/status/category, missing
package/fixed_version, ambiguous fixed_version, unparseable/empty diff)
yield a FAILED no-op plan, never a fabricated fix.
29 new pytest tests under agent-team/tests/test_fixer.py.
* feat(agent-team): run-team.py 'fix --dry-run' subcommand for the Plane-1 fixer
Adds the fixer front door to the operator CLI: load one confirmed
dependency-cve finding from a dependency-cve.json report (--report
--finding-id), plan the fix, and in --dry-run print the spec + patch + the
org-CI workflow_dispatch inputs WITHOUT dispatching anything.
Opt-in/inert: the command binds NO workflow dispatcher and holds no write
token, so even an ok plan only prints; live dispatch is provisioning-gated
(refuses to run without --dry-run). A non-fixable finding prints the
fail-safe reason and exits 1.
4 new pytest tests under agent-team/tests/test_run_team.py.
* feat(secrev): plan-groomer Plane-1 Phase 4 planner (report-only)
Aggregates the OTHER Plane-1 checkers' latest reports (compliance-drift,
dependency-cve, doc-drift, confluence-doc) into one prioritized, deduped
"groomed weekly plan" written into the mode-600 report. REPORT-ONLY per
decision D3: posts NOTHING to Slack; auto-write to Notion/Jira is a later
toggle (inert --notify seam). Reuses lib/sweep_substrate.sh redact().
Offline --canary asserts the groomed-plan item count (5) against a fixture
report set, exercising latest-date selection, dedup, multi-source aggregation,
and no-data discipline (a missing source is noted, never invented as work).
shellcheck-clean (only the shared SC1091 substrate-source info, at parity with
compliance-drift/dependency-cve). PROVISIONING (auto-write toggle, systemd
wiring, coordinator registry) deferred — gated.
* feat(secrev): confluence-doc Plane-1 Phase 4 doc-gap detector (recommend-only)
Scheduled, read-only documentation gap detector. Diffs the org repo set + an
optional read-only AWS inventory + the IT page-ID map (project_confluence_
migration) against Confluence and REPORTS doc gaps / stale pages / missing
runbooks into the mode-600 report. RECOMMEND-ONLY per D3/D7: NEVER auto-writes
Confluence; the on-demand SSH-invoked write path (incl. Mermaid edits via
~/.claude/scripts/confluence_mermaid.py) is a separate, gated provisioning path.
LIVE Confluence API reads need the gated confluence-bot service-account token
(D6); when creds are absent OR --no-api/--canary, the API checks are SKIPPED and
noted, NEVER reported as a gap on missing data (mirrors compliance-drift's
status-code-aware API-skip pattern: 200 parse, 404 real gap, else skip).
Offline --canary asserts the doc-gap count (3) against a fixture (repo list +
mock page-map + mock AWS inventory): a repo with no IT page, an AWS resource not
in the map, and a missing required runbook page; precision non-gaps (matched
repos/resources, doc-exempt repo, present required pages, skipped API) must not
inflate the count. shellcheck-clean (only the shared SC1091 substrate-source
info). PROVISIONING (confluence-bot account + 90-day rotation, page-1540098 live
dry-run expecting 16 weweave macros, systemd wiring, coordinator registry)
documented in the footer, deferred — gated.
* fix(secrev): commit compliance-drift secret fixture as dotenv.fixture (canary broke on fresh clone)
The compliance-drift canary's planted tracked-secret fixture was BadName_repo/.env,
but the repo root .gitignore lists '.env' — so it was never committed. On a fresh
clone of main the file is absent, the secrets-committed check stops firing, and the
canary FAILS (expected 6, got 5). It only passed where a gitignored, untracked
'.env' happened to exist locally. Verified the failure reproduces in a clean clone
of origin/main (3d97139) and in a fresh worktree.
Fix (in-convention, mirrors the dependency-cve .fixture-suffix trick): ship the
secret as BadName_repo/dotenv.fixture (committable, not gitignored); the --canary
materialization renames dotenv.fixture -> .env in its temp work area. The dotgit/
index already TRACKS .env, so git ls-files still reports it and the drift fires.
Restores the documented 6/6 canary on any fresh checkout. shellcheck stays clean.
* feat(agent-team): P5 checker-finding intake module + tests
Add agent_team.transport.checker_intake: turn a confirmed, at/above-threshold
Plane-1 checker FINDING into one Plane-2 pipeline remediation task via the
committed coordinator intake entry (start_task), mirroring github_intake.
- select_findings: status==confirmed AND severity>=threshold (default high);
unverified/suppressed/below-threshold dropped; unknown threshold rejected.
- finding_identity: stable de-dup key (finding id, else content-hash). In-memory
set, best-effort, NOT durable across restart (ledger table is the follow-up).
- finding_task_text/_sanitize: every repo-controlled field (title, proof, repo)
is newline/control-char neutralised and length-bounded before it reaches the
task text or operator log (log-injection hygiene).
- load_report_findings/ingest_reports: read the exact checker report JSON shape
(top-level object with findings[]; bare array and dir-of-*.json also accepted).
28 hermetic unit tests (stub coordinator, in-memory findings / temp reports).
* feat(agent-team): wire opt-in intake-checker run-team subcommand
Expose the P5 cross-plane loop only as a manual run-team subcommand
(intake-checker --report PATH [--threshold] [--transport] [--dry-run]),
mirroring how intake-github is exposed. NOT wired into the always-on serve
path: the loop stays opt-in/inert by default.