* fix(agent-team): handle real slack_bolt event envelope + map thread replies to open questions
The Socket Mode inbound listener was unit-tested against a SYNTHETIC payload
shape that does not match what slack_bolt actually delivers, so the suite was
green while a real Slack thread reply was silently dropped (the clarifier
question stayed `open`). Real slack_bolt delivers an Events API message /
app_mention as `{"type":"event_callback","event":{"type":"message",...}}` and
a free-text thread reply carries NO callback_id/question_id/metadata.
Three breaks fixed (all on the free-text reply path):
1. Type gate — handle_event gated on the OUTER `type`, which is
"event_callback" for a real message/app_mention, so the event fell outside
_ANSWER_BEARING_TYPES and was dropped. Now collapsed to the discriminating
INNER `event.type` via _discriminating_type / _inner_event.
2. question_id recovery — a real reply has no callback_id/question_id/metadata
(the bot's metadata is on the QUESTION message, not the reply). When explicit
id recovery fails, the listener now resolves the question by the inner
event's `thread_ts` against the OPEN ledger row whose `channel_ref` equals it
(new schema helper find_open_question_by_channel_ref, constrained to
status='open' as anti-replay). Explicit id recovery still takes precedence.
3. answer extraction — a real message event carries its text at `event.text`,
not a top-level `answer`/`text`. The thread-reply path now takes the inner
`event.text` (stripped) as the answer value.
AUTHZ-01 is unchanged and still runs FIRST: authorization gates on the sender's
Slack user id (`event.user` for the Events API shape) and fails closed on an
empty/unknown allowlist or unrecoverable sender. The new mapping only resolves
WHICH question is answered, never WHO may answer. Answers stay opaque DATA
(parameterized SQL + json.dumps; never eval/exec/interpolate).
Tests: replaced the synthetic events-API fixtures with REAL Bolt envelopes and
added regression coverage — real thread reply maps via channel_ref and is
accepted, text is stripped, non-owner reply rejected (row stays open), thread_ts
matching no open row is a no-op, reply to an already-answered row is a no-op
(anti-replay), and app_mention is normalized identically. block_actions /
slash_command paths retained.
* fix(agent-team): bind subscription invoker in the start CLI
`run-team.py start` runs the clarifier graph to the first human gate IN the CLI
process, and the clarifier calls Claude (assess_confidence). The invoker is a
process-local binding that only `serve` set, so `start` failed with
"claude_invoke has no invoker bound". Bind the real subscription invoker here,
mirroring Coordinator.serve(). Found during the live R720 P1 bring-up.
The Socket Mode inbound listener crashed at registration time on first live
run: `@app.action({})` raised `BoltError: action ({}) must be any of str,
Pattern, and dict` under slack_bolt 1.28.0, killing the listener thread (the
whole inbound answer path — message/app_mention/block_actions — went down,
caught only by the coordinator's respawn watchdog). serve() is marked
`# pragma: no cover - live socket`, so this was never exercised until the R720
bring-up. Replace the unsupported empty-dict matcher with a catch-all
`re.compile(r".*")` action_id regex; handle_event still does the real filtering
+ AUTHZ-01 owner-allowlist gate, so over-matching is safe.
Verified on sh-secrev: listener connects (live Socket Mode WebSocket), outbound
chat.postMessage works, 0 errors. /sh-security-review PASS (no confirmed
critical/high; matcher change introduces no new findings).
Also adds the dedicated Slack app (manifest + README) backing the clarifier
gate — "Sea Haven agent-team" (A0BC7AT8NUD), workspace-scoped install to avoid
the Enterprise-Grid `scope_not_allowed_on_enterprise` org-install trap — and
patches the provisioning runbook's stale langgraph pin (1.1.10 -> 1.2.5).
* fix(agent-team): serve() starts the inbound Slack listener (D-1)
Coordinator.serve() now constructs and starts the SlackListener concurrently
with the tick/drain loop on a background daemon thread, but ONLY when the live
transport is a SlackTransport AND SLACK_APP_TOKEN is configured. When Slack is
not the transport or the app token is absent, serve() behaves exactly as before
(tick/recover only) — Slack is never made mandatory.
- New injectable build_listener seam + default_slack_listener_factory sharing
the coordinator's own transport, ledger db_path, and resume_queue put.
- AUTHZ-01 owner-allowlist + open-status CAS untouched: serve() sources
AGENT_TEAM_SLACK_OWNER_IDS in SlackListener.serve, which still fails closed.
- SlackListener.close() added for clean Socket Mode teardown on shutdown;
serve() stops the listener + joins the thread in a finally.
- Tests: start-when-Slack+app-token, no-start otherwise, clean shutdown,
idempotent start, serve start/stop around the loop, listener close().
* fix(agent-team): systemd unit loads ~/orchestrator/.env + uses venv python (D-2/D-7)
D-2: add EnvironmentFile=-/home/adam/orchestrator/.env (optional '-') so the P2
GPT-4.1 review loop's cross_reviewer sub-process can read the non-Claude provider
key once a task reaches REVIEW. Mirrors the sea-haven-secrev unit.
D-7: point ExecStart at the agent-team venv interpreter
(/home/adam/orchestrator/agent-team/.venv/bin/python) instead of
/usr/bin/env python3, which resolved the system interpreter without the
installed deps under systemd's PATH.
All hardening (NoNewPrivileges / ProtectSystem=full / ProtectHome=read-only /
ReadWritePaths) is retained unchanged (locked decision).
* docs(agent-team): land provisioning + operator runbooks under docs/provisioning
- PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5
intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub
App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked
FIXED so the demo can use the live Slack answer path.
- P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the
Slack and operator-CLI answer paths documented for all four exit criteria.
- DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the
operator-CLI divergence kept as provisioning notes.
- OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks,
failed HITL resumes, budget exhaustion, transport outages, and
COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the
re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6).
* fix(agent-team): supervise the Slack listener thread — recurring ALARM + respawn
sh-security-review (logic) MEDIUM: a crashed listener thread was logged once,
then the daemon ran on 'deaf' — posting clarifier questions but receiving no
answers, every gate silently parking, process never exiting so systemd
Restart=on-failure never fired. serve() now calls _supervise_slack_listener()
each pass: when the listener is enabled but its thread is dead, it emits a
recurring ERROR ALARM and respawns via the idempotent starter (self-heal).
No-op when alive or disabled. +3 tests. (authz detector: wiring clean — AUTHZ-01
fail-closed allowlist + open-status CAS intact, dead listener fails SAFE.)
* feat(agent-team): P5 checker-finding intake module + tests
Add agent_team.transport.checker_intake: turn a confirmed, at/above-threshold
Plane-1 checker FINDING into one Plane-2 pipeline remediation task via the
committed coordinator intake entry (start_task), mirroring github_intake.
- select_findings: status==confirmed AND severity>=threshold (default high);
unverified/suppressed/below-threshold dropped; unknown threshold rejected.
- finding_identity: stable de-dup key (finding id, else content-hash). In-memory
set, best-effort, NOT durable across restart (ledger table is the follow-up).
- finding_task_text/_sanitize: every repo-controlled field (title, proof, repo)
is newline/control-char neutralised and length-bounded before it reaches the
task text or operator log (log-injection hygiene).
- load_report_findings/ingest_reports: read the exact checker report JSON shape
(top-level object with findings[]; bare array and dir-of-*.json also accepted).
28 hermetic unit tests (stub coordinator, in-memory findings / temp reports).
* feat(agent-team): wire opt-in intake-checker run-team subcommand
Expose the P5 cross-plane loop only as a manual run-team subcommand
(intake-checker --report PATH [--threshold] [--transport] [--dry-run]),
mirroring how intake-github is exposed. NOT wired into the always-on serve
path: the loop stays opt-in/inert by default.
clarifier_llm: sha1 -> sha256 for the non-security cache-discriminator (CWE-327 false positive). github_adapter + github_intake: inline nosemgrep on the urlopen lines (dynamic-urllib-use-detected) — the URL is built from a fixed https GitHub API base, dynamic part is the path only, no SSRF/file:// surface (extends the existing noqa:S310 trusted-host judgment to semgrep). Scanner now reports 0 mediums on the agent-team scope.
Live github (issue-comment poster) and claude_code (file-drop) transports, plus GithubIntake (labeled issue -> coordinator.start_task, de-duped). run-team _build_transport now wires github/claude_code live (was SystemExit) + adds the intake-github subcommand. claude_code drop-path also neutralizes backslash (defense-in-depth).
slack_live: real slack_sdk poster. slack_listener: Socket Mode inbound; trust boundary = app-token auth + an explicit owner allowlist on the sender (fail-closed, rejects all if AGENT_TEAM_SLACK_OWNER_IDS unset) + the open-status CAS as anti-replay. Closes the AUTHZ-01 missing-sender-authz finding from the security review.