WIP: feat(agent-team): one Slack thread per task + 👍 ack on received answers #49

Closed
amoussa1229 wants to merge 25 commits from feat/agent-team-slack-threading into main

25 commits

Author SHA1 Message Date
0a761c7ead feat(agent-team): 👍-acknowledge received Slack answers (reactions:write)
WS Slack-UX Feature 2. When the inbound listener acts on an answer in a task
thread, it adds a 👍 reaction to that reply so the human sees the machine
received it.

- SlackListener gains an optional reactor seam; handle_event reacts to the
  inbound reply message (channel + event ts) AFTER the AUTHZ-01 owner check
  passes — a non-owner message is rejected and never reacted to. Best-effort:
  any reaction failure (notably a missing scope) is swallowed and never breaks
  handle_event or the listen loop.
- /new-task is NOT reacted to (a slash command has no reactable message); its
  "📥 Task received" root post is the acknowledgement.
- build_slack_reactor wraps WebClient.reactions_add(name="thumbsup"); the
  default listener factory wires it best-effort from SLACK_BOT_TOKEN.
- Adds reactions:write to the bot scopes in agent-team-manifest.json.

NOTE: the new reactions:write scope requires Adam to re-apply the manifest to
app A0BCC7TTU66 and reinstall the app. Until then reactions.add returns
missing_scope, which the listener swallows (the reaction silently no-ops) —
answer handling is unaffected.

AUTHZ-01 and the first-answer-wins CAS remain unchanged.
2026-06-23 15:11:30 -04:00
8d31ef183d feat(agent-team): one Slack thread per task — root "Task received" message + threaded questions/milestones
WS Slack-UX Feature 1. A /new-task task now maps to ONE Slack thread instead of
several top-level messages.

- /new-task posts an immediate root "📥 Task received: …" ack and captures its
  ts (root_ts); this is the instant acknowledgement.
- root_ts is plumbed into start: new PipelineState/TaskRecord channel
  slack_thread_ts, seeded by graph.start_task and threaded through
  Coordinator.start_task. The NewTaskCallback is now (task_text, via, root_ts).
- All clarifier questions for the task post as THREADED REPLIES under root_ts
  (chat.postMessage thread_ts=root_ts), and each question's ledger channel_ref
  is set to root_ts (NOT the reply's own ts). Because answer-mapping resolves a
  reply via find_open_question_by_channel_ref(thread_ts), a reply in the root
  thread (thread_ts==root_ts) maps to the task's currently-open question with NO
  change to the mapping logic or the first-answer-wins CAS. The open-only
  partial-unique index still holds (one open question per task at a time).
- Lifecycle milestones (parked / plan-ready / needs-input) and follow-up
  questions thread under root_ts too; the notify sink gained an optional
  thread_ts kwarg (degrades to top-level on a sink that doesn't accept it).
  notify failures still never break tick.
- SlackTransport.post_question + the live poster accept/forward thread_ts.
- No root_ts (non-/new-task origin) ⇒ top-level posts exactly as before.

AUTHZ-01 (owner-allowlist-first, fail-closed) and the atomic open→answered
compare-and-set are unchanged.

Adds plumbing for the inbound-ack reactor seam used by Feature 2 (dormant until
a reactor is injected). Tests cover thread_ts forwarding, channel_ref=root_ts,
graph seeding, and coordinator threading.
2026-06-23 15:11:19 -04:00
c749d87518 fix(pipeline): single-shot Claude calls + planner actually reads review findings
Two root causes behind 'every task parks, and slowly':
1. SINGLE-SHOT INVOKER: subscription Claude calls ran as 40-turn, tool-enabled
   agentic sessions (--max-turns 40, $2 budget) for what are pure reasoning->JSON
   completions — minutes-long, and Claude wandered/returned unparseable output.
   _DEFAULT_MAX_TURNS 40->1 + allowed_tools=[] -> fast deterministic single turn.
2. PLANNER FEEDBACK KEY MISMATCH: the review stage writes verdict/findings, but
   _format_review_feedback read decision/notes/comment (never present) -> the
   planner re-planned with EMPTY feedback, re-introduced the rejected flaw
   ('assumptions persist'), hit the review cap, parked. Now reads verdict/findings
   (old keys kept as fallback) so GPT-4.1's objections reach the re-plan.
Plus: park notifications infer the phase the task was IN (review/plan/clarify)
instead of the terminal 'parked'.
Tests: planner real-verdict-keys regression + coordinator phase-inference. 1149 pass.
2026-06-23 14:53:21 -04:00
6c0152c2e5 feat(notify): richer park/lifecycle messages — task description + phase + blocker
Park notifications were opaque ('Task 6c3c3202 parked — needs your attention'):
no idea what the task is, where it got to, or what's blocking it. Now each
message names the task DESCRIPTION (not just the short id), the PHASE it reached,
and the actual BLOCKER — _summarize_blocker() pulls the last review_verdict's
findings (the GPT-4.1 REQUEST_CHANGES text, collapsed + truncated), falling back
to 'no plan built' / 'revision cap hit'. Applies to parked + needs-more-input +
plan-ready emits. +1 test (description + phase + blocker present). 1148 passed.
2026-06-23 14:35:21 -04:00
ebc10649a2 feat(notify): Slack lifecycle notifications + deliver multi-turn questions
The bot only ever posted clarifier questions; parks/completions were silent and
multi-turn follow-up questions were never posted during normal operation (only
the startup recover sweep posted them). So an answered task was a black box.

- Coordinator gains a 'notify' sink + _emit() (guarded, never breaks the loop).
- tick() now runs _post_resume_followups(results) after the drain: for each
  resumed thread it (a) POSTS a newly-pending clarifier question (fixes silent
  multi-turn — the drain path left it unposted) and emits 'needs more input';
  else emits 'parked — needs attention' or 'plan ready for review' from the
  settled state.
- run-team serve wires notify -> Slack channel (build_slack_poster) and an
  alarm_hook that logs the deadline-park WARNING AND posts a parked notice.
  Live-Slack only; dry-run/non-Slack/no-channel = silent (None), no token needed.

Tests: +4 (needs-input/parked/plan-ready emits + notify-failure swallow);
_FakeCoordinator gains notify/alarm_hook. 1147 passed, ruff clean.
2026-06-23 14:17:08 -04:00
6d76194474 fix(intake): seed the task description into graph state (was silently dropped)
/new-task (and every intake: GitHub issue, /sh-assign-task) reached the clarifier
with NO description -> the clarifier asked 'no task description provided'. Root
cause: coordinator.start_task only LOGGED task_text (a P1-era decision when the
deterministic clarifier didn't consume a description), graph.start_task took no
task arg, and PipelineState/TaskRecord had no 'task' channel at all.

Fix: add a first-class 'task' field to PipelineState + TaskRecord (+ round-trip
in task_from_dict); graph.start_task seeds task into the initial invoke (persists
through intake_node's partial-state return into CLARIFY); coordinator.start_task
passes task=task_text. The clarifier already reads state['task'] via
_task_description, so it now sees the real description.

Test: start_task(task='build a login form') -> suspended CLARIFY state carries
task. 1143 passed, ruff clean.
2026-06-23 13:53:39 -04:00
b9a3fc0d57 fix(ws2): register @app.command(/new-task) so Slack slash command is acked
'the app did not respond': serve() registered @app.action/@app.event but NO
@app.command handler, so Bolt never acked the /new-task slash command within
Slack's ~3s deadline. Add an @app.command(/new-task) handler that ack()s first,
re-stamps type:slash_commands onto Bolt's inner command body (Bolt strips the
Socket Mode envelope type that _is_new_task_command/_discriminating_type expect),
then forwards to handle_event (AUTHZ-01 + new-task dispatch). The resulting
payload shape is the one already covered by test_new_task_calls_callback_*.
2026-06-23 13:43:42 -04:00
f14a365604 fix(ops): deploy installs full requirements.txt (incl non-Claude model stack)
The agent-team venv was missing langchain-anthropic/-openai/-google-genai/
-community, so models.py failed to import and the in-process GPT-4.1 review /
Gemini scan / DeepSeek build silently fail-closed to REQUEST_CHANGES (the
non-Claude models never ran on the box). Step 3 now installs the full pinned
requirements.txt into the venv instead of just fastapi/uvicorn. Installed +
verified live on the box: all three model factories construct.
2026-06-23 13:37:11 -04:00
85553ac917 fix(ws1/ws5): bootstrap orchestrator root in in-process invokers (G1); thread handbook into clarifier (G4)
Gap-audit findings:
- G1 (blocks-feature): make_cross_reviewer_invoker / make_fast_coder_invoker did
  'from models import' without putting the orchestrator root on sys.path. The
  run-team serve daemon only bootstraps agent-team/, so on the live box every
  GPT-4.1 plan review hit ModuleNotFoundError -> review_plan's blanket except
  silently fail-closed to REQUEST_CHANGES (GPT-4.1 never actually ran). Both
  in-process invokers now call invoker_multi._ensure_orchestrator_on_path()
  before the deferred import. WS1 introduced this when it swapped the review
  default from the subprocess invoker to in-process.
- G4 (degrades): the handbook context_provider was wired into the planner only;
  default_clarify_node_factory now accepts + forwards it, and run-team wires it
  into build_clarify_node too, so clarifying questions are handbook-aware.

Tests: +2 regression tests (path-bootstrap, clarifier threading); _FakeCoordinator
gains build_clarify_node. 1142 passed, ruff clean.
2026-06-23 13:30:13 -04:00
5f40418dd7 fix(ws2): register /new-task slash command + commands scope in Slack manifest
The slack_listener handles {type:slash_commands, command:/new-task} but the app
manifest declared no slash commands and no 'commands' scope, so Slack never
offered /new-task (the command can't be invoked). Add the slash command +
commands bot scope. Socket Mode delivers it over the socket (no request URL).
APPLY: update the app A0BCC7TTU66 from this manifest + reinstall to pick up the
new scope.
2026-06-23 13:20:08 -04:00
8dda4fbffe docs(integration): document WS0-WS5 components + WS-rollout deploy 2026-06-23 12:56:08 -04:00
861d26106a feat(ops): R720 WS-rollout deploy/update script (G)
Idempotent attended update of the live coordinator to WS0-WS5: rsync repo +
handbook, install fastapi/uvicorn, append AGENT_TEAM_API_TOKEN/SEA_HAVEN_HANDBOOK_DIR
to secrev.env if absent, restart the daemon, smoke tests. HTTP API is an opt-in
separate step; P3 dispatch stays inert. Snapshot-first + confirm before restart.
2026-06-23 12:43:39 -04:00
e6d2f3efb0 feat(integration): wire WS5 context_provider + WS2 /new-task into serve
Integration branch combining WS0-WS5 (PRs #43-#46) + the activation wiring that
flips the safe seams ON in the run-team serve path:

- WS5 (D10): inject the Sea Haven handbook conventions into the planner prompt
  via context_provider (zero-arg handbook loader; fail-safe to '' when absent).
- WS2: an allowlisted Slack /new-task starts a task on this coordinator
  (set_new_task_callback adapter -> start_task; AUTHZ-01 gates it upstream).
- WS1 bind_multi_invoker() is already wired in _cmd_serve.

Coordinator gains a new_task_callback param + set_new_task_callback() (resolves
the constructor chicken-and-egg of referencing the coordinator's own start_task);
default_slack_listener_factory forwards it to the SlackListener.

NOT wired (deliberately): the P3 dispatch_node / build_verify path. Activating
it correctly needs a per-task expected_run_id bound into gated_build_verify_wiring
(plumbing that does not exist yet) AND the CI trust-boundary security re-review.
It stays inert pending that work.

Tests: +5 activation-wiring tests; _FakeCoordinator stub gains
set_new_task_callback. Full agent-team suite: 1140 passed, ruff clean.
2026-06-23 12:38:43 -04:00
c0bc46405f Merge remote-tracking branch 'origin/feat/ws0-plugin-slack-intake-autodelegate' into feat/ws-activation-wiring 2026-06-23 12:31:24 -04:00
23cc412884 Merge remote-tracking branch 'origin/feat/ws3-auto-dispatch-remove-env' into feat/ws-activation-wiring 2026-06-23 12:31:24 -04:00
e3a10a6260 Merge remote-tracking branch 'origin/feat/ws1-inprocess-models-http-api' into feat/ws-activation-wiring 2026-06-23 12:31:24 -04:00
b6c66877f5 fix(ws1): harden HTTP API + declare fastapi/uvicorn deps
Security follow-ups from the per-PR review (non-blocking MEDIUMs):
- Eager _get_token() at make_app build time so a missing AGENT_TEAM_API_TOKEN
  fails fast instead of serving requests first (matches the docstring contract).
- Disable /docs, /redoc, /openapi.json (no auth dependency in FastAPI) — the
  API is VPN-only/127.0.0.1 and should not expose its schema unauthenticated.
- Scrub raw exception text and subprocess stderr from 500 response bodies;
  log server-side instead (avoid internal-path/state disclosure).
- Bound /orchestrator/invoke concurrency with a semaphore (429 over the cap)
  so an authenticated caller cannot exhaust the box via many 600s subprocesses.

Also pin fastapi/uvicorn in requirements.txt (WS1 dep). With fastapi now
installed in CI, the previously skip-guarded TestClient tests run for real;
the importorskip guard stays as a no-op safety net.

Tests: 23 pass (adds docs-disabled + concurrency-429 cases).
2026-06-23 12:27:23 -04:00
b2684231b1 fix(ws3): keep agent-apply environment gate; drop fail-open Slack step
Reworked per GPT-4.1 cross-family review BLOCK. The original PR removed the
agent-apply GitHub Environment (the live required-reviewer human gate) and
replaced it with a Slack notice that FAILS OPEN when its webhook secret is
absent (which it is) plus an audit-log line. The cross-review correctly
flagged this as trading a preventive control for detective controls, one of
which silently no-ops.

This commit:
- Restores environment: agent-apply on gate-and-pr (the human approval pause).
- Drops the fail-open Slack notify step.
- Keeps the unconditional audit-log step as an additive detective control.
- Restores the MANDATORY-INVARIANT assertion (env must be present) and adds
  an assertion that the audit step is retained.

WS3's auto-dispatch (dispatch_invoker.py + graph/coordinator wiring) is
unchanged: it fires workflow_dispatch, which now pauses at the restored gate
for human approval — auto-dispatch up to the approval, then one click.
2026-06-23 12:17:19 -04:00
366d07a84e fix(ws1): skip fastapi TestClient tests when fastapi absent + ruff format
CI has no fastapi (box-only dependency); api.py imports it lazily. Guard the
7 TestClient smoke tests with skipif(find_spec('fastapi') is None) so they
skip in CI instead of failing collection, leaving the 14 invoker_multi tests
running. Also apply ruff format to the 5 WS1 files CI flagged.
2026-06-23 11:40:46 -04:00
83d17a5b7e style(ws0+ws2+ws4): ruff format slack_listener, hook, test (CI ruff format --check) 2026-06-23 11:38:58 -04:00
bfec8cc4d1 style(ws3): ruff format dispatch_invoker + test (CI ruff format --check) 2026-06-23 11:38:31 -04:00
Claude
80c04fc487
feat(ws3): add default_dispatch_node_factory + 25-test dispatch_invoker suite
- Add default_dispatch_node_factory() to coordinator.py: reads
  AGENT_TEAM_REPO_OWNER / AGENT_TEAM_REPO_NAME / AGENT_TEAM_BASE_BRANCH
  from env; fails closed (RuntimeError) if required vars absent; delegates
  to make_dispatch_node with owner/repo fixed at factory time. Added to __all__.
- Add tests/test_ws3_dispatch_invoker.py (25 tests): make_dispatch_node
  happy path + fail-closed paths (missing thread_id/diff/scope/plan, DispatcherError,
  unexpected exception); owner/repo injection from factory args; scope list
  flattening; DispatchNodeFactory export; graph DISPATCH_NODE constant;
  build_graph ValueError when dispatch_node given without build_verify; env-var
  binding for default_dispatch_node_factory; Coordinator.dispatch_node_wiring
  seam (ValueError when wired without build_verify_wiring).

All 1069 tests pass; ruff clean.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155wSD9kKFjnTMNNsDT3tiW
2026-06-23 01:40:32 +00:00
Claude
2aa72d73f4
feat(ws0+ws2+ws4): plugin scaffold, Slack /new-task, auto-delegate hook
WS0 — sea-haven-claude-plugin/ scaffold:
  - CLAUDE.md: Sea Haven engineering context (pipeline overview, rules,
    /new-task + /delegate usage, available phases)
  - settings.template.json: UserPromptSubmit hook wiring template
  - hooks/user_prompt_submit.py: standalone script (WS4)

WS2 — SlackListener /new-task intake:
  - Add NewTaskCallback type alias (Callable[[str, str], str])
  - Add new_task_callback param to SlackListener.__init__
  - _is_new_task_command() helper for slash_commands+/new-task detection
  - _handle_new_task_command() method: authorized-only, calls callback,
    exception-safe (listen loop stays alive on callback errors)
  - handle_event() routes /new-task BEFORE the answer path (post-AUTHZ-01)

WS4 — UserPromptSubmit auto-delegate hook:
  - /delegate <text> and DELEGATE: <text> prefixes trigger delegation
  - Calls POST /tasks on the agent-team HTTP API (WS1)
  - Blocks the Claude Code prompt; shows thread_id + next-steps message
  - Graceful degradation: missing token, HTTP error, network error all
    produce a block with a human-readable reason
  - run() is a pure function for testability (no stdin/stdout in tests)

Tests: 22 new tests in test_ws0_ws2_ws4_plugin_slack_hook.py.
Full suite: 1066 passed. ruff clean.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QYp761G9HojkmLqLZrASVi
2026-06-23 01:39:12 +00:00
Claude
72076b7874
feat(ws3): add dispatch node + remove agent-apply env gate
- Add agent_team/nodes/dispatch_invoker.py: make_dispatch_node() wraps
  dispatcher.dispatch_apply_verify() as a LangGraph node; fail-safe
  (parks task on any error); INERT unless wired by coordinator.
- graph.py: add DISPATCH_NODE constant; add dispatch_node param to
  build_graph; when provided, repoint APPROVED_ROUTE → DISPATCH_NODE →
  END (P3+). Validates dispatch_node requires build_verify.
- coordinator.py: add DispatchNodeFactory type; thread dispatch_node_wiring
  through __init__ and setup(); default None = inert (no auto-dispatch).
- agent-team-apply-verify.yml: remove agent-apply environment gate from
  gate-and-pr; add Slack-notify + audit-log compensating control steps.
- test_apply_verify_workflow_hardening.py: flip env assertion → assert env
  REMOVED and compensating steps present.

All 1044 tests passing. ruff clean.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QYp761G9HojkmLqLZrASVi
2026-06-23 01:31:29 +00:00
Claude
e5a8949d0d
feat(ws1): non-Claude in-process invokers + HTTP API (WS1, code)
- invoker_multi.py: multi_invoke(prompt, *, model) dispatches to GPT-4.1
  (cross_reviewer), DeepSeek (fast_coder), or Gemini (scanner) in-process;
  bind_multi_invoker() wires the review seam; lazy models import
- review_loop_llm: swap default_plan_reviewer from make_run_py_invoker()
  to make_cross_reviewer_invoker() (in-process GPT-4.1); subprocess path
  retained as make_run_py_invoker() for opt-in use
- builders_llm: add make_fast_coder_invoker() in-process DeepSeek path;
  rename subprocess path to subprocess_build (opt-in fallback); default_build
  now delegates to make_fast_coder_invoker()
- api.py: FastAPI app with bearer-token auth (AGENT_TEAM_API_TOKEN env),
  POST /tasks, GET /tasks/{id}, POST /orchestrator/invoke; binds 127.0.0.1;
  code only — not started here
- run-team.py: bind_multi_invoker() called in serve() alongside
  bind_subscription_invoker()
- tests: 21 new WS1 tests + updated builders_llm tests (1065 total, all passing)
2026-06-23 01:19:59 +00:00