Hermetic e2e tests driving the WHOLE stack composed together — real
build_graph(plan_gate=True) + real Coordinator + real SQLite ledger + real
review_loop router, with only the LLM nodes stubbed — through the daemon API
(start_task/submit_answer/tick), never nodes directly.
Flows: approve settles at BUILD; request-changes via RAW PROSE through the real
SlackListener -> _resolve_payload -> map_plan_decision proves the prose maps to
request_changes (NOT FAILED) and the notes reach the planner; abandon -> FAILED;
repeated request_changes terminates at MAX_PLAN_GATE_VISITS -> PARKED; and a
legacy (no-kind) ledger migrates in place then routes clarify vs plan_decision
correctly. 1459 passed.
Turn a human's Slack interaction at the plan gate into a structured decision the
graph can route, with the kind-aware mapping that closes a silent-FAIL hazard.
- KIND-AWARE NORMALIZATION (load-bearing): map_plan_decision() in slack_adapter
maps a reply to {"decision","notes"} — approve ∈ {approve,approved,yes,ok,lgtm,
ship}; abandon ∈ {abandon,reject,cancel,stop,kill}; EVERYTHING ELSE →
request_changes with the full reply as notes (never accidental abandon). Wired
in the listener's _resolve_payload for plan_decision rows ONLY (clarify passes
through). Without this, arbitrary change-notes hit the graph's
unrecognized-verb→FAILED path and silently fail the task. Anti-FAIL tests
assert prose → request_changes (!= abandon) at both the mapper and the
end-to-end listener seam; a regression test guards clarify pass-through.
New find_open_question_kind_by_channel_ref (anti-replay, status='open') powers
the thread-reply fallback's kind lookup.
- BUTTONS + MODAL: build_plan_decision_blocks() renders Approve (primary) /
Request changes / Abandon (danger+confirm); question_id double-anchored in
message metadata AND each button value ("<verb>:<question_id>"). Approve/abandon
submit via the existing @app.action(.*); request_changes has a dedicated
handler that AUTHORIZES before views_open (proven by test) and opens a notes
modal (private_metadata carries the id) → view_submission → request_changes +
notes. Free-text reply stays the always-available equal path. AUTHZ-01 ordering
preserved.
- No manifest change (views.open needs no extra scope).
1454 passed (1412 + 42).
Connect the graph plan-gate (B2a) to the durable ledger + Slack presentation.
- setup() passes build_graph(plan_gate=True) only on the wired review path
(plan_gate flag ANDed with review_node present); P1/stub paths force it off.
- _post_resume_followups detects a settled interrupt by payload
kind == PLAN_DECISION_KIND (NOT status, which still reads 'parked' at the
gate per B2a) and posts the decision gate: opens a pending_questions row with
kind='plan_decision' (24h deadline, threaded, channel_ref = root ts) and
presents _summarize_plan + findings + reply instructions, truncated to a
~2700-char Slack budget. A clarify/legacy interrupt keeps the existing path.
- Single-open-gate invariant: the opener skips if any open row already exists
for the thread (one row, one presentation).
- Expiry: _park posts a plan-decision-specific recovery notice (re-assign /
force-resume) for an expired gate row.
- Resume path unchanged: the decision answer flows through submit_answer →
ResumeWorker → plan_gate_node with no resume-worker special-casing.
notify_question has no kind param in this tree, so the opener calls
ledger.post_question(kind=...) + ledger.set_channel_ref directly; the clarifier
path still uses notify_question unchanged.
Tests: gate row+presentation+threading, approve/request_changes/abandon via
submit_answer, single-open-gate skip, expiry notice. 1412 passed.
Replace the terminal review-cap PARK with a resumable human decision gate,
opt-in via build_graph(plan_gate=True) (default False → all existing P1/P2/P3
wiring unchanged).
- plan_gate_node interrupt()s mirroring the clarifier contract (same payload
keys → existing pending_question() extractor + turn-guarded ResumeWorker drive
it with zero special-casing) plus a kind="plan_decision" discriminator and the
plan + latest review findings as context.
- Decision contract {"decision": approve|request_changes|abandon, "notes": ...}:
approve → the same terminal state an auto-approved plan reaches (ACTIVE/BUILD);
request_changes → append a synthetic human verdict to review_verdicts (so the
planner's _format_review_feedback surfaces the notes) and loop back to PLAN;
abandon / unrecognized → terminal FAILED (safe default, never accidental
approve).
- Bounded termination: MAX_PLAN_GATE_VISITS=3 combined ceiling on plan_gate_visits
(new channel on PipelineState + TaskRecord); on exhaustion the gate goes
terminal PARKED ("revision ceiling reached") WITHOUT interrupting. Proven by a
loop-past-ceiling test.
Notes for the coordinator wiring (B2b): while suspended at the gate the status
channel still reads 'parked' (carried over from review_node's escalate branch) —
the load-bearing "awaiting decision, not terminal" signal is the live pending
interrupt + kind="plan_decision", NOT the status channel.
Tests: interrupt-at-cap, approve/request_changes(notes)/abandon routing,
ceiling-terminates, auto-approve still bypasses the gate. 1404 passed.
The plan-review gate (coming next) needs to tell its decision questions apart
from clarifier questions in the durable ledger. Add a `kind` column to
pending_questions (values 'clarify' | 'plan_decision').
- Fresh DBs: `kind TEXT NOT NULL DEFAULT 'clarify'` (+ CHECK) in the DDL.
- Live ledger: idempotent additive migration (SCHEMA_VERSION 3→4) — a guarded
ALTER (PRAGMA table_info) run from both migrate() and init_db; legacy rows
take the 'clarify' default, never null. (SQLite can't add a CHECK via ALTER,
so the migrated column is NOT NULL DEFAULT only; value constraint is enforced
on fresh DBs by the CHECK and on all writes by the typed helper.)
- ledger.post_question gains a keyword-only `kind="clarify"` (backward
compatible — existing callers unchanged); PendingQuestion.from_row reads it.
Tests: fresh-DB column+default, idempotent init_db, legacy-DB backfill to
'clarify', plan_decision round-trip. 1396 passed.
The planner's single-shot Claude call intermittently failed with "Reached
maximum number of turns (1)" — it needs slightly more headroom than the
clarifier to finish emitting its JSON. Phase A of the planner-reliability plan:
- plan_node now invokes with max_turns=4 (allowed_tools stays []; the extra
turns buy completion, not exploration).
- build_plan_prompt instructs the model to use no tools and return only JSON
(a tool_use would consume the single turn before the plan is emitted).
- plan_node auto-retries the model call exactly once on a TRANSIENT failure
(turn-cap exhaustion or an empty reply), and fails fast on DETERMINISTIC ones
(malformed JSON, missing/blank phases) — a retry would just reproduce those.
review_loop_llm.py is intentionally GPT-4.1 cross-family (no claude_invoke), so
it gets no turn-budget change. verifier_llm.py does use claude_invoke but is P3
build/verify scope — left for a follow-up.
Tests: max_turns passthrough; retry on turn-cap and on empty; no retry on
malformed JSON; the no-tools prompt line. 1387 passed.
- run-team.py: add the 'dispatch <thread_id>' operator command (P3 option-b).
The read-only box parks at DISPATCH; this completes it with a just-in-time
WRITE token: reads candidate_diff + scope from the checkpoint (or --diff/--scope
files), pushes the head branch + fires workflow_dispatch via dispatch_apply_verify,
prints the located run_id, and (--write-back) writes it into the task checkpoint
so VERIFY binds. +2 tests.
- OPERATOR-RUNBOOK: fix the misleading 'systemctl show -p Environment' check (it
does NOT show EnvironmentFile= vars) -> use /proc/<MainPID>/environ +
_p3_env_is_configured(); document the operator-initiated dispatch flow + the
fine-grained-token write-probe caveat.
Suite green, ruff clean. Branch only; not merged.
The no-write-token detector's test fixtures + a doc comment contained contiguous
'-----BEGIN ... PRIVATE KEY-----' literals that tripped the repo's gitleaks
pre-push backstop (a false positive on a secret-DETECTOR's own test data). Build
the PEM markers at runtime so the source carries no contiguous literal; the
runtime values are still full PEM blocks (what the detector under test sees).
No behavior change; 39 no-write-token tests pass.
High-recall /sh-security-review fan-out + proof-or-kill verifier found two
confirmed HIGH; both now closed (verified empirically against the working tree):
- LOGIC-RACE-01 (HIGH, CWE-835): the build-loop budget was structurally dead
(verifier read a shared wiring-time VerifierConfig.build_loops, always 0, so
the max_build_loops park never fired -> a perpetually-failing task looped
BUILD->DISPATCH->VERIFY forever, force-pushing + firing a CI run each round).
Threaded build_loops through durable PipelineState/TaskRecord; verifier reads
state.get('build_loops',0), writes the incremented count back on each FAIL, and
PARKS at max_build_loops. Parks after exactly N failures, never unbounded.
- SEC-01 (HIGH, CWE-532) + SEC-02 (MED, CWE-214): p3_rollback.sh echoed the live
App JWT to stdout in default dry-run and passed it as a gh argv literal. Added
redact_secrets (Bearer/Authorization/ghX_/PEM masking) through run_or_plan; the
App uninstall now uses curl -H @<0600 tempfile> (JWT never on argv), shredded
after. Empirical: app/incident/all dry-runs leak 0 JWT occurrences.
- SEC-03 (MED, CWE-798): assert_no_write_token now applies the PEM regex + the
configured App-ID to env/config VALUES (not just files) — an App private key
under a benign env name is caught.
- SEC-04 (LOW) + P3-IAC-08 (LOW): tightened the box GITHUB_TOKEN fallback /
value-scan; staged-only WARN on the live workflow revert.
Suite: 1382 passed, ruff clean. Branch only; not merged/deployed.
NOTE: re-verifier flagged SEC-01 as open by grepping COMMITTED blobs (the fix was
uncommitted working-tree state); independently confirmed closed empirically.
GPT-4.1 cross-family review (APPROVE, no critical/high) raised two MEDIUMs on the
rollback tooling; addressed both:
- require_keys: each restore_* asserts its required baseline keys up front and
refuses a PARTIAL (silently-weaker) restore. environment accepts ids OR logins
(equivalent); a missing protection.full now REFUSES the enforce_admins-only
degrade unless P3_ROLLBACK_ALLOW_PARTIAL=1 is set (loud DEGRADED warning).
- out-of-band ACK: an --apply that needs a MANUAL App neutralise (no APP JWT, or
action=out-of-band) refuses unless P3_ROLLBACK_OOB_ACK=1 — so the App is never
left un-neutralised without a conscious operator sign-off; with the ack the
other surfaces still restore.
Documented both env vars in usage. +3 tests (required-key refuse, partial-protection
ack, oob ack). Suite: 1362 passed, ruff clean.
Wire the box-side build->dispatch->verify run identity so the verifier gate
can bind to the CI run the dispatcher triggered:
- task_model: add run_id / ci_correlation_tag / dispatched_at to TaskRecord +
PipelineState (+ dict round-trip).
- dispatcher: RunLocator seam + DispatchResult; dispatch_apply_verify stamps a
dispatched-at watermark, fires, then resolves the run via the workflow
run-name (gh run list; the per-task_id concurrency group makes it
unambiguous). Fails closed to run_id=None.
- dispatch_invoker: persist run_id/dispatched_at/ci_correlation_tag into state.
- workflow: additive run-name surfacing inputs.task_id as the correlation key
(flagged for the C1 /sh-security-review + GPT-4.1 cross-review re-run).
- docs: P3-PHASE0-DESIGN.md records the async-resume design decision.
Part of Phase 0 (feat/agent-team-p3-box-integration). No behavior change on the
default path: P3 wiring is still opt-in/inert.
/sh-security-review confirmed SEC-DASH-001 (low): build_snapshot embedded str(exc)
of a sqlite/OS error into the /api/state payload, leaking the absolute DB path /
table names to the unauthenticated LAN surface. Return only type(exc).__name__
(matching the task_detail hardening); apply the same to /api/topology's error
branch (SEC-DASH-003). Full exception detail stays in server-side logs.
- read_transitions: catch sqlite3.Error (not just OperationalError) so a corrupt
ledger degrades to empty rather than raising into callers
- recorder open-row lookup: order by the monotonic transition_id (drop the
timestamp-format dependency)
- TransitionRecorder.close(): release the retained in-memory test connection
- dashboard task_detail: return only the exception TYPE, never str(exc) (a SQLite
message can carry the DB path)
- _instrument: coerce a status enum to .value defensively before the terminal check
- schema.migrate: document the ordering constraint for future ALTERs vs the
unconditional idempotent tail
Switch agent-team-status.service ExecStart from status_page.serve to
dashboard.serve (uvicorn serving web/dist + the JSON API). Document the WebUI in
the README: live auto-laid pipeline map, click-through task history, endpoints,
and the Mac-side npm build + rsync flow.
New web/ SPA (React 18 + Vite 5 + TypeScript + React Flow + dagre, all pinned, no
CDN). Dark, three-pane layout: summary top bar, filterable task list, center
auto-laid pipeline map, and a right drawer showing a selected task's history
through each node (timeline with timestamps/duration/cost + Q&A/verdicts/plan).
Map nodes color by live state, loop-backs render dashed, clicking a node filters
the list, selecting a task highlights its path. Polls /api/state every 4s;
topology fetched once. dist/ is gitignored (built on the Mac, rsynced).
topology.py derives the pipeline map (nodes/edges/trees) from the compiled
LangGraph via get_graph() + a NODE_META display sidecar, so new agent nodes
appear automatically and group into trees branching off intake. dashboard.py is a
new read-only FastAPI app (0.0.0.0:8770) serving /api/state (contract preserved +
per-node live state), /api/topology, and /api/task/{id} (validated, timeline +
cost join + partial fallback) plus the built SPA — kept SEPARATE from the authed
api.py. status_page.py is trimmed to the /api/state data layer; the inline
HTML/SVG renderer + stdlib server are retired.
Add a v3 task_transitions table (additive migration + startup assertion) and a
fail-soft, idempotent TransitionRecorder. build_graph gains an injected
transition_recorder that wraps every node via a functools.wraps'd _instrument
(signature-preserving so LangGraph still injects RunnableConfig); the coordinator
wires it. Records one row per node entry (idempotent under resume replay) and
closes the open row on terminal status. Backs the dashboard task-history view.
Two gaps surfaced by a live /new-task (task 9bce78ad): the plan was produced
and approved, but the "plan ready" notice posted top-level (not in the task
thread) and contained no plan to review.
1. THREADING — `run-team.py` `_build_notifiers` exposed `notify(message)` with
no `thread_ts`. The coordinator's `_emit` calls `notify(message,
thread_ts=root)`; that raised TypeError, and `_emit`'s fallback re-posted
TOP-LEVEL. So every lifecycle milestone (plan-ready / parked / failed) landed
unthreaded, despite the coordinator computing the root ts. (The clarifier
QUESTION threaded fine — different path.) Fix: the sink now accepts and
forwards `thread_ts` into the chat.postMessage payload (build_slack_poster
already forwards the key).
2. PRESENTATION — the plan-ready milestone was a bare one-liner. It now posts a
CONDENSED plan (summary + numbered phase names; step detail stays on the
status dashboard) via new `Coordinator._summarize_plan`, so the plan is
actually reviewable in-thread.
Tests: notify sink forwards thread_ts (and omits it for top-level); condensed
plan renders summary + phase names (not steps); malformed plan falls back;
plan-ready milestone threads under the task root AND carries the plan. 1193 pass.
A live task (d30b697c) on the R720 crashed the coordinator: the planner's
single-shot Claude call raised "Reached maximum number of turns (1)", the
exception propagated out of `drain_resumes` through the serve loop, and systemd
restarted the daemon — with no Slack notice, so the failure was silent.
Defense in depth:
1. invoker: `_collect_subscription_text` now tolerates the single-shot turn cap.
When the Agent SDK raises "Reached maximum number of turns" mid-stream it
salvages the assistant text already collected (the JSON the planner needs)
instead of propagating. An empty salvage or any non-turn error still raises.
2. coordinator: `drain_resumes` wraps the per-job `resume()` so an unhandled
node exception fails THAT task instead of the daemon — it supersedes the
answered question (so the startup recovery sweep cannot re-drive it into the
same crash on reboot), marks the task FAILED via `update_state` with a short
`failure_reason`, and surfaces an honest "❌ FAILED" line to Slack.
3. task_model: add `failure_reason` to PipelineState + TaskRecord (kept in sync)
so the terminal-failure detail persists as a real graph channel.
4. resume_worker: add `ResumeOutcome.FAILED`.
Tests: invoker salvage/re-raise/propagate paths; drain_resumes fails-not-crashes,
supersedes the question, notifies, and one failing task does not block others.
1189 passed.
- render_html includes the inline SVG map + inline fetch('/api/state') poller
+ tooltip mount + live clock, stays offline (no CDN), keeps the <noscript>
meta-refresh fallback, and renders the map even on a not-ok snapshot.
- snapshot_to_dict: expected top-level keys, one entry per STAGE in order,
per-phase grouping, awaiting-human classification (open pending_question),
per-node model/agent role labels, summary counts, not-ok serializability.
- Descriptions escaped in BOTH the HTML and the JSON-in-script seed
(</script><script> breakout neutralised; \u003c form present); the raw
/api/state JSON round-trips the description for textContent rendering.
- End-to-end snapshot_to_dict over the seeded ledger.
Turn the read-only status page into an auto-updating visual map of the
agent-team DAG (INTAKE -> CLARIFY <-> gate -> PLAN <-> REVIEW ->
[BUILD -> VERIFY -> DISPATCH] -> DONE), rendered as hand-rolled inline
SVG (no CDN/D3 — the R720 is offline/LAN-only).
- Stage model (STAGES) with per-node model/agent role labels (Claude/
GPT-4.1/Gemini/DeepSeek/Slack owner) and gated/role-node flags.
- Per-stage live state (idle/active/awaiting-human/parked) + count badge,
grouped by current_phase; awaiting-human = OPEN pending_question.
- GET /api/state JSON sidecar (snapshot_to_dict); inline vanilla-JS poller
fetches it every 4s and repaints node states/counts/cards/clock/tooltip
in place (no reload, hover/scroll survive). <noscript> meta-refresh
fallback retained.
- Hover/focus tooltip per node: short_id, description, status, waiting age.
- Existing table view kept as a detail section below the map.
- Read-only (mode=ro), fail-safe, no secrets; descriptions escaped for both
HTML and the JSON-in-script seed (< / > -> \uXXXX), DOM via textContent.
The cfn-lint template-discovery step in review.sh piped the repo's whole
matched-file list into 'xargs -I{} sh -c "grep -l {}"'. On the orchestrator
monorepo — especially from a deep worktree path, where every matched path is a
long absolute path — xargs -I{} packs all paths into one assembled command and
aborts with 'xargs: command line cannot be assembled, too long'. The subprocess
exits non-zero having emitted ZERO findings, so the global pre-push hook BLOCKS
every push (agents were working around it with --no-verify).
Fix: switch the grep stage to NUL-delimited, un-batched xargs
(find ... -print0 | xargs -0 grep -lE ...). xargs -0 (no -I) splits the input
across multiple grep invocations, so the argv never exceeds ARG_MAX; grep -l
reports the same matching files as the old per-file grep, and -print0/-0 is safe
for paths with spaces/newlines. The first 'xargs -I{} find {}' is kept (find
needs the start path before its expression) and is bounded by the scope-path
count, so it is not an overflow source. Trailing '|| true' preserves the old
no-match/no-files semantics (TPLS = list-of-templates or empty, never fails).
Purely an argv-batching fix: the scanned file set, findings, and exit codes are
unchanged. Verified exit-code-identical: clean tree -> exit 0 (PASS); planted
GitHub PAT + RSA private key -> exit 1 (BLOCK, gitleaks high); planted CFN
template with a cfn-lint error on a deeply-nested path -> exit 1 (cfn-lint flags
it, no overflow). The previously-overflowing command now completes clean.
Codifies the hardened deploy-before-merge procedure for ongoing agent-team
code changes, generalizing the one-off deploy-r720-ws-rollout.sh. Prevents the
two self-inflicted live crash-loops:
- Whole agent_team/ package rsync (never per-file, which misplaces e.g.
nodes/planner.py at the package root -> ImportError/TypeError crash-loop).
- Snapshot HARD GATE (Adam's Hyper-V step; Claude ssh reaches only the guest)
+ ledger backup before any change.
- Pre-restart import sanity, then mandatory verify-after (is-active==active,
NRestarts didn't climb, ~6 threads, clean journal) with rollback guidance
on failure.
Idempotent, fails loudly. Optional SYNC_DEPS / SYNC_HANDBOOK / RESTART_STATUS.
Drives the new /sh-deploy-r720 skill.
WS Slack-UX Feature 2. When the inbound listener acts on an answer in a task
thread, it adds a 👍 reaction to that reply so the human sees the machine
received it.
- SlackListener gains an optional reactor seam; handle_event reacts to the
inbound reply message (channel + event ts) AFTER the AUTHZ-01 owner check
passes — a non-owner message is rejected and never reacted to. Best-effort:
any reaction failure (notably a missing scope) is swallowed and never breaks
handle_event or the listen loop.
- /new-task is NOT reacted to (a slash command has no reactable message); its
"📥 Task received" root post is the acknowledgement.
- build_slack_reactor wraps WebClient.reactions_add(name="thumbsup"); the
default listener factory wires it best-effort from SLACK_BOT_TOKEN.
- Adds reactions:write to the bot scopes in agent-team-manifest.json.
NOTE: the new reactions:write scope requires Adam to re-apply the manifest to
app A0BCC7TTU66 and reinstall the app. Until then reactions.add returns
missing_scope, which the listener swallows (the reaction silently no-ops) —
answer handling is unaffected.
AUTHZ-01 and the first-answer-wins CAS remain unchanged.
WS Slack-UX Feature 1. A /new-task task now maps to ONE Slack thread instead of
several top-level messages.
- /new-task posts an immediate root "📥 Task received: …" ack and captures its
ts (root_ts); this is the instant acknowledgement.
- root_ts is plumbed into start: new PipelineState/TaskRecord channel
slack_thread_ts, seeded by graph.start_task and threaded through
Coordinator.start_task. The NewTaskCallback is now (task_text, via, root_ts).
- All clarifier questions for the task post as THREADED REPLIES under root_ts
(chat.postMessage thread_ts=root_ts), and each question's ledger channel_ref
is set to root_ts (NOT the reply's own ts). Because answer-mapping resolves a
reply via find_open_question_by_channel_ref(thread_ts), a reply in the root
thread (thread_ts==root_ts) maps to the task's currently-open question with NO
change to the mapping logic or the first-answer-wins CAS. The open-only
partial-unique index still holds (one open question per task at a time).
- Lifecycle milestones (parked / plan-ready / needs-input) and follow-up
questions thread under root_ts too; the notify sink gained an optional
thread_ts kwarg (degrades to top-level on a sink that doesn't accept it).
notify failures still never break tick.
- SlackTransport.post_question + the live poster accept/forward thread_ts.
- No root_ts (non-/new-task origin) ⇒ top-level posts exactly as before.
AUTHZ-01 (owner-allowlist-first, fail-closed) and the atomic open→answered
compare-and-set are unchanged.
Adds plumbing for the inbound-ack reactor seam used by Feature 2 (dormant until
a reactor is injected). Tests cover thread_ts forwarding, channel_ref=root_ts,
graph seeding, and coordinator threading.
Two root causes behind 'every task parks, and slowly':
1. SINGLE-SHOT INVOKER: subscription Claude calls ran as 40-turn, tool-enabled
agentic sessions (--max-turns 40, $2 budget) for what are pure reasoning->JSON
completions — minutes-long, and Claude wandered/returned unparseable output.
_DEFAULT_MAX_TURNS 40->1 + allowed_tools=[] -> fast deterministic single turn.
2. PLANNER FEEDBACK KEY MISMATCH: the review stage writes verdict/findings, but
_format_review_feedback read decision/notes/comment (never present) -> the
planner re-planned with EMPTY feedback, re-introduced the rejected flaw
('assumptions persist'), hit the review cap, parked. Now reads verdict/findings
(old keys kept as fallback) so GPT-4.1's objections reach the re-plan.
Plus: park notifications infer the phase the task was IN (review/plan/clarify)
instead of the terminal 'parked'.
Tests: planner real-verdict-keys regression + coordinator phase-inference. 1149 pass.
Park notifications were opaque ('Task 6c3c3202 parked — needs your attention'):
no idea what the task is, where it got to, or what's blocking it. Now each
message names the task DESCRIPTION (not just the short id), the PHASE it reached,
and the actual BLOCKER — _summarize_blocker() pulls the last review_verdict's
findings (the GPT-4.1 REQUEST_CHANGES text, collapsed + truncated), falling back
to 'no plan built' / 'revision cap hit'. Applies to parked + needs-more-input +
plan-ready emits. +1 test (description + phase + blocker present). 1148 passed.
The bot only ever posted clarifier questions; parks/completions were silent and
multi-turn follow-up questions were never posted during normal operation (only
the startup recover sweep posted them). So an answered task was a black box.
- Coordinator gains a 'notify' sink + _emit() (guarded, never breaks the loop).
- tick() now runs _post_resume_followups(results) after the drain: for each
resumed thread it (a) POSTS a newly-pending clarifier question (fixes silent
multi-turn — the drain path left it unposted) and emits 'needs more input';
else emits 'parked — needs attention' or 'plan ready for review' from the
settled state.
- run-team serve wires notify -> Slack channel (build_slack_poster) and an
alarm_hook that logs the deadline-park WARNING AND posts a parked notice.
Live-Slack only; dry-run/non-Slack/no-channel = silent (None), no token needed.
Tests: +4 (needs-input/parked/plan-ready emits + notify-failure swallow);
_FakeCoordinator gains notify/alarm_hook. 1147 passed, ruff clean.
/new-task (and every intake: GitHub issue, /sh-assign-task) reached the clarifier
with NO description -> the clarifier asked 'no task description provided'. Root
cause: coordinator.start_task only LOGGED task_text (a P1-era decision when the
deterministic clarifier didn't consume a description), graph.start_task took no
task arg, and PipelineState/TaskRecord had no 'task' channel at all.
Fix: add a first-class 'task' field to PipelineState + TaskRecord (+ round-trip
in task_from_dict); graph.start_task seeds task into the initial invoke (persists
through intake_node's partial-state return into CLARIFY); coordinator.start_task
passes task=task_text. The clarifier already reads state['task'] via
_task_description, so it now sees the real description.
Test: start_task(task='build a login form') -> suspended CLARIFY state carries
task. 1143 passed, ruff clean.
'the app did not respond': serve() registered @app.action/@app.event but NO
@app.command handler, so Bolt never acked the /new-task slash command within
Slack's ~3s deadline. Add an @app.command(/new-task) handler that ack()s first,
re-stamps type:slash_commands onto Bolt's inner command body (Bolt strips the
Socket Mode envelope type that _is_new_task_command/_discriminating_type expect),
then forwards to handle_event (AUTHZ-01 + new-task dispatch). The resulting
payload shape is the one already covered by test_new_task_calls_callback_*.
The agent-team venv was missing langchain-anthropic/-openai/-google-genai/
-community, so models.py failed to import and the in-process GPT-4.1 review /
Gemini scan / DeepSeek build silently fail-closed to REQUEST_CHANGES (the
non-Claude models never ran on the box). Step 3 now installs the full pinned
requirements.txt into the venv instead of just fastapi/uvicorn. Installed +
verified live on the box: all three model factories construct.
Gap-audit findings:
- G1 (blocks-feature): make_cross_reviewer_invoker / make_fast_coder_invoker did
'from models import' without putting the orchestrator root on sys.path. The
run-team serve daemon only bootstraps agent-team/, so on the live box every
GPT-4.1 plan review hit ModuleNotFoundError -> review_plan's blanket except
silently fail-closed to REQUEST_CHANGES (GPT-4.1 never actually ran). Both
in-process invokers now call invoker_multi._ensure_orchestrator_on_path()
before the deferred import. WS1 introduced this when it swapped the review
default from the subprocess invoker to in-process.
- G4 (degrades): the handbook context_provider was wired into the planner only;
default_clarify_node_factory now accepts + forwards it, and run-team wires it
into build_clarify_node too, so clarifying questions are handbook-aware.
Tests: +2 regression tests (path-bootstrap, clarifier threading); _FakeCoordinator
gains build_clarify_node. 1142 passed, ruff clean.
The slack_listener handles {type:slash_commands, command:/new-task} but the app
manifest declared no slash commands and no 'commands' scope, so Slack never
offered /new-task (the command can't be invoked). Add the slash command +
commands bot scope. Socket Mode delivers it over the socket (no request URL).
APPLY: update the app A0BCC7TTU66 from this manifest + reinstall to pick up the
new scope.