Commit graph

178 commits

Author SHA1 Message Date
022befa796 feat(agent-team): Tailwind + shadcn foundation for dashboard redesign
Adds the styling foundation for the UI redesign (vibe-kanban / Magentic-UI /
Langflow references): Tailwind CSS + shadcn/ui primitives, HSL design tokens
(index.css) ported from theme.css, @/ alias, and the App.tsx Board|Pipeline tabs
shell. Extends StateResponse with the backend's existing stages[] for the Board.

theme.css kept transiently until components migrate to Tailwind. No backend
changes. Component fan-out (Board, MapNode, TaskDrawer, TopBar) follows.
2026-06-24 13:35:56 -04:00
21f2fe54c1 fix(agent-team): builder uses agentic invoker config (read-only tools + turn headroom)
The Plane-2 builder (default_diff_builder) called claude_invoke with no
overrides, inheriting the subscription invoker's single-shot defaults
(max_turns=1, allowed_tools=[]). Diff synthesis is agentic, so the call
died with 'Reached maximum number of turns (1)' and every task failed at
phase=build.

- invoker.py: thread allowed_tools through subscription_invoker and
  _collect_subscription_text (default None -> []), so callers can opt in;
  single-shot reasoning nodes are unchanged.
- builders.py: default_diff_builder now passes max_turns=8, a read-only
  tool allowlist (Read/Grep/Glob), and budget_usd=4.0. No write tools --
  the builder returns the diff as data and performs no repo writes (D2/D11).
- Tests: builder agentic-config passthrough; invoker allowed_tools thread +
  tool-less default guard (so future nodes must opt in explicitly).

Closes #60
2026-06-24 13:13:28 -04:00
Adam Moussa
0301b8e4e7
Merge pull request #59 from Sea-Haven-Industries/feature/agent-team-dashboard-polish
feat(agent-team): cleaner retry-loop display + human-readable task history
2026-06-24 12:50:11 -04:00
34c5f1d75e feat(agent-team): cleaner retry-loop display + human-readable task history
Two dashboard-SPA refinements (frontend only):

1. Retry loops (plan<->review, build<->verify) no longer draw a backward arc over
   the forward edge (the 'circular arrows'). Loop-backs are excluded from the
   default render and from the dagre layout; instead the source node shows a small
   ↺ chip ('can send work back to ...'), and the actual return arc is drawn only
   when a selected task ACTUALLY looped it (computed from its timeline), highlighted
   on that task's path.

2. Q&A / Review Verdicts / Plan render human-readably instead of JSON blobs:
   react-markdown (no rehype-raw -> raw HTML escaped, XSS-safe) renders findings/
   answers/summary; verdicts as cards (badge + round + outcome), Q&A as per-turn
   cards (string + dict shapes), plan as summary + phase/step lists.

16 frontend tests pass (incl. loopback default-off/on-when-looped, markdown bold,
and a no-raw-HTML XSS guard); typecheck + build clean.
2026-06-24 12:42:55 -04:00
Adam Moussa
7d53d6b1f3
Merge pull request #58 from Sea-Haven-Industries/feat/agent-team-plan-gate
feat(agent-team): planner reliability + resumable plan-review human gate
2026-06-24 12:40:52 -04:00
3ebeacdf31 fix(agent-team): init_db drives migrate() so the version stamp actually advances
init_db's own schema_meta write was ON CONFLICT DO NOTHING, and the daemon
(Coordinator.setup) calls init_db, never migrate() — so on an existing ledger
the column was ensured but schema_version was never advanced (observed live:
kind column present, schema_meta stuck at 3). migrate() already upserts the
version correctly but was effectively dead code (no production caller).

init_db now ends by calling migrate(conn), which steps the version and runs any
version-gated steps. Idempotent — re-running the create/ensure statements is
harmless. Regression test: an existing v3-stamped DB run through init_db now
reports schema_version == SCHEMA_VERSION (4) and has the kind column. 1491 passed.
2026-06-24 12:18:22 -04:00
71edeb3f3f fix(agent-team): review-round cap counts only reviewer verdicts (LOGIC-04)
_review_round_index counted EVERY review_verdicts entry, including the synthetic
human-gate verdict graph._apply_plan_decision folds in on a "request changes"
(reviewer == "human_plan_gate"). That inflated the count so a revised plan could
escalate prematurely without a fresh adversarial review.

Now counts only reviewer-authored verdicts: a new _is_reviewer_verdict excludes
entries tagged reviewer=="human_plan_gate" (read from the verdict dict's own
field — no graph.py import). A human request_changes now grants the revised plan
a fresh reviewer-round budget. Termination still bounded by MAX_PLAN_GATE_VISITS
(each request_changes consumes one gate visit). 1490 passed.
2026-06-24 11:57:03 -04:00
42f2438d0c fix(agent-team): make planner convergence cap count correctly (LOGIC-03)
_revision_count read verdict.get("decision"), but verdicts are keyed "verdict"
(both reviewer and synthetic human-gate), so the count was always 0 and the
plan_node MAX_PLAN_REVISIONS self-park was dead code. Now reads "verdict" first
(fallback "decision"), matching _format_review_feedback's precedence; the
existing .strip().upper()==_REQUEST_CHANGES compare covers both request_changes
and REQUEST_CHANGES. Counts reviewer + human request_changes.

Cap composition: the review-loop round cap and MAX_PLAN_GATE_VISITS govern the
live loops; the planner MAX_PLAN_REVISIONS is now a correct backstop (was inert),
not a behavior change to the gate. Tests drive the real "verdict" key and prove
the previously-dead park fires. 1487 passed.
2026-06-24 11:53:12 -04:00
48a81c0802 fix(agent-team): centralize safe decision mapping + remove free-text abandon hair-trigger
Security-review follow-up (LOGIC-01/02/05, all confirmed correctness).

- New transport-neutral `decisions.normalize_decision(raw, *, allow_abandon)` is
  the single source of truth: approve-allowlist→approve; abandon-allowlist→abandon
  ONLY when allow_abandon; everything else (prose, empty, abandon-verbs when
  disallowed) → request_changes with the full reply as notes; idempotent on an
  already-formed decision dict. slack_adapter.map_plan_decision is now a thin
  wrapper (default allow_abandon=True, no caller churn).
- LOGIC-01/02: graph._parse_decision now delegates to normalize_decision (was:
  any unrecognized verb → abandon → FAILED). The graph is now the universal safe
  backstop, so EVERY writer that bypassed the listener mapping — operator CLI
  answer_on_behalf (raw), Coordinator.submit_answer (raw), the recovery sweep —
  loops back on prose instead of silently FAILing the task. Explicit abandon
  still abandons (preserves the confirmed-button path).
- LOGIC-05: the Slack FREE-TEXT reply path maps with allow_abandon=False, so a
  bare "cancel"/"stop"/"abandon" typed in-thread → request_changes (never
  terminal abandon); abandon stays reachable only via the confirm-guarded button.

Tests: graph unrecognized→loops-back (not FAILED), operator raw-prose→request_
changes, free-text destructive verbs→request_changes vs button→abandon,
normalizer idempotency. 1484 passed.
2026-06-24 11:48:04 -04:00
f46d691e36 docs(agent-team): document the plan-review decision gate (Phase C)
README: two human gates (clarifier + plan-decision), the approve/request-changes/
abandon verbs, free-text-defaults-to-request-changes, MAX_PLAN_GATE_VISITS, and
the planner max_turns reliability fix.
OPERATOR-RUNBOOK: how the gate appears in Slack, the three decision paths
(buttons/modal/free-text), the single-open-gate invariant, ceiling→PARKED, and
the 24h expiry→PARKED→recovery (re-assign / force-resume).
DEPLOY-R720: the pending_questions.kind ledger migration (SCHEMA_VERSION→4,
idempotent additive ALTER on startup) + rollback (restore the ledger backup
before restart if the migration fails).
2026-06-24 11:30:54 -04:00
082e45bf88 test(agent-team): end-to-end plan-review gate composition (Phase C)
Hermetic e2e tests driving the WHOLE stack composed together — real
build_graph(plan_gate=True) + real Coordinator + real SQLite ledger + real
review_loop router, with only the LLM nodes stubbed — through the daemon API
(start_task/submit_answer/tick), never nodes directly.

Flows: approve settles at BUILD; request-changes via RAW PROSE through the real
SlackListener -> _resolve_payload -> map_plan_decision proves the prose maps to
request_changes (NOT FAILED) and the notes reach the planner; abandon -> FAILED;
repeated request_changes terminates at MAX_PLAN_GATE_VISITS -> PARKED; and a
legacy (no-kind) ledger migrates in place then routes clarify vs plan_decision
correctly. 1459 passed.
2026-06-23 21:03:54 -04:00
f4de957915 feat(agent-team): Slack decision surface for the plan-review gate (Phase B3)
Turn a human's Slack interaction at the plan gate into a structured decision the
graph can route, with the kind-aware mapping that closes a silent-FAIL hazard.

- KIND-AWARE NORMALIZATION (load-bearing): map_plan_decision() in slack_adapter
  maps a reply to {"decision","notes"} — approve ∈ {approve,approved,yes,ok,lgtm,
  ship}; abandon ∈ {abandon,reject,cancel,stop,kill}; EVERYTHING ELSE →
  request_changes with the full reply as notes (never accidental abandon). Wired
  in the listener's _resolve_payload for plan_decision rows ONLY (clarify passes
  through). Without this, arbitrary change-notes hit the graph's
  unrecognized-verb→FAILED path and silently fail the task. Anti-FAIL tests
  assert prose → request_changes (!= abandon) at both the mapper and the
  end-to-end listener seam; a regression test guards clarify pass-through.
  New find_open_question_kind_by_channel_ref (anti-replay, status='open') powers
  the thread-reply fallback's kind lookup.
- BUTTONS + MODAL: build_plan_decision_blocks() renders Approve (primary) /
  Request changes / Abandon (danger+confirm); question_id double-anchored in
  message metadata AND each button value ("<verb>:<question_id>"). Approve/abandon
  submit via the existing @app.action(.*); request_changes has a dedicated
  handler that AUTHORIZES before views_open (proven by test) and opens a notes
  modal (private_metadata carries the id) → view_submission → request_changes +
  notes. Free-text reply stays the always-available equal path. AUTHZ-01 ordering
  preserved.
- No manifest change (views.open needs no extra scope).

1454 passed (1412 + 42).
2026-06-23 21:03:54 -04:00
ba4fe68fdb feat(agent-team): wire the plan-review gate into the coordinator (Phase B2b)
Connect the graph plan-gate (B2a) to the durable ledger + Slack presentation.

- setup() passes build_graph(plan_gate=True) only on the wired review path
  (plan_gate flag ANDed with review_node present); P1/stub paths force it off.
- _post_resume_followups detects a settled interrupt by payload
  kind == PLAN_DECISION_KIND (NOT status, which still reads 'parked' at the
  gate per B2a) and posts the decision gate: opens a pending_questions row with
  kind='plan_decision' (24h deadline, threaded, channel_ref = root ts) and
  presents _summarize_plan + findings + reply instructions, truncated to a
  ~2700-char Slack budget. A clarify/legacy interrupt keeps the existing path.
- Single-open-gate invariant: the opener skips if any open row already exists
  for the thread (one row, one presentation).
- Expiry: _park posts a plan-decision-specific recovery notice (re-assign /
  force-resume) for an expired gate row.
- Resume path unchanged: the decision answer flows through submit_answer →
  ResumeWorker → plan_gate_node with no resume-worker special-casing.

notify_question has no kind param in this tree, so the opener calls
ledger.post_question(kind=...) + ledger.set_channel_ref directly; the clarifier
path still uses notify_question unchanged.

Tests: gate row+presentation+threading, approve/request_changes/abandon via
submit_answer, single-open-gate skip, expiry notice. 1412 passed.
2026-06-23 21:03:53 -04:00
67b0f4c6ae feat(agent-team): resumable plan-review gate in the graph (Phase B2a)
Replace the terminal review-cap PARK with a resumable human decision gate,
opt-in via build_graph(plan_gate=True) (default False → all existing P1/P2/P3
wiring unchanged).

- plan_gate_node interrupt()s mirroring the clarifier contract (same payload
  keys → existing pending_question() extractor + turn-guarded ResumeWorker drive
  it with zero special-casing) plus a kind="plan_decision" discriminator and the
  plan + latest review findings as context.
- Decision contract {"decision": approve|request_changes|abandon, "notes": ...}:
  approve → the same terminal state an auto-approved plan reaches (ACTIVE/BUILD);
  request_changes → append a synthetic human verdict to review_verdicts (so the
  planner's _format_review_feedback surfaces the notes) and loop back to PLAN;
  abandon / unrecognized → terminal FAILED (safe default, never accidental
  approve).
- Bounded termination: MAX_PLAN_GATE_VISITS=3 combined ceiling on plan_gate_visits
  (new channel on PipelineState + TaskRecord); on exhaustion the gate goes
  terminal PARKED ("revision ceiling reached") WITHOUT interrupting. Proven by a
  loop-past-ceiling test.

Notes for the coordinator wiring (B2b): while suspended at the gate the status
channel still reads 'parked' (carried over from review_node's escalate branch) —
the load-bearing "awaiting decision, not terminal" signal is the live pending
interrupt + kind="plan_decision", NOT the status channel.

Tests: interrupt-at-cap, approve/request_changes(notes)/abandon routing,
ceiling-terminates, auto-approve still bypasses the gate. 1404 passed.
2026-06-23 21:03:53 -04:00
1b4d30e47f feat(agent-team): add pending_questions.kind discriminator + migration (Phase B1)
The plan-review gate (coming next) needs to tell its decision questions apart
from clarifier questions in the durable ledger. Add a `kind` column to
pending_questions (values 'clarify' | 'plan_decision').

- Fresh DBs: `kind TEXT NOT NULL DEFAULT 'clarify'` (+ CHECK) in the DDL.
- Live ledger: idempotent additive migration (SCHEMA_VERSION 3→4) — a guarded
  ALTER (PRAGMA table_info) run from both migrate() and init_db; legacy rows
  take the 'clarify' default, never null. (SQLite can't add a CHECK via ALTER,
  so the migrated column is NOT NULL DEFAULT only; value constraint is enforced
  on fresh DBs by the CHECK and on all writes by the typed helper.)
- ledger.post_question gains a keyword-only `kind="clarify"` (backward
  compatible — existing callers unchanged); PendingQuestion.from_row reads it.

Tests: fresh-DB column+default, idempotent init_db, legacy-DB backfill to
'clarify', plan_decision round-trip. 1396 passed.
2026-06-23 21:03:53 -04:00
5332df60d9 fix(agent-team): planner turn headroom + classified retry-once (Phase A)
The planner's single-shot Claude call intermittently failed with "Reached
maximum number of turns (1)" — it needs slightly more headroom than the
clarifier to finish emitting its JSON. Phase A of the planner-reliability plan:

- plan_node now invokes with max_turns=4 (allowed_tools stays []; the extra
  turns buy completion, not exploration).
- build_plan_prompt instructs the model to use no tools and return only JSON
  (a tool_use would consume the single turn before the plan is emitted).
- plan_node auto-retries the model call exactly once on a TRANSIENT failure
  (turn-cap exhaustion or an empty reply), and fails fast on DETERMINISTIC ones
  (malformed JSON, missing/blank phases) — a retry would just reproduce those.

review_loop_llm.py is intentionally GPT-4.1 cross-family (no claude_invoke), so
it gets no turn-budget change. verifier_llm.py does use claude_invoke but is P3
build/verify scope — left for a follow-up.

Tests: max_turns passthrough; retry on turn-cap and on empty; no retry on
malformed JSON; the no-tools prompt line. 1387 passed.
2026-06-23 21:03:53 -04:00
Adam Moussa
addf23e883
Merge pull request #56 from Sea-Haven-Industries/feat/agent-team-p3-box-integration
feat(agent-team): P3 box-side build→dispatch→verify integration
2026-06-23 21:01:00 -04:00
a8f00ff676 feat(agent-team): operator dispatch command + runbook fixes
- run-team.py: add the 'dispatch <thread_id>' operator command (P3 option-b).
  The read-only box parks at DISPATCH; this completes it with a just-in-time
  WRITE token: reads candidate_diff + scope from the checkpoint (or --diff/--scope
  files), pushes the head branch + fires workflow_dispatch via dispatch_apply_verify,
  prints the located run_id, and (--write-back) writes it into the task checkpoint
  so VERIFY binds. +2 tests.
- OPERATOR-RUNBOOK: fix the misleading 'systemctl show -p Environment' check (it
  does NOT show EnvironmentFile= vars) -> use /proc/<MainPID>/environ +
  _p3_env_is_configured(); document the operator-initiated dispatch flow + the
  fine-grained-token write-probe caveat.

Suite green, ruff clean. Branch only; not merged.
2026-06-23 20:49:09 -04:00
ee97a69154 test(agent-team): assemble PEM test data at runtime (avoid gitleaks FP)
The no-write-token detector's test fixtures + a doc comment contained contiguous
'-----BEGIN ... PRIVATE KEY-----' literals that tripped the repo's gitleaks
pre-push backstop (a false positive on a secret-DETECTOR's own test data). Build
the PEM markers at runtime so the source carries no contiguous literal; the
runtime values are still full PEM blocks (what the detector under test sees).
No behavior change; 39 no-write-token tests pass.
2026-06-23 19:52:04 -04:00
00c51192c8 fix(agent-team): remediate C1 security-review BLOCK (2 HIGH + MED/LOW)
High-recall /sh-security-review fan-out + proof-or-kill verifier found two
confirmed HIGH; both now closed (verified empirically against the working tree):

- LOGIC-RACE-01 (HIGH, CWE-835): the build-loop budget was structurally dead
  (verifier read a shared wiring-time VerifierConfig.build_loops, always 0, so
  the max_build_loops park never fired -> a perpetually-failing task looped
  BUILD->DISPATCH->VERIFY forever, force-pushing + firing a CI run each round).
  Threaded build_loops through durable PipelineState/TaskRecord; verifier reads
  state.get('build_loops',0), writes the incremented count back on each FAIL, and
  PARKS at max_build_loops. Parks after exactly N failures, never unbounded.
- SEC-01 (HIGH, CWE-532) + SEC-02 (MED, CWE-214): p3_rollback.sh echoed the live
  App JWT to stdout in default dry-run and passed it as a gh argv literal. Added
  redact_secrets (Bearer/Authorization/ghX_/PEM masking) through run_or_plan; the
  App uninstall now uses curl -H @<0600 tempfile> (JWT never on argv), shredded
  after. Empirical: app/incident/all dry-runs leak 0 JWT occurrences.
- SEC-03 (MED, CWE-798): assert_no_write_token now applies the PEM regex + the
  configured App-ID to env/config VALUES (not just files) — an App private key
  under a benign env name is caught.
- SEC-04 (LOW) + P3-IAC-08 (LOW): tightened the box GITHUB_TOKEN fallback /
  value-scan; staged-only WARN on the live workflow revert.

Suite: 1382 passed, ruff clean. Branch only; not merged/deployed.
NOTE: re-verifier flagged SEC-01 as open by grepping COMMITTED blobs (the fix was
uncommitted working-tree state); independently confirmed closed empirically.
2026-06-23 19:52:04 -04:00
c4bea7270b harden(agent-team): fold C1 cross-review MEDIUMs into p3_rollback.sh
GPT-4.1 cross-family review (APPROVE, no critical/high) raised two MEDIUMs on the
rollback tooling; addressed both:
- require_keys: each restore_* asserts its required baseline keys up front and
  refuses a PARTIAL (silently-weaker) restore. environment accepts ids OR logins
  (equivalent); a missing protection.full now REFUSES the enforce_admins-only
  degrade unless P3_ROLLBACK_ALLOW_PARTIAL=1 is set (loud DEGRADED warning).
- out-of-band ACK: an --apply that needs a MANUAL App neutralise (no APP JWT, or
  action=out-of-band) refuses unless P3_ROLLBACK_OOB_ACK=1 — so the App is never
  left un-neutralised without a conscious operator sign-off; with the ack the
  other surfaces still restore.
Documented both env vars in usage. +3 tests (required-key refuse, partial-protection
ack, oob ack). Suite: 1362 passed, ruff clean.
2026-06-23 19:52:04 -04:00
cb84629c2b feat(agent-team): P3 Phases A/B/E — safety tooling, wiring, docs
Phase A (safety):
- scripts/p3_rollback.sh (+test): restore all privileged P3 surfaces from a
  recorded baseline; --dry-run default, --apply gated. Correct App-uninstall
  (App JWT) model; per-task env-reviewer restore by numeric id; real
  protection post-restore assert (normalize reads argv, fails loud, divergent
  state exits non-zero — regression-tested). KNOWN-LIMITATIONS header flags the
  branch-protection GET->PUT transform + live-validation for the C1 gate.
- scripts/assert_no_write_token.py (+test): box/CI audit that no write token
  (incl. ghu_/ghr_ prefixes + App PEM) lives on the box.
- draft_pr_monitor.py (+test): runaway (>3/15min) + stale (7d) draft-PR sweep,
  wired into tick() and bound a read-only provider in serve.

Phase B (wiring): systemd EnvironmentFile P3 vars + verification; new-draft-PR
lifecycle notice.

Phase E (docs): P3-LIVE-FLIP-PLAN/README/ci-README reflect CI-live-since-6/22 +
box-integration; runbook consolidated (rollback Incident 7 + box-env wiring);
removed a stray duplicate runbook.

Suite: 1360 passed, ruff clean. Branch only; not merged/deployed.
REMAINING HUMAN GATES: C1 /sh-security-review + GPT-4.1 cross-review on the
enabled workflow + rollback script; D box deploy + smoke + merge.
2026-06-23 19:52:04 -04:00
1f8c7e1ee3 fix(agent-team): close P3 async-resume BLOCKs (durable CI-watcher wiring)
Remediates the Phase-0 adversarial BLOCKs:
- Durable ci_pending_provider (_enumerate_ci_pending) walks the LangGraph
  SQLite checkpointer to enumerate threads suspended at VERIFY awaiting CI;
  re-derives across restart. Excludes human-clarify gates + advanced threads.
- run-team serve wires ci_pending_provider + ci_poller + ci_timeout ONLY on a
  configured box; inert path unchanged. Closes the 'VERIFY suspended forever'
  defect: tick()->_ci_watch resumes on terminal CI or timeout-parks.
- CI resume routes through the single-flight, turn-guarded ResumeWorker.
- FIXes: run-locator skips cancelled/stale runs on rapid re-dispatch; inert-mode
  wording matches behavior; added node-level fail-closed + spurious-resume tests.
- end-to-end async-resume proof (test_p3_async_resume.py, real checkpointer).

Suite: 1270 passed, ruff clean. Branch only; not merged/deployed.
2026-06-23 19:52:04 -04:00
f0c5cfe57f feat(agent-team): P3 Phase-0 box-side build->dispatch->verify (WIP)
0c-binding: per-task expected_run_id bound from state (gate rejects substituted
  run_id; None -> BLOCK, never vacuous pass).
0e: fail-safe serve default (failsafe_production_p3_wiring) — inert on
  unprovisioned env (one WARNING + one #agent-team notice), never crash-loops.
0a: reorder P3 subgraph BUILD -> DISPATCH -> VERIFY (preserves _instrument).
0d: ci_watcher engine + VERIFY interrupt()-wait (async resume-on-CI-complete).

KNOWN-OPEN (adversarial review BLOCKs, to remediate next):
- CI-watcher not wired into run-team serve (ci_pending_provider/ci_poller None)
  -> a VERIFY-suspended task never resumes/parks.
- no durable ci_pending_provider enumerating threads suspended at VERIFY.
Branch only; not merged, not deployed.
2026-06-23 19:52:04 -04:00
c3e935c904 feat(agent-team): capture dispatched run_id for P3 box-side verify
Wire the box-side build->dispatch->verify run identity so the verifier gate
can bind to the CI run the dispatcher triggered:

- task_model: add run_id / ci_correlation_tag / dispatched_at to TaskRecord +
  PipelineState (+ dict round-trip).
- dispatcher: RunLocator seam + DispatchResult; dispatch_apply_verify stamps a
  dispatched-at watermark, fires, then resolves the run via the workflow
  run-name (gh run list; the per-task_id concurrency group makes it
  unambiguous). Fails closed to run_id=None.
- dispatch_invoker: persist run_id/dispatched_at/ci_correlation_tag into state.
- workflow: additive run-name surfacing inputs.task_id as the correlation key
  (flagged for the C1 /sh-security-review + GPT-4.1 cross-review re-run).
- docs: P3-PHASE0-DESIGN.md records the async-resume design decision.

Part of Phase 0 (feat/agent-team-p3-box-integration). No behavior change on the
default path: P3 wiring is still opt-in/inert.
2026-06-23 19:52:04 -04:00
Adam Moussa
b7b9b92bfe
Merge pull request #55 from Sea-Haven-Industries/feature/agent-team-webui-makeover
feat(agent-team): WebUI makeover — branching pipeline map + click-through task history
2026-06-23 19:50:17 -04:00
a632a7df7d fix(agent-team): scrub exception text from /api/state + /api/topology errors
/sh-security-review confirmed SEC-DASH-001 (low): build_snapshot embedded str(exc)
of a sqlite/OS error into the /api/state payload, leaking the absolute DB path /
table names to the unauthenticated LAN surface. Return only type(exc).__name__
(matching the task_detail hardening); apply the same to /api/topology's error
branch (SEC-DASH-003). Full exception detail stays in server-side logs.
2026-06-23 17:28:22 -04:00
6951bf2fc6 fix(agent-team): address GPT-4.1 cross-review findings on the dashboard surface
- read_transitions: catch sqlite3.Error (not just OperationalError) so a corrupt
  ledger degrades to empty rather than raising into callers
- recorder open-row lookup: order by the monotonic transition_id (drop the
  timestamp-format dependency)
- TransitionRecorder.close(): release the retained in-memory test connection
- dashboard task_detail: return only the exception TYPE, never str(exc) (a SQLite
  message can carry the DB path)
- _instrument: coerce a status enum to .value defensively before the terminal check
- schema.migrate: document the ordering constraint for future ALTERs vs the
  unconditional idempotent tail
2026-06-23 17:19:04 -04:00
4f4db6ed03 docs(agent-team): point status systemd unit at the SPA dashboard + document WebUI
Switch agent-team-status.service ExecStart from status_page.serve to
dashboard.serve (uvicorn serving web/dist + the JSON API). Document the WebUI in
the README: live auto-laid pipeline map, click-through task history, endpoints,
and the Mac-side npm build + rsync flow.
2026-06-23 17:15:38 -04:00
e3137e33e6 test(agent-team): cover transitions, topology, dashboard API + rework status_page tests
Add test_transitions.py (recorder idempotency under replay, terminal close,
fail-soft, the B4 _instrument signature-preservation guarantee, end-to-end graph
drive), test_topology.py (meta coverage, tree grouping, edge classification,
phase->node map), and test_dashboard.py (/api/state contract + per-node state,
/api/topology, /api/task timeline+cost+partial+thread_id validation, all via
TestClient, fastapi-skipif guarded). Rework test_status_page.py to the data layer
(drop retired render_html/SVG tests). conftest exposes tests/ for sibling imports.
Frontend: Vitest+RTL for layout, TaskList filters/selection, TaskDrawer timeline.
2026-06-23 17:15:31 -04:00
677f3d3c71 feat(agent-team): React/Vite status-dashboard SPA
New web/ SPA (React 18 + Vite 5 + TypeScript + React Flow + dagre, all pinned, no
CDN). Dark, three-pane layout: summary top bar, filterable task list, center
auto-laid pipeline map, and a right drawer showing a selected task's history
through each node (timeline with timestamps/duration/cost + Q&A/verdicts/plan).
Map nodes color by live state, loop-backs render dashed, clicking a node filters
the list, selecting a task highlights its path. Polls /api/state every 4s;
topology fetched once. dist/ is gitignored (built on the Mac, rsynced).
2026-06-23 17:15:23 -04:00
322a1f6922 feat(agent-team): LangGraph-introspected topology + read-only dashboard API
topology.py derives the pipeline map (nodes/edges/trees) from the compiled
LangGraph via get_graph() + a NODE_META display sidecar, so new agent nodes
appear automatically and group into trees branching off intake. dashboard.py is a
new read-only FastAPI app (0.0.0.0:8770) serving /api/state (contract preserved +
per-node live state), /api/topology, and /api/task/{id} (validated, timeline +
cost join + partial fallback) plus the built SPA — kept SEPARATE from the authed
api.py. status_page.py is trimmed to the /api/state data layer; the inline
HTML/SVG renderer + stdlib server are retired.
2026-06-23 17:15:16 -04:00
4656ca64b6 feat(agent-team): task_transitions ledger + graph node instrumentation
Add a v3 task_transitions table (additive migration + startup assertion) and a
fail-soft, idempotent TransitionRecorder. build_graph gains an injected
transition_recorder that wraps every node via a functools.wraps'd _instrument
(signature-preserving so LangGraph still injects RunnableConfig); the coordinator
wires it. Records one row per node entry (idempotent under resume replay) and
closes the open row on terminal status. Backs the dashboard task-history view.
2026-06-23 17:15:08 -04:00
Adam Moussa
778737fa9f
Merge pull request #54 from Sea-Haven-Industries/fix/agent-team-plan-presentation-threading
fix(agent-team): thread lifecycle milestones + present the plan in Slack
2026-06-23 16:44:53 -04:00
403b913287 fix(agent-team): thread lifecycle milestones + present the plan in Slack
Two gaps surfaced by a live /new-task (task 9bce78ad): the plan was produced
and approved, but the "plan ready" notice posted top-level (not in the task
thread) and contained no plan to review.

1. THREADING — `run-team.py` `_build_notifiers` exposed `notify(message)` with
   no `thread_ts`. The coordinator's `_emit` calls `notify(message,
   thread_ts=root)`; that raised TypeError, and `_emit`'s fallback re-posted
   TOP-LEVEL. So every lifecycle milestone (plan-ready / parked / failed) landed
   unthreaded, despite the coordinator computing the root ts. (The clarifier
   QUESTION threaded fine — different path.) Fix: the sink now accepts and
   forwards `thread_ts` into the chat.postMessage payload (build_slack_poster
   already forwards the key).

2. PRESENTATION — the plan-ready milestone was a bare one-liner. It now posts a
   CONDENSED plan (summary + numbered phase names; step detail stays on the
   status dashboard) via new `Coordinator._summarize_plan`, so the plan is
   actually reviewable in-thread.

Tests: notify sink forwards thread_ts (and omits it for top-level); condensed
plan renders summary + phase names (not steps); malformed plan falls back;
plan-ready milestone threads under the task root AND carries the plan. 1193 pass.
2026-06-23 16:42:27 -04:00
Adam Moussa
7fbfdea03f
Merge pull request #53 from Sea-Haven-Industries/fix/coordinator-crash-on-node-exception
fix(agent-team): a crashing pipeline node fails one task, not the whole daemon
2026-06-23 16:22:53 -04:00
f12dfecd95 fix(agent-team): a crashing pipeline node fails one task, not the whole daemon
A live task (d30b697c) on the R720 crashed the coordinator: the planner's
single-shot Claude call raised "Reached maximum number of turns (1)", the
exception propagated out of `drain_resumes` through the serve loop, and systemd
restarted the daemon — with no Slack notice, so the failure was silent.

Defense in depth:

1. invoker: `_collect_subscription_text` now tolerates the single-shot turn cap.
   When the Agent SDK raises "Reached maximum number of turns" mid-stream it
   salvages the assistant text already collected (the JSON the planner needs)
   instead of propagating. An empty salvage or any non-turn error still raises.

2. coordinator: `drain_resumes` wraps the per-job `resume()` so an unhandled
   node exception fails THAT task instead of the daemon — it supersedes the
   answered question (so the startup recovery sweep cannot re-drive it into the
   same crash on reboot), marks the task FAILED via `update_state` with a short
   `failure_reason`, and surfaces an honest "❌ FAILED" line to Slack.

3. task_model: add `failure_reason` to PipelineState + TaskRecord (kept in sync)
   so the terminal-failure detail persists as a real graph channel.

4. resume_worker: add `ResumeOutcome.FAILED`.

Tests: invoker salvage/re-raise/propagate paths; drain_resumes fails-not-crashes,
supersedes the question, notifies, and one failing task does not block others.
1189 passed.
2026-06-23 16:20:45 -04:00
Adam Moussa
dd6aef0f16
Merge pull request #48 from Sea-Haven-Industries/feat/agent-team-status-page
feat(agent-team): LAN-only read-only status dashboard for the coordinator
2026-06-23 15:55:01 -04:00
505fdeebd3 docs(agent-team): describe the live pipeline map + /api/state endpoint 2026-06-23 15:53:04 -04:00
4e75a0bf93 test(agent-team): cover live pipeline map + /api/state JSON contract
- render_html includes the inline SVG map + inline fetch('/api/state') poller
  + tooltip mount + live clock, stays offline (no CDN), keeps the <noscript>
  meta-refresh fallback, and renders the map even on a not-ok snapshot.
- snapshot_to_dict: expected top-level keys, one entry per STAGE in order,
  per-phase grouping, awaiting-human classification (open pending_question),
  per-node model/agent role labels, summary counts, not-ok serializability.
- Descriptions escaped in BOTH the HTML and the JSON-in-script seed
  (</script><script> breakout neutralised; \u003c form present); the raw
  /api/state JSON round-trips the description for textContent rendering.
- End-to-end snapshot_to_dict over the seeded ledger.
2026-06-23 15:53:04 -04:00
72098ba26e feat(agent-team): live visual pipeline map for the status dashboard
Turn the read-only status page into an auto-updating visual map of the
agent-team DAG (INTAKE -> CLARIFY <-> gate -> PLAN <-> REVIEW ->
[BUILD -> VERIFY -> DISPATCH] -> DONE), rendered as hand-rolled inline
SVG (no CDN/D3 — the R720 is offline/LAN-only).

- Stage model (STAGES) with per-node model/agent role labels (Claude/
  GPT-4.1/Gemini/DeepSeek/Slack owner) and gated/role-node flags.
- Per-stage live state (idle/active/awaiting-human/parked) + count badge,
  grouped by current_phase; awaiting-human = OPEN pending_question.
- GET /api/state JSON sidecar (snapshot_to_dict); inline vanilla-JS poller
  fetches it every 4s and repaints node states/counts/cards/clock/tooltip
  in place (no reload, hover/scroll survive). <noscript> meta-refresh
  fallback retained.
- Hover/focus tooltip per node: short_id, description, status, waiting age.
- Existing table view kept as a detail section below the map.
- Read-only (mode=ro), fail-safe, no secrets; descriptions escaped for both
  HTML and the JSON-in-script seed (< / > -> \uXXXX), DOM via textContent.
2026-06-23 15:53:04 -04:00
bd5bec2f9a docs(agent-team): install/run notes for the read-only status dashboard 2026-06-23 15:53:04 -04:00
de215dfb64 test(agent-team): unit tests for status_page render + read-only reader 2026-06-23 15:53:04 -04:00
97a4befd24 feat(agent-team): systemd unit for the read-only status dashboard 2026-06-23 15:53:04 -04:00
3ddd9ca08c feat(agent-team): read-only LAN status dashboard module (status_page.py) 2026-06-23 15:53:04 -04:00
Adam Moussa
96b02e05f1
Merge pull request #52 from Sea-Haven-Industries/fix/security-review-xargs-overflow
fix(security-review): batch template-discovery grep so review.sh --scanners-only stops overflowing argv
2026-06-23 15:52:47 -04:00
Adam Moussa
e8369e3c1d
Merge pull request #50 from Sea-Haven-Industries/feat/r720-deploy-script
feat(agent-team): generalized SAFE deploy-r720.sh (whole-package rsync, snapshot-gate, verify-after)
2026-06-23 15:52:08 -04:00
bdafb7bbf1 fix(security-review): batch template-discovery grep so review.sh --scanners-only stops overflowing argv on the monorepo
The cfn-lint template-discovery step in review.sh piped the repo's whole
matched-file list into 'xargs -I{} sh -c "grep -l {}"'. On the orchestrator
monorepo — especially from a deep worktree path, where every matched path is a
long absolute path — xargs -I{} packs all paths into one assembled command and
aborts with 'xargs: command line cannot be assembled, too long'. The subprocess
exits non-zero having emitted ZERO findings, so the global pre-push hook BLOCKS
every push (agents were working around it with --no-verify).

Fix: switch the grep stage to NUL-delimited, un-batched xargs
(find ... -print0 | xargs -0 grep -lE ...). xargs -0 (no -I) splits the input
across multiple grep invocations, so the argv never exceeds ARG_MAX; grep -l
reports the same matching files as the old per-file grep, and -print0/-0 is safe
for paths with spaces/newlines. The first 'xargs -I{} find {}' is kept (find
needs the start path before its expression) and is bounded by the scope-path
count, so it is not an overflow source. Trailing '|| true' preserves the old
no-match/no-files semantics (TPLS = list-of-templates or empty, never fails).

Purely an argv-batching fix: the scanned file set, findings, and exit codes are
unchanged. Verified exit-code-identical: clean tree -> exit 0 (PASS); planted
GitHub PAT + RSA private key -> exit 1 (BLOCK, gitleaks high); planted CFN
template with a cfn-lint error on a deeply-nested path -> exit 1 (cfn-lint flags
it, no overflow). The previously-overflowing command now completes clean.
2026-06-23 15:51:43 -04:00
d89ca9ae7f feat(agent-team): generalized SAFE deploy-r720.sh (whole-package rsync, snapshot-gate, verify-after)
Codifies the hardened deploy-before-merge procedure for ongoing agent-team
code changes, generalizing the one-off deploy-r720-ws-rollout.sh. Prevents the
two self-inflicted live crash-loops:

- Whole agent_team/ package rsync (never per-file, which misplaces e.g.
  nodes/planner.py at the package root -> ImportError/TypeError crash-loop).
- Snapshot HARD GATE (Adam's Hyper-V step; Claude ssh reaches only the guest)
  + ledger backup before any change.
- Pre-restart import sanity, then mandatory verify-after (is-active==active,
  NRestarts didn't climb, ~6 threads, clean journal) with rollback guidance
  on failure.

Idempotent, fails loudly. Optional SYNC_DEPS / SYNC_HANDBOOK / RESTART_STATUS.
Drives the new /sh-deploy-r720 skill.
2026-06-23 15:51:28 -04:00
Adam Moussa
f7bfa5baf0
Merge pull request #47 from Sea-Haven-Industries/feat/ws-activation-wiring
feat(integration): WS0-WS5 + activation wiring (context_provider + /new-task)
2026-06-23 15:50:28 -04:00