Commit graph

13 commits

Author SHA1 Message Date
48a81c0802 fix(agent-team): centralize safe decision mapping + remove free-text abandon hair-trigger
Security-review follow-up (LOGIC-01/02/05, all confirmed correctness).

- New transport-neutral `decisions.normalize_decision(raw, *, allow_abandon)` is
  the single source of truth: approve-allowlist→approve; abandon-allowlist→abandon
  ONLY when allow_abandon; everything else (prose, empty, abandon-verbs when
  disallowed) → request_changes with the full reply as notes; idempotent on an
  already-formed decision dict. slack_adapter.map_plan_decision is now a thin
  wrapper (default allow_abandon=True, no caller churn).
- LOGIC-01/02: graph._parse_decision now delegates to normalize_decision (was:
  any unrecognized verb → abandon → FAILED). The graph is now the universal safe
  backstop, so EVERY writer that bypassed the listener mapping — operator CLI
  answer_on_behalf (raw), Coordinator.submit_answer (raw), the recovery sweep —
  loops back on prose instead of silently FAILing the task. Explicit abandon
  still abandons (preserves the confirmed-button path).
- LOGIC-05: the Slack FREE-TEXT reply path maps with allow_abandon=False, so a
  bare "cancel"/"stop"/"abandon" typed in-thread → request_changes (never
  terminal abandon); abandon stays reachable only via the confirm-guarded button.

Tests: graph unrecognized→loops-back (not FAILED), operator raw-prose→request_
changes, free-text destructive verbs→request_changes vs button→abandon,
normalizer idempotency. 1484 passed.
2026-06-24 11:48:04 -04:00
67b0f4c6ae feat(agent-team): resumable plan-review gate in the graph (Phase B2a)
Replace the terminal review-cap PARK with a resumable human decision gate,
opt-in via build_graph(plan_gate=True) (default False → all existing P1/P2/P3
wiring unchanged).

- plan_gate_node interrupt()s mirroring the clarifier contract (same payload
  keys → existing pending_question() extractor + turn-guarded ResumeWorker drive
  it with zero special-casing) plus a kind="plan_decision" discriminator and the
  plan + latest review findings as context.
- Decision contract {"decision": approve|request_changes|abandon, "notes": ...}:
  approve → the same terminal state an auto-approved plan reaches (ACTIVE/BUILD);
  request_changes → append a synthetic human verdict to review_verdicts (so the
  planner's _format_review_feedback surfaces the notes) and loop back to PLAN;
  abandon / unrecognized → terminal FAILED (safe default, never accidental
  approve).
- Bounded termination: MAX_PLAN_GATE_VISITS=3 combined ceiling on plan_gate_visits
  (new channel on PipelineState + TaskRecord); on exhaustion the gate goes
  terminal PARKED ("revision ceiling reached") WITHOUT interrupting. Proven by a
  loop-past-ceiling test.

Notes for the coordinator wiring (B2b): while suspended at the gate the status
channel still reads 'parked' (carried over from review_node's escalate branch) —
the load-bearing "awaiting decision, not terminal" signal is the live pending
interrupt + kind="plan_decision", NOT the status channel.

Tests: interrupt-at-cap, approve/request_changes(notes)/abandon routing,
ceiling-terminates, auto-approve still bypasses the gate. 1404 passed.
2026-06-23 21:03:53 -04:00
f0c5cfe57f feat(agent-team): P3 Phase-0 box-side build->dispatch->verify (WIP)
0c-binding: per-task expected_run_id bound from state (gate rejects substituted
  run_id; None -> BLOCK, never vacuous pass).
0e: fail-safe serve default (failsafe_production_p3_wiring) — inert on
  unprovisioned env (one WARNING + one #agent-team notice), never crash-loops.
0a: reorder P3 subgraph BUILD -> DISPATCH -> VERIFY (preserves _instrument).
0d: ci_watcher engine + VERIFY interrupt()-wait (async resume-on-CI-complete).

KNOWN-OPEN (adversarial review BLOCKs, to remediate next):
- CI-watcher not wired into run-team serve (ci_pending_provider/ci_poller None)
  -> a VERIFY-suspended task never resumes/parks.
- no durable ci_pending_provider enumerating threads suspended at VERIFY.
Branch only; not merged, not deployed.
2026-06-23 19:52:04 -04:00
6951bf2fc6 fix(agent-team): address GPT-4.1 cross-review findings on the dashboard surface
- read_transitions: catch sqlite3.Error (not just OperationalError) so a corrupt
  ledger degrades to empty rather than raising into callers
- recorder open-row lookup: order by the monotonic transition_id (drop the
  timestamp-format dependency)
- TransitionRecorder.close(): release the retained in-memory test connection
- dashboard task_detail: return only the exception TYPE, never str(exc) (a SQLite
  message can carry the DB path)
- _instrument: coerce a status enum to .value defensively before the terminal check
- schema.migrate: document the ordering constraint for future ALTERs vs the
  unconditional idempotent tail
2026-06-23 17:19:04 -04:00
4656ca64b6 feat(agent-team): task_transitions ledger + graph node instrumentation
Add a v3 task_transitions table (additive migration + startup assertion) and a
fail-soft, idempotent TransitionRecorder. build_graph gains an injected
transition_recorder that wraps every node via a functools.wraps'd _instrument
(signature-preserving so LangGraph still injects RunnableConfig); the coordinator
wires it. Records one row per node entry (idempotent under resume replay) and
closes the open row on terminal status. Backs the dashboard task-history view.
2026-06-23 17:15:08 -04:00
3847e43ba3 feat(agent-team): one Slack thread per task — root "Task received" message + threaded questions/milestones
WS Slack-UX Feature 1. A /new-task task now maps to ONE Slack thread instead of
several top-level messages.

- /new-task posts an immediate root "📥 Task received: …" ack and captures its
  ts (root_ts); this is the instant acknowledgement.
- root_ts is plumbed into start: new PipelineState/TaskRecord channel
  slack_thread_ts, seeded by graph.start_task and threaded through
  Coordinator.start_task. The NewTaskCallback is now (task_text, via, root_ts).
- All clarifier questions for the task post as THREADED REPLIES under root_ts
  (chat.postMessage thread_ts=root_ts), and each question's ledger channel_ref
  is set to root_ts (NOT the reply's own ts). Because answer-mapping resolves a
  reply via find_open_question_by_channel_ref(thread_ts), a reply in the root
  thread (thread_ts==root_ts) maps to the task's currently-open question with NO
  change to the mapping logic or the first-answer-wins CAS. The open-only
  partial-unique index still holds (one open question per task at a time).
- Lifecycle milestones (parked / plan-ready / needs-input) and follow-up
  questions thread under root_ts too; the notify sink gained an optional
  thread_ts kwarg (degrades to top-level on a sink that doesn't accept it).
  notify failures still never break tick.
- SlackTransport.post_question + the live poster accept/forward thread_ts.
- No root_ts (non-/new-task origin) ⇒ top-level posts exactly as before.

AUTHZ-01 (owner-allowlist-first, fail-closed) and the atomic open→answered
compare-and-set are unchanged.

Adds plumbing for the inbound-ack reactor seam used by Feature 2 (dormant until
a reactor is injected). Tests cover thread_ts forwarding, channel_ref=root_ts,
graph seeding, and coordinator threading.
2026-06-23 15:49:40 -04:00
a6fd1cfc75 fix(intake): seed the task description into graph state (was silently dropped)
/new-task (and every intake: GitHub issue, /sh-assign-task) reached the clarifier
with NO description -> the clarifier asked 'no task description provided'. Root
cause: coordinator.start_task only LOGGED task_text (a P1-era decision when the
deterministic clarifier didn't consume a description), graph.start_task took no
task arg, and PipelineState/TaskRecord had no 'task' channel at all.

Fix: add a first-class 'task' field to PipelineState + TaskRecord (+ round-trip
in task_from_dict); graph.start_task seeds task into the initial invoke (persists
through intake_node's partial-state return into CLARIFY); coordinator.start_task
passes task=task_text. The clarifier already reads state['task'] via
_task_description, so it now sees the real description.

Test: start_task(task='build a login form') -> suspended CLARIFY state carries
task. 1143 passed, ruff clean.
2026-06-23 15:49:40 -04:00
Claude
72076b7874
feat(ws3): add dispatch node + remove agent-apply env gate
- Add agent_team/nodes/dispatch_invoker.py: make_dispatch_node() wraps
  dispatcher.dispatch_apply_verify() as a LangGraph node; fail-safe
  (parks task on any error); INERT unless wired by coordinator.
- graph.py: add DISPATCH_NODE constant; add dispatch_node param to
  build_graph; when provided, repoint APPROVED_ROUTE → DISPATCH_NODE →
  END (P3+). Validates dispatch_node requires build_verify.
- coordinator.py: add DispatchNodeFactory type; thread dispatch_node_wiring
  through __init__ and setup(); default None = inert (no auto-dispatch).
- agent-team-apply-verify.yml: remove agent-apply environment gate from
  gate-and-pr; add Slack-notify + audit-log compensating control steps.
- test_apply_verify_workflow_hardening.py: flip env assertion → assert env
  REMOVED and compensating steps present.

All 1044 tests passing. ruff clean.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QYp761G9HojkmLqLZrASVi
2026-06-23 01:31:29 +00:00
Adam Moussa
c671e6bdbb
fix(agent-team): register QuestionSet with the langgraph checkpoint serializer (silence/avoid msgpack block) (#30) 2026-06-22 16:14:27 -04:00
e1208ee563 feat(agent-team): P3-inert build/verify subgraph topology (opt-in, no live CI)
build_verify_subgraph: BUILD->VERIFY nodes + route_after_verify. build_graph gains an opt-in build_verify param that repoints the review 'build' route at the subgraph (BUILD->VERIFY->{approved->END | loop->PLAN | parked->END}); default unchanged (P2). Coordinator build_verify_wiring composes it INERT (no ci_result -> ci_gate BLOCK -> PARKED; LLM is fix-proposer only, never declares green). NOT enabled in production: the live CI apply/verify + OIDC stays held for its /sh-security-review + GPT-4.1 cross-review gate. Also escapes untrusted intake text in logs (log-injection hygiene).
2026-06-18 13:23:03 -04:00
22edd4143a feat(agent-team): P1/P2 graph wiring + coordinator daemon + run-team start/serve
build_graph gains injected live_plan_node/review_node/route_review: P1 = plan->END, P2 = clarify->plan->review->{build|loop-back|parked}. Coordinator composes clarifier->graph->ResumeWorker, wraps planner fail-safe, binds the GPT-4.1 review loop; run-team start/serve opt production into P2. Re-delivery uses a guarded CAS so a concurrently-answered row is never clobbered (closes RACE-REDELIVER).
2026-06-18 12:56:42 -04:00
0eae5dbfc3 Prove P1 exit criteria against the real LangGraph graph; fix question_id stability
Reworks the P1 sim so the four §7.1 exit criteria are demonstrated against the
ACTUAL mechanic, not a model (resolves the verifier's "sim models the ledger,
not the LangGraph integration" finding).

- New tests/sim/test_p1_graph_integration.py drives the real agent_team.graph
  StateGraph (interrupt/Command(resume)) + the real langgraph SqliteSaver
  checkpointer + the committed pending_questions compare-and-set, proving:
  (a) suspend survives a simulated restart (drop saver/conn, rebuild over the
  same checkpoint DB) and resumes; (b) duplicate answer loses the CAS and the
  graph never double-advances; (c) a post-deadline answer loses to expire and
  the task is not resumed; (d) two concurrent tasks resume to the correct
  thread, with a turn-guarded no-double-apply check.
- graph.py: derive a STABLE question_id from uuid5(thread_id, turn). The
  clarifier node replays on resume, so the prior fresh-uuid id changed between
  the delivered/ledgered question and the qa_history entry — breaking the
  §3.3.1 identity contract. Now the delivered id == ledger key == history entry
  (unit-tested in test_graph.py).
- harness._connect() now uses the committed schema.connect() (WAL + busy_timeout)
  instead of a raw sqlite3.connect, so concurrent responders genuinely serialize;
  the criterion-(d) concurrency test no longer swallows OperationalError (it
  asserts zero errors + exactly one CAS winner).
- requirements.txt: pin langgraph-checkpoint-sqlite==3.1.0 (design D9 durable
  checkpointer), now exercised by the integration test.

Full suite: 564 passed; ruff + format clean.
2026-06-17 15:16:12 -04:00
15a416d31a Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.

Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled

KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven

Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 15:16:12 -04:00