Commit graph

44 commits

Author SHA1 Message Date
3847e43ba3 feat(agent-team): one Slack thread per task — root "Task received" message + threaded questions/milestones
WS Slack-UX Feature 1. A /new-task task now maps to ONE Slack thread instead of
several top-level messages.

- /new-task posts an immediate root "📥 Task received: …" ack and captures its
  ts (root_ts); this is the instant acknowledgement.
- root_ts is plumbed into start: new PipelineState/TaskRecord channel
  slack_thread_ts, seeded by graph.start_task and threaded through
  Coordinator.start_task. The NewTaskCallback is now (task_text, via, root_ts).
- All clarifier questions for the task post as THREADED REPLIES under root_ts
  (chat.postMessage thread_ts=root_ts), and each question's ledger channel_ref
  is set to root_ts (NOT the reply's own ts). Because answer-mapping resolves a
  reply via find_open_question_by_channel_ref(thread_ts), a reply in the root
  thread (thread_ts==root_ts) maps to the task's currently-open question with NO
  change to the mapping logic or the first-answer-wins CAS. The open-only
  partial-unique index still holds (one open question per task at a time).
- Lifecycle milestones (parked / plan-ready / needs-input) and follow-up
  questions thread under root_ts too; the notify sink gained an optional
  thread_ts kwarg (degrades to top-level on a sink that doesn't accept it).
  notify failures still never break tick.
- SlackTransport.post_question + the live poster accept/forward thread_ts.
- No root_ts (non-/new-task origin) ⇒ top-level posts exactly as before.

AUTHZ-01 (owner-allowlist-first, fail-closed) and the atomic open→answered
compare-and-set are unchanged.

Adds plumbing for the inbound-ack reactor seam used by Feature 2 (dormant until
a reactor is injected). Tests cover thread_ts forwarding, channel_ref=root_ts,
graph seeding, and coordinator threading.
2026-06-23 15:49:40 -04:00
e7ef4387e6 fix(pipeline): single-shot Claude calls + planner actually reads review findings
Two root causes behind 'every task parks, and slowly':
1. SINGLE-SHOT INVOKER: subscription Claude calls ran as 40-turn, tool-enabled
   agentic sessions (--max-turns 40, $2 budget) for what are pure reasoning->JSON
   completions — minutes-long, and Claude wandered/returned unparseable output.
   _DEFAULT_MAX_TURNS 40->1 + allowed_tools=[] -> fast deterministic single turn.
2. PLANNER FEEDBACK KEY MISMATCH: the review stage writes verdict/findings, but
   _format_review_feedback read decision/notes/comment (never present) -> the
   planner re-planned with EMPTY feedback, re-introduced the rejected flaw
   ('assumptions persist'), hit the review cap, parked. Now reads verdict/findings
   (old keys kept as fallback) so GPT-4.1's objections reach the re-plan.
Plus: park notifications infer the phase the task was IN (review/plan/clarify)
instead of the terminal 'parked'.
Tests: planner real-verdict-keys regression + coordinator phase-inference. 1149 pass.
2026-06-23 15:49:40 -04:00
655f6d80f4 feat(notify): richer park/lifecycle messages — task description + phase + blocker
Park notifications were opaque ('Task 6c3c3202 parked — needs your attention'):
no idea what the task is, where it got to, or what's blocking it. Now each
message names the task DESCRIPTION (not just the short id), the PHASE it reached,
and the actual BLOCKER — _summarize_blocker() pulls the last review_verdict's
findings (the GPT-4.1 REQUEST_CHANGES text, collapsed + truncated), falling back
to 'no plan built' / 'revision cap hit'. Applies to parked + needs-more-input +
plan-ready emits. +1 test (description + phase + blocker present). 1148 passed.
2026-06-23 15:49:40 -04:00
f0a2dc27a5 feat(notify): Slack lifecycle notifications + deliver multi-turn questions
The bot only ever posted clarifier questions; parks/completions were silent and
multi-turn follow-up questions were never posted during normal operation (only
the startup recover sweep posted them). So an answered task was a black box.

- Coordinator gains a 'notify' sink + _emit() (guarded, never breaks the loop).
- tick() now runs _post_resume_followups(results) after the drain: for each
  resumed thread it (a) POSTS a newly-pending clarifier question (fixes silent
  multi-turn — the drain path left it unposted) and emits 'needs more input';
  else emits 'parked — needs attention' or 'plan ready for review' from the
  settled state.
- run-team serve wires notify -> Slack channel (build_slack_poster) and an
  alarm_hook that logs the deadline-park WARNING AND posts a parked notice.
  Live-Slack only; dry-run/non-Slack/no-channel = silent (None), no token needed.

Tests: +4 (needs-input/parked/plan-ready emits + notify-failure swallow);
_FakeCoordinator gains notify/alarm_hook. 1147 passed, ruff clean.
2026-06-23 15:49:40 -04:00
a6fd1cfc75 fix(intake): seed the task description into graph state (was silently dropped)
/new-task (and every intake: GitHub issue, /sh-assign-task) reached the clarifier
with NO description -> the clarifier asked 'no task description provided'. Root
cause: coordinator.start_task only LOGGED task_text (a P1-era decision when the
deterministic clarifier didn't consume a description), graph.start_task took no
task arg, and PipelineState/TaskRecord had no 'task' channel at all.

Fix: add a first-class 'task' field to PipelineState + TaskRecord (+ round-trip
in task_from_dict); graph.start_task seeds task into the initial invoke (persists
through intake_node's partial-state return into CLARIFY); coordinator.start_task
passes task=task_text. The clarifier already reads state['task'] via
_task_description, so it now sees the real description.

Test: start_task(task='build a login form') -> suspended CLARIFY state carries
task. 1143 passed, ruff clean.
2026-06-23 15:49:40 -04:00
0945d60338 fix(ws2): register @app.command(/new-task) so Slack slash command is acked
'the app did not respond': serve() registered @app.action/@app.event but NO
@app.command handler, so Bolt never acked the /new-task slash command within
Slack's ~3s deadline. Add an @app.command(/new-task) handler that ack()s first,
re-stamps type:slash_commands onto Bolt's inner command body (Bolt strips the
Socket Mode envelope type that _is_new_task_command/_discriminating_type expect),
then forwards to handle_event (AUTHZ-01 + new-task dispatch). The resulting
payload shape is the one already covered by test_new_task_calls_callback_*.
2026-06-23 15:49:40 -04:00
75f1dc6f75 fix(ws1/ws5): bootstrap orchestrator root in in-process invokers (G1); thread handbook into clarifier (G4)
Gap-audit findings:
- G1 (blocks-feature): make_cross_reviewer_invoker / make_fast_coder_invoker did
  'from models import' without putting the orchestrator root on sys.path. The
  run-team serve daemon only bootstraps agent-team/, so on the live box every
  GPT-4.1 plan review hit ModuleNotFoundError -> review_plan's blanket except
  silently fail-closed to REQUEST_CHANGES (GPT-4.1 never actually ran). Both
  in-process invokers now call invoker_multi._ensure_orchestrator_on_path()
  before the deferred import. WS1 introduced this when it swapped the review
  default from the subprocess invoker to in-process.
- G4 (degrades): the handbook context_provider was wired into the planner only;
  default_clarify_node_factory now accepts + forwards it, and run-team wires it
  into build_clarify_node too, so clarifying questions are handbook-aware.

Tests: +2 regression tests (path-bootstrap, clarifier threading); _FakeCoordinator
gains build_clarify_node. 1142 passed, ruff clean.
2026-06-23 15:49:40 -04:00
a96a5b487a feat(integration): wire WS5 context_provider + WS2 /new-task into serve
Integration branch combining WS0-WS5 (PRs #43-#46) + the activation wiring that
flips the safe seams ON in the run-team serve path:

- WS5 (D10): inject the Sea Haven handbook conventions into the planner prompt
  via context_provider (zero-arg handbook loader; fail-safe to '' when absent).
- WS2: an allowlisted Slack /new-task starts a task on this coordinator
  (set_new_task_callback adapter -> start_task; AUTHZ-01 gates it upstream).
- WS1 bind_multi_invoker() is already wired in _cmd_serve.

Coordinator gains a new_task_callback param + set_new_task_callback() (resolves
the constructor chicken-and-egg of referencing the coordinator's own start_task);
default_slack_listener_factory forwards it to the SlackListener.

NOT wired (deliberately): the P3 dispatch_node / build_verify path. Activating
it correctly needs a per-task expected_run_id bound into gated_build_verify_wiring
(plumbing that does not exist yet) AND the CI trust-boundary security re-review.
It stays inert pending that work.

Tests: +5 activation-wiring tests; _FakeCoordinator stub gains
set_new_task_callback. Full agent-team suite: 1140 passed, ruff clean.
2026-06-23 15:49:40 -04:00
Adam Moussa
e920ffc92c
Merge pull request #46 from Sea-Haven-Industries/feat/ws0-plugin-slack-intake-autodelegate
feat(ws0+ws2+ws4): Sea Haven plugin scaffold + Slack /new-task + auto-delegate hook
2026-06-23 15:48:29 -04:00
Adam Moussa
45a6560512
Merge pull request #45 from Sea-Haven-Industries/feat/ws3-auto-dispatch-remove-env
feat(ws3): auto-dispatch node wiring (keeps agent-apply human gate)
2026-06-23 15:48:24 -04:00
Adam Moussa
462f44d9de
Merge pull request #44 from Sea-Haven-Industries/feat/ws1-inprocess-models-http-api
feat(ws1): non-Claude in-process invokers + FastAPI HTTP API (WS1, code only)
2026-06-23 15:48:18 -04:00
b6c66877f5 fix(ws1): harden HTTP API + declare fastapi/uvicorn deps
Security follow-ups from the per-PR review (non-blocking MEDIUMs):
- Eager _get_token() at make_app build time so a missing AGENT_TEAM_API_TOKEN
  fails fast instead of serving requests first (matches the docstring contract).
- Disable /docs, /redoc, /openapi.json (no auth dependency in FastAPI) — the
  API is VPN-only/127.0.0.1 and should not expose its schema unauthenticated.
- Scrub raw exception text and subprocess stderr from 500 response bodies;
  log server-side instead (avoid internal-path/state disclosure).
- Bound /orchestrator/invoke concurrency with a semaphore (429 over the cap)
  so an authenticated caller cannot exhaust the box via many 600s subprocesses.

Also pin fastapi/uvicorn in requirements.txt (WS1 dep). With fastapi now
installed in CI, the previously skip-guarded TestClient tests run for real;
the importorskip guard stays as a no-op safety net.

Tests: 23 pass (adds docs-disabled + concurrency-429 cases).
2026-06-23 12:27:23 -04:00
366d07a84e fix(ws1): skip fastapi TestClient tests when fastapi absent + ruff format
CI has no fastapi (box-only dependency); api.py imports it lazily. Guard the
7 TestClient smoke tests with skipif(find_spec('fastapi') is None) so they
skip in CI instead of failing collection, leaving the 14 invoker_multi tests
running. Also apply ruff format to the 5 WS1 files CI flagged.
2026-06-23 11:40:46 -04:00
83d17a5b7e style(ws0+ws2+ws4): ruff format slack_listener, hook, test (CI ruff format --check) 2026-06-23 11:38:58 -04:00
bfec8cc4d1 style(ws3): ruff format dispatch_invoker + test (CI ruff format --check) 2026-06-23 11:38:31 -04:00
47d7a81b6b style(ws5): ruff format handbook.py (CI ruff format --check) 2026-06-23 11:37:52 -04:00
Claude
80c04fc487
feat(ws3): add default_dispatch_node_factory + 25-test dispatch_invoker suite
- Add default_dispatch_node_factory() to coordinator.py: reads
  AGENT_TEAM_REPO_OWNER / AGENT_TEAM_REPO_NAME / AGENT_TEAM_BASE_BRANCH
  from env; fails closed (RuntimeError) if required vars absent; delegates
  to make_dispatch_node with owner/repo fixed at factory time. Added to __all__.
- Add tests/test_ws3_dispatch_invoker.py (25 tests): make_dispatch_node
  happy path + fail-closed paths (missing thread_id/diff/scope/plan, DispatcherError,
  unexpected exception); owner/repo injection from factory args; scope list
  flattening; DispatchNodeFactory export; graph DISPATCH_NODE constant;
  build_graph ValueError when dispatch_node given without build_verify; env-var
  binding for default_dispatch_node_factory; Coordinator.dispatch_node_wiring
  seam (ValueError when wired without build_verify_wiring).

All 1069 tests pass; ruff clean.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155wSD9kKFjnTMNNsDT3tiW
2026-06-23 01:40:32 +00:00
Claude
2aa72d73f4
feat(ws0+ws2+ws4): plugin scaffold, Slack /new-task, auto-delegate hook
WS0 — sea-haven-claude-plugin/ scaffold:
  - CLAUDE.md: Sea Haven engineering context (pipeline overview, rules,
    /new-task + /delegate usage, available phases)
  - settings.template.json: UserPromptSubmit hook wiring template
  - hooks/user_prompt_submit.py: standalone script (WS4)

WS2 — SlackListener /new-task intake:
  - Add NewTaskCallback type alias (Callable[[str, str], str])
  - Add new_task_callback param to SlackListener.__init__
  - _is_new_task_command() helper for slash_commands+/new-task detection
  - _handle_new_task_command() method: authorized-only, calls callback,
    exception-safe (listen loop stays alive on callback errors)
  - handle_event() routes /new-task BEFORE the answer path (post-AUTHZ-01)

WS4 — UserPromptSubmit auto-delegate hook:
  - /delegate <text> and DELEGATE: <text> prefixes trigger delegation
  - Calls POST /tasks on the agent-team HTTP API (WS1)
  - Blocks the Claude Code prompt; shows thread_id + next-steps message
  - Graceful degradation: missing token, HTTP error, network error all
    produce a block with a human-readable reason
  - run() is a pure function for testability (no stdin/stdout in tests)

Tests: 22 new tests in test_ws0_ws2_ws4_plugin_slack_hook.py.
Full suite: 1066 passed. ruff clean.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QYp761G9HojkmLqLZrASVi
2026-06-23 01:39:12 +00:00
Claude
72076b7874
feat(ws3): add dispatch node + remove agent-apply env gate
- Add agent_team/nodes/dispatch_invoker.py: make_dispatch_node() wraps
  dispatcher.dispatch_apply_verify() as a LangGraph node; fail-safe
  (parks task on any error); INERT unless wired by coordinator.
- graph.py: add DISPATCH_NODE constant; add dispatch_node param to
  build_graph; when provided, repoint APPROVED_ROUTE → DISPATCH_NODE →
  END (P3+). Validates dispatch_node requires build_verify.
- coordinator.py: add DispatchNodeFactory type; thread dispatch_node_wiring
  through __init__ and setup(); default None = inert (no auto-dispatch).
- agent-team-apply-verify.yml: remove agent-apply environment gate from
  gate-and-pr; add Slack-notify + audit-log compensating control steps.
- test_apply_verify_workflow_hardening.py: flip env assertion → assert env
  REMOVED and compensating steps present.

All 1044 tests passing. ruff clean.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QYp761G9HojkmLqLZrASVi
2026-06-23 01:31:29 +00:00
Claude
e5a8949d0d
feat(ws1): non-Claude in-process invokers + HTTP API (WS1, code)
- invoker_multi.py: multi_invoke(prompt, *, model) dispatches to GPT-4.1
  (cross_reviewer), DeepSeek (fast_coder), or Gemini (scanner) in-process;
  bind_multi_invoker() wires the review seam; lazy models import
- review_loop_llm: swap default_plan_reviewer from make_run_py_invoker()
  to make_cross_reviewer_invoker() (in-process GPT-4.1); subprocess path
  retained as make_run_py_invoker() for opt-in use
- builders_llm: add make_fast_coder_invoker() in-process DeepSeek path;
  rename subprocess path to subprocess_build (opt-in fallback); default_build
  now delegates to make_fast_coder_invoker()
- api.py: FastAPI app with bearer-token auth (AGENT_TEAM_API_TOKEN env),
  POST /tasks, GET /tasks/{id}, POST /orchestrator/invoke; binds 127.0.0.1;
  code only — not started here
- run-team.py: bind_multi_invoker() called in serve() alongside
  bind_subscription_invoker()
- tests: 21 new WS1 tests + updated builders_llm tests (1065 total, all passing)
2026-06-23 01:19:59 +00:00
Claude
a12de32a34
feat(ws5): add memory/handbook injection seams for context-provider pattern
- retriever: add save_memory() writing to _box-drafts/ review queue, add
  memory_dir param to retrieve() for isolated test routing
- handbook: new load_handbook_conventions() with safe no-op contract (returns
  "" when dir absent/empty/unreadable, never raises)
- clarifier_llm: add context_provider=None seam to ClaudeClarifier and
  build_claude_clarifier_callables(); failure in provider is silent
- planner: add context_provider=None seam to build_plan_prompt() and plan_node()
- coordinator: thread context_provider through default_plan_node_factory(),
  only pass kwarg when non-None to preserve stub-monkeypatching in tests
- tests: 19 new WS5 tests covering all seams (1063 total, all passing)
2026-06-23 01:08:34 +00:00
Adam Moussa
5756178d62
feat(agent-team): wire apply/verify into .github/workflows (make it a live GitHub Actions workflow) (#37)
GitHub Actions only runs workflows under .github/workflows/, so the apply/verify
workflow at agent-team/ci/ was never registered (workflow_dispatch 404'd). Move it
to .github/workflows/agent-team-apply-verify.yml so it is a real, dispatchable
workflow. Its only trigger is workflow_dispatch + it is gated by the agent-apply
required-reviewer environment, so it never auto-runs and nothing privileged runs
unapproved. Updated the two workflow test files' path refs (parents[2]/.github/
workflows) and the ci/README pointer. Dispatcher push gains --no-verify: the apply
path is scanned CI-side (guard + the PR's checks), so it must not be blocked by the
operator's LOCAL human-commit pre-push dev hook (which flags pre-existing whole-repo
FPs like .env.example). 1044 tests, ruff clean.
2026-06-22 19:12:17 -04:00
Adam Moussa
03b9a94881
feat(agent-team): P3-flip Phase 1 — CI trust-boundary hardening (WIP, gated) (#34)
* feat(agent-team): P3-flip Phase 1 — expand denylist vectors (§4.2) + runner-trust assertion (§4.1)

First controls of the P3-live-flip Phase-1 CI hardening (workflow stays INERT;
this only tightens the trust boundary). Whole Phase-1 surface is gated by
/sh-security-review + GPT-4.1 cross-review before any flip.

§4.2 — expand the trust-control denylist with direct code-execution / supply-chain
vectors, kept byte-identical across all three copies (ci_gate.DENYLIST_GLOBS + the
guard + post-build inline DENY_GLOBS), drift-guarded:
  .gitmodules, .husky/**, .githooks/**, .gitattributes, .npmrc, and generated/build
  artifacts (__generated__, *.generated.*, dist/**, build/**, *.min.js).
Deliberate: lockfiles are NOT wholesale denied — lockfile-postinstall RCE is already
contained by the credential-less egress-blocked build sandbox, and the Tier-3 dep-CVE
fixer rewrites lockfiles to produce its draft PRs; a blanket deny would make it
un-shippable. Flagged in-code for the security gate. Direct code-execution config
(hooks/filters/npmrc/submodules) is the actual §4.2 RCE surface.

§4.1 — runner-trust: assert no job (esp. the privileged gate-and-pr) can run on a
self-hosted/user-provided runner; all must be GitHub-hosted.

998 tests pass, ruff clean.

* feat(agent-team): P3-flip Phase 1 — gate-weakening detector (§4.5)

A diff that ADDS a lint/type/coverage/security suppression (noqa, type: ignore,
pragma: no cover, nosec, nosemgrep), a test skip/xfail, or a hook bypass
(--no-verify) could make CI pass falsely. The pure-code gate now flags these via
gate_weakening_violations() and BLOCKs in evaluate_ci_gate as a top-priority trust
violation (step 1b, alongside the denylist) — regardless of the authenticated CI
conclusion. A build cannot pass itself by disabling its own checks; flagged diffs
escalate to a human. Only ADDED lines are inspected (removing a suppression is fine).

1015 tests pass, ruff clean.

* feat(agent-team): P3-flip — diff transport (§4.3) + flip privileged apply path live

Completes the box->CI diff handoff and flips the apply/verify privileged job
live (gated behind the agent-apply environment's required reviewer).

Transport (§4.3): the read-only box (D2) emits a diff but holds no write token.
- New credential-less `materialize` job decodes the untrusted `diff_b64`
  dispatch input via env (CWE-94), fail-closed re-hashes it against
  `expected_diff_hash`, and uploads it as the named artifact so guard/build-test
  download it same-run. guard now `needs: materialize`.
- New `dispatcher.py` (the trusted apply path, operator/Mac-side — never the
  box): pushes the diff as a head branch then `gh workflow run`s the workflow.
  Pure input-assembly (sha256 == sha256sum, b64 round-trip, head ref) is
  unit-tested; git/gh are injected seams. Push-before-dispatch; fail-closed on
  empty diff/scope, unsafe task_id/owner/repo.

Flip: gate-and-pr binds `environment: agent-apply` (required reviewer
amoussa1229) + grants exactly `pull-requests: write`; the App-token + draft-PR
steps run only on `steps.gate.outputs.gate == 'pass'` (no more if:false); the
draft PR opens with an explicit `--head`; task_id/head_branch charset-validated
(§4.6). Updated the hardening tests from inert-state to live-state assertions +
added transport tests. 1039 tests, ruff clean, workflow YAML valid.

NOTE: workflow only runs on manual workflow_dispatch and the privileged job is
held at the required-reviewer gate, so nothing privileged runs unapproved.

* fix(agent-team): P3-flip — address GPT-4.1 cross-review (size bound, ref-traversal guard)

- BLOCK: cap candidate diff at 40 KB in the dispatcher (the diff rides a base64
  workflow_dispatch input; GitHub caps inputs at ~64 KB so an oversized diff
  cannot dispatch at all) + a defense-in-depth decoded-size bound in materialize.
- FIX: harden the draft-PR HEAD_BRANCH guard to reject leading/trailing slash,
  '..' segments, and '//' (CWE-88 git ref-traversal), not just bad charset.
- NIT: document the mandatory invariants on gate-and-pr (required-reviewer
  environment must stay; runs-on must stay GitHub-hosted).
- QUESTION (lockfiles): answered in-code — the build-test sandbox is
  credential-less + egress-blocked, so lockfile-postinstall RCE is contained.
Tests added for all guards. 1042 tests, ruff clean, YAML valid.

* fix(agent-team): P3-flip — resolve /sh-security-review findings (LOGIC-1/2/3)

High-recall fan-out (injection/logic/iac+secrets) + proof-or-kill on the LIVE
apply path found 3 real issues the cross-review missed; all fixed:

- LOGIC-2 (HIGH, was a live hole): build-test ran `ruff check . || echo` /
  `pytest -q || echo`, swallowing failures so the job was always 'success' and
  the gate would open draft PRs on RED builds. ruff/pytest now run
  authoritatively under set -e (pytest exit 5 'no tests' is the only non-fatal
  case); the exit code IS the build-test conclusion the gate keys on.
- LOGIC-1 (verified!=shipped): the dispatcher used `git apply` + `git add -A`,
  staging stray untracked content into the pushed PR head. Now `git apply
  --index` stages exactly the diff, so the head tree is precisely base+diff —
  bound to the bytes CI hash-verified.
- LOGIC-3 (§4.5 on the live path): gate-weakening was enforced only box-side;
  added a gate-weakening check to the guard job so the live PR-opening path
  rejects a diff that adds suppressions/skips, even on a green build.

Injection / secrets / least-privilege / flip-correctness / no-untrusted-checkout
all came back clean. 1044 tests, ruff clean, YAML valid.
2026-06-22 18:51:52 -04:00
Adam Moussa
f916c03818
feat(agent-team): durable GitHub-issue intake de-dup + intake hardening (#32)
* feat(agent-team): durable GitHub-issue intake de-dup (schema v2)

The intake poller de-duped ingested issues in an in-memory set that does
not survive a process restart. A scheduled/cron intake (each run a fresh
process) would therefore re-ingest every still-open labeled issue on every
run and spawn duplicate pipeline tasks. Since the box is read-only (no write
token to remove the intake label), durable de-dup is the only correct guard.

- schema v2: new ingested_issues(source, issue_id, ingested_at) table +
  issue_already_ingested / record_issue_ingested helpers; migrate() adds the
  table to a legacy v1 DB and restamps; init_db creates it.
- github_intake: pluggable IngestStore seam (in-memory default preserved for
  tests/one-off; durable build_ledger_ingest_store for production). Record is
  after start_task succeeds, so a failed intake stays retryable.
- run-team.py intake-github wires the ledger store keyed by github:owner/repo,
  making a scheduled timer idempotent across runs.

978 tests pass (ruff clean).

* fix(agent-team): harden github intake per /sh-security-review (CWE-918, idempotency)

Fixes from the high-recall detector fan-out on the durable-dedup change:

- INTAKE-LOGIC-01 (idempotency): switch the IngestStore seam from
  check-then-record (seen/mark) to claim-then-do (claim/release). The id is
  now reserved BEFORE the non-idempotent start_task side effect, so a crash in
  that window cannot re-spawn a duplicate task on the next run; a raising
  start_task releases the claim so transient failures stay retryable. Adds
  delete_issue_ingested to the schema layer for the release path.
- INTAKE-SSRF-001 / INTAKE-PATHSPLICE-002 (CWE-918) in build_default_issue_client:
  drop the caller-overridable api_root (hardcode GITHUB_API_ROOT) and validate
  owner/repo against an anchored charset before splicing them into the
  token-bearing API URL — mirrors the sibling ci_fetcher BLOCK-3/FIX-3 fixes.

Tests cover cross-process duplicate prevention, release-on-failure retry, and
the owner/repo + api_root rejection. 982 tests pass, ruff clean.

Follow-up (pre-existing, not introduced here): the label-only intake has no
author allowlist (cf. AGENT_TEAM_SLACK_OWNER_IDS on the Slack listener); the
Slack answer gate bounds the blast radius. Track as separate hardening.
2026-06-22 17:20:43 -04:00
Adam Moussa
c671e6bdbb
fix(agent-team): register QuestionSet with the langgraph checkpoint serializer (silence/avoid msgpack block) (#30) 2026-06-22 16:14:27 -04:00
Adam Moussa
a6275b4000
fix(agent-team): harden listener respawn/close, broaden handle_event guard, channel_ref partial-unique (#29)
Three robustness/hardening fixes surfaced by /sh-security-review on the
agent-team listener/coordinator surface. None alter AUTHZ-01 allowlist
behavior or the first-answer-wins compare-and-set semantics.

1. Respawn close() leak (CWE-772). The watchdog _supervise_slack_listener
   respawned the inbound Slack listener without tearing down the dead one,
   leaking a Socket Mode WebSocket / SDK thread set per flap. Now the dead
   listener is closed before respawn (new _close_dead_listener, idempotent),
   AND _run_listener has a finally that always closes the listener so a
   crashed serve() releases its socket. SlackListener.close() is idempotent,
   so the belt-and-braces close stays a safe no-op.

2. Broadened exception guard in handle_event (CWE-248). The submit block
   only caught ValueError; the accept path (submit_answer ->
   _question_turn/_question_thread) can raise KeyError on a concurrently
   mutated row, and the CAS can raise sqlite3.Error. An uncaught exception
   would escape into the Bolt dispatch. Added a separate `except Exception`
   that logs at WARNING (not silent, not debug) and returns None. The
   existing ValueError-as-debug behavior is unchanged; authorization still
   runs first, so the trust boundary is not widened.

3. channel_ref partial-unique index (defense-in-depth). Added
   uq_pending_questions_open_channel_ref — a PARTIAL UNIQUE index on
   (channel_ref) WHERE channel_ref IS NOT NULL AND status='open' — so two
   OPEN rows can never share a non-null channel_ref (a thread_ts can never
   map to two open questions). Installed in init_db AND unconditionally in
   migrate (idempotent IF NOT EXISTS) so existing v1 DBs gain it. NULLs and
   closed rows are excluded; mirrored verbatim into schema.sql.

Tests: +8 (was 960, now 968). New: schema partial-unique reject/null/closed/
migrate cases; handle_event KeyError + sqlite3.Error swallow cases;
coordinator close-before-respawn + run_listener-closes-on-crash. Fixed the
operator-cli test fixture to use a per-question channel_ref (it previously
inserted multiple open rows sharing one ref, which the new index correctly
rejects).
2026-06-22 16:14:22 -04:00
Adam Moussa
e85a56148e
fix(agent-team): handle real slack_bolt event envelope + map thread replies; bind start invoker (#26)
* fix(agent-team): handle real slack_bolt event envelope + map thread replies to open questions

The Socket Mode inbound listener was unit-tested against a SYNTHETIC payload
shape that does not match what slack_bolt actually delivers, so the suite was
green while a real Slack thread reply was silently dropped (the clarifier
question stayed `open`). Real slack_bolt delivers an Events API message /
app_mention as `{"type":"event_callback","event":{"type":"message",...}}` and
a free-text thread reply carries NO callback_id/question_id/metadata.

Three breaks fixed (all on the free-text reply path):

1. Type gate — handle_event gated on the OUTER `type`, which is
   "event_callback" for a real message/app_mention, so the event fell outside
   _ANSWER_BEARING_TYPES and was dropped. Now collapsed to the discriminating
   INNER `event.type` via _discriminating_type / _inner_event.

2. question_id recovery — a real reply has no callback_id/question_id/metadata
   (the bot's metadata is on the QUESTION message, not the reply). When explicit
   id recovery fails, the listener now resolves the question by the inner
   event's `thread_ts` against the OPEN ledger row whose `channel_ref` equals it
   (new schema helper find_open_question_by_channel_ref, constrained to
   status='open' as anti-replay). Explicit id recovery still takes precedence.

3. answer extraction — a real message event carries its text at `event.text`,
   not a top-level `answer`/`text`. The thread-reply path now takes the inner
   `event.text` (stripped) as the answer value.

AUTHZ-01 is unchanged and still runs FIRST: authorization gates on the sender's
Slack user id (`event.user` for the Events API shape) and fails closed on an
empty/unknown allowlist or unrecoverable sender. The new mapping only resolves
WHICH question is answered, never WHO may answer. Answers stay opaque DATA
(parameterized SQL + json.dumps; never eval/exec/interpolate).

Tests: replaced the synthetic events-API fixtures with REAL Bolt envelopes and
added regression coverage — real thread reply maps via channel_ref and is
accepted, text is stripped, non-owner reply rejected (row stays open), thread_ts
matching no open row is a no-op, reply to an already-answered row is a no-op
(anti-replay), and app_mention is normalized identically. block_actions /
slash_command paths retained.

* fix(agent-team): bind subscription invoker in the start CLI

`run-team.py start` runs the clarifier graph to the first human gate IN the CLI
process, and the clarifier calls Claude (assess_confidence). The invoker is a
process-local binding that only `serve` set, so `start` failed with
"claude_invoke has no invoker bound". Bind the real subscription invoker here,
mirroring Coordinator.serve(). Found during the live R720 P1 bring-up.
2026-06-22 15:30:48 -04:00
Adam Moussa
f59b293022
fix(agent-team): repair Slack listener block_actions matcher; add dedicated Slack app (#25)
The Socket Mode inbound listener crashed at registration time on first live
run: `@app.action({})` raised `BoltError: action ({}) must be any of str,
Pattern, and dict` under slack_bolt 1.28.0, killing the listener thread (the
whole inbound answer path — message/app_mention/block_actions — went down,
caught only by the coordinator's respawn watchdog). serve() is marked
`# pragma: no cover - live socket`, so this was never exercised until the R720
bring-up. Replace the unsupported empty-dict matcher with a catch-all
`re.compile(r".*")` action_id regex; handle_event still does the real filtering
+ AUTHZ-01 owner-allowlist gate, so over-matching is safe.

Verified on sh-secrev: listener connects (live Socket Mode WebSocket), outbound
chat.postMessage works, 0 errors. /sh-security-review PASS (no confirmed
critical/high; matcher change introduces no new findings).

Also adds the dedicated Slack app (manifest + README) backing the clarifier
gate — "Sea Haven agent-team" (A0BC7AT8NUD), workspace-scoped install to avoid
the Enterprise-Grid `scope_not_allowed_on_enterprise` org-install trap — and
patches the provisioning runbook's stale langgraph pin (1.1.10 -> 1.2.5).
2026-06-22 13:16:35 -04:00
Adam Moussa
2d1dca0804
feat(agent-team): deploy-readiness — serve starts Slack listener + systemd + provisioning docs (#23)
* fix(agent-team): serve() starts the inbound Slack listener (D-1)

Coordinator.serve() now constructs and starts the SlackListener concurrently
with the tick/drain loop on a background daemon thread, but ONLY when the live
transport is a SlackTransport AND SLACK_APP_TOKEN is configured. When Slack is
not the transport or the app token is absent, serve() behaves exactly as before
(tick/recover only) — Slack is never made mandatory.

- New injectable build_listener seam + default_slack_listener_factory sharing
  the coordinator's own transport, ledger db_path, and resume_queue put.
- AUTHZ-01 owner-allowlist + open-status CAS untouched: serve() sources
  AGENT_TEAM_SLACK_OWNER_IDS in SlackListener.serve, which still fails closed.
- SlackListener.close() added for clean Socket Mode teardown on shutdown;
  serve() stops the listener + joins the thread in a finally.
- Tests: start-when-Slack+app-token, no-start otherwise, clean shutdown,
  idempotent start, serve start/stop around the loop, listener close().

* fix(agent-team): systemd unit loads ~/orchestrator/.env + uses venv python (D-2/D-7)

D-2: add EnvironmentFile=-/home/adam/orchestrator/.env (optional '-') so the P2
GPT-4.1 review loop's cross_reviewer sub-process can read the non-Claude provider
key once a task reaches REVIEW. Mirrors the sea-haven-secrev unit.

D-7: point ExecStart at the agent-team venv interpreter
(/home/adam/orchestrator/agent-team/.venv/bin/python) instead of
/usr/bin/env python3, which resolved the system interpreter without the
installed deps under systemd's PATH.

All hardening (NoNewPrivileges / ProtectSystem=full / ProtectHome=read-only /
ReadWritePaths) is retained unchanged (locked decision).

* docs(agent-team): land provisioning + operator runbooks under docs/provisioning

- PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5
  intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub
  App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked
  FIXED so the demo can use the live Slack answer path.
- P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the
  Slack and operator-CLI answer paths documented for all four exit criteria.
- DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the
  operator-CLI divergence kept as provisioning notes.
- OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks,
  failed HITL resumes, budget exhaustion, transport outages, and
  COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the
  re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6).

* fix(agent-team): supervise the Slack listener thread — recurring ALARM + respawn

sh-security-review (logic) MEDIUM: a crashed listener thread was logged once,
then the daemon ran on 'deaf' — posting clarifier questions but receiving no
answers, every gate silently parking, process never exiting so systemd
Restart=on-failure never fired. serve() now calls _supervise_slack_listener()
each pass: when the listener is enabled but its thread is dead, it emits a
recurring ERROR ALARM and respawns via the idempotent starter (self-heal).
No-op when alive or disabled. +3 tests. (authz detector: wiring clean — AUTHZ-01
fail-closed allowlist + open-status CAS intact, dead listener fails SAFE.)
2026-06-18 16:56:21 -04:00
Adam Moussa
c49f97d316
feat(agent-team): Plane-1 fixer — finding→patch→CI draft-PR (dep-bumps, opt-in/inert) (#21)
* feat(agent-team): Plane-1 Tier-3 fixer — dependency-cve finding -> patch + CI dispatch (opt-in/inert)

The fixer (design §4 fixer row, §7 Phase 5, §3.3.2) takes a CONFIRMED,
low-risk dependency-cve finding (the narrowest fix class) and produces:

  * a fix SPEC (Claude, via the §3.1 billing seam), and
  * a minimal bump PATCH (DeepSeek fast_coder, via the orchestrator run.py
    path that builders_llm uses),

records the candidate diff + its content-hash, and emits the org-CI
workflow_dispatch inputs (task_id / diff_artifact_name / expected_diff_hash /
declared_scope) for the gate-passed P3-live apply/verify surface.

INERT / opt-in / fail-safe, mirroring build_verify_wiring:
  * plan_fix dispatches NOTHING; dispatch_fix has NO default dispatcher
    (the box holds no write token, D2) so an un-wired call can never fire a
    workflow.
  * no git/patch/subprocess/fs-write in executable code — the patch is emitted
    as diff TEXT only; CI applies it and opens a DRAFT PR, the box never
    applies/pushes/merges.
  * untrusted-patch hygiene: the generated diff is confined box-side to the
    single dependency manifest (declared_scope) and rejected via
    ci_gate.denylist_violations if it escapes scope or touches the
    trust-control surface — defense-in-depth with the CI guard.
  * bad/ambiguous findings (wrong check/status/category, missing
    package/fixed_version, ambiguous fixed_version, unparseable/empty diff)
    yield a FAILED no-op plan, never a fabricated fix.

29 new pytest tests under agent-team/tests/test_fixer.py.

* feat(agent-team): run-team.py 'fix --dry-run' subcommand for the Plane-1 fixer

Adds the fixer front door to the operator CLI: load one confirmed
dependency-cve finding from a dependency-cve.json report (--report
--finding-id), plan the fix, and in --dry-run print the spec + patch + the
org-CI workflow_dispatch inputs WITHOUT dispatching anything.

Opt-in/inert: the command binds NO workflow dispatcher and holds no write
token, so even an ok plan only prints; live dispatch is provisioning-gated
(refuses to run without --dry-run). A non-fixable finding prints the
fail-safe reason and exits 1.

4 new pytest tests under agent-team/tests/test_run_team.py.
2026-06-18 16:17:33 -04:00
Adam Moussa
02ff708593
feat(agent-team): P5 cross-plane loop — checker finding -> pipeline task (opt-in) (#18)
* feat(agent-team): P5 checker-finding intake module + tests

Add agent_team.transport.checker_intake: turn a confirmed, at/above-threshold
Plane-1 checker FINDING into one Plane-2 pipeline remediation task via the
committed coordinator intake entry (start_task), mirroring github_intake.

- select_findings: status==confirmed AND severity>=threshold (default high);
  unverified/suppressed/below-threshold dropped; unknown threshold rejected.
- finding_identity: stable de-dup key (finding id, else content-hash). In-memory
  set, best-effort, NOT durable across restart (ledger table is the follow-up).
- finding_task_text/_sanitize: every repo-controlled field (title, proof, repo)
  is newline/control-char neutralised and length-bounded before it reaches the
  task text or operator log (log-injection hygiene).
- load_report_findings/ingest_reports: read the exact checker report JSON shape
  (top-level object with findings[]; bare array and dir-of-*.json also accepted).

28 hermetic unit tests (stub coordinator, in-memory findings / temp reports).

* feat(agent-team): wire opt-in intake-checker run-team subcommand

Expose the P5 cross-plane loop only as a manual run-team subcommand
(intake-checker --report PATH [--threshold] [--transport] [--dry-run]),
mirroring how intake-github is exposed. NOT wired into the always-on serve
path: the loop stays opt-in/inert by default.
2026-06-18 15:56:52 -04:00
Adam Moussa
3d97139300
feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17)
* feat(agent-team): read-only CI-result fetcher for P3 verify gate (opt-in, inert)

ci_fetcher.py: fail-closed CiResultFetcher reading the GitHub Actions run
conclusion via a read-only PAT (AGENT_TEAM_CI_READ_TOKEN→GITHUB_TOKEN), returns
{run_id,conclusion,diff_hash} or None on any error. Data-fetcher only — ci_gate
owns the verdict; never writes, no OIDC/AWS, never reads patch artifacts.
coordinator gains opt-in gated_build_verify_wiring() composing it via
bind_ci_result_fetcher; NOT wired into the default run-team.py path. 20 tests.

* harden(agent-team): P3 apply/verify workflow — GitHub App token, CWE-94, fail-closed

Decision-1 auth model: gate-and-pr uses a GitHub App installation token
(pull-requests:write) behind the agent-apply environment; ALL OIDC/id-token/AWS
removed. Hardening: task_id env-indirection (CWE-94 — GitHub expands ${{ }} into
the run shell before exec, so %s/quoting is insufficient); run-id pinning on both
download-artifact; post-build denied-path check (build-hook writes into denied
paths fail the job); empty-hash fail-closed in BOTH the embedded gate (fixed a
real ''=='' pass bug) and ci_gate.py. App-token + draft-PR steps stay if:${{ false }}
until provisioning (App + environment + branch protection). +17 tests.

* harden(agent-team): apply P3-live security-gate fixes (GPT-4.1 xreview + sh-security-review)

BLOCK-1/FIX-4: gate-and-pr re-comments pull-requests:write + environment:agent-apply
(provisioning-time uncomment) and gains needs.guard/build-test=='success' job guard —
zero privilege until provisioning. BLOCK-2/3+FIX-5: ci_fetcher validates run_id (^[0-9]{1,20}$),
owner/repo (^[A-Za-z0-9_.-]{1,100}$), and fetched_id (int) — fail closed, no SSRF/path
injection. FIX-1: conclusion allowlist. FIX-3: api_root removed from public builder (no
injectable endpoint). INJ-02: post-build denied-path check uses NUL-delimited git output +
explicit rename parsing, no backslash mangling, non-UTF8=violation. INJ-03: all three trust-
control denylists unified to one 22-entry union + drift-guard test. Q1: documented run_id/
diff_hash trust source (dispatcher/ledger only). 884 tests, ruff clean. Privileged steps stay
if:${{ false }} until provisioning.

* build(security-review): prune .claude worktrees from deterministic scanners

Agent worktrees under .claude/worktrees/ are full repo copies; the cfn-lint
find|xargs template scan overflowed ('command line cannot be assembled') and the
pre-push hook fail-closed to BLOCK whenever a worktree was present. Prune .claude
in the cfn-lint find + semgrep/checkov excludes, and gitignore .claude/ so it is
never scanned or committed. Unblocks main-tree pushes during parallel agent work.
2026-06-18 15:53:26 -04:00
035c6e57e5 fix(agent-team): clear 3 non-blocking SAST mediums on main
clarifier_llm: sha1 -> sha256 for the non-security cache-discriminator (CWE-327 false positive). github_adapter + github_intake: inline nosemgrep on the urlopen lines (dynamic-urllib-use-detected) — the URL is built from a fixed https GitHub API base, dynamic part is the path only, no SSRF/file:// surface (extends the existing noqa:S310 trusted-host judgment to semgrep). Scanner now reports 0 mediums on the agent-team scope.
2026-06-18 14:16:31 -04:00
e1208ee563 feat(agent-team): P3-inert build/verify subgraph topology (opt-in, no live CI)
build_verify_subgraph: BUILD->VERIFY nodes + route_after_verify. build_graph gains an opt-in build_verify param that repoints the review 'build' route at the subgraph (BUILD->VERIFY->{approved->END | loop->PLAN | parked->END}); default unchanged (P2). Coordinator build_verify_wiring composes it INERT (no ci_result -> ci_gate BLOCK -> PARKED; LLM is fix-proposer only, never declares green). NOT enabled in production: the live CI apply/verify + OIDC stays held for its /sh-security-review + GPT-4.1 cross-review gate. Also escapes untrusted intake text in logs (log-injection hygiene).
2026-06-18 13:23:03 -04:00
0842ff778b feat(agent-team): P4 live github/claude_code transports + GitHub-issue intake
Live github (issue-comment poster) and claude_code (file-drop) transports, plus GithubIntake (labeled issue -> coordinator.start_task, de-duped). run-team _build_transport now wires github/claude_code live (was SystemExit) + adds the intake-github subcommand. claude_code drop-path also neutralizes backslash (defense-in-depth).
2026-06-18 13:23:02 -04:00
22edd4143a feat(agent-team): P1/P2 graph wiring + coordinator daemon + run-team start/serve
build_graph gains injected live_plan_node/review_node/route_review: P1 = plan->END, P2 = clarify->plan->review->{build|loop-back|parked}. Coordinator composes clarifier->graph->ResumeWorker, wraps planner fail-safe, binds the GPT-4.1 review loop; run-team start/serve opt production into P2. Re-delivery uses a guarded CAS so a concurrently-answered row is never clobbered (closes RACE-REDELIVER).
2026-06-18 12:56:42 -04:00
253e31b0e8 feat(agent-team): live Slack transport + Socket Mode listener with owner allowlist
slack_live: real slack_sdk poster. slack_listener: Socket Mode inbound; trust boundary = app-token auth + an explicit owner allowlist on the sender (fail-closed, rejects all if AGENT_TEAM_SLACK_OWNER_IDS unset) + the open-status CAS as anti-replay. Closes the AUTHZ-01 missing-sender-authz finding from the security review.
2026-06-18 12:56:42 -04:00
4b17e8ebd4 feat(agent-team): bind planner/review/builder/verifier nodes to their models
review_loop_llm -> GPT-4.1 cross_reviewer (orchestrator run.py); builders_llm -> DeepSeek fast_coder (INERT, proposes diff text only); verifier_llm -> ci_gate is sole PASS authority, Claude is fix-proposer only. Hardens review_loop.parse_verdict to word-boundary matching, adds a fail-closed subprocess timeout, and bind_review_node (single-arg, no LangGraph config injection). All fail safe on untrusted model output.
2026-06-18 12:56:42 -04:00
270ce93b2a feat(agent-team): bind P1 clarifier to real Claude via subscription-OAuth invoker
Adds the billing-seam invoker (claude_agent_sdk subscription-OAuth, deferred import, API/Bedrock paths) and the Claude-backed clarifier callables (ConfidenceAssessor/QuestionGenerator, one call/turn memoized on (thread_id,len,content-hash), fail-safe to 0.0 so garbage never clears the 98% human gate).
2026-06-18 12:56:42 -04:00
721cec5315 Resolve security-review BLOCK: CI-guard bypasses, denylist parity, force-resume
Addresses the confirmed findings from /sh-security-review + the GPT-4.1
cross-review of the Plane-2 scaffold. Full suite: 589 passed; ruff clean.

FIXED (proven-exploitable):
- CI-guard denylist bypass (HIGH): Python fnmatch '**/' is non-recursive, so
  root-level template.yaml/*.tf/cdk.json/*.pem/*.key/*-stack.* evaded the
  trust-control surface. Replaced fnmatch with a recursive, case-insensitive
  glob->regex matcher. (verified: fnmatch('template.yaml','**/template.yaml')==False)
- CI-guard scope bypass (HIGH): a '**' declared_scope made every path in-scope.
  Scope is now concrete-prefix confinement (reduces a glob to its leading
  metacharacter-free segments; '**' -> empty -> dropped -> unscoped reject).
- Box-side vs CI denylist divergence (MED): builders.py _DENY_PATTERNS now covers
  Terraform, *.pem/*.key, CDK stack files, .github/actions, *iam*, bare policy*.json
  (case-insensitive), matching the CI surface.
- force-resume was backwards (MED): it superseded the answered row recovery
  resumes from, making a stuck task permanently un-resumable while printing
  success. Now re-opens an EXPIRED (parked) question via a new reopen_question
  CAS helper; never supersedes an answered row; honest exit codes.
- operator attribution (MED): run-team.py --operator defaulted to "" -> now the
  OS login, so destructive actions are always attributable.
- audit-log append race (MED): replaced read-modify-rewrite (lost records under
  concurrent operators) with an O_APPEND single-line write, mode 600 enforced.
- lstrip("ab/") path-mangling in the symlink error path -> regex prefix strip.

Regression tests added across test_ci_gate_workflow / test_builders / test_run_team
/ test_schema. Design-level findings (resume-worker durability, egress breadth,
answered_at ordering, DB-swap TOCTOU, diff-hash threat-model) are pre-deployment
/ P1-build-proper and recorded with written justification in
agent-team/.security-review/suppressions.json; CI README diff-hash wording made
honest.
2026-06-17 15:16:12 -04:00
0eae5dbfc3 Prove P1 exit criteria against the real LangGraph graph; fix question_id stability
Reworks the P1 sim so the four §7.1 exit criteria are demonstrated against the
ACTUAL mechanic, not a model (resolves the verifier's "sim models the ledger,
not the LangGraph integration" finding).

- New tests/sim/test_p1_graph_integration.py drives the real agent_team.graph
  StateGraph (interrupt/Command(resume)) + the real langgraph SqliteSaver
  checkpointer + the committed pending_questions compare-and-set, proving:
  (a) suspend survives a simulated restart (drop saver/conn, rebuild over the
  same checkpoint DB) and resumes; (b) duplicate answer loses the CAS and the
  graph never double-advances; (c) a post-deadline answer loses to expire and
  the task is not resumed; (d) two concurrent tasks resume to the correct
  thread, with a turn-guarded no-double-apply check.
- graph.py: derive a STABLE question_id from uuid5(thread_id, turn). The
  clarifier node replays on resume, so the prior fresh-uuid id changed between
  the delivered/ledgered question and the qa_history entry — breaking the
  §3.3.1 identity contract. Now the delivered id == ledger key == history entry
  (unit-tested in test_graph.py).
- harness._connect() now uses the committed schema.connect() (WAL + busy_timeout)
  instead of a raw sqlite3.connect, so concurrent responders genuinely serialize;
  the criterion-(d) concurrency test no longer swallows OperationalError (it
  asserts zero errors + exactly one CAS winner).
- requirements.txt: pin langgraph-checkpoint-sqlite==3.1.0 (design D9 durable
  checkpointer), now exercised by the integration test.

Full suite: 564 passed; ruff + format clean.
2026-06-17 15:16:12 -04:00
3c29ce3fdb Fix verified P1 findings: denylist bypasses, CAS concurrency, operator CLI
Resolves three execution-proven verifier findings from the scaffold review.
Full suite: 548 passed, 1 skipped (stable across repeated runs); ruff clean.

builders denylist (§3.3.2 #2): scan was +++-only and missed header-only
sections. Now section-driven off `diff --git a/<src> b/<dest>`, catching the 4
proven bypasses — delete of a denied path, mode-change-only, `copy to` a denied
path, out-of-scope delete (regression tests for each).

§3.3.1 compare-and-set concurrency: BEGIN IMMEDIATE moved inside guarded retry;
each CAS now runs on its own connection (shared sqlite3.Connection cannot hold
two transactions, and is unsafe for concurrent use even for reads). connect()
stashes the db path on a Connection subclass so the path is derived by a
thread-safe attribute read, not a PRAGMA on the shared conn; busy_timeout set
before the WAL pragma. Added shared-connection concurrent regression tests
(distinct + same question) — previously raised "transaction within a
transaction".

operator CLI (run-team.py): added the design-named re-deliver and force-resume
verbs (were missing); audit now records the attempt BEFORE the mutation and the
outcome after, so a ledger mutation can never land without a trail; main()
catches OSError instead of leaving an uncaught traceback on audit-write failure.
2026-06-17 15:16:12 -04:00
15a416d31a Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.

Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled

KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven

Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 15:16:12 -04:00
dc2da449fd Plane 2 foundation: interfaces, SQLite schemas, state-store, billing seam 2026-06-17 15:16:12 -04:00