The README still described 'FOUNDATION modules only — no coordinator, no transports, no CI workflow'. Updated to reflect the built+merged pipeline (P1 human gate, P2 planner+review, live coordinator+transports+intake, P3-inert build/verify subgraph, P4) with the current layout, the pipeline diagram, key design points, and the deploy-gated/not-built items (P3 live CI, provisioning).
clarifier_llm: sha1 -> sha256 for the non-security cache-discriminator (CWE-327 false positive). github_adapter + github_intake: inline nosemgrep on the urlopen lines (dynamic-urllib-use-detected) — the URL is built from a fixed https GitHub API base, dynamic part is the path only, no SSRF/file:// surface (extends the existing noqa:S310 trusted-host judgment to semgrep). Scanner now reports 0 mediums on the agent-team scope.
build_verify_subgraph: BUILD->VERIFY nodes + route_after_verify. build_graph gains an opt-in build_verify param that repoints the review 'build' route at the subgraph (BUILD->VERIFY->{approved->END | loop->PLAN | parked->END}); default unchanged (P2). Coordinator build_verify_wiring composes it INERT (no ci_result -> ci_gate BLOCK -> PARKED; LLM is fix-proposer only, never declares green). NOT enabled in production: the live CI apply/verify + OIDC stays held for its /sh-security-review + GPT-4.1 cross-review gate. Also escapes untrusted intake text in logs (log-injection hygiene).
Live github (issue-comment poster) and claude_code (file-drop) transports, plus GithubIntake (labeled issue -> coordinator.start_task, de-duped). run-team _build_transport now wires github/claude_code live (was SystemExit) + adds the intake-github subcommand. claude_code drop-path also neutralizes backslash (defense-in-depth).
CI lacks slack_sdk, so the deferred-import-missing error fired before the no-token check and masked it. Stub slack_sdk into sys.modules so the token branch is deterministically exercised in both environments.
P1c deploy artifacts: DEPLOY-R720.md (snapshot-first, rsync, venv deps, ~/secrev.env tokens incl. AGENT_TEAM_SLACK_OWNER_IDS, init-db, systemd, the 4-criteria live demo, rollback) and the long-running coordinator service unit.
build_graph gains injected live_plan_node/review_node/route_review: P1 = plan->END, P2 = clarify->plan->review->{build|loop-back|parked}. Coordinator composes clarifier->graph->ResumeWorker, wraps planner fail-safe, binds the GPT-4.1 review loop; run-team start/serve opt production into P2. Re-delivery uses a guarded CAS so a concurrently-answered row is never clobbered (closes RACE-REDELIVER).
slack_live: real slack_sdk poster. slack_listener: Socket Mode inbound; trust boundary = app-token auth + an explicit owner allowlist on the sender (fail-closed, rejects all if AGENT_TEAM_SLACK_OWNER_IDS unset) + the open-status CAS as anti-replay. Closes the AUTHZ-01 missing-sender-authz finding from the security review.
review_loop_llm -> GPT-4.1 cross_reviewer (orchestrator run.py); builders_llm -> DeepSeek fast_coder (INERT, proposes diff text only); verifier_llm -> ci_gate is sole PASS authority, Claude is fix-proposer only. Hardens review_loop.parse_verdict to word-boundary matching, adds a fail-closed subprocess timeout, and bind_review_node (single-arg, no LangGraph config injection). All fail safe on untrusted model output.
Adds the billing-seam invoker (claude_agent_sdk subscription-OAuth, deferred import, API/Bedrock paths) and the Claude-backed clarifier callables (ConfidenceAssessor/QuestionGenerator, one call/turn memoized on (thread_id,len,content-hash), fail-safe to 0.0 so garbage never clears the 98% human gate).
The prior exclusion (--skip-path cdk.out) stopped the CDK-repo stall but also
silenced checkov on cdk.out/<stack>.template.json — the actual deploy artifact —
losing real S3/IAM IaC coverage (CKV_AWS_53-56, CKV_AWS_111). Switch to skipping
only the cdk.out asset.<hash>/ dependency bundles (the stall cause) so the
synthesized templates are still scanned.
- --skip-path 'cdk\.out/asset\.' anchors to cdk.out so a source file literally
named asset.* is not also excluded; keeps cdk.out/*.template.json scanned.
- venv/dist/build kept as bare names (match anywhere); .venv/.aws-sam escaped.
cfn-lint + semgrep already prune these trees (prior commit). gitleaks runs in
git-mode and respects .gitignore, so cdk.out is already skipped there.
Verified: shellcheck clean; synthetic cdk.out test confirms the stack template is
scanned while asset.* is skipped; the orchestrator's own --scanners-only gate
still exits 0 with suppressions (inert on non-CDK repos: A1==A2 findings here).
Repo-root 'pytest --collect-only' failed with ImportPathMismatchError because
agent-team/tests/ and the root tests/ are both the 'tests' package. Add a root
conftest.py that excludes agent-team from the root collection, and a dedicated
agent-team-tests CI job that runs the (API-key-free) agent-team suite in its own
working dir.
Addresses the confirmed findings from /sh-security-review + the GPT-4.1
cross-review of the Plane-2 scaffold. Full suite: 589 passed; ruff clean.
FIXED (proven-exploitable):
- CI-guard denylist bypass (HIGH): Python fnmatch '**/' is non-recursive, so
root-level template.yaml/*.tf/cdk.json/*.pem/*.key/*-stack.* evaded the
trust-control surface. Replaced fnmatch with a recursive, case-insensitive
glob->regex matcher. (verified: fnmatch('template.yaml','**/template.yaml')==False)
- CI-guard scope bypass (HIGH): a '**' declared_scope made every path in-scope.
Scope is now concrete-prefix confinement (reduces a glob to its leading
metacharacter-free segments; '**' -> empty -> dropped -> unscoped reject).
- Box-side vs CI denylist divergence (MED): builders.py _DENY_PATTERNS now covers
Terraform, *.pem/*.key, CDK stack files, .github/actions, *iam*, bare policy*.json
(case-insensitive), matching the CI surface.
- force-resume was backwards (MED): it superseded the answered row recovery
resumes from, making a stuck task permanently un-resumable while printing
success. Now re-opens an EXPIRED (parked) question via a new reopen_question
CAS helper; never supersedes an answered row; honest exit codes.
- operator attribution (MED): run-team.py --operator defaulted to "" -> now the
OS login, so destructive actions are always attributable.
- audit-log append race (MED): replaced read-modify-rewrite (lost records under
concurrent operators) with an O_APPEND single-line write, mode 600 enforced.
- lstrip("ab/") path-mangling in the symlink error path -> regex prefix strip.
Regression tests added across test_ci_gate_workflow / test_builders / test_run_team
/ test_schema. Design-level findings (resume-worker durability, egress breadth,
answered_at ordering, DB-swap TOCTOU, diff-hash threat-model) are pre-deployment
/ P1-build-proper and recorded with written justification in
agent-team/.security-review/suppressions.json; CI README diff-hash wording made
honest.
Reworks the P1 sim so the four §7.1 exit criteria are demonstrated against the
ACTUAL mechanic, not a model (resolves the verifier's "sim models the ledger,
not the LangGraph integration" finding).
- New tests/sim/test_p1_graph_integration.py drives the real agent_team.graph
StateGraph (interrupt/Command(resume)) + the real langgraph SqliteSaver
checkpointer + the committed pending_questions compare-and-set, proving:
(a) suspend survives a simulated restart (drop saver/conn, rebuild over the
same checkpoint DB) and resumes; (b) duplicate answer loses the CAS and the
graph never double-advances; (c) a post-deadline answer loses to expire and
the task is not resumed; (d) two concurrent tasks resume to the correct
thread, with a turn-guarded no-double-apply check.
- graph.py: derive a STABLE question_id from uuid5(thread_id, turn). The
clarifier node replays on resume, so the prior fresh-uuid id changed between
the delivered/ledgered question and the qa_history entry — breaking the
§3.3.1 identity contract. Now the delivered id == ledger key == history entry
(unit-tested in test_graph.py).
- harness._connect() now uses the committed schema.connect() (WAL + busy_timeout)
instead of a raw sqlite3.connect, so concurrent responders genuinely serialize;
the criterion-(d) concurrency test no longer swallows OperationalError (it
asserts zero errors + exactly one CAS winner).
- requirements.txt: pin langgraph-checkpoint-sqlite==3.1.0 (design D9 durable
checkpointer), now exercised by the integration test.
Full suite: 564 passed; ruff + format clean.
Resolves the GPT-4.1 cross-review FIX items on the §3.3.2 CI apply/verify guard:
- Symlink-escape (Medium-High): reject any candidate diff that introduces a
symlink (git mode 120000). A symlink can redirect a later in-diff write into a
denied path that textual canonicalization cannot see; auto-built diffs have no
legitimate symlinks, so this fails closed (exit 7).
- Diff-parse robustness (Medium): decode the diff as strict UTF-8 and fail closed
(exit 8) instead of errors='replace', closing homoglyph/encoding evasion.
- Declared-scope canonicalization (Medium): drop parent-escaping scope globs so a
malformed scope can only shrink coverage, never widen it past repo root.
- Egress allowlist (Low-Med): explicit DEPLOY marker to parameterize the
build-test registries per target repo before enabling.
Backs the gate's correctness claim with a committed, runnable suite
(tests/test_ci_gate_workflow.py) that extracts the inline guard from the YAML and
exercises good + adversarial diffs (clean, hash mismatch, workflow delete,
copy-into-denied, symlink, non-UTF-8, out-of-scope, unscoped, escaping scope).
Corrects the README "Tests" section that claimed coverage that did not exist.
Full suite: 557 passed, 1 skipped; ruff clean.
Resolves three execution-proven verifier findings from the scaffold review.
Full suite: 548 passed, 1 skipped (stable across repeated runs); ruff clean.
builders denylist (§3.3.2 #2): scan was +++-only and missed header-only
sections. Now section-driven off `diff --git a/<src> b/<dest>`, catching the 4
proven bypasses — delete of a denied path, mode-change-only, `copy to` a denied
path, out-of-scope delete (regression tests for each).
§3.3.1 compare-and-set concurrency: BEGIN IMMEDIATE moved inside guarded retry;
each CAS now runs on its own connection (shared sqlite3.Connection cannot hold
two transactions, and is unsafe for concurrent use even for reads). connect()
stashes the db path on a Connection subclass so the path is derived by a
thread-safe attribute read, not a PRAGMA on the shared conn; busy_timeout set
before the WAL pragma. Added shared-connection concurrent regression tests
(distinct + same question) — previously raised "transaction within a
transaction".
operator CLI (run-team.py): added the design-named re-deliver and force-resume
verbs (were missing); audit now records the attempt BEFORE the mutation and the
outcome after, so a ledger mutation can never land without a trail; main()
catches OSError instead of leaving an uncaught traceback on audit-write failure.
cfn-lint, semgrep, and checkov were scanning synthesized/vendored output
(cdk.out, node_modules, .venv/venv, .aws-sam, dist, build). On CDK repos this
explodes the find/xargs arg list and stalls the scan, and flagging synthesized
templates is wrong. Prune those trees in the cfn-lint find, and pass
--exclude / --skip-path to semgrep / checkov.
First-run data (2026-06-17) showed the $20 ceiling covered only the
canary + 5 of ~22 scannable repos before pausing the rotation, leaving
16 repos un-deep-scanned that night. Raise TOTAL_BUDGET_USD default to
$120 so every repo gets a deep agentic pass each night (~22 x ~$5 +
canary, with headroom). Spend draws on the Max subscription pool; the
per-target cap ($12) and round-robin rotation are unchanged, so this is
a ceiling raise, not a per-repo cost change. Updates README, DEPLOY, and
the systemd Environment example to match.
First full VM run: canary recall 18 with the expanded Node/.NET corpus, and a real
agentic repo cost ~$3.50 (vs the $1.4 testbed). Raise CANARY_FLOOR 10->14 (catches a
language-blindness recall collapse to ~9 while keeping margin under 18) and MAX_CYCLE_NIGHTS
4->6 (at ~4-6 agentic repos/night a full 22-repo rotation takes ~4-5 nights; 4 would
false-fire the coverage alarm). Both stay env-overridable and re-tunable as data accrues.
scan_scanners communicates results via globals (T1_BLOCK/T1_CRIT/T1_HIGH); its last
statement was a bare [ $rc -eq 1 ] && T1_BLOCK=1. On a PASS (rc=0) that test is false,
so the function returned non-zero and set -e killed the whole sweep at the first passing
repo in tier 1 (right after the unbound-variable fix let it get that far). Add an explicit
return 0. Audited the rest of the tier1/tier2/summary path; scan_agentic and the others
already end on a zero-status command.
scan_scanners referenced ${slug} inside the SAME local statement that defines it
(local target=... slug=$2 result_json=...${slug}...). bash expands the local's
arguments before the builtin assigns them, so under set -u ${slug} is unbound and
the sweep died right after the canary, before tier 1 ever ran. Split the local so
slug exists first. Fix the same latent self-reference in mirror_repo (dir=...$name),
which only worked by accident because the discovery loop left a global $name.
Per the machine-level-suppressions policy, the .env.example FP suppression now lives at
~/.config/sea-haven/security-review/orchestrator/suppressions.json (resolved first by the
hooks), not a committed .security-review/suppressions.json. Gate still passes via the
machine-level file.
The root README covered only the router and omitted the security-review/ subsystem
entirely (Path A skill, review.sh gate, Path B nightly sweep, global hooks). Add a
Security Review section summarizing both paths and linking the subsystem README +
DEPLOY-R720 runbook.
The live global pre-push hook was hand-edited to resolve suppressions from a
machine-level file (${SH_SECURITY_SUPPRESSIONS_DIR:-~/.config/sea-haven/security-review}/
<repo-basename>/suppressions.json) kept out of repo history, falling back to a
repo-local .security-review/suppressions.json. The repo-sourced hooks lacked it, so
install-hooks.sh --global would overwrite the live hook and lose the feature.
Port the prefer-machine/fallback-repo-local block into hooks/pre-push, align
hooks/pre-commit to the same (else a machine-suppressed finding passes at push but
blocks at commit), and document the path + SH_SECURITY_SUPPRESSIONS_DIR override +
basename-collision caveat in the README.
run_headless.py is the Path B engine the nightly sweep calls: the 6 fresh-context
detectors + proof-or-kill verifier from /sh-security-review, run unattended via the
Claude Agent SDK on subscription OAuth (pops ANTHROPIC_API_KEY so the API key can't
silently win). Read-only tools, hermetic, per-call + total budget caps, fails toward
over-reporting. Emits the finding schema that review.sh --agent-findings consumes.
Rewrite README + DEPLOY-R720 for the clean-clone mirror model, the read-only GH_TOKEN
PAT recipe, the global hook installer, and the no-CI-by-design decision.
gitleaks git-mode scans committed history, so the .env.example placeholder flags from
history and would block the orchestrator's own pre-push gate. Suppress it with a written
justification. Also gitignore stray .adf_final*.json left by an unrelated tool.
Replace the opt-in sweep-targets allowlist with zero-wiring discovery: enumerate org
repos via the GitHub REST API (curl + read-only GH_TOKEN, no gh dependency) and mirror
each as a shallow clean clone (git clone --depth=1, default branch from the API) into
~/repo-mirrors. Scanning server-side clones keeps local .env secrets out of scope.
Tier 1 runs deterministic scanners over every repo nightly ($0 Claude); tier 2 runs the
agentic pass over a budget-bounded round-robin rotation with a persistent cycle pointer,
so the draw on the shared Max limits stays bounded and coverage never goes silently
incomplete (COVERAGE ALARM if the rotation falls behind). Skip = committed marker or
central list (marker-skips logged). Redact secrets from Slack; reports mode 600. Raise
the systemd timeout to 6h for the longer two-tier run.
The global pre-push hook, the /sh-security-review prompt, and finding.schema.json
previously lived only in ~/.config/git and ~/.claude (untracked) — unreproducible.
Source them here: add hooks/pre-push, rewrite install-hooks.sh with a --global mode
(lays down both hooks, sets core.hooksPath, links skill+schema into ~/.claude) and a
per-repo mode. Align pre-commit with pre-push (honor skip marker + suppressions). Add
semgrep p/javascript so the scanners cover the org's Node/.NET repos.