* feat(agent-team): Plane-1 Tier-3 fixer — dependency-cve finding -> patch + CI dispatch (opt-in/inert)
The fixer (design §4 fixer row, §7 Phase 5, §3.3.2) takes a CONFIRMED,
low-risk dependency-cve finding (the narrowest fix class) and produces:
* a fix SPEC (Claude, via the §3.1 billing seam), and
* a minimal bump PATCH (DeepSeek fast_coder, via the orchestrator run.py
path that builders_llm uses),
records the candidate diff + its content-hash, and emits the org-CI
workflow_dispatch inputs (task_id / diff_artifact_name / expected_diff_hash /
declared_scope) for the gate-passed P3-live apply/verify surface.
INERT / opt-in / fail-safe, mirroring build_verify_wiring:
* plan_fix dispatches NOTHING; dispatch_fix has NO default dispatcher
(the box holds no write token, D2) so an un-wired call can never fire a
workflow.
* no git/patch/subprocess/fs-write in executable code — the patch is emitted
as diff TEXT only; CI applies it and opens a DRAFT PR, the box never
applies/pushes/merges.
* untrusted-patch hygiene: the generated diff is confined box-side to the
single dependency manifest (declared_scope) and rejected via
ci_gate.denylist_violations if it escapes scope or touches the
trust-control surface — defense-in-depth with the CI guard.
* bad/ambiguous findings (wrong check/status/category, missing
package/fixed_version, ambiguous fixed_version, unparseable/empty diff)
yield a FAILED no-op plan, never a fabricated fix.
29 new pytest tests under agent-team/tests/test_fixer.py.
* feat(agent-team): run-team.py 'fix --dry-run' subcommand for the Plane-1 fixer
Adds the fixer front door to the operator CLI: load one confirmed
dependency-cve finding from a dependency-cve.json report (--report
--finding-id), plan the fix, and in --dry-run print the spec + patch + the
org-CI workflow_dispatch inputs WITHOUT dispatching anything.
Opt-in/inert: the command binds NO workflow dispatcher and holds no write
token, so even an ok plan only prints; live dispatch is provisioning-gated
(refuses to run without --dry-run). A non-fixable finding prints the
fail-safe reason and exits 1.
4 new pytest tests under agent-team/tests/test_run_team.py.
* feat(agent-team): P5 checker-finding intake module + tests
Add agent_team.transport.checker_intake: turn a confirmed, at/above-threshold
Plane-1 checker FINDING into one Plane-2 pipeline remediation task via the
committed coordinator intake entry (start_task), mirroring github_intake.
- select_findings: status==confirmed AND severity>=threshold (default high);
unverified/suppressed/below-threshold dropped; unknown threshold rejected.
- finding_identity: stable de-dup key (finding id, else content-hash). In-memory
set, best-effort, NOT durable across restart (ledger table is the follow-up).
- finding_task_text/_sanitize: every repo-controlled field (title, proof, repo)
is newline/control-char neutralised and length-bounded before it reaches the
task text or operator log (log-injection hygiene).
- load_report_findings/ingest_reports: read the exact checker report JSON shape
(top-level object with findings[]; bare array and dir-of-*.json also accepted).
28 hermetic unit tests (stub coordinator, in-memory findings / temp reports).
* feat(agent-team): wire opt-in intake-checker run-team subcommand
Expose the P5 cross-plane loop only as a manual run-team subcommand
(intake-checker --report PATH [--threshold] [--transport] [--dry-run]),
mirroring how intake-github is exposed. NOT wired into the always-on serve
path: the loop stays opt-in/inert by default.
* feat(agent-team): read-only CI-result fetcher for P3 verify gate (opt-in, inert)
ci_fetcher.py: fail-closed CiResultFetcher reading the GitHub Actions run
conclusion via a read-only PAT (AGENT_TEAM_CI_READ_TOKEN→GITHUB_TOKEN), returns
{run_id,conclusion,diff_hash} or None on any error. Data-fetcher only — ci_gate
owns the verdict; never writes, no OIDC/AWS, never reads patch artifacts.
coordinator gains opt-in gated_build_verify_wiring() composing it via
bind_ci_result_fetcher; NOT wired into the default run-team.py path. 20 tests.
* harden(agent-team): P3 apply/verify workflow — GitHub App token, CWE-94, fail-closed
Decision-1 auth model: gate-and-pr uses a GitHub App installation token
(pull-requests:write) behind the agent-apply environment; ALL OIDC/id-token/AWS
removed. Hardening: task_id env-indirection (CWE-94 — GitHub expands ${{ }} into
the run shell before exec, so %s/quoting is insufficient); run-id pinning on both
download-artifact; post-build denied-path check (build-hook writes into denied
paths fail the job); empty-hash fail-closed in BOTH the embedded gate (fixed a
real ''=='' pass bug) and ci_gate.py. App-token + draft-PR steps stay if:${{ false }}
until provisioning (App + environment + branch protection). +17 tests.
* harden(agent-team): apply P3-live security-gate fixes (GPT-4.1 xreview + sh-security-review)
BLOCK-1/FIX-4: gate-and-pr re-comments pull-requests:write + environment:agent-apply
(provisioning-time uncomment) and gains needs.guard/build-test=='success' job guard —
zero privilege until provisioning. BLOCK-2/3+FIX-5: ci_fetcher validates run_id (^[0-9]{1,20}$),
owner/repo (^[A-Za-z0-9_.-]{1,100}$), and fetched_id (int) — fail closed, no SSRF/path
injection. FIX-1: conclusion allowlist. FIX-3: api_root removed from public builder (no
injectable endpoint). INJ-02: post-build denied-path check uses NUL-delimited git output +
explicit rename parsing, no backslash mangling, non-UTF8=violation. INJ-03: all three trust-
control denylists unified to one 22-entry union + drift-guard test. Q1: documented run_id/
diff_hash trust source (dispatcher/ledger only). 884 tests, ruff clean. Privileged steps stay
if:${{ false }} until provisioning.
* build(security-review): prune .claude worktrees from deterministic scanners
Agent worktrees under .claude/worktrees/ are full repo copies; the cfn-lint
find|xargs template scan overflowed ('command line cannot be assembled') and the
pre-push hook fail-closed to BLOCK whenever a worktree was present. Prune .claude
in the cfn-lint find + semgrep/checkov excludes, and gitignore .claude/ so it is
never scanned or committed. Unblocks main-tree pushes during parallel agent work.
* fix(agent-team): read SLACK_CHANNEL_ID, aligning code with deploy doc + systemd
run-team.py read os.environ['SLACK_CHANNEL'] while DEPLOY-R720.md and the
coordinator systemd unit both document SLACK_CHANNEL_ID; the mismatch would
silently default the live Slack transport channel to empty. Standardize on
SLACK_CHANNEL_ID (decision locked 2026-06-18).
* feat(secrev): dependency-cve Plane-1 Tier-1 checker (OSV, ALARM-only)
Read-only checker on the Phase-0 substrate: scans $MIRROR_DIR mirrors for
pinned deps (requirements/poetry/Pipfile/package-lock/yarn/csproj across
PyPI/npm/NuGet), cross-refs OSV querybatch (live) or an offline advisory
fixture (canary). Mode-600 reports, ALARM-only, --canary asserts 2 planted
vulns (jinja2 2.11.2, lodash 4.17.15). Complements Dependabot. Not provisioned.
* feat(secrev): Plane-1 checker coordinator (shared budget, rotation, dedup)
Coordinator (design §5/§6.7) orchestrating Tier-1 checkers under one shared
budget ledger + versioned rotation/coverage state (atomic write + schema/hash/
logical-consistency integrity, park-on-corrupt). Canary-suite-first
(COMPLACENCY skip), fan-out under the shared cap with defer-not-drop, COVERAGE
alarm past MAX_CYCLE_NIGHTS, cross-checker dedup/prioritize, ALARM-only routing.
--squeeze-dry-run proves deferral-not-drop + COVERAGE alarm. Not provisioned.
* fix(secrev): hide dependency-cve canary manifests from dependency-review
The canary fixtures intentionally pin known-vulnerable deps (jinja2 2.11.2,
lodash 4.17.15) so the checker has something to detect. GitHub's dependency
graph parsed those fixture manifests as real project deps, failing the
dependency-review PR gate (fail-on-severity: high). Store the manifests with a
.fixture suffix so the dependency graph ignores them; the --canary materializer
strips the suffix in its temp work area before scanning, so detection is
unchanged (still 2/2). No advisory allowlist, no change to the shared org
reusable workflow — the real gate stays strict for actual deps.
The README still described 'FOUNDATION modules only — no coordinator, no transports, no CI workflow'. Updated to reflect the built+merged pipeline (P1 human gate, P2 planner+review, live coordinator+transports+intake, P3-inert build/verify subgraph, P4) with the current layout, the pipeline diagram, key design points, and the deploy-gated/not-built items (P3 live CI, provisioning).
clarifier_llm: sha1 -> sha256 for the non-security cache-discriminator (CWE-327 false positive). github_adapter + github_intake: inline nosemgrep on the urlopen lines (dynamic-urllib-use-detected) — the URL is built from a fixed https GitHub API base, dynamic part is the path only, no SSRF/file:// surface (extends the existing noqa:S310 trusted-host judgment to semgrep). Scanner now reports 0 mediums on the agent-team scope.
build_verify_subgraph: BUILD->VERIFY nodes + route_after_verify. build_graph gains an opt-in build_verify param that repoints the review 'build' route at the subgraph (BUILD->VERIFY->{approved->END | loop->PLAN | parked->END}); default unchanged (P2). Coordinator build_verify_wiring composes it INERT (no ci_result -> ci_gate BLOCK -> PARKED; LLM is fix-proposer only, never declares green). NOT enabled in production: the live CI apply/verify + OIDC stays held for its /sh-security-review + GPT-4.1 cross-review gate. Also escapes untrusted intake text in logs (log-injection hygiene).
Live github (issue-comment poster) and claude_code (file-drop) transports, plus GithubIntake (labeled issue -> coordinator.start_task, de-duped). run-team _build_transport now wires github/claude_code live (was SystemExit) + adds the intake-github subcommand. claude_code drop-path also neutralizes backslash (defense-in-depth).
CI lacks slack_sdk, so the deferred-import-missing error fired before the no-token check and masked it. Stub slack_sdk into sys.modules so the token branch is deterministically exercised in both environments.
P1c deploy artifacts: DEPLOY-R720.md (snapshot-first, rsync, venv deps, ~/secrev.env tokens incl. AGENT_TEAM_SLACK_OWNER_IDS, init-db, systemd, the 4-criteria live demo, rollback) and the long-running coordinator service unit.
build_graph gains injected live_plan_node/review_node/route_review: P1 = plan->END, P2 = clarify->plan->review->{build|loop-back|parked}. Coordinator composes clarifier->graph->ResumeWorker, wraps planner fail-safe, binds the GPT-4.1 review loop; run-team start/serve opt production into P2. Re-delivery uses a guarded CAS so a concurrently-answered row is never clobbered (closes RACE-REDELIVER).
slack_live: real slack_sdk poster. slack_listener: Socket Mode inbound; trust boundary = app-token auth + an explicit owner allowlist on the sender (fail-closed, rejects all if AGENT_TEAM_SLACK_OWNER_IDS unset) + the open-status CAS as anti-replay. Closes the AUTHZ-01 missing-sender-authz finding from the security review.
review_loop_llm -> GPT-4.1 cross_reviewer (orchestrator run.py); builders_llm -> DeepSeek fast_coder (INERT, proposes diff text only); verifier_llm -> ci_gate is sole PASS authority, Claude is fix-proposer only. Hardens review_loop.parse_verdict to word-boundary matching, adds a fail-closed subprocess timeout, and bind_review_node (single-arg, no LangGraph config injection). All fail safe on untrusted model output.
Adds the billing-seam invoker (claude_agent_sdk subscription-OAuth, deferred import, API/Bedrock paths) and the Claude-backed clarifier callables (ConfidenceAssessor/QuestionGenerator, one call/turn memoized on (thread_id,len,content-hash), fail-safe to 0.0 so garbage never clears the 98% human gate).
Addresses the confirmed findings from /sh-security-review + the GPT-4.1
cross-review of the Plane-2 scaffold. Full suite: 589 passed; ruff clean.
FIXED (proven-exploitable):
- CI-guard denylist bypass (HIGH): Python fnmatch '**/' is non-recursive, so
root-level template.yaml/*.tf/cdk.json/*.pem/*.key/*-stack.* evaded the
trust-control surface. Replaced fnmatch with a recursive, case-insensitive
glob->regex matcher. (verified: fnmatch('template.yaml','**/template.yaml')==False)
- CI-guard scope bypass (HIGH): a '**' declared_scope made every path in-scope.
Scope is now concrete-prefix confinement (reduces a glob to its leading
metacharacter-free segments; '**' -> empty -> dropped -> unscoped reject).
- Box-side vs CI denylist divergence (MED): builders.py _DENY_PATTERNS now covers
Terraform, *.pem/*.key, CDK stack files, .github/actions, *iam*, bare policy*.json
(case-insensitive), matching the CI surface.
- force-resume was backwards (MED): it superseded the answered row recovery
resumes from, making a stuck task permanently un-resumable while printing
success. Now re-opens an EXPIRED (parked) question via a new reopen_question
CAS helper; never supersedes an answered row; honest exit codes.
- operator attribution (MED): run-team.py --operator defaulted to "" -> now the
OS login, so destructive actions are always attributable.
- audit-log append race (MED): replaced read-modify-rewrite (lost records under
concurrent operators) with an O_APPEND single-line write, mode 600 enforced.
- lstrip("ab/") path-mangling in the symlink error path -> regex prefix strip.
Regression tests added across test_ci_gate_workflow / test_builders / test_run_team
/ test_schema. Design-level findings (resume-worker durability, egress breadth,
answered_at ordering, DB-swap TOCTOU, diff-hash threat-model) are pre-deployment
/ P1-build-proper and recorded with written justification in
agent-team/.security-review/suppressions.json; CI README diff-hash wording made
honest.
Reworks the P1 sim so the four §7.1 exit criteria are demonstrated against the
ACTUAL mechanic, not a model (resolves the verifier's "sim models the ledger,
not the LangGraph integration" finding).
- New tests/sim/test_p1_graph_integration.py drives the real agent_team.graph
StateGraph (interrupt/Command(resume)) + the real langgraph SqliteSaver
checkpointer + the committed pending_questions compare-and-set, proving:
(a) suspend survives a simulated restart (drop saver/conn, rebuild over the
same checkpoint DB) and resumes; (b) duplicate answer loses the CAS and the
graph never double-advances; (c) a post-deadline answer loses to expire and
the task is not resumed; (d) two concurrent tasks resume to the correct
thread, with a turn-guarded no-double-apply check.
- graph.py: derive a STABLE question_id from uuid5(thread_id, turn). The
clarifier node replays on resume, so the prior fresh-uuid id changed between
the delivered/ledgered question and the qa_history entry — breaking the
§3.3.1 identity contract. Now the delivered id == ledger key == history entry
(unit-tested in test_graph.py).
- harness._connect() now uses the committed schema.connect() (WAL + busy_timeout)
instead of a raw sqlite3.connect, so concurrent responders genuinely serialize;
the criterion-(d) concurrency test no longer swallows OperationalError (it
asserts zero errors + exactly one CAS winner).
- requirements.txt: pin langgraph-checkpoint-sqlite==3.1.0 (design D9 durable
checkpointer), now exercised by the integration test.
Full suite: 564 passed; ruff + format clean.
Resolves the GPT-4.1 cross-review FIX items on the §3.3.2 CI apply/verify guard:
- Symlink-escape (Medium-High): reject any candidate diff that introduces a
symlink (git mode 120000). A symlink can redirect a later in-diff write into a
denied path that textual canonicalization cannot see; auto-built diffs have no
legitimate symlinks, so this fails closed (exit 7).
- Diff-parse robustness (Medium): decode the diff as strict UTF-8 and fail closed
(exit 8) instead of errors='replace', closing homoglyph/encoding evasion.
- Declared-scope canonicalization (Medium): drop parent-escaping scope globs so a
malformed scope can only shrink coverage, never widen it past repo root.
- Egress allowlist (Low-Med): explicit DEPLOY marker to parameterize the
build-test registries per target repo before enabling.
Backs the gate's correctness claim with a committed, runnable suite
(tests/test_ci_gate_workflow.py) that extracts the inline guard from the YAML and
exercises good + adversarial diffs (clean, hash mismatch, workflow delete,
copy-into-denied, symlink, non-UTF-8, out-of-scope, unscoped, escaping scope).
Corrects the README "Tests" section that claimed coverage that did not exist.
Full suite: 557 passed, 1 skipped; ruff clean.
Resolves three execution-proven verifier findings from the scaffold review.
Full suite: 548 passed, 1 skipped (stable across repeated runs); ruff clean.
builders denylist (§3.3.2 #2): scan was +++-only and missed header-only
sections. Now section-driven off `diff --git a/<src> b/<dest>`, catching the 4
proven bypasses — delete of a denied path, mode-change-only, `copy to` a denied
path, out-of-scope delete (regression tests for each).
§3.3.1 compare-and-set concurrency: BEGIN IMMEDIATE moved inside guarded retry;
each CAS now runs on its own connection (shared sqlite3.Connection cannot hold
two transactions, and is unsafe for concurrent use even for reads). connect()
stashes the db path on a Connection subclass so the path is derived by a
thread-safe attribute read, not a PRAGMA on the shared conn; busy_timeout set
before the WAL pragma. Added shared-connection concurrent regression tests
(distinct + same question) — previously raised "transaction within a
transaction".
operator CLI (run-team.py): added the design-named re-deliver and force-resume
verbs (were missing); audit now records the attempt BEFORE the mutation and the
outcome after, so a ledger mutation can never land without a trail; main()
catches OSError instead of leaving an uncaught traceback on audit-write failure.