Gap-audit findings:
- G1 (blocks-feature): make_cross_reviewer_invoker / make_fast_coder_invoker did
'from models import' without putting the orchestrator root on sys.path. The
run-team serve daemon only bootstraps agent-team/, so on the live box every
GPT-4.1 plan review hit ModuleNotFoundError -> review_plan's blanket except
silently fail-closed to REQUEST_CHANGES (GPT-4.1 never actually ran). Both
in-process invokers now call invoker_multi._ensure_orchestrator_on_path()
before the deferred import. WS1 introduced this when it swapped the review
default from the subprocess invoker to in-process.
- G4 (degrades): the handbook context_provider was wired into the planner only;
default_clarify_node_factory now accepts + forwards it, and run-team wires it
into build_clarify_node too, so clarifying questions are handbook-aware.
Tests: +2 regression tests (path-bootstrap, clarifier threading); _FakeCoordinator
gains build_clarify_node. 1142 passed, ruff clean.
Integration branch combining WS0-WS5 (PRs #43-#46) + the activation wiring that
flips the safe seams ON in the run-team serve path:
- WS5 (D10): inject the Sea Haven handbook conventions into the planner prompt
via context_provider (zero-arg handbook loader; fail-safe to '' when absent).
- WS2: an allowlisted Slack /new-task starts a task on this coordinator
(set_new_task_callback adapter -> start_task; AUTHZ-01 gates it upstream).
- WS1 bind_multi_invoker() is already wired in _cmd_serve.
Coordinator gains a new_task_callback param + set_new_task_callback() (resolves
the constructor chicken-and-egg of referencing the coordinator's own start_task);
default_slack_listener_factory forwards it to the SlackListener.
NOT wired (deliberately): the P3 dispatch_node / build_verify path. Activating
it correctly needs a per-task expected_run_id bound into gated_build_verify_wiring
(plumbing that does not exist yet) AND the CI trust-boundary security re-review.
It stays inert pending that work.
Tests: +5 activation-wiring tests; _FakeCoordinator stub gains
set_new_task_callback. Full agent-team suite: 1140 passed, ruff clean.
Security follow-up from the per-PR review (non-blocking, defense-in-depth):
- Replace the save_memory name blocklist with an allowlist regex
(^[A-Za-z0-9][A-Za-z0-9._-]*$, max 128) so dot-only/hidden/backslash/NUL/
over-long names are rejected outright, not written as malformed-but-contained
files.
- Write via os.open(..., O_NOFOLLOW): the open fails (ELOOP) if the final
path component is a pre-planted symlink, closing the TOCTOU where a symlink
in _box-drafts/ could redirect the write outside the dir. O_CREAT|O_TRUNC
keeps overwrite-on-resave for regular files.
Tests: adds allowlist-rejection + symlink-refusal cases (21 pass).
Security follow-ups from the per-PR review (non-blocking MEDIUMs):
- Eager _get_token() at make_app build time so a missing AGENT_TEAM_API_TOKEN
fails fast instead of serving requests first (matches the docstring contract).
- Disable /docs, /redoc, /openapi.json (no auth dependency in FastAPI) — the
API is VPN-only/127.0.0.1 and should not expose its schema unauthenticated.
- Scrub raw exception text and subprocess stderr from 500 response bodies;
log server-side instead (avoid internal-path/state disclosure).
- Bound /orchestrator/invoke concurrency with a semaphore (429 over the cap)
so an authenticated caller cannot exhaust the box via many 600s subprocesses.
Also pin fastapi/uvicorn in requirements.txt (WS1 dep). With fastapi now
installed in CI, the previously skip-guarded TestClient tests run for real;
the importorskip guard stays as a no-op safety net.
Tests: 23 pass (adds docs-disabled + concurrency-429 cases).
Reworked per GPT-4.1 cross-family review BLOCK. The original PR removed the
agent-apply GitHub Environment (the live required-reviewer human gate) and
replaced it with a Slack notice that FAILS OPEN when its webhook secret is
absent (which it is) plus an audit-log line. The cross-review correctly
flagged this as trading a preventive control for detective controls, one of
which silently no-ops.
This commit:
- Restores environment: agent-apply on gate-and-pr (the human approval pause).
- Drops the fail-open Slack notify step.
- Keeps the unconditional audit-log step as an additive detective control.
- Restores the MANDATORY-INVARIANT assertion (env must be present) and adds
an assertion that the audit step is retained.
WS3's auto-dispatch (dispatch_invoker.py + graph/coordinator wiring) is
unchanged: it fires workflow_dispatch, which now pauses at the restored gate
for human approval — auto-dispatch up to the approval, then one click.
CI has no fastapi (box-only dependency); api.py imports it lazily. Guard the
7 TestClient smoke tests with skipif(find_spec('fastapi') is None) so they
skip in CI instead of failing collection, leaving the 14 invoker_multi tests
running. Also apply ruff format to the 5 WS1 files CI flagged.
- invoker_multi.py: multi_invoke(prompt, *, model) dispatches to GPT-4.1
(cross_reviewer), DeepSeek (fast_coder), or Gemini (scanner) in-process;
bind_multi_invoker() wires the review seam; lazy models import
- review_loop_llm: swap default_plan_reviewer from make_run_py_invoker()
to make_cross_reviewer_invoker() (in-process GPT-4.1); subprocess path
retained as make_run_py_invoker() for opt-in use
- builders_llm: add make_fast_coder_invoker() in-process DeepSeek path;
rename subprocess path to subprocess_build (opt-in fallback); default_build
now delegates to make_fast_coder_invoker()
- api.py: FastAPI app with bearer-token auth (AGENT_TEAM_API_TOKEN env),
POST /tasks, GET /tasks/{id}, POST /orchestrator/invoke; binds 127.0.0.1;
code only — not started here
- run-team.py: bind_multi_invoker() called in serve() alongside
bind_subscription_invoker()
- tests: 21 new WS1 tests + updated builders_llm tests (1065 total, all passing)
- retriever: add save_memory() writing to _box-drafts/ review queue, add
memory_dir param to retrieve() for isolated test routing
- handbook: new load_handbook_conventions() with safe no-op contract (returns
"" when dir absent/empty/unreadable, never raises)
- clarifier_llm: add context_provider=None seam to ClaudeClarifier and
build_claude_clarifier_callables(); failure in provider is silent
- planner: add context_provider=None seam to build_plan_prompt() and plan_node()
- coordinator: thread context_provider through default_plan_node_factory(),
only pass kwarg when non-None to preserve stub-monkeypatching in tests
- tests: 19 new WS5 tests covering all seams (1063 total, all passing)
Live smoke exposed a self-pollution bug (pre-existing from PR #17): the post-build
denied-path check wrote its own _build_diff_z.bin / _build_status_z.bin into the
working tree, then its own `git status --untracked-files=all` flagged them as
out-of-scope writes — failing any run with a narrow declared_scope (the smoke's
docs/**). Write them to $RUNNER_TEMP instead (read via $_DIFF_Z/$_STATUS_Z), so
the check no longer sees its own temp files. The agent-team pytest artifacts were
already correctly gitignored; only the check's own files tripped it. Test harness
updated to pass the env paths. 1044 tests, ruff clean.
GitHub Actions only runs workflows under .github/workflows/, so the apply/verify
workflow at agent-team/ci/ was never registered (workflow_dispatch 404'd). Move it
to .github/workflows/agent-team-apply-verify.yml so it is a real, dispatchable
workflow. Its only trigger is workflow_dispatch + it is gated by the agent-apply
required-reviewer environment, so it never auto-runs and nothing privileged runs
unapproved. Updated the two workflow test files' path refs (parents[2]/.github/
workflows) and the ci/README pointer. Dispatcher push gains --no-verify: the apply
path is scanned CI-side (guard + the PR's checks), so it must not be blocked by the
operator's LOCAL human-commit pre-push dev hook (which flags pre-existing whole-repo
FPs like .env.example). 1044 tests, ruff clean.
* feat(agent-team): P3-flip Phase 1 — expand denylist vectors (§4.2) + runner-trust assertion (§4.1)
First controls of the P3-live-flip Phase-1 CI hardening (workflow stays INERT;
this only tightens the trust boundary). Whole Phase-1 surface is gated by
/sh-security-review + GPT-4.1 cross-review before any flip.
§4.2 — expand the trust-control denylist with direct code-execution / supply-chain
vectors, kept byte-identical across all three copies (ci_gate.DENYLIST_GLOBS + the
guard + post-build inline DENY_GLOBS), drift-guarded:
.gitmodules, .husky/**, .githooks/**, .gitattributes, .npmrc, and generated/build
artifacts (__generated__, *.generated.*, dist/**, build/**, *.min.js).
Deliberate: lockfiles are NOT wholesale denied — lockfile-postinstall RCE is already
contained by the credential-less egress-blocked build sandbox, and the Tier-3 dep-CVE
fixer rewrites lockfiles to produce its draft PRs; a blanket deny would make it
un-shippable. Flagged in-code for the security gate. Direct code-execution config
(hooks/filters/npmrc/submodules) is the actual §4.2 RCE surface.
§4.1 — runner-trust: assert no job (esp. the privileged gate-and-pr) can run on a
self-hosted/user-provided runner; all must be GitHub-hosted.
998 tests pass, ruff clean.
* feat(agent-team): P3-flip Phase 1 — gate-weakening detector (§4.5)
A diff that ADDS a lint/type/coverage/security suppression (noqa, type: ignore,
pragma: no cover, nosec, nosemgrep), a test skip/xfail, or a hook bypass
(--no-verify) could make CI pass falsely. The pure-code gate now flags these via
gate_weakening_violations() and BLOCKs in evaluate_ci_gate as a top-priority trust
violation (step 1b, alongside the denylist) — regardless of the authenticated CI
conclusion. A build cannot pass itself by disabling its own checks; flagged diffs
escalate to a human. Only ADDED lines are inspected (removing a suppression is fine).
1015 tests pass, ruff clean.
* feat(agent-team): P3-flip — diff transport (§4.3) + flip privileged apply path live
Completes the box->CI diff handoff and flips the apply/verify privileged job
live (gated behind the agent-apply environment's required reviewer).
Transport (§4.3): the read-only box (D2) emits a diff but holds no write token.
- New credential-less `materialize` job decodes the untrusted `diff_b64`
dispatch input via env (CWE-94), fail-closed re-hashes it against
`expected_diff_hash`, and uploads it as the named artifact so guard/build-test
download it same-run. guard now `needs: materialize`.
- New `dispatcher.py` (the trusted apply path, operator/Mac-side — never the
box): pushes the diff as a head branch then `gh workflow run`s the workflow.
Pure input-assembly (sha256 == sha256sum, b64 round-trip, head ref) is
unit-tested; git/gh are injected seams. Push-before-dispatch; fail-closed on
empty diff/scope, unsafe task_id/owner/repo.
Flip: gate-and-pr binds `environment: agent-apply` (required reviewer
amoussa1229) + grants exactly `pull-requests: write`; the App-token + draft-PR
steps run only on `steps.gate.outputs.gate == 'pass'` (no more if:false); the
draft PR opens with an explicit `--head`; task_id/head_branch charset-validated
(§4.6). Updated the hardening tests from inert-state to live-state assertions +
added transport tests. 1039 tests, ruff clean, workflow YAML valid.
NOTE: workflow only runs on manual workflow_dispatch and the privileged job is
held at the required-reviewer gate, so nothing privileged runs unapproved.
* fix(agent-team): P3-flip — address GPT-4.1 cross-review (size bound, ref-traversal guard)
- BLOCK: cap candidate diff at 40 KB in the dispatcher (the diff rides a base64
workflow_dispatch input; GitHub caps inputs at ~64 KB so an oversized diff
cannot dispatch at all) + a defense-in-depth decoded-size bound in materialize.
- FIX: harden the draft-PR HEAD_BRANCH guard to reject leading/trailing slash,
'..' segments, and '//' (CWE-88 git ref-traversal), not just bad charset.
- NIT: document the mandatory invariants on gate-and-pr (required-reviewer
environment must stay; runs-on must stay GitHub-hosted).
- QUESTION (lockfiles): answered in-code — the build-test sandbox is
credential-less + egress-blocked, so lockfile-postinstall RCE is contained.
Tests added for all guards. 1042 tests, ruff clean, YAML valid.
* fix(agent-team): P3-flip — resolve /sh-security-review findings (LOGIC-1/2/3)
High-recall fan-out (injection/logic/iac+secrets) + proof-or-kill on the LIVE
apply path found 3 real issues the cross-review missed; all fixed:
- LOGIC-2 (HIGH, was a live hole): build-test ran `ruff check . || echo` /
`pytest -q || echo`, swallowing failures so the job was always 'success' and
the gate would open draft PRs on RED builds. ruff/pytest now run
authoritatively under set -e (pytest exit 5 'no tests' is the only non-fatal
case); the exit code IS the build-test conclusion the gate keys on.
- LOGIC-1 (verified!=shipped): the dispatcher used `git apply` + `git add -A`,
staging stray untracked content into the pushed PR head. Now `git apply
--index` stages exactly the diff, so the head tree is precisely base+diff —
bound to the bytes CI hash-verified.
- LOGIC-3 (§4.5 on the live path): gate-weakening was enforced only box-side;
added a gate-weakening check to the guard job so the live PR-opening path
rejects a diff that adds suppressions/skips, even on a green build.
Injection / secrets / least-privilege / flip-correctness / no-untrusted-checkout
all came back clean. 1044 tests, ruff clean, YAML valid.
* feat(agent-team): durable GitHub-issue intake de-dup (schema v2)
The intake poller de-duped ingested issues in an in-memory set that does
not survive a process restart. A scheduled/cron intake (each run a fresh
process) would therefore re-ingest every still-open labeled issue on every
run and spawn duplicate pipeline tasks. Since the box is read-only (no write
token to remove the intake label), durable de-dup is the only correct guard.
- schema v2: new ingested_issues(source, issue_id, ingested_at) table +
issue_already_ingested / record_issue_ingested helpers; migrate() adds the
table to a legacy v1 DB and restamps; init_db creates it.
- github_intake: pluggable IngestStore seam (in-memory default preserved for
tests/one-off; durable build_ledger_ingest_store for production). Record is
after start_task succeeds, so a failed intake stays retryable.
- run-team.py intake-github wires the ledger store keyed by github:owner/repo,
making a scheduled timer idempotent across runs.
978 tests pass (ruff clean).
* fix(agent-team): harden github intake per /sh-security-review (CWE-918, idempotency)
Fixes from the high-recall detector fan-out on the durable-dedup change:
- INTAKE-LOGIC-01 (idempotency): switch the IngestStore seam from
check-then-record (seen/mark) to claim-then-do (claim/release). The id is
now reserved BEFORE the non-idempotent start_task side effect, so a crash in
that window cannot re-spawn a duplicate task on the next run; a raising
start_task releases the claim so transient failures stay retryable. Adds
delete_issue_ingested to the schema layer for the release path.
- INTAKE-SSRF-001 / INTAKE-PATHSPLICE-002 (CWE-918) in build_default_issue_client:
drop the caller-overridable api_root (hardcode GITHUB_API_ROOT) and validate
owner/repo against an anchored charset before splicing them into the
token-bearing API URL — mirrors the sibling ci_fetcher BLOCK-3/FIX-3 fixes.
Tests cover cross-process duplicate prevention, release-on-failure retry, and
the owner/repo + api_root rejection. 982 tests pass, ruff clean.
Follow-up (pre-existing, not introduced here): the label-only intake has no
author allowlist (cf. AGENT_TEAM_SLACK_OWNER_IDS on the Slack listener); the
Slack answer gate bounds the blast radius. Track as separate hardening.
Three robustness/hardening fixes surfaced by /sh-security-review on the
agent-team listener/coordinator surface. None alter AUTHZ-01 allowlist
behavior or the first-answer-wins compare-and-set semantics.
1. Respawn close() leak (CWE-772). The watchdog _supervise_slack_listener
respawned the inbound Slack listener without tearing down the dead one,
leaking a Socket Mode WebSocket / SDK thread set per flap. Now the dead
listener is closed before respawn (new _close_dead_listener, idempotent),
AND _run_listener has a finally that always closes the listener so a
crashed serve() releases its socket. SlackListener.close() is idempotent,
so the belt-and-braces close stays a safe no-op.
2. Broadened exception guard in handle_event (CWE-248). The submit block
only caught ValueError; the accept path (submit_answer ->
_question_turn/_question_thread) can raise KeyError on a concurrently
mutated row, and the CAS can raise sqlite3.Error. An uncaught exception
would escape into the Bolt dispatch. Added a separate `except Exception`
that logs at WARNING (not silent, not debug) and returns None. The
existing ValueError-as-debug behavior is unchanged; authorization still
runs first, so the trust boundary is not widened.
3. channel_ref partial-unique index (defense-in-depth). Added
uq_pending_questions_open_channel_ref — a PARTIAL UNIQUE index on
(channel_ref) WHERE channel_ref IS NOT NULL AND status='open' — so two
OPEN rows can never share a non-null channel_ref (a thread_ts can never
map to two open questions). Installed in init_db AND unconditionally in
migrate (idempotent IF NOT EXISTS) so existing v1 DBs gain it. NULLs and
closed rows are excluded; mirrored verbatim into schema.sql.
Tests: +8 (was 960, now 968). New: schema partial-unique reject/null/closed/
migrate cases; handle_event KeyError + sqlite3.Error swallow cases;
coordinator close-before-respawn + run_listener-closes-on-crash. Fixed the
operator-cli test fixture to use a per-question channel_ref (it previously
inserted multiple open rows sharing one ref, which the new index correctly
rejects).
* fix(agent-team): handle real slack_bolt event envelope + map thread replies to open questions
The Socket Mode inbound listener was unit-tested against a SYNTHETIC payload
shape that does not match what slack_bolt actually delivers, so the suite was
green while a real Slack thread reply was silently dropped (the clarifier
question stayed `open`). Real slack_bolt delivers an Events API message /
app_mention as `{"type":"event_callback","event":{"type":"message",...}}` and
a free-text thread reply carries NO callback_id/question_id/metadata.
Three breaks fixed (all on the free-text reply path):
1. Type gate — handle_event gated on the OUTER `type`, which is
"event_callback" for a real message/app_mention, so the event fell outside
_ANSWER_BEARING_TYPES and was dropped. Now collapsed to the discriminating
INNER `event.type` via _discriminating_type / _inner_event.
2. question_id recovery — a real reply has no callback_id/question_id/metadata
(the bot's metadata is on the QUESTION message, not the reply). When explicit
id recovery fails, the listener now resolves the question by the inner
event's `thread_ts` against the OPEN ledger row whose `channel_ref` equals it
(new schema helper find_open_question_by_channel_ref, constrained to
status='open' as anti-replay). Explicit id recovery still takes precedence.
3. answer extraction — a real message event carries its text at `event.text`,
not a top-level `answer`/`text`. The thread-reply path now takes the inner
`event.text` (stripped) as the answer value.
AUTHZ-01 is unchanged and still runs FIRST: authorization gates on the sender's
Slack user id (`event.user` for the Events API shape) and fails closed on an
empty/unknown allowlist or unrecoverable sender. The new mapping only resolves
WHICH question is answered, never WHO may answer. Answers stay opaque DATA
(parameterized SQL + json.dumps; never eval/exec/interpolate).
Tests: replaced the synthetic events-API fixtures with REAL Bolt envelopes and
added regression coverage — real thread reply maps via channel_ref and is
accepted, text is stripped, non-owner reply rejected (row stays open), thread_ts
matching no open row is a no-op, reply to an already-answered row is a no-op
(anti-replay), and app_mention is normalized identically. block_actions /
slash_command paths retained.
* fix(agent-team): bind subscription invoker in the start CLI
`run-team.py start` runs the clarifier graph to the first human gate IN the CLI
process, and the clarifier calls Claude (assess_confidence). The invoker is a
process-local binding that only `serve` set, so `start` failed with
"claude_invoke has no invoker bound". Bind the real subscription invoker here,
mirroring Coordinator.serve(). Found during the live R720 P1 bring-up.
* fix(agent-team): serve() starts the inbound Slack listener (D-1)
Coordinator.serve() now constructs and starts the SlackListener concurrently
with the tick/drain loop on a background daemon thread, but ONLY when the live
transport is a SlackTransport AND SLACK_APP_TOKEN is configured. When Slack is
not the transport or the app token is absent, serve() behaves exactly as before
(tick/recover only) — Slack is never made mandatory.
- New injectable build_listener seam + default_slack_listener_factory sharing
the coordinator's own transport, ledger db_path, and resume_queue put.
- AUTHZ-01 owner-allowlist + open-status CAS untouched: serve() sources
AGENT_TEAM_SLACK_OWNER_IDS in SlackListener.serve, which still fails closed.
- SlackListener.close() added for clean Socket Mode teardown on shutdown;
serve() stops the listener + joins the thread in a finally.
- Tests: start-when-Slack+app-token, no-start otherwise, clean shutdown,
idempotent start, serve start/stop around the loop, listener close().
* fix(agent-team): systemd unit loads ~/orchestrator/.env + uses venv python (D-2/D-7)
D-2: add EnvironmentFile=-/home/adam/orchestrator/.env (optional '-') so the P2
GPT-4.1 review loop's cross_reviewer sub-process can read the non-Claude provider
key once a task reaches REVIEW. Mirrors the sea-haven-secrev unit.
D-7: point ExecStart at the agent-team venv interpreter
(/home/adam/orchestrator/agent-team/.venv/bin/python) instead of
/usr/bin/env python3, which resolved the system interpreter without the
installed deps under systemd's PATH.
All hardening (NoNewPrivileges / ProtectSystem=full / ProtectHome=read-only /
ReadWritePaths) is retained unchanged (locked decision).
* docs(agent-team): land provisioning + operator runbooks under docs/provisioning
- PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5
intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub
App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked
FIXED so the demo can use the live Slack answer path.
- P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the
Slack and operator-CLI answer paths documented for all four exit criteria.
- DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the
operator-CLI divergence kept as provisioning notes.
- OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks,
failed HITL resumes, budget exhaustion, transport outages, and
COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the
re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6).
* fix(agent-team): supervise the Slack listener thread — recurring ALARM + respawn
sh-security-review (logic) MEDIUM: a crashed listener thread was logged once,
then the daemon ran on 'deaf' — posting clarifier questions but receiving no
answers, every gate silently parking, process never exiting so systemd
Restart=on-failure never fired. serve() now calls _supervise_slack_listener()
each pass: when the listener is enabled but its thread is dead, it emits a
recurring ERROR ALARM and respawns via the idempotent starter (self-heal).
No-op when alive or disabled. +3 tests. (authz detector: wiring clean — AUTHZ-01
fail-closed allowlist + open-status CAS intact, dead listener fails SAFE.)
* feat(agent-team): Plane-1 Tier-3 fixer — dependency-cve finding -> patch + CI dispatch (opt-in/inert)
The fixer (design §4 fixer row, §7 Phase 5, §3.3.2) takes a CONFIRMED,
low-risk dependency-cve finding (the narrowest fix class) and produces:
* a fix SPEC (Claude, via the §3.1 billing seam), and
* a minimal bump PATCH (DeepSeek fast_coder, via the orchestrator run.py
path that builders_llm uses),
records the candidate diff + its content-hash, and emits the org-CI
workflow_dispatch inputs (task_id / diff_artifact_name / expected_diff_hash /
declared_scope) for the gate-passed P3-live apply/verify surface.
INERT / opt-in / fail-safe, mirroring build_verify_wiring:
* plan_fix dispatches NOTHING; dispatch_fix has NO default dispatcher
(the box holds no write token, D2) so an un-wired call can never fire a
workflow.
* no git/patch/subprocess/fs-write in executable code — the patch is emitted
as diff TEXT only; CI applies it and opens a DRAFT PR, the box never
applies/pushes/merges.
* untrusted-patch hygiene: the generated diff is confined box-side to the
single dependency manifest (declared_scope) and rejected via
ci_gate.denylist_violations if it escapes scope or touches the
trust-control surface — defense-in-depth with the CI guard.
* bad/ambiguous findings (wrong check/status/category, missing
package/fixed_version, ambiguous fixed_version, unparseable/empty diff)
yield a FAILED no-op plan, never a fabricated fix.
29 new pytest tests under agent-team/tests/test_fixer.py.
* feat(agent-team): run-team.py 'fix --dry-run' subcommand for the Plane-1 fixer
Adds the fixer front door to the operator CLI: load one confirmed
dependency-cve finding from a dependency-cve.json report (--report
--finding-id), plan the fix, and in --dry-run print the spec + patch + the
org-CI workflow_dispatch inputs WITHOUT dispatching anything.
Opt-in/inert: the command binds NO workflow dispatcher and holds no write
token, so even an ok plan only prints; live dispatch is provisioning-gated
(refuses to run without --dry-run). A non-fixable finding prints the
fail-safe reason and exits 1.
4 new pytest tests under agent-team/tests/test_run_team.py.
* feat(agent-team): P5 checker-finding intake module + tests
Add agent_team.transport.checker_intake: turn a confirmed, at/above-threshold
Plane-1 checker FINDING into one Plane-2 pipeline remediation task via the
committed coordinator intake entry (start_task), mirroring github_intake.
- select_findings: status==confirmed AND severity>=threshold (default high);
unverified/suppressed/below-threshold dropped; unknown threshold rejected.
- finding_identity: stable de-dup key (finding id, else content-hash). In-memory
set, best-effort, NOT durable across restart (ledger table is the follow-up).
- finding_task_text/_sanitize: every repo-controlled field (title, proof, repo)
is newline/control-char neutralised and length-bounded before it reaches the
task text or operator log (log-injection hygiene).
- load_report_findings/ingest_reports: read the exact checker report JSON shape
(top-level object with findings[]; bare array and dir-of-*.json also accepted).
28 hermetic unit tests (stub coordinator, in-memory findings / temp reports).
* feat(agent-team): wire opt-in intake-checker run-team subcommand
Expose the P5 cross-plane loop only as a manual run-team subcommand
(intake-checker --report PATH [--threshold] [--transport] [--dry-run]),
mirroring how intake-github is exposed. NOT wired into the always-on serve
path: the loop stays opt-in/inert by default.
* feat(agent-team): read-only CI-result fetcher for P3 verify gate (opt-in, inert)
ci_fetcher.py: fail-closed CiResultFetcher reading the GitHub Actions run
conclusion via a read-only PAT (AGENT_TEAM_CI_READ_TOKEN→GITHUB_TOKEN), returns
{run_id,conclusion,diff_hash} or None on any error. Data-fetcher only — ci_gate
owns the verdict; never writes, no OIDC/AWS, never reads patch artifacts.
coordinator gains opt-in gated_build_verify_wiring() composing it via
bind_ci_result_fetcher; NOT wired into the default run-team.py path. 20 tests.
* harden(agent-team): P3 apply/verify workflow — GitHub App token, CWE-94, fail-closed
Decision-1 auth model: gate-and-pr uses a GitHub App installation token
(pull-requests:write) behind the agent-apply environment; ALL OIDC/id-token/AWS
removed. Hardening: task_id env-indirection (CWE-94 — GitHub expands ${{ }} into
the run shell before exec, so %s/quoting is insufficient); run-id pinning on both
download-artifact; post-build denied-path check (build-hook writes into denied
paths fail the job); empty-hash fail-closed in BOTH the embedded gate (fixed a
real ''=='' pass bug) and ci_gate.py. App-token + draft-PR steps stay if:${{ false }}
until provisioning (App + environment + branch protection). +17 tests.
* harden(agent-team): apply P3-live security-gate fixes (GPT-4.1 xreview + sh-security-review)
BLOCK-1/FIX-4: gate-and-pr re-comments pull-requests:write + environment:agent-apply
(provisioning-time uncomment) and gains needs.guard/build-test=='success' job guard —
zero privilege until provisioning. BLOCK-2/3+FIX-5: ci_fetcher validates run_id (^[0-9]{1,20}$),
owner/repo (^[A-Za-z0-9_.-]{1,100}$), and fetched_id (int) — fail closed, no SSRF/path
injection. FIX-1: conclusion allowlist. FIX-3: api_root removed from public builder (no
injectable endpoint). INJ-02: post-build denied-path check uses NUL-delimited git output +
explicit rename parsing, no backslash mangling, non-UTF8=violation. INJ-03: all three trust-
control denylists unified to one 22-entry union + drift-guard test. Q1: documented run_id/
diff_hash trust source (dispatcher/ledger only). 884 tests, ruff clean. Privileged steps stay
if:${{ false }} until provisioning.
* build(security-review): prune .claude worktrees from deterministic scanners
Agent worktrees under .claude/worktrees/ are full repo copies; the cfn-lint
find|xargs template scan overflowed ('command line cannot be assembled') and the
pre-push hook fail-closed to BLOCK whenever a worktree was present. Prune .claude
in the cfn-lint find + semgrep/checkov excludes, and gitignore .claude/ so it is
never scanned or committed. Unblocks main-tree pushes during parallel agent work.
build_verify_subgraph: BUILD->VERIFY nodes + route_after_verify. build_graph gains an opt-in build_verify param that repoints the review 'build' route at the subgraph (BUILD->VERIFY->{approved->END | loop->PLAN | parked->END}); default unchanged (P2). Coordinator build_verify_wiring composes it INERT (no ci_result -> ci_gate BLOCK -> PARKED; LLM is fix-proposer only, never declares green). NOT enabled in production: the live CI apply/verify + OIDC stays held for its /sh-security-review + GPT-4.1 cross-review gate. Also escapes untrusted intake text in logs (log-injection hygiene).
Live github (issue-comment poster) and claude_code (file-drop) transports, plus GithubIntake (labeled issue -> coordinator.start_task, de-duped). run-team _build_transport now wires github/claude_code live (was SystemExit) + adds the intake-github subcommand. claude_code drop-path also neutralizes backslash (defense-in-depth).
CI lacks slack_sdk, so the deferred-import-missing error fired before the no-token check and masked it. Stub slack_sdk into sys.modules so the token branch is deterministically exercised in both environments.
build_graph gains injected live_plan_node/review_node/route_review: P1 = plan->END, P2 = clarify->plan->review->{build|loop-back|parked}. Coordinator composes clarifier->graph->ResumeWorker, wraps planner fail-safe, binds the GPT-4.1 review loop; run-team start/serve opt production into P2. Re-delivery uses a guarded CAS so a concurrently-answered row is never clobbered (closes RACE-REDELIVER).
slack_live: real slack_sdk poster. slack_listener: Socket Mode inbound; trust boundary = app-token auth + an explicit owner allowlist on the sender (fail-closed, rejects all if AGENT_TEAM_SLACK_OWNER_IDS unset) + the open-status CAS as anti-replay. Closes the AUTHZ-01 missing-sender-authz finding from the security review.
review_loop_llm -> GPT-4.1 cross_reviewer (orchestrator run.py); builders_llm -> DeepSeek fast_coder (INERT, proposes diff text only); verifier_llm -> ci_gate is sole PASS authority, Claude is fix-proposer only. Hardens review_loop.parse_verdict to word-boundary matching, adds a fail-closed subprocess timeout, and bind_review_node (single-arg, no LangGraph config injection). All fail safe on untrusted model output.
Adds the billing-seam invoker (claude_agent_sdk subscription-OAuth, deferred import, API/Bedrock paths) and the Claude-backed clarifier callables (ConfidenceAssessor/QuestionGenerator, one call/turn memoized on (thread_id,len,content-hash), fail-safe to 0.0 so garbage never clears the 98% human gate).
Addresses the confirmed findings from /sh-security-review + the GPT-4.1
cross-review of the Plane-2 scaffold. Full suite: 589 passed; ruff clean.
FIXED (proven-exploitable):
- CI-guard denylist bypass (HIGH): Python fnmatch '**/' is non-recursive, so
root-level template.yaml/*.tf/cdk.json/*.pem/*.key/*-stack.* evaded the
trust-control surface. Replaced fnmatch with a recursive, case-insensitive
glob->regex matcher. (verified: fnmatch('template.yaml','**/template.yaml')==False)
- CI-guard scope bypass (HIGH): a '**' declared_scope made every path in-scope.
Scope is now concrete-prefix confinement (reduces a glob to its leading
metacharacter-free segments; '**' -> empty -> dropped -> unscoped reject).
- Box-side vs CI denylist divergence (MED): builders.py _DENY_PATTERNS now covers
Terraform, *.pem/*.key, CDK stack files, .github/actions, *iam*, bare policy*.json
(case-insensitive), matching the CI surface.
- force-resume was backwards (MED): it superseded the answered row recovery
resumes from, making a stuck task permanently un-resumable while printing
success. Now re-opens an EXPIRED (parked) question via a new reopen_question
CAS helper; never supersedes an answered row; honest exit codes.
- operator attribution (MED): run-team.py --operator defaulted to "" -> now the
OS login, so destructive actions are always attributable.
- audit-log append race (MED): replaced read-modify-rewrite (lost records under
concurrent operators) with an O_APPEND single-line write, mode 600 enforced.
- lstrip("ab/") path-mangling in the symlink error path -> regex prefix strip.
Regression tests added across test_ci_gate_workflow / test_builders / test_run_team
/ test_schema. Design-level findings (resume-worker durability, egress breadth,
answered_at ordering, DB-swap TOCTOU, diff-hash threat-model) are pre-deployment
/ P1-build-proper and recorded with written justification in
agent-team/.security-review/suppressions.json; CI README diff-hash wording made
honest.
Reworks the P1 sim so the four §7.1 exit criteria are demonstrated against the
ACTUAL mechanic, not a model (resolves the verifier's "sim models the ledger,
not the LangGraph integration" finding).
- New tests/sim/test_p1_graph_integration.py drives the real agent_team.graph
StateGraph (interrupt/Command(resume)) + the real langgraph SqliteSaver
checkpointer + the committed pending_questions compare-and-set, proving:
(a) suspend survives a simulated restart (drop saver/conn, rebuild over the
same checkpoint DB) and resumes; (b) duplicate answer loses the CAS and the
graph never double-advances; (c) a post-deadline answer loses to expire and
the task is not resumed; (d) two concurrent tasks resume to the correct
thread, with a turn-guarded no-double-apply check.
- graph.py: derive a STABLE question_id from uuid5(thread_id, turn). The
clarifier node replays on resume, so the prior fresh-uuid id changed between
the delivered/ledgered question and the qa_history entry — breaking the
§3.3.1 identity contract. Now the delivered id == ledger key == history entry
(unit-tested in test_graph.py).
- harness._connect() now uses the committed schema.connect() (WAL + busy_timeout)
instead of a raw sqlite3.connect, so concurrent responders genuinely serialize;
the criterion-(d) concurrency test no longer swallows OperationalError (it
asserts zero errors + exactly one CAS winner).
- requirements.txt: pin langgraph-checkpoint-sqlite==3.1.0 (design D9 durable
checkpointer), now exercised by the integration test.
Full suite: 564 passed; ruff + format clean.
Resolves the GPT-4.1 cross-review FIX items on the §3.3.2 CI apply/verify guard:
- Symlink-escape (Medium-High): reject any candidate diff that introduces a
symlink (git mode 120000). A symlink can redirect a later in-diff write into a
denied path that textual canonicalization cannot see; auto-built diffs have no
legitimate symlinks, so this fails closed (exit 7).
- Diff-parse robustness (Medium): decode the diff as strict UTF-8 and fail closed
(exit 8) instead of errors='replace', closing homoglyph/encoding evasion.
- Declared-scope canonicalization (Medium): drop parent-escaping scope globs so a
malformed scope can only shrink coverage, never widen it past repo root.
- Egress allowlist (Low-Med): explicit DEPLOY marker to parameterize the
build-test registries per target repo before enabling.
Backs the gate's correctness claim with a committed, runnable suite
(tests/test_ci_gate_workflow.py) that extracts the inline guard from the YAML and
exercises good + adversarial diffs (clean, hash mismatch, workflow delete,
copy-into-denied, symlink, non-UTF-8, out-of-scope, unscoped, escaping scope).
Corrects the README "Tests" section that claimed coverage that did not exist.
Full suite: 557 passed, 1 skipped; ruff clean.
Resolves three execution-proven verifier findings from the scaffold review.
Full suite: 548 passed, 1 skipped (stable across repeated runs); ruff clean.
builders denylist (§3.3.2 #2): scan was +++-only and missed header-only
sections. Now section-driven off `diff --git a/<src> b/<dest>`, catching the 4
proven bypasses — delete of a denied path, mode-change-only, `copy to` a denied
path, out-of-scope delete (regression tests for each).
§3.3.1 compare-and-set concurrency: BEGIN IMMEDIATE moved inside guarded retry;
each CAS now runs on its own connection (shared sqlite3.Connection cannot hold
two transactions, and is unsafe for concurrent use even for reads). connect()
stashes the db path on a Connection subclass so the path is derived by a
thread-safe attribute read, not a PRAGMA on the shared conn; busy_timeout set
before the WAL pragma. Added shared-connection concurrent regression tests
(distinct + same question) — previously raised "transaction within a
transaction".
operator CLI (run-team.py): added the design-named re-deliver and force-resume
verbs (were missing); audit now records the attempt BEFORE the mutation and the
outcome after, so a ledger mutation can never land without a trail; main()
catches OSError instead of leaving an uncaught traceback on audit-write failure.