The cfn-lint template-discovery step in review.sh piped the repo's whole
matched-file list into 'xargs -I{} sh -c "grep -l {}"'. On the orchestrator
monorepo — especially from a deep worktree path, where every matched path is a
long absolute path — xargs -I{} packs all paths into one assembled command and
aborts with 'xargs: command line cannot be assembled, too long'. The subprocess
exits non-zero having emitted ZERO findings, so the global pre-push hook BLOCKS
every push (agents were working around it with --no-verify).
Fix: switch the grep stage to NUL-delimited, un-batched xargs
(find ... -print0 | xargs -0 grep -lE ...). xargs -0 (no -I) splits the input
across multiple grep invocations, so the argv never exceeds ARG_MAX; grep -l
reports the same matching files as the old per-file grep, and -print0/-0 is safe
for paths with spaces/newlines. The first 'xargs -I{} find {}' is kept (find
needs the start path before its expression) and is bounded by the scope-path
count, so it is not an overflow source. Trailing '|| true' preserves the old
no-match/no-files semantics (TPLS = list-of-templates or empty, never fails).
Purely an argv-batching fix: the scanned file set, findings, and exit codes are
unchanged. Verified exit-code-identical: clean tree -> exit 0 (PASS); planted
GitHub PAT + RSA private key -> exit 1 (BLOCK, gitleaks high); planted CFN
template with a cfn-lint error on a deeply-nested path -> exit 1 (cfn-lint flags
it, no overflow). The previously-overflowing command now completes clean.
Codifies the hardened deploy-before-merge procedure for ongoing agent-team
code changes, generalizing the one-off deploy-r720-ws-rollout.sh. Prevents the
two self-inflicted live crash-loops:
- Whole agent_team/ package rsync (never per-file, which misplaces e.g.
nodes/planner.py at the package root -> ImportError/TypeError crash-loop).
- Snapshot HARD GATE (Adam's Hyper-V step; Claude ssh reaches only the guest)
+ ledger backup before any change.
- Pre-restart import sanity, then mandatory verify-after (is-active==active,
NRestarts didn't climb, ~6 threads, clean journal) with rollback guidance
on failure.
Idempotent, fails loudly. Optional SYNC_DEPS / SYNC_HANDBOOK / RESTART_STATUS.
Drives the new /sh-deploy-r720 skill.
WS Slack-UX Feature 2. When the inbound listener acts on an answer in a task
thread, it adds a 👍 reaction to that reply so the human sees the machine
received it.
- SlackListener gains an optional reactor seam; handle_event reacts to the
inbound reply message (channel + event ts) AFTER the AUTHZ-01 owner check
passes — a non-owner message is rejected and never reacted to. Best-effort:
any reaction failure (notably a missing scope) is swallowed and never breaks
handle_event or the listen loop.
- /new-task is NOT reacted to (a slash command has no reactable message); its
"📥 Task received" root post is the acknowledgement.
- build_slack_reactor wraps WebClient.reactions_add(name="thumbsup"); the
default listener factory wires it best-effort from SLACK_BOT_TOKEN.
- Adds reactions:write to the bot scopes in agent-team-manifest.json.
NOTE: the new reactions:write scope requires Adam to re-apply the manifest to
app A0BCC7TTU66 and reinstall the app. Until then reactions.add returns
missing_scope, which the listener swallows (the reaction silently no-ops) —
answer handling is unaffected.
AUTHZ-01 and the first-answer-wins CAS remain unchanged.
WS Slack-UX Feature 1. A /new-task task now maps to ONE Slack thread instead of
several top-level messages.
- /new-task posts an immediate root "📥 Task received: …" ack and captures its
ts (root_ts); this is the instant acknowledgement.
- root_ts is plumbed into start: new PipelineState/TaskRecord channel
slack_thread_ts, seeded by graph.start_task and threaded through
Coordinator.start_task. The NewTaskCallback is now (task_text, via, root_ts).
- All clarifier questions for the task post as THREADED REPLIES under root_ts
(chat.postMessage thread_ts=root_ts), and each question's ledger channel_ref
is set to root_ts (NOT the reply's own ts). Because answer-mapping resolves a
reply via find_open_question_by_channel_ref(thread_ts), a reply in the root
thread (thread_ts==root_ts) maps to the task's currently-open question with NO
change to the mapping logic or the first-answer-wins CAS. The open-only
partial-unique index still holds (one open question per task at a time).
- Lifecycle milestones (parked / plan-ready / needs-input) and follow-up
questions thread under root_ts too; the notify sink gained an optional
thread_ts kwarg (degrades to top-level on a sink that doesn't accept it).
notify failures still never break tick.
- SlackTransport.post_question + the live poster accept/forward thread_ts.
- No root_ts (non-/new-task origin) ⇒ top-level posts exactly as before.
AUTHZ-01 (owner-allowlist-first, fail-closed) and the atomic open→answered
compare-and-set are unchanged.
Adds plumbing for the inbound-ack reactor seam used by Feature 2 (dormant until
a reactor is injected). Tests cover thread_ts forwarding, channel_ref=root_ts,
graph seeding, and coordinator threading.
Two root causes behind 'every task parks, and slowly':
1. SINGLE-SHOT INVOKER: subscription Claude calls ran as 40-turn, tool-enabled
agentic sessions (--max-turns 40, $2 budget) for what are pure reasoning->JSON
completions — minutes-long, and Claude wandered/returned unparseable output.
_DEFAULT_MAX_TURNS 40->1 + allowed_tools=[] -> fast deterministic single turn.
2. PLANNER FEEDBACK KEY MISMATCH: the review stage writes verdict/findings, but
_format_review_feedback read decision/notes/comment (never present) -> the
planner re-planned with EMPTY feedback, re-introduced the rejected flaw
('assumptions persist'), hit the review cap, parked. Now reads verdict/findings
(old keys kept as fallback) so GPT-4.1's objections reach the re-plan.
Plus: park notifications infer the phase the task was IN (review/plan/clarify)
instead of the terminal 'parked'.
Tests: planner real-verdict-keys regression + coordinator phase-inference. 1149 pass.
Park notifications were opaque ('Task 6c3c3202 parked — needs your attention'):
no idea what the task is, where it got to, or what's blocking it. Now each
message names the task DESCRIPTION (not just the short id), the PHASE it reached,
and the actual BLOCKER — _summarize_blocker() pulls the last review_verdict's
findings (the GPT-4.1 REQUEST_CHANGES text, collapsed + truncated), falling back
to 'no plan built' / 'revision cap hit'. Applies to parked + needs-more-input +
plan-ready emits. +1 test (description + phase + blocker present). 1148 passed.
The bot only ever posted clarifier questions; parks/completions were silent and
multi-turn follow-up questions were never posted during normal operation (only
the startup recover sweep posted them). So an answered task was a black box.
- Coordinator gains a 'notify' sink + _emit() (guarded, never breaks the loop).
- tick() now runs _post_resume_followups(results) after the drain: for each
resumed thread it (a) POSTS a newly-pending clarifier question (fixes silent
multi-turn — the drain path left it unposted) and emits 'needs more input';
else emits 'parked — needs attention' or 'plan ready for review' from the
settled state.
- run-team serve wires notify -> Slack channel (build_slack_poster) and an
alarm_hook that logs the deadline-park WARNING AND posts a parked notice.
Live-Slack only; dry-run/non-Slack/no-channel = silent (None), no token needed.
Tests: +4 (needs-input/parked/plan-ready emits + notify-failure swallow);
_FakeCoordinator gains notify/alarm_hook. 1147 passed, ruff clean.
/new-task (and every intake: GitHub issue, /sh-assign-task) reached the clarifier
with NO description -> the clarifier asked 'no task description provided'. Root
cause: coordinator.start_task only LOGGED task_text (a P1-era decision when the
deterministic clarifier didn't consume a description), graph.start_task took no
task arg, and PipelineState/TaskRecord had no 'task' channel at all.
Fix: add a first-class 'task' field to PipelineState + TaskRecord (+ round-trip
in task_from_dict); graph.start_task seeds task into the initial invoke (persists
through intake_node's partial-state return into CLARIFY); coordinator.start_task
passes task=task_text. The clarifier already reads state['task'] via
_task_description, so it now sees the real description.
Test: start_task(task='build a login form') -> suspended CLARIFY state carries
task. 1143 passed, ruff clean.
'the app did not respond': serve() registered @app.action/@app.event but NO
@app.command handler, so Bolt never acked the /new-task slash command within
Slack's ~3s deadline. Add an @app.command(/new-task) handler that ack()s first,
re-stamps type:slash_commands onto Bolt's inner command body (Bolt strips the
Socket Mode envelope type that _is_new_task_command/_discriminating_type expect),
then forwards to handle_event (AUTHZ-01 + new-task dispatch). The resulting
payload shape is the one already covered by test_new_task_calls_callback_*.
The agent-team venv was missing langchain-anthropic/-openai/-google-genai/
-community, so models.py failed to import and the in-process GPT-4.1 review /
Gemini scan / DeepSeek build silently fail-closed to REQUEST_CHANGES (the
non-Claude models never ran on the box). Step 3 now installs the full pinned
requirements.txt into the venv instead of just fastapi/uvicorn. Installed +
verified live on the box: all three model factories construct.
Gap-audit findings:
- G1 (blocks-feature): make_cross_reviewer_invoker / make_fast_coder_invoker did
'from models import' without putting the orchestrator root on sys.path. The
run-team serve daemon only bootstraps agent-team/, so on the live box every
GPT-4.1 plan review hit ModuleNotFoundError -> review_plan's blanket except
silently fail-closed to REQUEST_CHANGES (GPT-4.1 never actually ran). Both
in-process invokers now call invoker_multi._ensure_orchestrator_on_path()
before the deferred import. WS1 introduced this when it swapped the review
default from the subprocess invoker to in-process.
- G4 (degrades): the handbook context_provider was wired into the planner only;
default_clarify_node_factory now accepts + forwards it, and run-team wires it
into build_clarify_node too, so clarifying questions are handbook-aware.
Tests: +2 regression tests (path-bootstrap, clarifier threading); _FakeCoordinator
gains build_clarify_node. 1142 passed, ruff clean.
The slack_listener handles {type:slash_commands, command:/new-task} but the app
manifest declared no slash commands and no 'commands' scope, so Slack never
offered /new-task (the command can't be invoked). Add the slash command +
commands bot scope. Socket Mode delivers it over the socket (no request URL).
APPLY: update the app A0BCC7TTU66 from this manifest + reinstall to pick up the
new scope.
Idempotent attended update of the live coordinator to WS0-WS5: rsync repo +
handbook, install fastapi/uvicorn, append AGENT_TEAM_API_TOKEN/SEA_HAVEN_HANDBOOK_DIR
to secrev.env if absent, restart the daemon, smoke tests. HTTP API is an opt-in
separate step; P3 dispatch stays inert. Snapshot-first + confirm before restart.
Integration branch combining WS0-WS5 (PRs #43-#46) + the activation wiring that
flips the safe seams ON in the run-team serve path:
- WS5 (D10): inject the Sea Haven handbook conventions into the planner prompt
via context_provider (zero-arg handbook loader; fail-safe to '' when absent).
- WS2: an allowlisted Slack /new-task starts a task on this coordinator
(set_new_task_callback adapter -> start_task; AUTHZ-01 gates it upstream).
- WS1 bind_multi_invoker() is already wired in _cmd_serve.
Coordinator gains a new_task_callback param + set_new_task_callback() (resolves
the constructor chicken-and-egg of referencing the coordinator's own start_task);
default_slack_listener_factory forwards it to the SlackListener.
NOT wired (deliberately): the P3 dispatch_node / build_verify path. Activating
it correctly needs a per-task expected_run_id bound into gated_build_verify_wiring
(plumbing that does not exist yet) AND the CI trust-boundary security re-review.
It stays inert pending that work.
Tests: +5 activation-wiring tests; _FakeCoordinator stub gains
set_new_task_callback. Full agent-team suite: 1140 passed, ruff clean.
Security follow-up from the per-PR review (non-blocking, defense-in-depth):
- Replace the save_memory name blocklist with an allowlist regex
(^[A-Za-z0-9][A-Za-z0-9._-]*$, max 128) so dot-only/hidden/backslash/NUL/
over-long names are rejected outright, not written as malformed-but-contained
files.
- Write via os.open(..., O_NOFOLLOW): the open fails (ELOOP) if the final
path component is a pre-planted symlink, closing the TOCTOU where a symlink
in _box-drafts/ could redirect the write outside the dir. O_CREAT|O_TRUNC
keeps overwrite-on-resave for regular files.
Tests: adds allowlist-rejection + symlink-refusal cases (21 pass).
Security follow-ups from the per-PR review (non-blocking MEDIUMs):
- Eager _get_token() at make_app build time so a missing AGENT_TEAM_API_TOKEN
fails fast instead of serving requests first (matches the docstring contract).
- Disable /docs, /redoc, /openapi.json (no auth dependency in FastAPI) — the
API is VPN-only/127.0.0.1 and should not expose its schema unauthenticated.
- Scrub raw exception text and subprocess stderr from 500 response bodies;
log server-side instead (avoid internal-path/state disclosure).
- Bound /orchestrator/invoke concurrency with a semaphore (429 over the cap)
so an authenticated caller cannot exhaust the box via many 600s subprocesses.
Also pin fastapi/uvicorn in requirements.txt (WS1 dep). With fastapi now
installed in CI, the previously skip-guarded TestClient tests run for real;
the importorskip guard stays as a no-op safety net.
Tests: 23 pass (adds docs-disabled + concurrency-429 cases).
Reworked per GPT-4.1 cross-family review BLOCK. The original PR removed the
agent-apply GitHub Environment (the live required-reviewer human gate) and
replaced it with a Slack notice that FAILS OPEN when its webhook secret is
absent (which it is) plus an audit-log line. The cross-review correctly
flagged this as trading a preventive control for detective controls, one of
which silently no-ops.
This commit:
- Restores environment: agent-apply on gate-and-pr (the human approval pause).
- Drops the fail-open Slack notify step.
- Keeps the unconditional audit-log step as an additive detective control.
- Restores the MANDATORY-INVARIANT assertion (env must be present) and adds
an assertion that the audit step is retained.
WS3's auto-dispatch (dispatch_invoker.py + graph/coordinator wiring) is
unchanged: it fires workflow_dispatch, which now pauses at the restored gate
for human approval — auto-dispatch up to the approval, then one click.
CI has no fastapi (box-only dependency); api.py imports it lazily. Guard the
7 TestClient smoke tests with skipif(find_spec('fastapi') is None) so they
skip in CI instead of failing collection, leaving the 14 invoker_multi tests
running. Also apply ruff format to the 5 WS1 files CI flagged.
- invoker_multi.py: multi_invoke(prompt, *, model) dispatches to GPT-4.1
(cross_reviewer), DeepSeek (fast_coder), or Gemini (scanner) in-process;
bind_multi_invoker() wires the review seam; lazy models import
- review_loop_llm: swap default_plan_reviewer from make_run_py_invoker()
to make_cross_reviewer_invoker() (in-process GPT-4.1); subprocess path
retained as make_run_py_invoker() for opt-in use
- builders_llm: add make_fast_coder_invoker() in-process DeepSeek path;
rename subprocess path to subprocess_build (opt-in fallback); default_build
now delegates to make_fast_coder_invoker()
- api.py: FastAPI app with bearer-token auth (AGENT_TEAM_API_TOKEN env),
POST /tasks, GET /tasks/{id}, POST /orchestrator/invoke; binds 127.0.0.1;
code only — not started here
- run-team.py: bind_multi_invoker() called in serve() alongside
bind_subscription_invoker()
- tests: 21 new WS1 tests + updated builders_llm tests (1065 total, all passing)
- retriever: add save_memory() writing to _box-drafts/ review queue, add
memory_dir param to retrieve() for isolated test routing
- handbook: new load_handbook_conventions() with safe no-op contract (returns
"" when dir absent/empty/unreadable, never raises)
- clarifier_llm: add context_provider=None seam to ClaudeClarifier and
build_claude_clarifier_callables(); failure in provider is silent
- planner: add context_provider=None seam to build_plan_prompt() and plan_node()
- coordinator: thread context_provider through default_plan_node_factory(),
only pass kwarg when non-None to preserve stub-monkeypatching in tests
- tests: 19 new WS5 tests covering all seams (1063 total, all passing)
confluence-doc is provisioned (OAuth 2.0 service-account creds in ~/secrev.env +
PAGE_MAP_FILE from the live IT space). Drop it from COORDINATOR_SKIP_ROLES (now
just aws-posture) and add Environment=PAGE_MAP_FILE. Verified live on the box:
'Confluence API: oauth auth ready', 21 recommend-only doc gaps reported (D7, no
writes). aws-posture stays skipped (IAM/step-ca not stood up).
The public /_edge/tenant_info endpoint returns an HTML 'Page Unavailable' on the
seahavenind site, so cloudId auto-resolution failed. Fetch the token first, then
resolve cloudId from the OAuth-native https://api.atlassian.com/oauth/token/
accessible-resources (Bearer), preferring the resource whose url matches the
configured site, else the first. CONFLUENCE_CLOUD_ID override still honored.
Canary 3/3 (offline), shellcheck clean.
Atlassian org service accounts have no classic API token — they authenticate via
OAuth 2.0 client-credentials (2LO). Add a dual-mode auth seam to confluence-doc:
- OAuth (preferred when CONFLUENCE_OAUTH_CLIENT_ID/_SECRET set): POST
auth.atlassian.com/oauth/token (client_id+client_secret+grant_type=client_credentials)
→ 60-min Bearer; calls go to api.atlassian.com/ex/confluence/<cloudId>/wiki/api/v2/...
cloudId auto-resolves from the site's public /_edge/tenant_info (no input needed).
- Basic (email+API token) retained as a fallback.
conf_api_init() picks the mode once; conf_get() does the authenticated GET. Any
failure (no cloudId / token request fails) → RUN_API=0, live checks SKIPPED, NO
false alarm (matches the existing no-data discipline). Secret passed in the request
body (--data-urlencode), never logged. Still read + recommend-only (D7); canary
unaffected (offline) — 3/3. shellcheck clean (accepted SC1091).
Live smoke exposed a self-pollution bug (pre-existing from PR #17): the post-build
denied-path check wrote its own _build_diff_z.bin / _build_status_z.bin into the
working tree, then its own `git status --untracked-files=all` flagged them as
out-of-scope writes — failing any run with a narrow declared_scope (the smoke's
docs/**). Write them to $RUNNER_TEMP instead (read via $_DIFF_Z/$_STATUS_Z), so
the check no longer sees its own temp files. The agent-team pytest artifacts were
already correctly gitignored; only the check's own files tripped it. Test harness
updated to pass the env paths. 1044 tests, ruff clean.
GitHub Actions only runs workflows under .github/workflows/, so the apply/verify
workflow at agent-team/ci/ was never registered (workflow_dispatch 404'd). Move it
to .github/workflows/agent-team-apply-verify.yml so it is a real, dispatchable
workflow. Its only trigger is workflow_dispatch + it is gated by the agent-apply
required-reviewer environment, so it never auto-runs and nothing privileged runs
unapproved. Updated the two workflow test files' path refs (parents[2]/.github/
workflows) and the ci/README pointer. Dispatcher push gains --no-verify: the apply
path is scanned CI-side (guard + the PR's checks), so it must not be blocked by the
operator's LOCAL human-commit pre-push dev hook (which flags pre-existing whole-repo
FPs like .env.example). 1044 tests, ruff clean.
- checker_coordinator.sh: add COORDINATOR_SKIP_ROLES env (comma-separated) to drop
roles whose credentials are not provisioned from the registry entirely (never
canaried/run/ALARMed). Fail-safe: empty/unset = run all.
- systemd: sea-haven-checkers.{service,timer} run the coordinator nightly at ~03:30
UTC (90 min after the secrev sweep so they don't contend on $MIRROR_DIR / the
Claude pool). The unit sets COORDINATOR_SKIP_ROLES=aws-posture,confluence-doc
(aws-posture needs IAM Roles Anywhere; confluence-doc needs the confluence-bot
token — both intentionally unprovisioned).
Deployed + enabled on the box (deploy-before-merge): canary 4/4 with the skip,
timer scheduled for 2026-06-23 03:35 UTC.
The build-test step ran the generic root `pytest -q`, but this repo's root
suite needs live ANTHROPIC/COMPOSIO keys (collect-only in CI) — so on the
orchestrator target it would fail in the credential-less build sandbox and block
every apply. Retarget the authoritative test to the self-contained agent-team/
subproject (exactly what the repo's green `ci / subproject-tests` runs). ruff
stays repo-wide. Command is AUTHOR-FIXED (not a dispatch input) so no injection
surface is added; multi-target parameterization remains a provisioning item.
This unblocks the first live smoke test.
* feat(agent-team): P3-flip Phase 1 — expand denylist vectors (§4.2) + runner-trust assertion (§4.1)
First controls of the P3-live-flip Phase-1 CI hardening (workflow stays INERT;
this only tightens the trust boundary). Whole Phase-1 surface is gated by
/sh-security-review + GPT-4.1 cross-review before any flip.
§4.2 — expand the trust-control denylist with direct code-execution / supply-chain
vectors, kept byte-identical across all three copies (ci_gate.DENYLIST_GLOBS + the
guard + post-build inline DENY_GLOBS), drift-guarded:
.gitmodules, .husky/**, .githooks/**, .gitattributes, .npmrc, and generated/build
artifacts (__generated__, *.generated.*, dist/**, build/**, *.min.js).
Deliberate: lockfiles are NOT wholesale denied — lockfile-postinstall RCE is already
contained by the credential-less egress-blocked build sandbox, and the Tier-3 dep-CVE
fixer rewrites lockfiles to produce its draft PRs; a blanket deny would make it
un-shippable. Flagged in-code for the security gate. Direct code-execution config
(hooks/filters/npmrc/submodules) is the actual §4.2 RCE surface.
§4.1 — runner-trust: assert no job (esp. the privileged gate-and-pr) can run on a
self-hosted/user-provided runner; all must be GitHub-hosted.
998 tests pass, ruff clean.
* feat(agent-team): P3-flip Phase 1 — gate-weakening detector (§4.5)
A diff that ADDS a lint/type/coverage/security suppression (noqa, type: ignore,
pragma: no cover, nosec, nosemgrep), a test skip/xfail, or a hook bypass
(--no-verify) could make CI pass falsely. The pure-code gate now flags these via
gate_weakening_violations() and BLOCKs in evaluate_ci_gate as a top-priority trust
violation (step 1b, alongside the denylist) — regardless of the authenticated CI
conclusion. A build cannot pass itself by disabling its own checks; flagged diffs
escalate to a human. Only ADDED lines are inspected (removing a suppression is fine).
1015 tests pass, ruff clean.
* feat(agent-team): P3-flip — diff transport (§4.3) + flip privileged apply path live
Completes the box->CI diff handoff and flips the apply/verify privileged job
live (gated behind the agent-apply environment's required reviewer).
Transport (§4.3): the read-only box (D2) emits a diff but holds no write token.
- New credential-less `materialize` job decodes the untrusted `diff_b64`
dispatch input via env (CWE-94), fail-closed re-hashes it against
`expected_diff_hash`, and uploads it as the named artifact so guard/build-test
download it same-run. guard now `needs: materialize`.
- New `dispatcher.py` (the trusted apply path, operator/Mac-side — never the
box): pushes the diff as a head branch then `gh workflow run`s the workflow.
Pure input-assembly (sha256 == sha256sum, b64 round-trip, head ref) is
unit-tested; git/gh are injected seams. Push-before-dispatch; fail-closed on
empty diff/scope, unsafe task_id/owner/repo.
Flip: gate-and-pr binds `environment: agent-apply` (required reviewer
amoussa1229) + grants exactly `pull-requests: write`; the App-token + draft-PR
steps run only on `steps.gate.outputs.gate == 'pass'` (no more if:false); the
draft PR opens with an explicit `--head`; task_id/head_branch charset-validated
(§4.6). Updated the hardening tests from inert-state to live-state assertions +
added transport tests. 1039 tests, ruff clean, workflow YAML valid.
NOTE: workflow only runs on manual workflow_dispatch and the privileged job is
held at the required-reviewer gate, so nothing privileged runs unapproved.
* fix(agent-team): P3-flip — address GPT-4.1 cross-review (size bound, ref-traversal guard)
- BLOCK: cap candidate diff at 40 KB in the dispatcher (the diff rides a base64
workflow_dispatch input; GitHub caps inputs at ~64 KB so an oversized diff
cannot dispatch at all) + a defense-in-depth decoded-size bound in materialize.
- FIX: harden the draft-PR HEAD_BRANCH guard to reject leading/trailing slash,
'..' segments, and '//' (CWE-88 git ref-traversal), not just bad charset.
- NIT: document the mandatory invariants on gate-and-pr (required-reviewer
environment must stay; runs-on must stay GitHub-hosted).
- QUESTION (lockfiles): answered in-code — the build-test sandbox is
credential-less + egress-blocked, so lockfile-postinstall RCE is contained.
Tests added for all guards. 1042 tests, ruff clean, YAML valid.
* fix(agent-team): P3-flip — resolve /sh-security-review findings (LOGIC-1/2/3)
High-recall fan-out (injection/logic/iac+secrets) + proof-or-kill on the LIVE
apply path found 3 real issues the cross-review missed; all fixed:
- LOGIC-2 (HIGH, was a live hole): build-test ran `ruff check . || echo` /
`pytest -q || echo`, swallowing failures so the job was always 'success' and
the gate would open draft PRs on RED builds. ruff/pytest now run
authoritatively under set -e (pytest exit 5 'no tests' is the only non-fatal
case); the exit code IS the build-test conclusion the gate keys on.
- LOGIC-1 (verified!=shipped): the dispatcher used `git apply` + `git add -A`,
staging stray untracked content into the pushed PR head. Now `git apply
--index` stages exactly the diff, so the head tree is precisely base+diff —
bound to the bytes CI hash-verified.
- LOGIC-3 (§4.5 on the live path): gate-weakening was enforced only box-side;
added a gate-weakening check to the guard job so the live PR-opening path
rejects a diff that adds suppressions/skips, even on a green build.
Injection / secrets / least-privilege / flip-correctness / no-untrusted-checkout
all came back clean. 1044 tests, ruff clean, YAML valid.
* docs(agent-team): fold round-2 GPT-4.1 plan-review findings into P3-live-flip plan
Round-2 cross-review (REQUEST CHANGES) folded:
- B1 deploy-before-merge made a concrete CI-enforced gate (required status check
fed by a box-exercised dispatch dry-run), not prose.
- B2 added rollback for a prematurely-flipped privileged job that actually RAN
(token rotate, revert opened PR/branch, audit the window) — distinct from an
accidental merge.
- B3 rollback is a tested, re-runnable script over ALL privileged surfaces
(workflow, environment, App perms, branch protection), not a one-time manual run.
- B4 Phase 3 explicitly gated on Phase 2 being fully provisioned + verified.
- B5 /sh-security-review + GPT-4.1 cross-review RE-RUN on the actual enabled
workflow before the flip, not only the inert version.
- B6 docs/memory updated incrementally at each privileged step; Phase 6 is the
final reconciliation pass.
Plus FIX/NIT/QUESTION: gate-weakening pattern review, check-name discovery test,
memory update on denylist change, snapshot retention in Phase 0, draft-PR
notification (Phase 4) + conservative-rollout controls (Phase 5) made concrete.
GH_TOKEN->GITHUB_TOKEN gate marked done (box alias added).
* docs(agent-team): close B1 — deploy-before-merge is a committed hard gate, not a manual fallback
Round-3 re-review resolved B2-B6 but flagged B1 still-open: the prior wording left
a 'enforced by hand until the check exists' escape hatch. Reframe B1 as a REQUIRED
Phase-1 build deliverable that blocks the flip (no manual fallback), with the
required-status-check + admin-bypass-disabled branch protection, and a dry-run
faithfulness note (same workflow file/jobs as live, only the privileged if: differs).
* docs(agent-team): close B1 ordering — admin-bypass disabled before any flip via the B4 precondition
Round-4 confirmatory re-review noted (c) admin-bypass-disable sits in Phase 2 while
the gate is framed Phase-1. Clarify there is no ordering window: the flip (Phase 3)
is gated on Phase 2 completion (B4), so branch protection incl. admin-bypass-disable
is necessarily in place before any flip. The check's implementation being a Phase-1
build task (not yet physically built) is expected for a pre-build plan; it is
non-optional and flip-blocking, enforced at the Phase-1 hard stop. Stopping the
plan-review cycle here per the project's '3 cycles, residual is build-time' rule.
* feat(agent-team): durable GitHub-issue intake de-dup (schema v2)
The intake poller de-duped ingested issues in an in-memory set that does
not survive a process restart. A scheduled/cron intake (each run a fresh
process) would therefore re-ingest every still-open labeled issue on every
run and spawn duplicate pipeline tasks. Since the box is read-only (no write
token to remove the intake label), durable de-dup is the only correct guard.
- schema v2: new ingested_issues(source, issue_id, ingested_at) table +
issue_already_ingested / record_issue_ingested helpers; migrate() adds the
table to a legacy v1 DB and restamps; init_db creates it.
- github_intake: pluggable IngestStore seam (in-memory default preserved for
tests/one-off; durable build_ledger_ingest_store for production). Record is
after start_task succeeds, so a failed intake stays retryable.
- run-team.py intake-github wires the ledger store keyed by github:owner/repo,
making a scheduled timer idempotent across runs.
978 tests pass (ruff clean).
* fix(agent-team): harden github intake per /sh-security-review (CWE-918, idempotency)
Fixes from the high-recall detector fan-out on the durable-dedup change:
- INTAKE-LOGIC-01 (idempotency): switch the IngestStore seam from
check-then-record (seen/mark) to claim-then-do (claim/release). The id is
now reserved BEFORE the non-idempotent start_task side effect, so a crash in
that window cannot re-spawn a duplicate task on the next run; a raising
start_task releases the claim so transient failures stay retryable. Adds
delete_issue_ingested to the schema layer for the release path.
- INTAKE-SSRF-001 / INTAKE-PATHSPLICE-002 (CWE-918) in build_default_issue_client:
drop the caller-overridable api_root (hardcode GITHUB_API_ROOT) and validate
owner/repo against an anchored charset before splicing them into the
token-bearing API URL — mirrors the sibling ci_fetcher BLOCK-3/FIX-3 fixes.
Tests cover cross-process duplicate prevention, release-on-failure retry, and
the owner/repo + api_root rejection. 982 tests pass, ruff clean.
Follow-up (pre-existing, not introduced here): the label-only intake has no
author allowlist (cf. AGENT_TEAM_SLACK_OWNER_IDS on the Slack listener); the
Slack answer gate bounds the blast radius. Track as separate hardening.
Three robustness/hardening fixes surfaced by /sh-security-review on the
agent-team listener/coordinator surface. None alter AUTHZ-01 allowlist
behavior or the first-answer-wins compare-and-set semantics.
1. Respawn close() leak (CWE-772). The watchdog _supervise_slack_listener
respawned the inbound Slack listener without tearing down the dead one,
leaking a Socket Mode WebSocket / SDK thread set per flap. Now the dead
listener is closed before respawn (new _close_dead_listener, idempotent),
AND _run_listener has a finally that always closes the listener so a
crashed serve() releases its socket. SlackListener.close() is idempotent,
so the belt-and-braces close stays a safe no-op.
2. Broadened exception guard in handle_event (CWE-248). The submit block
only caught ValueError; the accept path (submit_answer ->
_question_turn/_question_thread) can raise KeyError on a concurrently
mutated row, and the CAS can raise sqlite3.Error. An uncaught exception
would escape into the Bolt dispatch. Added a separate `except Exception`
that logs at WARNING (not silent, not debug) and returns None. The
existing ValueError-as-debug behavior is unchanged; authorization still
runs first, so the trust boundary is not widened.
3. channel_ref partial-unique index (defense-in-depth). Added
uq_pending_questions_open_channel_ref — a PARTIAL UNIQUE index on
(channel_ref) WHERE channel_ref IS NOT NULL AND status='open' — so two
OPEN rows can never share a non-null channel_ref (a thread_ts can never
map to two open questions). Installed in init_db AND unconditionally in
migrate (idempotent IF NOT EXISTS) so existing v1 DBs gain it. NULLs and
closed rows are excluded; mirrored verbatim into schema.sql.
Tests: +8 (was 960, now 968). New: schema partial-unique reject/null/closed/
migrate cases; handle_event KeyError + sqlite3.Error swallow cases;
coordinator close-before-respawn + run_listener-closes-on-crash. Fixed the
operator-cli test fixture to use a per-question channel_ref (it previously
inserted multiple open rows sharing one ref, which the new index correctly
rejects).