* feat(agent-team): read-only CI-result fetcher for P3 verify gate (opt-in, inert)
ci_fetcher.py: fail-closed CiResultFetcher reading the GitHub Actions run
conclusion via a read-only PAT (AGENT_TEAM_CI_READ_TOKEN→GITHUB_TOKEN), returns
{run_id,conclusion,diff_hash} or None on any error. Data-fetcher only — ci_gate
owns the verdict; never writes, no OIDC/AWS, never reads patch artifacts.
coordinator gains opt-in gated_build_verify_wiring() composing it via
bind_ci_result_fetcher; NOT wired into the default run-team.py path. 20 tests.
* harden(agent-team): P3 apply/verify workflow — GitHub App token, CWE-94, fail-closed
Decision-1 auth model: gate-and-pr uses a GitHub App installation token
(pull-requests:write) behind the agent-apply environment; ALL OIDC/id-token/AWS
removed. Hardening: task_id env-indirection (CWE-94 — GitHub expands ${{ }} into
the run shell before exec, so %s/quoting is insufficient); run-id pinning on both
download-artifact; post-build denied-path check (build-hook writes into denied
paths fail the job); empty-hash fail-closed in BOTH the embedded gate (fixed a
real ''=='' pass bug) and ci_gate.py. App-token + draft-PR steps stay if:${{ false }}
until provisioning (App + environment + branch protection). +17 tests.
* harden(agent-team): apply P3-live security-gate fixes (GPT-4.1 xreview + sh-security-review)
BLOCK-1/FIX-4: gate-and-pr re-comments pull-requests:write + environment:agent-apply
(provisioning-time uncomment) and gains needs.guard/build-test=='success' job guard —
zero privilege until provisioning. BLOCK-2/3+FIX-5: ci_fetcher validates run_id (^[0-9]{1,20}$),
owner/repo (^[A-Za-z0-9_.-]{1,100}$), and fetched_id (int) — fail closed, no SSRF/path
injection. FIX-1: conclusion allowlist. FIX-3: api_root removed from public builder (no
injectable endpoint). INJ-02: post-build denied-path check uses NUL-delimited git output +
explicit rename parsing, no backslash mangling, non-UTF8=violation. INJ-03: all three trust-
control denylists unified to one 22-entry union + drift-guard test. Q1: documented run_id/
diff_hash trust source (dispatcher/ledger only). 884 tests, ruff clean. Privileged steps stay
if:${{ false }} until provisioning.
* build(security-review): prune .claude worktrees from deterministic scanners
Agent worktrees under .claude/worktrees/ are full repo copies; the cfn-lint
find|xargs template scan overflowed ('command line cannot be assembled') and the
pre-push hook fail-closed to BLOCK whenever a worktree was present. Prune .claude
in the cfn-lint find + semgrep/checkov excludes, and gitignore .claude/ so it is
never scanned or committed. Unblocks main-tree pushes during parallel agent work.
gitleaks git-mode scans committed history, so the .env.example placeholder flags from
history and would block the orchestrator's own pre-push gate. Suppress it with a written
justification. Also gitignore stray .adf_final*.json left by an unrelated tool.
Closes the no-CI gap surfaced in the post-merge retrospective. The org
reusable workflows live under Sea-Haven-Industries and assume SAM/CDK
projects — orchestrator is a personal CLI tool with neither, so this is
a standalone workflow on the same conventions (actions/checkout@v6,
Python 3.12, ruff).
- lint job: ruff check + ruff format --check.
- test-collect job: installs requirements + runs `pytest --collect-only`.
Catches import errors and golden-set test discovery regressions without
needing live ANTHROPIC/COMPOSIO secrets — full pytest stays a local
pre-push responsibility.
Also adds .DS_Store to the local .gitignore (also covered by the user
global gitignore, but belt-and-suspenders).
Plugs the orchestrator into Adam's existing memory store at
~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/. Every
run starts with a top-3 retrieval pass that is then surfaced in the CLI
output and injected as system context into the router and downstream agent.
- retriever.py: load *.md memories (skipping the MEMORY.md index), embed
with text-embedding-3-small, cache to .cache/embeddings.json keyed on
file mtime. Cosine similarity, top-k=3 default. Reads only — never
writes back to the memory store.
- state.py: add `retrieved: list[dict]` to OrchestratorState; relax to
total=False to match LangGraph's partial-update semantics.
- graph.py: new retriever_node wired as START -> retriever -> router.
router_node and connector_node now inject retrieved memories into their
SystemMessage. Retrieval failures are caught and the run continues with
empty memory context (logged).
- agents.py: make_agent_node injects retrieved memories into each agent's
system prompt.
- run.py: prints `[retrieved: name1, name2, name3]` (or `[retrieved: none]`)
before route/result for both --route-only and full-run modes, so bad
retrieval is visible at a glance.
- .gitignore: add .cache/, .pytest_cache/, .ruff_cache/.
Validated: golden-set still 21/21 passing; smoke tests retrieve plausible
memories ("Send a Slack message to ops about the new exec-aide deploy" ->
project_exec_aide, feedback_exec_aide_vip_management, project_seahaven_slack_bot).
Multi-model agent routing with 7 agent nodes (Sonnet, Haiku,
GPT-4.1, Gemini, DeepSeek) and 15 pre-loaded Composio tools
for Slack, Notion, GitHub, and Google Drive integration.