* feat(agent-team): P3-flip Phase 1 — expand denylist vectors (§4.2) + runner-trust assertion (§4.1) First controls of the P3-live-flip Phase-1 CI hardening (workflow stays INERT; this only tightens the trust boundary). Whole Phase-1 surface is gated by /sh-security-review + GPT-4.1 cross-review before any flip. §4.2 — expand the trust-control denylist with direct code-execution / supply-chain vectors, kept byte-identical across all three copies (ci_gate.DENYLIST_GLOBS + the guard + post-build inline DENY_GLOBS), drift-guarded: .gitmodules, .husky/**, .githooks/**, .gitattributes, .npmrc, and generated/build artifacts (__generated__, *.generated.*, dist/**, build/**, *.min.js). Deliberate: lockfiles are NOT wholesale denied — lockfile-postinstall RCE is already contained by the credential-less egress-blocked build sandbox, and the Tier-3 dep-CVE fixer rewrites lockfiles to produce its draft PRs; a blanket deny would make it un-shippable. Flagged in-code for the security gate. Direct code-execution config (hooks/filters/npmrc/submodules) is the actual §4.2 RCE surface. §4.1 — runner-trust: assert no job (esp. the privileged gate-and-pr) can run on a self-hosted/user-provided runner; all must be GitHub-hosted. 998 tests pass, ruff clean. * feat(agent-team): P3-flip Phase 1 — gate-weakening detector (§4.5) A diff that ADDS a lint/type/coverage/security suppression (noqa, type: ignore, pragma: no cover, nosec, nosemgrep), a test skip/xfail, or a hook bypass (--no-verify) could make CI pass falsely. The pure-code gate now flags these via gate_weakening_violations() and BLOCKs in evaluate_ci_gate as a top-priority trust violation (step 1b, alongside the denylist) — regardless of the authenticated CI conclusion. A build cannot pass itself by disabling its own checks; flagged diffs escalate to a human. Only ADDED lines are inspected (removing a suppression is fine). 1015 tests pass, ruff clean. * feat(agent-team): P3-flip — diff transport (§4.3) + flip privileged apply path live Completes the box->CI diff handoff and flips the apply/verify privileged job live (gated behind the agent-apply environment's required reviewer). Transport (§4.3): the read-only box (D2) emits a diff but holds no write token. - New credential-less `materialize` job decodes the untrusted `diff_b64` dispatch input via env (CWE-94), fail-closed re-hashes it against `expected_diff_hash`, and uploads it as the named artifact so guard/build-test download it same-run. guard now `needs: materialize`. - New `dispatcher.py` (the trusted apply path, operator/Mac-side — never the box): pushes the diff as a head branch then `gh workflow run`s the workflow. Pure input-assembly (sha256 == sha256sum, b64 round-trip, head ref) is unit-tested; git/gh are injected seams. Push-before-dispatch; fail-closed on empty diff/scope, unsafe task_id/owner/repo. Flip: gate-and-pr binds `environment: agent-apply` (required reviewer amoussa1229) + grants exactly `pull-requests: write`; the App-token + draft-PR steps run only on `steps.gate.outputs.gate == 'pass'` (no more if:false); the draft PR opens with an explicit `--head`; task_id/head_branch charset-validated (§4.6). Updated the hardening tests from inert-state to live-state assertions + added transport tests. 1039 tests, ruff clean, workflow YAML valid. NOTE: workflow only runs on manual workflow_dispatch and the privileged job is held at the required-reviewer gate, so nothing privileged runs unapproved. * fix(agent-team): P3-flip — address GPT-4.1 cross-review (size bound, ref-traversal guard) - BLOCK: cap candidate diff at 40 KB in the dispatcher (the diff rides a base64 workflow_dispatch input; GitHub caps inputs at ~64 KB so an oversized diff cannot dispatch at all) + a defense-in-depth decoded-size bound in materialize. - FIX: harden the draft-PR HEAD_BRANCH guard to reject leading/trailing slash, '..' segments, and '//' (CWE-88 git ref-traversal), not just bad charset. - NIT: document the mandatory invariants on gate-and-pr (required-reviewer environment must stay; runs-on must stay GitHub-hosted). - QUESTION (lockfiles): answered in-code — the build-test sandbox is credential-less + egress-blocked, so lockfile-postinstall RCE is contained. Tests added for all guards. 1042 tests, ruff clean, YAML valid. * fix(agent-team): P3-flip — resolve /sh-security-review findings (LOGIC-1/2/3) High-recall fan-out (injection/logic/iac+secrets) + proof-or-kill on the LIVE apply path found 3 real issues the cross-review missed; all fixed: - LOGIC-2 (HIGH, was a live hole): build-test ran `ruff check . || echo` / `pytest -q || echo`, swallowing failures so the job was always 'success' and the gate would open draft PRs on RED builds. ruff/pytest now run authoritatively under set -e (pytest exit 5 'no tests' is the only non-fatal case); the exit code IS the build-test conclusion the gate keys on. - LOGIC-1 (verified!=shipped): the dispatcher used `git apply` + `git add -A`, staging stray untracked content into the pushed PR head. Now `git apply --index` stages exactly the diff, so the head tree is precisely base+diff — bound to the bytes CI hash-verified. - LOGIC-3 (§4.5 on the live path): gate-weakening was enforced only box-side; added a gate-weakening check to the guard job so the live PR-opening path rejects a diff that adds suppressions/skips, even on a green build. Injection / secrets / least-privilege / flip-correctness / no-untrusted-checkout all came back clean. 1044 tests, ruff clean, YAML valid. |
||
|---|---|---|
| .github | ||
| agent-team | ||
| docs | ||
| scripts | ||
| security-review | ||
| tests | ||
| .env.example | ||
| .gitignore | ||
| agents.py | ||
| conftest.py | ||
| graph.py | ||
| models.py | ||
| README.md | ||
| requirements.txt | ||
| retriever.py | ||
| run.py | ||
| state.py | ||
| telemetry.py | ||
| tools.py | ||
orchestrator
Multi-model AI agent orchestration via LangGraph + Composio. Routes tasks to the best-fit model and connects to external services (Slack, Notion, GitHub, Google Drive). Memory-aware — each run is enriched with the top-3 most relevant notes from Adam's project/feedback/reference memory store.
Architecture
Claude Code ──► run.py ──► LangGraph StateGraph
│
▼
retriever ──► top-3 memories from
│ ~/.claude/projects/.../memory/
▼
router (Sonnet, structured output)
│
┌─────────────┼─────────────────────┐
▼ ▼ ▼
┌──────────────┐ ┌───────────┐ ┌──────────┐
│ implementer │ │ connector │ │ unknown │
│ reviewer │ │ (Composio)│ │ (no fit) │
│ researcher │ └───────────┘ └──────────┘
│ cross_reviewer│ │
│ scanner │ ▼
│ fast_coder │ tool_executor ──► summarizer
└──────────────┘
The retriever embeds Adam's memory files once and caches vectors to .cache/embeddings.json (mtime-keyed; only changed files re-embed). Each run picks the top-3 most relevant memories and surfaces them in the CLI output before the route line.
The router uses Pydantic structured output (RouteDecision) and returns an explicit "unknown" route when no agent fits — no silent fallback. All LLM invocations are wrapped with retry-on-transient-error.
Files
| File | Purpose |
|---|---|
run.py |
CLI entry point — python3 run.py "<task>" |
graph.py |
LangGraph graph: retriever, router, connector, summarizer, unknown nodes |
agents.py |
AGENTS registry (label → model_fn, prompt, description) + make_agent_node factory |
models.py |
LLM factories, model-ID constants, with_retries() helper |
state.py |
OrchestratorState TypedDict |
retriever.py |
Memory loader, embedder, cache, top-k retrieval |
tools.py |
Composio tool loading (Slack, Notion, GitHub, Google Drive) |
tests/test_routing_golden.py |
20-case golden-set regression test for the router |
Usage
# Full execution — retrieves memory, routes, and runs the task
python3 run.py "What is the LangGraph checkpoint API?"
# Route-only — retrieves memory and prints the agent that would handle the task
python3 run.py --route-only "Review this code for security issues"
Output shape:
[retrieved: project_seahaven_slack_bot, feedback_secrets_manager, reference_sea_haven_aws]
[reviewer]
<agent output>
From Claude Code (via CLAUDE.md hybrid delegation):
python3 ~/Documents/repositories/orchestrator/run.py "<task description>"
python3 ~/Documents/repositories/orchestrator/run.py --route-only "<task description>"
When Claude Code delegates vs. handles natively
Claude Code uses a hybrid model — it delegates to the orchestrator when a different model has a genuine advantage, and handles everything else natively:
| Delegate to orchestrator | Handle natively in Claude Code |
|---|---|
| Cross-family code review (GPT-4.1) | File editing, refactoring, bug fixes |
| Large codebase scanning (Gemini) | Git operations, PRs, merges |
| Quick bounded coding (DeepSeek) | AWS/SAM/CDK deployments |
| External service actions (Composio) | Shell commands, system admin |
| Interactive planning and conversation |
Agents
| Agent | Model | Use Case |
|---|---|---|
| implementer | Claude Sonnet | Write code with a clear spec |
| reviewer | Claude Sonnet | Code review (BLOCK/FIX/NIT/QUESTION) |
| researcher | Claude Haiku | Doc lookups, API research |
| cross_reviewer | GPT-4.1 | Independent second-opinion review |
| scanner | Gemini 2.5 Pro | Large codebase analysis |
| fast_coder | DeepSeek Coder | Quick, bounded coding tasks |
| connector | Sonnet + Composio | Slack, Notion, GitHub, Google Drive |
The router can also return done (no agent needed) or unknown (no clear fit). Model IDs are centralized as constants in models.py.
Memory retrieval
The retriever reads ~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/*.md (skipping the MEMORY.md index), embeds each file once with text-embedding-3-small, and caches the vectors to .cache/embeddings.json. On subsequent runs:
- Only files whose mtime changed are re-embedded.
- Top-3 memories by cosine similarity are injected as system context into both the router and the agent.
- Retrieved names are printed as the first line of every run so bad retrieval is visible.
- Retrieval is read-only. The orchestrator never writes back to the memory store.
If retrieval fails (network, missing key), the run continues with no memory context and logs the failure into the message trail.
Connectors (via Composio)
All connections authenticated under Composio user amoussa:
- Slack: send messages, read channels/threads, find users, add reactions
- Notion: search/read/create/update pages, add content
- GitHub: create issues, list issues, get repo info
- Google Drive: find files, get metadata
The connector node is restricted to one tool call per run — a load-bearing rule learned from a 1.9M-token incident with meta-tool routing.
Security Review
The security-review/ subsystem is a high-recall, anti-complacency security gate. It is separate from the router — it does not route through run.py or LangGraph. One pure-code script, review.sh, owns the block decision (confirmed critical/high → block); no agent decides.
- Path A — interactive: the
/sh-security-reviewClaude Code skill (Max-covered). Narrow fresh-context detector fan-out + a proof-or-kill verifier; emits the structured finding schema forreview.shto gate. - Path B — unattended: a nightly two-tier sweep on the
sh-secrevR720 VM. Tier 1 runs deterministic scanners (review.sh --scanners-only) over every Sea-Haven-Industries org repo; Tier 2 is a budget-bounded agentic pass (run_headless.py) on a round-robin rotation. Clean-clone auto-discovery via a read-only GitHub PAT; ALARM-only Slack (a clean night posts nothing). - Git hooks: global pre-commit / pre-push hooks (
install-hooks.sh --global) gate every local repo viareview.sh --scanners-only.
See security-review/README.md for full detail and security-review/DEPLOY-R720.md for the VM runbook.
Setup
- Install dependencies:
pip install -r requirements.txt - Copy
.env.exampleto.envand fill in API keys - Authenticate Composio integrations at app.composio.dev
Configuration
All API keys are stored in .env (gitignored):
ANTHROPIC_API_KEY— Claude models + routerOPENAI_API_KEY— GPT-4.1 cross-reviewer + text-embedding-3-smallGOOGLE_API_KEY— Gemini scannerDEEPSEEK_API_KEY— DeepSeek fast-coderCOMPOSIO_API_KEY— Composio connectorsLANGSMITH_API_KEY— LangSmith tracing
Tracing is enabled via LangSmith (project: orchestration).
Testing
pytest tests/test_routing_golden.py -v
20 labelled tasks → expected agent. Skipped cleanly if ANTHROPIC_API_KEY or COMPOSIO_API_KEY are unset.