Multi-model AI agent orchestration via LangGraph + Composio. Routes tasks to the best-fit model and connects to external services (Slack, Notion, GitHub, Google Drive). Memory-aware — each run is enriched with the top-3 most relevant notes from Adam's project/feedback/reference memory store.
The retriever embeds Adam's memory files once and caches vectors to `.cache/embeddings.json` (mtime-keyed; only changed files re-embed). Each run picks the top-3 most relevant memories and surfaces them in the CLI output before the route line.
The router uses Pydantic structured output (`RouteDecision`) and returns an explicit `"unknown"` route when no agent fits — no silent fallback. All LLM invocations are wrapped with retry-on-transient-error.
### When Claude Code delegates vs. handles natively
Claude Code uses a hybrid model — it delegates to the orchestrator when a different model has a genuine advantage, and handles everything else natively:
| Delegate to orchestrator | Handle natively in Claude Code |
The router can also return `done` (no agent needed) or `unknown` (no clear fit). Model IDs are centralized as constants in `models.py`.
## Memory retrieval
The retriever reads `~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/*.md` (skipping the `MEMORY.md` index), embeds each file once with `text-embedding-3-small`, and caches the vectors to `.cache/embeddings.json`. On subsequent runs:
- Only files whose mtime changed are re-embedded.
- Top-3 memories by cosine similarity are injected as system context into both the router and the agent.
- Retrieved names are printed as the first line of every run so bad retrieval is visible.
- Retrieval is **read-only.** The orchestrator never writes back to the memory store.
If retrieval fails (network, missing key), the run continues with no memory context and logs the failure into the message trail.
The `security-review/` subsystem is a high-recall, anti-complacency security gate. It is **separate from the router** — it does not route through `run.py` or LangGraph. One pure-code script, `review.sh`, owns the **block decision** (confirmed critical/high → block); no agent decides.
- **Path A — interactive:** the `/sh-security-review` Claude Code skill (Max-covered). Narrow fresh-context detector fan-out + a proof-or-kill verifier; emits the structured finding schema for `review.sh` to gate.
- **Path B — unattended:** a nightly two-tier sweep on the `sh-secrev` R720 VM. Tier 1 runs deterministic scanners (`review.sh --scanners-only`) over every Sea-Haven-Industries org repo; Tier 2 is a budget-bounded agentic pass (`run_headless.py`) on a round-robin rotation. Clean-clone auto-discovery via a read-only GitHub PAT; **ALARM-only** Slack (a clean night posts nothing).
- **Git hooks:** global pre-commit / pre-push hooks (`install-hooks.sh --global`) gate every local repo via `review.sh --scanners-only`.
See `security-review/README.md` for full detail and `security-review/DEPLOY-R720.md` for the VM runbook.
The `agent-team/` subsystem is a separate, durable, human-gated SDLC pipeline (LangGraph + SQLite ledger) that runs as an always-on coordinator daemon on the same `sh-secrev` R720 VM. It is distinct from this stateless router: it persists tasks across restarts and runs INTAKE → CLARIFY → PLAN → REVIEW (build/verify is deploy-gated and inert). It reuses this orchestrator's `models.py` for its non-Claude invokers, and exposes an opt-in FastAPI HTTP API (loopback, bearer auth) plus a `/delegate` Claude Code plugin hook (`sea-haven-claude-plugin/`). See `agent-team/README.md` and `agent-team/DEPLOY-R720.md`.