* fix(agent-team): serve() starts the inbound Slack listener (D-1) Coordinator.serve() now constructs and starts the SlackListener concurrently with the tick/drain loop on a background daemon thread, but ONLY when the live transport is a SlackTransport AND SLACK_APP_TOKEN is configured. When Slack is not the transport or the app token is absent, serve() behaves exactly as before (tick/recover only) — Slack is never made mandatory. - New injectable build_listener seam + default_slack_listener_factory sharing the coordinator's own transport, ledger db_path, and resume_queue put. - AUTHZ-01 owner-allowlist + open-status CAS untouched: serve() sources AGENT_TEAM_SLACK_OWNER_IDS in SlackListener.serve, which still fails closed. - SlackListener.close() added for clean Socket Mode teardown on shutdown; serve() stops the listener + joins the thread in a finally. - Tests: start-when-Slack+app-token, no-start otherwise, clean shutdown, idempotent start, serve start/stop around the loop, listener close(). * fix(agent-team): systemd unit loads ~/orchestrator/.env + uses venv python (D-2/D-7) D-2: add EnvironmentFile=-/home/adam/orchestrator/.env (optional '-') so the P2 GPT-4.1 review loop's cross_reviewer sub-process can read the non-Claude provider key once a task reaches REVIEW. Mirrors the sea-haven-secrev unit. D-7: point ExecStart at the agent-team venv interpreter (/home/adam/orchestrator/agent-team/.venv/bin/python) instead of /usr/bin/env python3, which resolved the system interpreter without the installed deps under systemd's PATH. All hardening (NoNewPrivileges / ProtectSystem=full / ProtectHome=read-only / ReadWritePaths) is retained unchanged (locked decision). * docs(agent-team): land provisioning + operator runbooks under docs/provisioning - PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5 intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked FIXED so the demo can use the live Slack answer path. - P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the Slack and operator-CLI answer paths documented for all four exit criteria. - DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the operator-CLI divergence kept as provisioning notes. - OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks, failed HITL resumes, budget exhaustion, transport outages, and COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6). * fix(agent-team): supervise the Slack listener thread — recurring ALARM + respawn sh-security-review (logic) MEDIUM: a crashed listener thread was logged once, then the daemon ran on 'deaf' — posting clarifier questions but receiving no answers, every gate silently parking, process never exiting so systemd Restart=on-failure never fired. serve() now calls _supervise_slack_listener() each pass: when the listener is enabled but its thread is dead, it emits a recurring ERROR ALARM and respawns via the idempotent starter (self-heal). No-op when alive or disabled. +3 tests. (authz detector: wiring clean — AUTHZ-01 fail-closed allowlist + open-status CAS intact, dead listener fails SAFE.) |
||
|---|---|---|
| .github | ||
| agent-team | ||
| docs | ||
| scripts | ||
| security-review | ||
| tests | ||
| .env.example | ||
| .gitignore | ||
| agents.py | ||
| conftest.py | ||
| graph.py | ||
| models.py | ||
| README.md | ||
| requirements.txt | ||
| retriever.py | ||
| run.py | ||
| state.py | ||
| telemetry.py | ||
| tools.py | ||
orchestrator
Multi-model AI agent orchestration via LangGraph + Composio. Routes tasks to the best-fit model and connects to external services (Slack, Notion, GitHub, Google Drive). Memory-aware — each run is enriched with the top-3 most relevant notes from Adam's project/feedback/reference memory store.
Architecture
Claude Code ──► run.py ──► LangGraph StateGraph
│
▼
retriever ──► top-3 memories from
│ ~/.claude/projects/.../memory/
▼
router (Sonnet, structured output)
│
┌─────────────┼─────────────────────┐
▼ ▼ ▼
┌──────────────┐ ┌───────────┐ ┌──────────┐
│ implementer │ │ connector │ │ unknown │
│ reviewer │ │ (Composio)│ │ (no fit) │
│ researcher │ └───────────┘ └──────────┘
│ cross_reviewer│ │
│ scanner │ ▼
│ fast_coder │ tool_executor ──► summarizer
└──────────────┘
The retriever embeds Adam's memory files once and caches vectors to .cache/embeddings.json (mtime-keyed; only changed files re-embed). Each run picks the top-3 most relevant memories and surfaces them in the CLI output before the route line.
The router uses Pydantic structured output (RouteDecision) and returns an explicit "unknown" route when no agent fits — no silent fallback. All LLM invocations are wrapped with retry-on-transient-error.
Files
| File | Purpose |
|---|---|
run.py |
CLI entry point — python3 run.py "<task>" |
graph.py |
LangGraph graph: retriever, router, connector, summarizer, unknown nodes |
agents.py |
AGENTS registry (label → model_fn, prompt, description) + make_agent_node factory |
models.py |
LLM factories, model-ID constants, with_retries() helper |
state.py |
OrchestratorState TypedDict |
retriever.py |
Memory loader, embedder, cache, top-k retrieval |
tools.py |
Composio tool loading (Slack, Notion, GitHub, Google Drive) |
tests/test_routing_golden.py |
20-case golden-set regression test for the router |
Usage
# Full execution — retrieves memory, routes, and runs the task
python3 run.py "What is the LangGraph checkpoint API?"
# Route-only — retrieves memory and prints the agent that would handle the task
python3 run.py --route-only "Review this code for security issues"
Output shape:
[retrieved: project_seahaven_slack_bot, feedback_secrets_manager, reference_sea_haven_aws]
[reviewer]
<agent output>
From Claude Code (via CLAUDE.md hybrid delegation):
python3 ~/Documents/repositories/orchestrator/run.py "<task description>"
python3 ~/Documents/repositories/orchestrator/run.py --route-only "<task description>"
When Claude Code delegates vs. handles natively
Claude Code uses a hybrid model — it delegates to the orchestrator when a different model has a genuine advantage, and handles everything else natively:
| Delegate to orchestrator | Handle natively in Claude Code |
|---|---|
| Cross-family code review (GPT-4.1) | File editing, refactoring, bug fixes |
| Large codebase scanning (Gemini) | Git operations, PRs, merges |
| Quick bounded coding (DeepSeek) | AWS/SAM/CDK deployments |
| External service actions (Composio) | Shell commands, system admin |
| Interactive planning and conversation |
Agents
| Agent | Model | Use Case |
|---|---|---|
| implementer | Claude Sonnet | Write code with a clear spec |
| reviewer | Claude Sonnet | Code review (BLOCK/FIX/NIT/QUESTION) |
| researcher | Claude Haiku | Doc lookups, API research |
| cross_reviewer | GPT-4.1 | Independent second-opinion review |
| scanner | Gemini 2.5 Pro | Large codebase analysis |
| fast_coder | DeepSeek Coder | Quick, bounded coding tasks |
| connector | Sonnet + Composio | Slack, Notion, GitHub, Google Drive |
The router can also return done (no agent needed) or unknown (no clear fit). Model IDs are centralized as constants in models.py.
Memory retrieval
The retriever reads ~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/*.md (skipping the MEMORY.md index), embeds each file once with text-embedding-3-small, and caches the vectors to .cache/embeddings.json. On subsequent runs:
- Only files whose mtime changed are re-embedded.
- Top-3 memories by cosine similarity are injected as system context into both the router and the agent.
- Retrieved names are printed as the first line of every run so bad retrieval is visible.
- Retrieval is read-only. The orchestrator never writes back to the memory store.
If retrieval fails (network, missing key), the run continues with no memory context and logs the failure into the message trail.
Connectors (via Composio)
All connections authenticated under Composio user amoussa:
- Slack: send messages, read channels/threads, find users, add reactions
- Notion: search/read/create/update pages, add content
- GitHub: create issues, list issues, get repo info
- Google Drive: find files, get metadata
The connector node is restricted to one tool call per run — a load-bearing rule learned from a 1.9M-token incident with meta-tool routing.
Security Review
The security-review/ subsystem is a high-recall, anti-complacency security gate. It is separate from the router — it does not route through run.py or LangGraph. One pure-code script, review.sh, owns the block decision (confirmed critical/high → block); no agent decides.
- Path A — interactive: the
/sh-security-reviewClaude Code skill (Max-covered). Narrow fresh-context detector fan-out + a proof-or-kill verifier; emits the structured finding schema forreview.shto gate. - Path B — unattended: a nightly two-tier sweep on the
sh-secrevR720 VM. Tier 1 runs deterministic scanners (review.sh --scanners-only) over every Sea-Haven-Industries org repo; Tier 2 is a budget-bounded agentic pass (run_headless.py) on a round-robin rotation. Clean-clone auto-discovery via a read-only GitHub PAT; ALARM-only Slack (a clean night posts nothing). - Git hooks: global pre-commit / pre-push hooks (
install-hooks.sh --global) gate every local repo viareview.sh --scanners-only.
See security-review/README.md for full detail and security-review/DEPLOY-R720.md for the VM runbook.
Setup
- Install dependencies:
pip install -r requirements.txt - Copy
.env.exampleto.envand fill in API keys - Authenticate Composio integrations at app.composio.dev
Configuration
All API keys are stored in .env (gitignored):
ANTHROPIC_API_KEY— Claude models + routerOPENAI_API_KEY— GPT-4.1 cross-reviewer + text-embedding-3-smallGOOGLE_API_KEY— Gemini scannerDEEPSEEK_API_KEY— DeepSeek fast-coderCOMPOSIO_API_KEY— Composio connectorsLANGSMITH_API_KEY— LangSmith tracing
Tracing is enabled via LangSmith (project: orchestration).
Testing
pytest tests/test_routing_golden.py -v
20 labelled tasks → expected agent. Skipped cleanly if ANTHROPIC_API_KEY or COMPOSIO_API_KEY are unset.