WS Slack-UX Feature 1. A /new-task task now maps to ONE Slack thread instead of
several top-level messages.
- /new-task posts an immediate root "📥 Task received: …" ack and captures its
ts (root_ts); this is the instant acknowledgement.
- root_ts is plumbed into start: new PipelineState/TaskRecord channel
slack_thread_ts, seeded by graph.start_task and threaded through
Coordinator.start_task. The NewTaskCallback is now (task_text, via, root_ts).
- All clarifier questions for the task post as THREADED REPLIES under root_ts
(chat.postMessage thread_ts=root_ts), and each question's ledger channel_ref
is set to root_ts (NOT the reply's own ts). Because answer-mapping resolves a
reply via find_open_question_by_channel_ref(thread_ts), a reply in the root
thread (thread_ts==root_ts) maps to the task's currently-open question with NO
change to the mapping logic or the first-answer-wins CAS. The open-only
partial-unique index still holds (one open question per task at a time).
- Lifecycle milestones (parked / plan-ready / needs-input) and follow-up
questions thread under root_ts too; the notify sink gained an optional
thread_ts kwarg (degrades to top-level on a sink that doesn't accept it).
notify failures still never break tick.
- SlackTransport.post_question + the live poster accept/forward thread_ts.
- No root_ts (non-/new-task origin) ⇒ top-level posts exactly as before.
AUTHZ-01 (owner-allowlist-first, fail-closed) and the atomic open→answered
compare-and-set are unchanged.
Adds plumbing for the inbound-ack reactor seam used by Feature 2 (dormant until
a reactor is injected). Tests cover thread_ts forwarding, channel_ref=root_ts,
graph seeding, and coordinator threading.
|
||
|---|---|---|
| .github | ||
| agent-team | ||
| docs | ||
| scripts | ||
| sea-haven-claude-plugin | ||
| security-review | ||
| tests | ||
| .env.example | ||
| .gitignore | ||
| agents.py | ||
| conftest.py | ||
| graph.py | ||
| models.py | ||
| README.md | ||
| requirements.txt | ||
| retriever.py | ||
| run.py | ||
| state.py | ||
| telemetry.py | ||
| tools.py | ||
orchestrator
Multi-model AI agent orchestration via LangGraph + Composio. Routes tasks to the best-fit model and connects to external services (Slack, Notion, GitHub, Google Drive). Memory-aware — each run is enriched with the top-3 most relevant notes from Adam's project/feedback/reference memory store.
Architecture
Claude Code ──► run.py ──► LangGraph StateGraph
│
▼
retriever ──► top-3 memories from
│ ~/.claude/projects/.../memory/
▼
router (Sonnet, structured output)
│
┌─────────────┼─────────────────────┐
▼ ▼ ▼
┌──────────────┐ ┌───────────┐ ┌──────────┐
│ implementer │ │ connector │ │ unknown │
│ reviewer │ │ (Composio)│ │ (no fit) │
│ researcher │ └───────────┘ └──────────┘
│ cross_reviewer│ │
│ scanner │ ▼
│ fast_coder │ tool_executor ──► summarizer
└──────────────┘
The retriever embeds Adam's memory files once and caches vectors to .cache/embeddings.json (mtime-keyed; only changed files re-embed). Each run picks the top-3 most relevant memories and surfaces them in the CLI output before the route line.
The router uses Pydantic structured output (RouteDecision) and returns an explicit "unknown" route when no agent fits — no silent fallback. All LLM invocations are wrapped with retry-on-transient-error.
Files
| File | Purpose |
|---|---|
run.py |
CLI entry point — python3 run.py "<task>" |
graph.py |
LangGraph graph: retriever, router, connector, summarizer, unknown nodes |
agents.py |
AGENTS registry (label → model_fn, prompt, description) + make_agent_node factory |
models.py |
LLM factories, model-ID constants, with_retries() helper |
state.py |
OrchestratorState TypedDict |
retriever.py |
Memory loader, embedder, cache, top-k retrieval |
tools.py |
Composio tool loading (Slack, Notion, GitHub, Google Drive) |
tests/test_routing_golden.py |
20-case golden-set regression test for the router |
Usage
# Full execution — retrieves memory, routes, and runs the task
python3 run.py "What is the LangGraph checkpoint API?"
# Route-only — retrieves memory and prints the agent that would handle the task
python3 run.py --route-only "Review this code for security issues"
Output shape:
[retrieved: project_seahaven_slack_bot, feedback_secrets_manager, reference_sea_haven_aws]
[reviewer]
<agent output>
From Claude Code (via CLAUDE.md hybrid delegation):
python3 ~/Documents/repositories/orchestrator/run.py "<task description>"
python3 ~/Documents/repositories/orchestrator/run.py --route-only "<task description>"
When Claude Code delegates vs. handles natively
Claude Code uses a hybrid model — it delegates to the orchestrator when a different model has a genuine advantage, and handles everything else natively:
| Delegate to orchestrator | Handle natively in Claude Code |
|---|---|
| Cross-family code review (GPT-4.1) | File editing, refactoring, bug fixes |
| Large codebase scanning (Gemini) | Git operations, PRs, merges |
| Quick bounded coding (DeepSeek) | AWS/SAM/CDK deployments |
| External service actions (Composio) | Shell commands, system admin |
| Interactive planning and conversation |
Agents
| Agent | Model | Use Case |
|---|---|---|
| implementer | Claude Sonnet | Write code with a clear spec |
| reviewer | Claude Sonnet | Code review (BLOCK/FIX/NIT/QUESTION) |
| researcher | Claude Haiku | Doc lookups, API research |
| cross_reviewer | GPT-4.1 | Independent second-opinion review |
| scanner | Gemini 2.5 Pro | Large codebase analysis |
| fast_coder | DeepSeek Coder | Quick, bounded coding tasks |
| connector | Sonnet + Composio | Slack, Notion, GitHub, Google Drive |
The router can also return done (no agent needed) or unknown (no clear fit). Model IDs are centralized as constants in models.py.
Memory retrieval
The retriever reads ~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/*.md (skipping the MEMORY.md index), embeds each file once with text-embedding-3-small, and caches the vectors to .cache/embeddings.json. On subsequent runs:
- Only files whose mtime changed are re-embedded.
- Top-3 memories by cosine similarity are injected as system context into both the router and the agent.
- Retrieved names are printed as the first line of every run so bad retrieval is visible.
- Retrieval is read-only. The orchestrator never writes back to the memory store.
If retrieval fails (network, missing key), the run continues with no memory context and logs the failure into the message trail.
Connectors (via Composio)
All connections authenticated under Composio user amoussa:
- Slack: send messages, read channels/threads, find users, add reactions
- Notion: search/read/create/update pages, add content
- GitHub: create issues, list issues, get repo info
- Google Drive: find files, get metadata
The connector node is restricted to one tool call per run — a load-bearing rule learned from a 1.9M-token incident with meta-tool routing.
Security Review
The security-review/ subsystem is a high-recall, anti-complacency security gate. It is separate from the router — it does not route through run.py or LangGraph. One pure-code script, review.sh, owns the block decision (confirmed critical/high → block); no agent decides.
- Path A — interactive: the
/sh-security-reviewClaude Code skill (Max-covered). Narrow fresh-context detector fan-out + a proof-or-kill verifier; emits the structured finding schema forreview.shto gate. - Path B — unattended: a nightly two-tier sweep on the
sh-secrevR720 VM. Tier 1 runs deterministic scanners (review.sh --scanners-only) over every Sea-Haven-Industries org repo; Tier 2 is a budget-bounded agentic pass (run_headless.py) on a round-robin rotation. Clean-clone auto-discovery via a read-only GitHub PAT; ALARM-only Slack (a clean night posts nothing). - Git hooks: global pre-commit / pre-push hooks (
install-hooks.sh --global) gate every local repo viareview.sh --scanners-only.
See security-review/README.md for full detail and security-review/DEPLOY-R720.md for the VM runbook.
agent-team (R720 durable SDLC pipeline)
The agent-team/ subsystem is a separate, durable, human-gated SDLC pipeline (LangGraph + SQLite ledger) that runs as an always-on coordinator daemon on the same sh-secrev R720 VM. It is distinct from this stateless router: it persists tasks across restarts and runs INTAKE → CLARIFY → PLAN → REVIEW (build/verify is deploy-gated and inert). It reuses this orchestrator's models.py for its non-Claude invokers, and exposes an opt-in FastAPI HTTP API (loopback, bearer auth) plus a /delegate Claude Code plugin hook (sea-haven-claude-plugin/). See agent-team/README.md and agent-team/DEPLOY-R720.md.
Setup
- Install dependencies:
pip install -r requirements.txt - Copy
.env.exampleto.envand fill in API keys - Authenticate Composio integrations at app.composio.dev
Configuration
All API keys are stored in .env (gitignored):
ANTHROPIC_API_KEY— Claude models + routerOPENAI_API_KEY— GPT-4.1 cross-reviewer + text-embedding-3-smallGOOGLE_API_KEY— Gemini scannerDEEPSEEK_API_KEY— DeepSeek fast-coderCOMPOSIO_API_KEY— Composio connectorsLANGSMITH_API_KEY— LangSmith tracing
Tracing is enabled via LangSmith (project: orchestration).
Testing
pytest tests/test_routing_golden.py -v
20 labelled tasks → expected agent. Skipped cleanly if ANTHROPIC_API_KEY or COMPOSIO_API_KEY are unset.