This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/README.md
2026-07-06 17:41:09 -04:00

155 lines
8.1 KiB
Markdown

# orchestrator
[![CI](https://github.com/Sea-Haven-Industries/orchestrator/actions/workflows/ci.yaml/badge.svg)](https://github.com/Sea-Haven-Industries/orchestrator/actions/workflows/ci.yaml)
[![Dependency Review](https://github.com/Sea-Haven-Industries/orchestrator/actions/workflows/dependency-review.yml/badge.svg)](https://github.com/Sea-Haven-Industries/orchestrator/actions/workflows/dependency-review.yml)
![Python](https://img.shields.io/badge/Python-3776AB?logo=python&logoColor=white)
![Slack](https://img.shields.io/badge/Slack-integration-4A154B?logo=slack&logoColor=white)
Multi-model AI agent orchestration via LangGraph + Composio. Routes tasks to the best-fit model and connects to external services (Slack, Notion, GitHub, Google Drive). Memory-aware — each run is enriched with the top-3 most relevant notes from Adam's project/feedback/reference memory store.
## Architecture
```
Claude Code ──► run.py ──► LangGraph StateGraph
│
▼
retriever ──► top-3 memories from
│ ~/.claude/projects/.../memory/
▼
router (Sonnet, structured output)
│
┌─────────────┼─────────────────────┐
▼ ▼ ▼
┌──────────────┐ ┌───────────┐ ┌──────────┐
│ implementer │ │ connector │ │ unknown │
│ reviewer │ │ (Composio)│ │ (no fit) │
│ researcher │ └───────────┘ └──────────┘
│ cross_reviewer│ │
│ scanner │ ▼
│ fast_coder │ tool_executor ──► summarizer
└──────────────┘
```
The retriever embeds Adam's memory files once and caches vectors to `.cache/embeddings.json` (mtime-keyed; only changed files re-embed). Each run picks the top-3 most relevant memories and surfaces them in the CLI output before the route line.
The router uses Pydantic structured output (`RouteDecision`) and returns an explicit `"unknown"` route when no agent fits — no silent fallback. All LLM invocations are wrapped with retry-on-transient-error.
## Files
| File | Purpose |
|---|---|
| `run.py` | CLI entry point — `python3 run.py "<task>"` |
| `graph.py` | LangGraph graph: retriever, router, connector, summarizer, unknown nodes |
| `agents.py` | `AGENTS` registry (label → model_fn, prompt, description) + `make_agent_node` factory |
| `models.py` | LLM factories, model-ID constants, `with_retries()` helper |
| `state.py` | `OrchestratorState` TypedDict |
| `retriever.py` | Memory loader, embedder, cache, top-k retrieval |
| `tools.py` | Composio tool loading (Slack, Notion, GitHub, Google Drive) |
| `tests/test_routing_golden.py` | 20-case golden-set regression test for the router |
## Usage
```bash
# Full execution — retrieves memory, routes, and runs the task
python3 run.py "What is the LangGraph checkpoint API?"
# Route-only — retrieves memory and prints the agent that would handle the task
python3 run.py --route-only "Review this code for security issues"
```
Output shape:
```
[retrieved: project_seahaven_slack_bot, feedback_secrets_manager, reference_sea_haven_aws]
[reviewer]
<agent output>
```
From Claude Code (via CLAUDE.md hybrid delegation):
```bash
python3 ~/Documents/repositories/orchestrator/run.py "<task description>"
python3 ~/Documents/repositories/orchestrator/run.py --route-only "<task description>"
```
### When Claude Code delegates vs. handles natively
Claude Code uses a hybrid model — it delegates to the orchestrator when a different model has a genuine advantage, and handles everything else natively:
| Delegate to orchestrator | Handle natively in Claude Code |
|---|---|
| Cross-family code review (GPT-4.1) | File editing, refactoring, bug fixes |
| Large codebase scanning (Gemini) | Git operations, PRs, merges |
| Quick bounded coding (DeepSeek) | AWS/SAM/CDK deployments |
| External service actions (Composio) | Shell commands, system admin |
| | Interactive planning and conversation |
## Agents
| Agent | Model | Use Case |
|---|---|---|
| implementer | Claude Sonnet | Write code with a clear spec |
| reviewer | Claude Sonnet | Code review (BLOCK/FIX/NIT/QUESTION) |
| researcher | Claude Haiku | Doc lookups, API research |
| cross_reviewer | GPT-4.1 | Independent second-opinion review |
| scanner | Gemini 2.5 Pro | Large codebase analysis |
| fast_coder | DeepSeek Coder | Quick, bounded coding tasks |
| connector | Sonnet + Composio | Slack, Notion, GitHub, Google Drive |
The router can also return `done` (no agent needed) or `unknown` (no clear fit). Model IDs are centralized as constants in `models.py`.
## Memory retrieval
The retriever reads `~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/*.md` (skipping the `MEMORY.md` index), embeds each file once with `text-embedding-3-small`, and caches the vectors to `.cache/embeddings.json`. On subsequent runs:
- Only files whose mtime changed are re-embedded.
- Top-3 memories by cosine similarity are injected as system context into both the router and the agent.
- Retrieved names are printed as the first line of every run so bad retrieval is visible.
- Retrieval is **read-only.** The orchestrator never writes back to the memory store.
If retrieval fails (network, missing key), the run continues with no memory context and logs the failure into the message trail.
## Connectors (via Composio)
All connections authenticated under Composio user `amoussa`:
- **Slack**: send messages, read channels/threads, find users, add reactions
- **Notion**: search/read/create/update pages, add content
- **GitHub**: create issues, list issues, get repo info
- **Google Drive**: find files, get metadata
The connector node is restricted to **one tool call per run** — a load-bearing rule learned from a 1.9M-token incident with meta-tool routing.
## Security Review
The `security-review/` subsystem is a high-recall, anti-complacency security gate. It is **separate from the router** — it does not route through `run.py` or LangGraph. One pure-code script, `review.sh`, owns the **block decision** (confirmed critical/high → block); no agent decides.
- **Path A — interactive:** the `/sh-security-review` Claude Code skill (Max-covered). Narrow fresh-context detector fan-out + a proof-or-kill verifier; emits the structured finding schema for `review.sh` to gate.
- **Path B — unattended:** a nightly two-tier sweep on the `sh-secrev` R720 VM. Tier 1 runs deterministic scanners (`review.sh --scanners-only`) over every Sea-Haven-Industries org repo; Tier 2 is a budget-bounded agentic pass (`run_headless.py`) on a round-robin rotation. Clean-clone auto-discovery via a read-only GitHub PAT; **ALARM-only** Slack (a clean night posts nothing).
- **Git hooks:** global pre-commit / pre-push hooks (`install-hooks.sh --global`) gate every local repo via `review.sh --scanners-only`.
See `security-review/README.md` for full detail and `security-review/DEPLOY-R720.md` for the VM runbook.
## Setup
1. Install dependencies: `pip install -r requirements.txt`
2. Copy `.env.example` to `.env` and fill in API keys
3. Authenticate Composio integrations at [app.composio.dev](https://app.composio.dev)
## Configuration
All API keys are stored in `.env` (gitignored):
- `ANTHROPIC_API_KEY` — Claude models + router
- `OPENAI_API_KEY` — GPT-4.1 cross-reviewer + text-embedding-3-small
- `GOOGLE_API_KEY` — Gemini scanner
- `DEEPSEEK_API_KEY` — DeepSeek fast-coder
- `COMPOSIO_API_KEY` — Composio connectors
- `LANGSMITH_API_KEY` — LangSmith tracing
Tracing is enabled via LangSmith (project: `orchestration`).
## Testing
```bash
pytest tests/test_routing_golden.py -v
```
20 labelled tasks → expected agent. Skipped cleanly if `ANTHROPIC_API_KEY` or `COMPOSIO_API_KEY` are unset.