Hybrid multi-model task orchestrator — routes coding/review/scan tasks across LLMs via LangGraph
Plugs the orchestrator into Adam's existing memory store at
~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/. Every
run starts with a top-3 retrieval pass that is then surfaced in the CLI
output and injected as system context into the router and downstream agent.
- retriever.py: load *.md memories (skipping the MEMORY.md index), embed
with text-embedding-3-small, cache to .cache/embeddings.json keyed on
file mtime. Cosine similarity, top-k=3 default. Reads only — never
writes back to the memory store.
- state.py: add `retrieved: list[dict]` to OrchestratorState; relax to
total=False to match LangGraph's partial-update semantics.
- graph.py: new retriever_node wired as START -> retriever -> router.
router_node and connector_node now inject retrieved memories into their
SystemMessage. Retrieval failures are caught and the run continues with
empty memory context (logged).
- agents.py: make_agent_node injects retrieved memories into each agent's
system prompt.
- run.py: prints `[retrieved: name1, name2, name3]` (or `[retrieved: none]`)
before route/result for both --route-only and full-run modes, so bad
retrieval is visible at a glance.
- .gitignore: add .cache/, .pytest_cache/, .ruff_cache/.
Validated: golden-set still 21/21 passing; smoke tests retrieve plausible
memories ("Send a Slack message to ops about the new exec-aide deploy" ->
project_exec_aide, feedback_exec_aide_vip_management, project_seahaven_slack_bot).
|
||
|---|---|---|
| tests | ||
| .env.example | ||
| .gitignore | ||
| agents.py | ||
| graph.py | ||
| models.py | ||
| README.md | ||
| requirements.txt | ||
| retriever.py | ||
| run.py | ||
| state.py | ||
| tools.py | ||
orchestrator
Multi-model AI agent orchestration via LangGraph + Composio. Routes tasks to the best-fit model and connects to external services (Slack, Notion, GitHub, Google Drive).
Architecture
Claude Code ──► run.py ──► LangGraph StateGraph
│
┌─────┤ router (Sonnet)
│ │
▼ ▼
┌──────────────────────────────┐
│ implementer (Sonnet) │
│ reviewer (Sonnet) │
│ researcher (Haiku) │
│ cross_reviewer (GPT-4.1) │
│ scanner (Gemini 2.5) │
│ fast_coder (DeepSeek) │
│ connector (Composio) │
└──────────────────────────────┘
The router node evaluates each task and routes to one of 7 agent nodes via conditional edges. The connector node uses Composio tools for external service interactions.
Files
| File | Purpose |
|---|---|
run.py |
CLI entry point — python3 run.py "<task>" |
graph.py |
LangGraph graph definition, router, connector, and summarizer nodes |
agents.py |
Agent node functions with system prompts |
models.py |
LLM factory functions for each provider |
state.py |
Graph state schema (OrchestratorState) |
tools.py |
Composio tool loading (Slack, Notion, GitHub, Google Drive) |
Usage
# Full execution — routes and runs the task
python3 run.py "What is the LangGraph checkpoint API?"
# Route-only — prints which agent would handle the task
python3 run.py --route-only "Review this code for security issues"
From Claude Code (via CLAUDE.md hybrid delegation):
# Claude Code delegates automatically when another model is better for the task
python3 ~/Documents/repositories/orchestrator/run.py "<task description>"
# Check routing without executing
python3 ~/Documents/repositories/orchestrator/run.py --route-only "<task description>"
When Claude Code delegates vs. handles natively
Claude Code uses a hybrid model — it delegates to the orchestrator when a different model has a genuine advantage, and handles everything else natively:
| Delegate to orchestrator | Handle natively in Claude Code |
|---|---|
| Cross-family code review (GPT-4.1) | File editing, refactoring, bug fixes |
| Large codebase scanning (Gemini) | Git operations, PRs, merges |
| Quick bounded coding (DeepSeek) | AWS/SAM/CDK deployments |
| External service actions (Composio) | Shell commands, system admin |
| Interactive planning and conversation |
Agents
| Agent | Model | Use Case |
|---|---|---|
| implementer | Claude Sonnet 4.6 | Write code with a clear spec |
| reviewer | Claude Sonnet 4.6 | Code review (BLOCK/FIX/NIT/QUESTION) |
| researcher | Claude Haiku 4.5 | Doc lookups, API research |
| cross_reviewer | GPT-4.1 | Independent second-opinion review |
| scanner | Gemini 2.5 Pro | Large codebase analysis |
| fast_coder | DeepSeek Coder | Quick, bounded coding tasks |
| connector | Sonnet + Composio | Slack, Notion, GitHub, Google Drive |
Connectors (via Composio)
All connections authenticated under Composio user amoussa:
- Slack: send messages, read channels/threads, find users, add reactions
- Notion: search/read/create/update pages, add content
- GitHub: create issues, list issues, get repo info
- Google Drive: find files, get metadata
Setup
- Install dependencies:
pip install -r requirements.txt - Copy
.env.exampleto.envand fill in API keys - Authenticate Composio integrations at app.composio.dev
Configuration
All API keys are stored in .env (gitignored):
ANTHROPIC_API_KEY— Claude modelsOPENAI_API_KEY— GPT-4.1 cross-reviewerGOOGLE_API_KEY— Gemini scannerDEEPSEEK_API_KEY— DeepSeek fast-coderCOMPOSIO_API_KEY— Composio connectorsLANGSMITH_API_KEY— LangSmith tracing
Tracing is enabled via LangSmith (project: orchestration).