Update README for Phases 1-2 architecture
Reflect retriever node, structured router with explicit unknown route, AGENTS registry, model-ID constants, memory-aware CLI output, and the golden-set test.
This commit is contained in:
parent
ac7101f3df
commit
67180e124e
1 changed files with 65 additions and 29 deletions
94
README.md
94
README.md
|
|
@ -1,55 +1,68 @@
|
|||
# orchestrator
|
||||
|
||||
Multi-model AI agent orchestration via LangGraph + Composio. Routes tasks to the best-fit model and connects to external services (Slack, Notion, GitHub, Google Drive).
|
||||
Multi-model AI agent orchestration via LangGraph + Composio. Routes tasks to the best-fit model and connects to external services (Slack, Notion, GitHub, Google Drive). Memory-aware — each run is enriched with the top-3 most relevant notes from Adam's project/feedback/reference memory store.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
Claude Code ──► run.py ──► LangGraph StateGraph
|
||||
│
|
||||
┌─────┤ router (Sonnet)
|
||||
│ │
|
||||
▼ ▼
|
||||
┌──────────────────────────────┐
|
||||
│ implementer (Sonnet) │
|
||||
│ reviewer (Sonnet) │
|
||||
│ researcher (Haiku) │
|
||||
│ cross_reviewer (GPT-4.1) │
|
||||
│ scanner (Gemini 2.5) │
|
||||
│ fast_coder (DeepSeek) │
|
||||
│ connector (Composio) │
|
||||
└──────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
retriever ──► top-3 memories from
|
||||
│ ~/.claude/projects/.../memory/
|
||||
▼
|
||||
router (Sonnet, structured output)
|
||||
│
|
||||
┌─────────────┼─────────────────────┐
|
||||
▼ ▼ ▼
|
||||
┌──────────────┐ ┌───────────┐ ┌──────────┐
|
||||
│ implementer │ │ connector │ │ unknown │
|
||||
│ reviewer │ │ (Composio)│ │ (no fit) │
|
||||
│ researcher │ └───────────┘ └──────────┘
|
||||
│ cross_reviewer│ │
|
||||
│ scanner │ ▼
|
||||
│ fast_coder │ tool_executor ──► summarizer
|
||||
└──────────────┘
|
||||
```
|
||||
|
||||
The router node evaluates each task and routes to one of 7 agent nodes via conditional edges. The connector node uses Composio tools for external service interactions.
|
||||
The retriever embeds Adam's memory files once and caches vectors to `.cache/embeddings.json` (mtime-keyed; only changed files re-embed). Each run picks the top-3 most relevant memories and surfaces them in the CLI output before the route line.
|
||||
|
||||
The router uses Pydantic structured output (`RouteDecision`) and returns an explicit `"unknown"` route when no agent fits — no silent fallback. All LLM invocations are wrapped with retry-on-transient-error.
|
||||
|
||||
## Files
|
||||
|
||||
| File | Purpose |
|
||||
|---|---|
|
||||
| `run.py` | CLI entry point — `python3 run.py "<task>"` |
|
||||
| `graph.py` | LangGraph graph definition, router, connector, and summarizer nodes |
|
||||
| `agents.py` | Agent node functions with system prompts |
|
||||
| `models.py` | LLM factory functions for each provider |
|
||||
| `state.py` | Graph state schema (`OrchestratorState`) |
|
||||
| `graph.py` | LangGraph graph: retriever, router, connector, summarizer, unknown nodes |
|
||||
| `agents.py` | `AGENTS` registry (label → model_fn, prompt, description) + `make_agent_node` factory |
|
||||
| `models.py` | LLM factories, model-ID constants, `with_retries()` helper |
|
||||
| `state.py` | `OrchestratorState` TypedDict |
|
||||
| `retriever.py` | Memory loader, embedder, cache, top-k retrieval |
|
||||
| `tools.py` | Composio tool loading (Slack, Notion, GitHub, Google Drive) |
|
||||
| `tests/test_routing_golden.py` | 20-case golden-set regression test for the router |
|
||||
|
||||
## Usage
|
||||
|
||||
```bash
|
||||
# Full execution — routes and runs the task
|
||||
# Full execution — retrieves memory, routes, and runs the task
|
||||
python3 run.py "What is the LangGraph checkpoint API?"
|
||||
|
||||
# Route-only — prints which agent would handle the task
|
||||
# Route-only — retrieves memory and prints the agent that would handle the task
|
||||
python3 run.py --route-only "Review this code for security issues"
|
||||
```
|
||||
|
||||
Output shape:
|
||||
```
|
||||
[retrieved: project_seahaven_slack_bot, feedback_secrets_manager, reference_sea_haven_aws]
|
||||
[reviewer]
|
||||
|
||||
<agent output>
|
||||
```
|
||||
|
||||
From Claude Code (via CLAUDE.md hybrid delegation):
|
||||
```bash
|
||||
# Claude Code delegates automatically when another model is better for the task
|
||||
python3 ~/Documents/repositories/orchestrator/run.py "<task description>"
|
||||
|
||||
# Check routing without executing
|
||||
python3 ~/Documents/repositories/orchestrator/run.py --route-only "<task description>"
|
||||
```
|
||||
|
||||
|
|
@ -69,14 +82,27 @@ Claude Code uses a hybrid model — it delegates to the orchestrator when a diff
|
|||
|
||||
| Agent | Model | Use Case |
|
||||
|---|---|---|
|
||||
| implementer | Claude Sonnet 4.6 | Write code with a clear spec |
|
||||
| reviewer | Claude Sonnet 4.6 | Code review (BLOCK/FIX/NIT/QUESTION) |
|
||||
| researcher | Claude Haiku 4.5 | Doc lookups, API research |
|
||||
| implementer | Claude Sonnet | Write code with a clear spec |
|
||||
| reviewer | Claude Sonnet | Code review (BLOCK/FIX/NIT/QUESTION) |
|
||||
| researcher | Claude Haiku | Doc lookups, API research |
|
||||
| cross_reviewer | GPT-4.1 | Independent second-opinion review |
|
||||
| scanner | Gemini 2.5 Pro | Large codebase analysis |
|
||||
| fast_coder | DeepSeek Coder | Quick, bounded coding tasks |
|
||||
| connector | Sonnet + Composio | Slack, Notion, GitHub, Google Drive |
|
||||
|
||||
The router can also return `done` (no agent needed) or `unknown` (no clear fit). Model IDs are centralized as constants in `models.py`.
|
||||
|
||||
## Memory retrieval
|
||||
|
||||
The retriever reads `~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/*.md` (skipping the `MEMORY.md` index), embeds each file once with `text-embedding-3-small`, and caches the vectors to `.cache/embeddings.json`. On subsequent runs:
|
||||
|
||||
- Only files whose mtime changed are re-embedded.
|
||||
- Top-3 memories by cosine similarity are injected as system context into both the router and the agent.
|
||||
- Retrieved names are printed as the first line of every run so bad retrieval is visible.
|
||||
- Retrieval is **read-only.** The orchestrator never writes back to the memory store.
|
||||
|
||||
If retrieval fails (network, missing key), the run continues with no memory context and logs the failure into the message trail.
|
||||
|
||||
## Connectors (via Composio)
|
||||
|
||||
All connections authenticated under Composio user `amoussa`:
|
||||
|
|
@ -85,6 +111,8 @@ All connections authenticated under Composio user `amoussa`:
|
|||
- **GitHub**: create issues, list issues, get repo info
|
||||
- **Google Drive**: find files, get metadata
|
||||
|
||||
The connector node is restricted to **one tool call per run** — a load-bearing rule learned from a 1.9M-token incident with meta-tool routing.
|
||||
|
||||
## Setup
|
||||
|
||||
1. Install dependencies: `pip install -r requirements.txt`
|
||||
|
|
@ -94,11 +122,19 @@ All connections authenticated under Composio user `amoussa`:
|
|||
## Configuration
|
||||
|
||||
All API keys are stored in `.env` (gitignored):
|
||||
- `ANTHROPIC_API_KEY` — Claude models
|
||||
- `OPENAI_API_KEY` — GPT-4.1 cross-reviewer
|
||||
- `ANTHROPIC_API_KEY` — Claude models + router
|
||||
- `OPENAI_API_KEY` — GPT-4.1 cross-reviewer + text-embedding-3-small
|
||||
- `GOOGLE_API_KEY` — Gemini scanner
|
||||
- `DEEPSEEK_API_KEY` — DeepSeek fast-coder
|
||||
- `COMPOSIO_API_KEY` — Composio connectors
|
||||
- `LANGSMITH_API_KEY` — LangSmith tracing
|
||||
|
||||
Tracing is enabled via LangSmith (project: `orchestration`).
|
||||
|
||||
## Testing
|
||||
|
||||
```bash
|
||||
pytest tests/test_routing_golden.py -v
|
||||
```
|
||||
|
||||
20 labelled tasks → expected agent. Skipped cleanly if `ANTHROPIC_API_KEY` or `COMPOSIO_API_KEY` are unset.
|
||||
|
|
|
|||
Reference in a new issue