Hybrid multi-model task orchestrator — routes coding/review/scan tasks across LLMs via LangGraph
This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
Find a file
Adam Moussa 3d97139300
feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17)
* feat(agent-team): read-only CI-result fetcher for P3 verify gate (opt-in, inert)

ci_fetcher.py: fail-closed CiResultFetcher reading the GitHub Actions run
conclusion via a read-only PAT (AGENT_TEAM_CI_READ_TOKEN→GITHUB_TOKEN), returns
{run_id,conclusion,diff_hash} or None on any error. Data-fetcher only — ci_gate
owns the verdict; never writes, no OIDC/AWS, never reads patch artifacts.
coordinator gains opt-in gated_build_verify_wiring() composing it via
bind_ci_result_fetcher; NOT wired into the default run-team.py path. 20 tests.

* harden(agent-team): P3 apply/verify workflow — GitHub App token, CWE-94, fail-closed

Decision-1 auth model: gate-and-pr uses a GitHub App installation token
(pull-requests:write) behind the agent-apply environment; ALL OIDC/id-token/AWS
removed. Hardening: task_id env-indirection (CWE-94 — GitHub expands ${{ }} into
the run shell before exec, so %s/quoting is insufficient); run-id pinning on both
download-artifact; post-build denied-path check (build-hook writes into denied
paths fail the job); empty-hash fail-closed in BOTH the embedded gate (fixed a
real ''=='' pass bug) and ci_gate.py. App-token + draft-PR steps stay if:${{ false }}
until provisioning (App + environment + branch protection). +17 tests.

* harden(agent-team): apply P3-live security-gate fixes (GPT-4.1 xreview + sh-security-review)

BLOCK-1/FIX-4: gate-and-pr re-comments pull-requests:write + environment:agent-apply
(provisioning-time uncomment) and gains needs.guard/build-test=='success' job guard —
zero privilege until provisioning. BLOCK-2/3+FIX-5: ci_fetcher validates run_id (^[0-9]{1,20}$),
owner/repo (^[A-Za-z0-9_.-]{1,100}$), and fetched_id (int) — fail closed, no SSRF/path
injection. FIX-1: conclusion allowlist. FIX-3: api_root removed from public builder (no
injectable endpoint). INJ-02: post-build denied-path check uses NUL-delimited git output +
explicit rename parsing, no backslash mangling, non-UTF8=violation. INJ-03: all three trust-
control denylists unified to one 22-entry union + drift-guard test. Q1: documented run_id/
diff_hash trust source (dispatcher/ledger only). 884 tests, ruff clean. Privileged steps stay
if:${{ false }} until provisioning.

* build(security-review): prune .claude worktrees from deterministic scanners

Agent worktrees under .claude/worktrees/ are full repo copies; the cfn-lint
find|xargs template scan overflowed ('command line cannot be assembled') and the
pre-push hook fail-closed to BLOCK whenever a worktree was present. Prune .claude
in the cfn-lint find + semgrep/checkov excludes, and gitignore .claude/ so it is
never scanned or committed. Unblocks main-tree pushes during parallel agent work.
2026-06-18 15:53:26 -04:00
.github ci: add workflow permissions from GHAS notes 2026-06-17 17:56:32 -04:00
agent-team feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17) 2026-06-18 15:53:26 -04:00
docs Plane 2 foundation: interfaces, SQLite schemas, state-store, billing seam 2026-06-17 15:16:12 -04:00
scripts Add retrieval-augmented routing tests, fix lazy loading and open items #2-3 2026-05-26 18:26:59 -04:00
security-review feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17) 2026-06-18 15:53:26 -04:00
tests Add retrieval-augmented routing tests, fix lazy loading and open items #2-3 2026-05-26 18:26:59 -04:00
.env.example Add README and fix .env.example project name 2026-05-08 13:12:02 -04:00
.gitignore feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17) 2026-06-18 15:53:26 -04:00
agents.py Add memory retriever node (Phase 2) 2026-05-15 11:36:03 -04:00
conftest.py Fix CI collection: isolate agent-team tests; add agent-team CI job 2026-06-17 15:19:51 -04:00
graph.py Add retrieval-augmented routing tests, fix lazy loading and open items #2-3 2026-05-26 18:26:59 -04:00
models.py Bump Sonnet model ID to claude-sonnet-4-6 2026-05-15 13:14:18 -04:00
README.md Document the security-review subsystem in the root README 2026-06-16 15:55:01 -04:00
requirements.txt build(deps): bump the minor-and-patch group across 1 directory with 6 updates (#11) 2026-06-18 14:44:46 -04:00
retriever.py Add retrieval-augmented routing tests, fix lazy loading and open items #2-3 2026-05-26 18:26:59 -04:00
run.py Add retrieval-augmented routing tests, fix lazy loading and open items #2-3 2026-05-26 18:26:59 -04:00
state.py Add memory retriever node (Phase 2) 2026-05-15 11:36:03 -04:00
telemetry.py Address code review FIX items from PR #1 2026-05-15 13:01:14 -04:00
tools.py Add CLI entry point and fix Composio user_id 2026-05-08 13:10:39 -04:00

orchestrator

Multi-model AI agent orchestration via LangGraph + Composio. Routes tasks to the best-fit model and connects to external services (Slack, Notion, GitHub, Google Drive). Memory-aware — each run is enriched with the top-3 most relevant notes from Adam's project/feedback/reference memory store.

Architecture

Claude Code ──► run.py ──► LangGraph StateGraph
                              │
                              ▼
                          retriever  ──► top-3 memories from
                              │           ~/.claude/projects/.../memory/
                              ▼
                          router (Sonnet, structured output)
                              │
                ┌─────────────┼─────────────────────┐
                ▼             ▼                     ▼
        ┌──────────────┐  ┌───────────┐      ┌──────────┐
        │ implementer  │  │ connector │      │ unknown  │
        │ reviewer     │  │ (Composio)│      │ (no fit) │
        │ researcher   │  └───────────┘      └──────────┘
        │ cross_reviewer│       │
        │ scanner       │       ▼
        │ fast_coder    │  tool_executor ──► summarizer
        └──────────────┘

The retriever embeds Adam's memory files once and caches vectors to .cache/embeddings.json (mtime-keyed; only changed files re-embed). Each run picks the top-3 most relevant memories and surfaces them in the CLI output before the route line.

The router uses Pydantic structured output (RouteDecision) and returns an explicit "unknown" route when no agent fits — no silent fallback. All LLM invocations are wrapped with retry-on-transient-error.

Files

File Purpose
run.py CLI entry point — python3 run.py "<task>"
graph.py LangGraph graph: retriever, router, connector, summarizer, unknown nodes
agents.py AGENTS registry (label → model_fn, prompt, description) + make_agent_node factory
models.py LLM factories, model-ID constants, with_retries() helper
state.py OrchestratorState TypedDict
retriever.py Memory loader, embedder, cache, top-k retrieval
tools.py Composio tool loading (Slack, Notion, GitHub, Google Drive)
tests/test_routing_golden.py 20-case golden-set regression test for the router

Usage

# Full execution — retrieves memory, routes, and runs the task
python3 run.py "What is the LangGraph checkpoint API?"

# Route-only — retrieves memory and prints the agent that would handle the task
python3 run.py --route-only "Review this code for security issues"

Output shape:

[retrieved: project_seahaven_slack_bot, feedback_secrets_manager, reference_sea_haven_aws]
[reviewer]

<agent output>

From Claude Code (via CLAUDE.md hybrid delegation):

python3 ~/Documents/repositories/orchestrator/run.py "<task description>"
python3 ~/Documents/repositories/orchestrator/run.py --route-only "<task description>"

When Claude Code delegates vs. handles natively

Claude Code uses a hybrid model — it delegates to the orchestrator when a different model has a genuine advantage, and handles everything else natively:

Delegate to orchestrator Handle natively in Claude Code
Cross-family code review (GPT-4.1) File editing, refactoring, bug fixes
Large codebase scanning (Gemini) Git operations, PRs, merges
Quick bounded coding (DeepSeek) AWS/SAM/CDK deployments
External service actions (Composio) Shell commands, system admin
Interactive planning and conversation

Agents

Agent Model Use Case
implementer Claude Sonnet Write code with a clear spec
reviewer Claude Sonnet Code review (BLOCK/FIX/NIT/QUESTION)
researcher Claude Haiku Doc lookups, API research
cross_reviewer GPT-4.1 Independent second-opinion review
scanner Gemini 2.5 Pro Large codebase analysis
fast_coder DeepSeek Coder Quick, bounded coding tasks
connector Sonnet + Composio Slack, Notion, GitHub, Google Drive

The router can also return done (no agent needed) or unknown (no clear fit). Model IDs are centralized as constants in models.py.

Memory retrieval

The retriever reads ~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/*.md (skipping the MEMORY.md index), embeds each file once with text-embedding-3-small, and caches the vectors to .cache/embeddings.json. On subsequent runs:

  • Only files whose mtime changed are re-embedded.
  • Top-3 memories by cosine similarity are injected as system context into both the router and the agent.
  • Retrieved names are printed as the first line of every run so bad retrieval is visible.
  • Retrieval is read-only. The orchestrator never writes back to the memory store.

If retrieval fails (network, missing key), the run continues with no memory context and logs the failure into the message trail.

Connectors (via Composio)

All connections authenticated under Composio user amoussa:

  • Slack: send messages, read channels/threads, find users, add reactions
  • Notion: search/read/create/update pages, add content
  • GitHub: create issues, list issues, get repo info
  • Google Drive: find files, get metadata

The connector node is restricted to one tool call per run — a load-bearing rule learned from a 1.9M-token incident with meta-tool routing.

Security Review

The security-review/ subsystem is a high-recall, anti-complacency security gate. It is separate from the router — it does not route through run.py or LangGraph. One pure-code script, review.sh, owns the block decision (confirmed critical/high → block); no agent decides.

  • Path A — interactive: the /sh-security-review Claude Code skill (Max-covered). Narrow fresh-context detector fan-out + a proof-or-kill verifier; emits the structured finding schema for review.sh to gate.
  • Path B — unattended: a nightly two-tier sweep on the sh-secrev R720 VM. Tier 1 runs deterministic scanners (review.sh --scanners-only) over every Sea-Haven-Industries org repo; Tier 2 is a budget-bounded agentic pass (run_headless.py) on a round-robin rotation. Clean-clone auto-discovery via a read-only GitHub PAT; ALARM-only Slack (a clean night posts nothing).
  • Git hooks: global pre-commit / pre-push hooks (install-hooks.sh --global) gate every local repo via review.sh --scanners-only.

See security-review/README.md for full detail and security-review/DEPLOY-R720.md for the VM runbook.

Setup

  1. Install dependencies: pip install -r requirements.txt
  2. Copy .env.example to .env and fill in API keys
  3. Authenticate Composio integrations at app.composio.dev

Configuration

All API keys are stored in .env (gitignored):

  • ANTHROPIC_API_KEY — Claude models + router
  • OPENAI_API_KEY — GPT-4.1 cross-reviewer + text-embedding-3-small
  • GOOGLE_API_KEY — Gemini scanner
  • DEEPSEEK_API_KEY — DeepSeek fast-coder
  • COMPOSIO_API_KEY — Composio connectors
  • LANGSMITH_API_KEY — LangSmith tracing

Tracing is enabled via LangSmith (project: orchestration).

Testing

pytest tests/test_routing_golden.py -v

20 labelled tasks → expected agent. Skipped cleanly if ANTHROPIC_API_KEY or COMPOSIO_API_KEY are unset.