Orchestrator modernization — Phases 1, 2, 4 #1

Merged
amoussa1229 merged 4 commits from phase-1-stabilize-foundations into main 2026-05-15 16:14:51 +00:00
amoussa1229 commented 2026-05-15 15:54:43 +00:00 (Migrated from github.com)

Summary

Executes Phases 1, 2, and 4 of the orchestrator modernization assessment (~/Desktop/orchestrator-modernization-assessment.md). Phase 3 deferred per assessment ("only build if drift pain shows up"); Phase 5+ deferred pending a week of telemetry data.

  • Phase 1 — stabilize router (366d724): centralized model IDs as models.py constants, collapsed six near-identical agent nodes into a single AGENTS registry + make_agent_node factory, replaced the silent route = "researcher" fallback with Pydantic structured output (RouteDecision) and an explicit unknown route, wrapped every LLM invoke in with_retries(), and added a 20-case pytest golden-set for routing regressions. Killed the four-copy prompt drift surface (graph.py / run.py --route-only / agent prompts / handbook descriptions) by generating the router prompt from AGENTS.
  • Phase 2 — memory retriever (ac7101f): new retriever.py embeds ~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/*.md (skipping MEMORY.md index) with text-embedding-3-small, caches vectors to .cache/embeddings.json keyed on file mtime so only changed memories re-embed. Wired retriever_node as START → retriever → router; top-3 memories injected into the router + downstream agent system prompts. [retrieved: name1, name2, name3] printed before route/result for visibility. Retrieval is read-only and failure-tolerant — if it errors the run continues with empty memory context.
  • Phase 4 — JSONL telemetry + weekly digest (f05a98d): every full run appends one JSON line to ~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl (timestamp, sha256-prefix task hash, retrieved names, route, runtime, tokens in/out, success/error; risk_class and confidence reserved as nulls for Phase 5). Telemetry failures are swallowed — never kills a run. scripts/weekly_summary.py prints a markdown digest of the last 7 days (route distribution, unknown rate, cross-review rate, success rate, total spend, mean tokens per route); intended for /schedule to pipe into Slack.
  • Docs (67180e1): README updated for the new architecture (retriever node, structured router, AGENTS registry, golden-set test).

Architectural intent (preserved)

  • One-shot router → agent → END. No planner above the router (Claude Code is the planner). No multi-step agent chains.
  • Connector still restricted to one tool call per run.
  • Memory is read-only — distillation deliberately deferred to Phase 6.
  • LangSmith remains primary per-call observability; JSONL covers aggregate trends only.

Test plan

  • ruff check . clean
  • ruff format --check . clean (12 files)
  • pytest tests/test_routing_golden.py -v — 21/21 pass, run twice (Phase 2 and Phase 4 validation)
  • Full-run smoke (route=researcher) writes a complete JSONL record with 928 tokens in / 109 tokens out captured from usage_metadata
  • Full-run smoke (route=done) writes a record with tokens_in/out=0 (router-only, expected — structured output doesn't surface usage)
  • python3 scripts/weekly_summary.py produces a clean markdown digest from the live log
  • --route-only smoke for a memory-relevant task surfaces plausible memories (project_exec_aide, feedback_exec_aide_vip_management, project_seahaven_slack_bot for a Slack deploy task)
  • Cache hit on second run — .cache/embeddings.json (2.6 MB, 87 entries) reused without re-embedding

Known limitations / deferred

  • Router tokens for with_structured_output calls aren't captured (Pydantic return path doesn't surface usage_metadata through the message trail). Documented in commit message; revisit in Phase 5.
  • Cost rates in scripts/weekly_summary.py are rough — intended for spotting runaway prompts, not finance.
  • claude-sonnet-4-20250514 EOL 2026-06-15. Centralized at models.CLAUDE_SONNET — one-line bump when ready (out of scope for this PR).

Rollback

Each phase is independently reversible:

  • Phase 4: revert f05a98d (deletes telemetry, restores plain app.invoke).
  • Phase 2: revert ac7101f (deletes retriever, restores START → router).
  • Phase 1: revert 366d724 (restores free-text router + six agent functions).

Next

Recommended: run on real workload for ~1 week, then revisit Phase 5 (risk class) using telemetry to decide.

## Summary Executes Phases 1, 2, and 4 of the orchestrator modernization assessment (`~/Desktop/orchestrator-modernization-assessment.md`). Phase 3 deferred per assessment ("only build if drift pain shows up"); Phase 5+ deferred pending a week of telemetry data. - **Phase 1 — stabilize router** (`366d724`): centralized model IDs as `models.py` constants, collapsed six near-identical agent nodes into a single `AGENTS` registry + `make_agent_node` factory, replaced the silent `route = "researcher"` fallback with Pydantic structured output (`RouteDecision`) and an explicit `unknown` route, wrapped every LLM invoke in `with_retries()`, and added a 20-case pytest golden-set for routing regressions. Killed the four-copy prompt drift surface (`graph.py` / `run.py --route-only` / agent prompts / handbook descriptions) by generating the router prompt from `AGENTS`. - **Phase 2 — memory retriever** (`ac7101f`): new `retriever.py` embeds `~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/*.md` (skipping `MEMORY.md` index) with `text-embedding-3-small`, caches vectors to `.cache/embeddings.json` keyed on file mtime so only changed memories re-embed. Wired `retriever_node` as `START → retriever → router`; top-3 memories injected into the router + downstream agent system prompts. `[retrieved: name1, name2, name3]` printed before route/result for visibility. Retrieval is read-only and failure-tolerant — if it errors the run continues with empty memory context. - **Phase 4 — JSONL telemetry + weekly digest** (`f05a98d`): every full run appends one JSON line to `~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl` (timestamp, sha256-prefix task hash, retrieved names, route, runtime, tokens in/out, success/error; `risk_class` and `confidence` reserved as nulls for Phase 5). Telemetry failures are swallowed — never kills a run. `scripts/weekly_summary.py` prints a markdown digest of the last 7 days (route distribution, unknown rate, cross-review rate, success rate, total spend, mean tokens per route); intended for `/schedule` to pipe into Slack. - **Docs** (`67180e1`): README updated for the new architecture (retriever node, structured router, AGENTS registry, golden-set test). ## Architectural intent (preserved) - One-shot router → agent → END. No planner above the router (Claude Code is the planner). No multi-step agent chains. - Connector still restricted to one tool call per run. - Memory is read-only — distillation deliberately deferred to Phase 6. - LangSmith remains primary per-call observability; JSONL covers aggregate trends only. ## Test plan - [x] `ruff check .` clean - [x] `ruff format --check .` clean (12 files) - [x] `pytest tests/test_routing_golden.py -v` — **21/21 pass**, run twice (Phase 2 and Phase 4 validation) - [x] Full-run smoke (route=`researcher`) writes a complete JSONL record with 928 tokens in / 109 tokens out captured from `usage_metadata` - [x] Full-run smoke (route=`done`) writes a record with `tokens_in/out=0` (router-only, expected — structured output doesn't surface usage) - [x] `python3 scripts/weekly_summary.py` produces a clean markdown digest from the live log - [x] `--route-only` smoke for a memory-relevant task surfaces plausible memories (`project_exec_aide`, `feedback_exec_aide_vip_management`, `project_seahaven_slack_bot` for a Slack deploy task) - [x] Cache hit on second run — `.cache/embeddings.json` (2.6 MB, 87 entries) reused without re-embedding ## Known limitations / deferred - Router tokens for `with_structured_output` calls aren't captured (Pydantic return path doesn't surface `usage_metadata` through the message trail). Documented in commit message; revisit in Phase 5. - Cost rates in `scripts/weekly_summary.py` are rough — intended for spotting runaway prompts, not finance. - `claude-sonnet-4-20250514` EOL **2026-06-15**. Centralized at `models.CLAUDE_SONNET` — one-line bump when ready (out of scope for this PR). ## Rollback Each phase is independently reversible: - Phase 4: revert `f05a98d` (deletes telemetry, restores plain `app.invoke`). - Phase 2: revert `ac7101f` (deletes retriever, restores `START → router`). - Phase 1: revert `366d724` (restores free-text router + six agent functions). ## Next Recommended: run on real workload for ~1 week, then revisit Phase 5 (risk class) using telemetry to decide.
This repo is archived. You cannot comment on pull requests.
No description provided.