- Add 6 retrieval-augmented routing tests (3 live retrieval, 3 off-topic fake memories) to unblock Phase 5 - Defer Composio tool loading and graph construction to first use so expired or missing keys don't crash imports - Atomic cache write in retriever via temp file (open item #2) - Log rotation in weekly_summary.py, pruning JSONL >90 days (open item #3)
9.9 KiB
Sea Haven Orchestrator — Session Handoff
Generated 2026-05-15 at end of modernization session. Companion to the original assessment at ~/Desktop/orchestrator-modernization-assessment.md.
TL;DR
Phases 1, 2, and 4 of the modernization assessment are merged to main (db34d7e is the latest commit). The orchestrator now has a memory-aware structured router, JSONL telemetry, CI, and is bumped to Sonnet 4.6. Three FIX-level review items were caught and fixed in a follow-up PR. One known follow-up to do before Phase 5 (retrieval-augmented routing test) and two low-priority items can ride along on the next time anyone touches the relevant files.
Current state
Repo: amoussa1229/orchestrator (personal, private) at ~/Documents/repositories/orchestrator/
Branch: all work merged into main via PRs #1, #2, #3. No open branches.
Architecture: START → retriever → router → agent → END. Six agent nodes (implementer, reviewer, researcher, cross_reviewer, scanner, fast_coder) via an AGENTS registry + factory. Connector node (one-tool-call rule preserved). Unknown route returns an explicit error string.
Memory: ~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/ (87 files) embedded with text-embedding-3-small, cached to .cache/embeddings.json. Top-3 injected into router + agent prompts. Read-only.
Telemetry: every full run writes one JSONL line to ~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl (UTC). scripts/weekly_summary.py produces a markdown digest.
CI: .github/workflows/ci.yaml runs ruff check + ruff format check + pytest --collect-only. Two jobs, both green.
Models: Sonnet 4.6 (claude-sonnet-4-6), Haiku 4.5 (claude-haiku-4-5-20251001), GPT-4.1, Gemini 2.5 Pro, DeepSeek Coder. All centralized as constants in models.py.
What shipped this session
| PR | Commits | What |
|---|---|---|
| #1 | 366d724, ac7101f, 67180e1, f05a98d |
Phases 1, 2, 4 + README update |
| #2 | 5378bac |
Code-review FIX items: Gemini retry coverage, batched embeddings, UTC log filenames, cost-includes-failures |
| #3 | fb5dbb2, 3638d6b, 2bb8365 |
CI workflow, Sonnet 4.6 bump, golden-set CI collection fix |
Memory updates:
project_orchestration_migration.md— rewritten for current architecturefeedback_executor_mode_workflow.md— new (executor-mode workflow rules)feedback_pre_action_smoke_test.md— new (smoke-test prompt builders)MEMORY.md— indexed both new feedback entries
System updates:
~/.config/git/ignore—.DS_Storeadded globally
Open items
1. Retrieval-augmented routing test — DO BEFORE PHASE 5
Status: DONE (2026-05-26). Added 6 tests: 3 with live retrieval, 3 with off-topic fake memories. Also fixed Composio lazy loading so tests don't crash on expired keys.
What's missing. The golden-set test (tests/test_routing_golden.py) calls router_node with no retrieved key, so all 21 cases exercise the router on a bare prompt without memory injection. We currently have zero test signal on whether retrieval flips a routing decision.
Why it blocks Phase 5. Phase 5 modifies RouteDecision (adds risk_class, confidence). That changes the router prompt path that retrieval also touches. Without a test that validates memory-augmented routing, we can't tell whether Phase 5 breaks routing only in the with-memory case.
Fix. Add 2–3 cases with populated retrieved:
from retriever import retrieve
@requires_keys
def test_router_unchanged_with_relevant_memory():
retrieved = retrieve("Send a Slack message to ops", k=3)
out = router_node({
"task": "Send a Slack message to ops about the deploy",
"messages": [],
"retrieved": retrieved,
})
assert out["route"] == "connector"
@requires_keys
def test_router_unchanged_with_offtopic_memory():
# Hand-built off-topic context — router should still pick the obvious route.
fake_retrieved = [
{"name": "feedback_emulator_coordinates", "score": 0.1,
"content": "ADB input uses landscape; getevent uses portrait"},
]
out = router_node({
"task": "Review this PR diff for security issues",
"messages": [],
"retrieved": fake_retrieved,
})
assert out["route"] == "reviewer"
Priority: medium. Not urgent for current usage, but it's the only gate on confidently merging Phase 5.
2. Non-atomic cache write — LOW
Status: DONE (2026-05-26). Atomic write via temp file in retriever._save_cache.
What's wrong. retriever._save_cache calls CACHE_FILE.write_text(...), which truncates the file to zero before writing. A crash, Ctrl-C, or kill during the ~50ms write window leaves .cache/embeddings.json empty or half-written. Next run, _load_cache hits JSONDecodeError, returns {}, and the retriever re-embeds all 87 memories (~10s + ~$0.001).
Impact severity. Latency-only, not correctness or data-loss. Source memory files are untouched; re-embedded vectors are identical.
Probability. Low. Vulnerable window is short and only occurs when memories actually changed (cache-hit-only path doesn't write).
Fix. Atomic write via temp file:
def _save_cache(cache: dict) -> None:
CACHE_DIR.mkdir(exist_ok=True)
tmp = CACHE_FILE.with_suffix(".json.tmp")
tmp.write_text(json.dumps(cache))
tmp.replace(CACHE_FILE)
Priority: low. Bundle with the next change that touches retriever.py. Not worth a dedicated PR.
3. Telemetry log rotation — LOW
Status: DONE (2026-05-26). prune_old_logs() added to weekly_summary.py; deletes JSONL files >90 days old, runs automatically after each digest.
What's wrong. ~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl files accumulate forever. No deletion, no compression. At current volume (20 runs/day) the directory grows ~2 MB/year — disk is not the issue, but file count climbs and the digest only reads the last 7 days, so older logs are noise.
Fix. One-liner in a scheduled routine:
find ~/.claude/logs/orchestrator -name "*.jsonl" -mtime +90 -delete
90 days well exceeds the 7-day digest window.
Priority: low. Bundle with the weekly digest /schedule — same job can emit the digest and then prune.
Scheduled-but-unscheduled
These were called out as "do it when ready" but never wired:
- Weekly digest →
python3 ~/Documents/repositories/orchestrator/scripts/weekly_summary.pypiped to Slack. Suggested cadence: Monday morning. Wire via/schedule. - Quarterly log prune (see item 3 above) — bundle with the digest schedule.
Telemetry collection period
The next decision point on the assessment roadmap (Phase 5 vs Phase 3 vs stop) was deliberately deferred until ~1 week of real-workload telemetry is in the JSONL log. As of 2026-05-15 the log has 2 entries (both from smoke testing this session). Real usage starts now.
Re-run python3 scripts/weekly_summary.py after ~7 days of normal use to get the first meaningful digest. Decisions that depend on this:
- Unknown rate — target <2%. If higher, router prompt or AGENTS descriptions need work.
- Top routes — confirms model-diversity hypothesis is real (i.e., non-Sonnet routes get used).
- Mean tokens per route — flags runaway prompts.
- Failed-run cost — flags transient provider issues that retries aren't catching.
Roadmap status
From ~/Desktop/orchestrator-modernization-assessment.md:
| Phase | Status | Notes |
|---|---|---|
| 1 — Stabilize foundations | ✅ DONE | PR #1, commit 366d724 |
| 2 — Memory retriever | ✅ DONE | PR #1, commit ac7101f |
| 3 — Governance compiler | HOLD | Only if prompt drift pain shows up. AGENTS dict + handbook links currently cover ~80%. |
| 4 — JSONL telemetry | ✅ DONE | PR #1, commit f05a98d |
| 5 — Risk class + connector audit log | GATED on item 1 above + telemetry data | Low-risk addition once item 1 lands and digest shows clean baseline. |
| 6 — Memory distillation | HOLD | Only if Phase 2 retrieval is paying off after a few weeks of real use. |
| 7 — Bounded connector chaining | HOLD | Only if connector-audit.jsonl (Phase 5) shows a recurring multi-call pattern. |
How to pick up next session
If continuing this work in a future session, paste this prompt template:
Context: orchestrator modernization handoff at
~/Documents/repositories/orchestrator/HANDOFF-2026-05-15.md. Read that first. Phases 1, 2, 4 are merged. Open items #1 (retrieval test), #2 (atomic cache), #3 (log rotation) documented there. I want to <pick one: tackle item #1 / schedule the weekly digest / start Phase 5 / something else>.
Memory entries the next session will load automatically (already in MEMORY.md):
project_orchestration_migration.mdfeedback_executor_mode_workflow.md— relevant if executing an approved planfeedback_pre_action_smoke_test.md— relevant for any prompt-builder workfeedback_composio_direct_tools.md— relevant if anyone proposes meta-tool routing
References
- Assessment:
~/Desktop/orchestrator-modernization-assessment.md— original analysis and phase plan (kept outside the repo as the originating artifact) - This handoff:
~/Documents/repositories/orchestrator/HANDOFF-2026-05-15.md(you're reading it) - Repo: https://github.com/amoussa1229/orchestrator
- PRs: #1, #2, #3
- Telemetry logs:
~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl(UTC date) - Embedding cache:
~/Documents/repositories/orchestrator/.cache/embeddings.json - Memory store:
~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/