This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/HANDOFF-2026-05-15.md
Adam Moussa c23ea679e4 Add retrieval-augmented routing tests, fix lazy loading and open items #2-3
- Add 6 retrieval-augmented routing tests (3 live retrieval, 3 off-topic
  fake memories) to unblock Phase 5
- Defer Composio tool loading and graph construction to first use so
  expired or missing keys don't crash imports
- Atomic cache write in retriever via temp file (open item #2)
- Log rotation in weekly_summary.py, pruning JSONL >90 days (open item #3)
2026-05-26 18:26:59 -04:00

9.9 KiB
Raw Blame History

Sea Haven Orchestrator — Session Handoff

Generated 2026-05-15 at end of modernization session. Companion to the original assessment at ~/Desktop/orchestrator-modernization-assessment.md.

TL;DR

Phases 1, 2, and 4 of the modernization assessment are merged to main (db34d7e is the latest commit). The orchestrator now has a memory-aware structured router, JSONL telemetry, CI, and is bumped to Sonnet 4.6. Three FIX-level review items were caught and fixed in a follow-up PR. One known follow-up to do before Phase 5 (retrieval-augmented routing test) and two low-priority items can ride along on the next time anyone touches the relevant files.


Current state

Repo: amoussa1229/orchestrator (personal, private) at ~/Documents/repositories/orchestrator/ Branch: all work merged into main via PRs #1, #2, #3. No open branches. Architecture: START → retriever → router → agent → END. Six agent nodes (implementer, reviewer, researcher, cross_reviewer, scanner, fast_coder) via an AGENTS registry + factory. Connector node (one-tool-call rule preserved). Unknown route returns an explicit error string. Memory: ~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/ (87 files) embedded with text-embedding-3-small, cached to .cache/embeddings.json. Top-3 injected into router + agent prompts. Read-only. Telemetry: every full run writes one JSONL line to ~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl (UTC). scripts/weekly_summary.py produces a markdown digest. CI: .github/workflows/ci.yaml runs ruff check + ruff format check + pytest --collect-only. Two jobs, both green. Models: Sonnet 4.6 (claude-sonnet-4-6), Haiku 4.5 (claude-haiku-4-5-20251001), GPT-4.1, Gemini 2.5 Pro, DeepSeek Coder. All centralized as constants in models.py.

What shipped this session

PR Commits What
#1 366d724, ac7101f, 67180e1, f05a98d Phases 1, 2, 4 + README update
#2 5378bac Code-review FIX items: Gemini retry coverage, batched embeddings, UTC log filenames, cost-includes-failures
#3 fb5dbb2, 3638d6b, 2bb8365 CI workflow, Sonnet 4.6 bump, golden-set CI collection fix

Memory updates:

  • project_orchestration_migration.md — rewritten for current architecture
  • feedback_executor_mode_workflow.md — new (executor-mode workflow rules)
  • feedback_pre_action_smoke_test.md — new (smoke-test prompt builders)
  • MEMORY.md — indexed both new feedback entries

System updates:

  • ~/.config/git/ignore — .DS_Store added globally

Open items

1. Retrieval-augmented routing test — DO BEFORE PHASE 5

Status: DONE (2026-05-26). Added 6 tests: 3 with live retrieval, 3 with off-topic fake memories. Also fixed Composio lazy loading so tests don't crash on expired keys.

What's missing. The golden-set test (tests/test_routing_golden.py) calls router_node with no retrieved key, so all 21 cases exercise the router on a bare prompt without memory injection. We currently have zero test signal on whether retrieval flips a routing decision.

Why it blocks Phase 5. Phase 5 modifies RouteDecision (adds risk_class, confidence). That changes the router prompt path that retrieval also touches. Without a test that validates memory-augmented routing, we can't tell whether Phase 5 breaks routing only in the with-memory case.

Fix. Add 2–3 cases with populated retrieved:

from retriever import retrieve

@requires_keys
def test_router_unchanged_with_relevant_memory():
    retrieved = retrieve("Send a Slack message to ops", k=3)
    out = router_node({
        "task": "Send a Slack message to ops about the deploy",
        "messages": [],
        "retrieved": retrieved,
    })
    assert out["route"] == "connector"


@requires_keys
def test_router_unchanged_with_offtopic_memory():
    # Hand-built off-topic context — router should still pick the obvious route.
    fake_retrieved = [
        {"name": "feedback_emulator_coordinates", "score": 0.1,
         "content": "ADB input uses landscape; getevent uses portrait"},
    ]
    out = router_node({
        "task": "Review this PR diff for security issues",
        "messages": [],
        "retrieved": fake_retrieved,
    })
    assert out["route"] == "reviewer"

Priority: medium. Not urgent for current usage, but it's the only gate on confidently merging Phase 5.


2. Non-atomic cache write — LOW

Status: DONE (2026-05-26). Atomic write via temp file in retriever._save_cache.

What's wrong. retriever._save_cache calls CACHE_FILE.write_text(...), which truncates the file to zero before writing. A crash, Ctrl-C, or kill during the ~50ms write window leaves .cache/embeddings.json empty or half-written. Next run, _load_cache hits JSONDecodeError, returns {}, and the retriever re-embeds all 87 memories (~10s + ~$0.001).

Impact severity. Latency-only, not correctness or data-loss. Source memory files are untouched; re-embedded vectors are identical.

Probability. Low. Vulnerable window is short and only occurs when memories actually changed (cache-hit-only path doesn't write).

Fix. Atomic write via temp file:

def _save_cache(cache: dict) -> None:
    CACHE_DIR.mkdir(exist_ok=True)
    tmp = CACHE_FILE.with_suffix(".json.tmp")
    tmp.write_text(json.dumps(cache))
    tmp.replace(CACHE_FILE)

Priority: low. Bundle with the next change that touches retriever.py. Not worth a dedicated PR.


3. Telemetry log rotation — LOW

Status: DONE (2026-05-26). prune_old_logs() added to weekly_summary.py; deletes JSONL files >90 days old, runs automatically after each digest.

What's wrong. ~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl files accumulate forever. No deletion, no compression. At current volume (20 runs/day) the directory grows ~2 MB/year — disk is not the issue, but file count climbs and the digest only reads the last 7 days, so older logs are noise.

Fix. One-liner in a scheduled routine:

find ~/.claude/logs/orchestrator -name "*.jsonl" -mtime +90 -delete

90 days well exceeds the 7-day digest window.

Priority: low. Bundle with the weekly digest /schedule — same job can emit the digest and then prune.


Scheduled-but-unscheduled

These were called out as "do it when ready" but never wired:

  • Weekly digest → python3 ~/Documents/repositories/orchestrator/scripts/weekly_summary.py piped to Slack. Suggested cadence: Monday morning. Wire via /schedule.
  • Quarterly log prune (see item 3 above) — bundle with the digest schedule.

Telemetry collection period

The next decision point on the assessment roadmap (Phase 5 vs Phase 3 vs stop) was deliberately deferred until ~1 week of real-workload telemetry is in the JSONL log. As of 2026-05-15 the log has 2 entries (both from smoke testing this session). Real usage starts now.

Re-run python3 scripts/weekly_summary.py after ~7 days of normal use to get the first meaningful digest. Decisions that depend on this:

  • Unknown rate — target <2%. If higher, router prompt or AGENTS descriptions need work.
  • Top routes — confirms model-diversity hypothesis is real (i.e., non-Sonnet routes get used).
  • Mean tokens per route — flags runaway prompts.
  • Failed-run cost — flags transient provider issues that retries aren't catching.

Roadmap status

From ~/Desktop/orchestrator-modernization-assessment.md:

Phase Status Notes
1 — Stabilize foundations ✅ DONE PR #1, commit 366d724
2 — Memory retriever ✅ DONE PR #1, commit ac7101f
3 — Governance compiler HOLD Only if prompt drift pain shows up. AGENTS dict + handbook links currently cover ~80%.
4 — JSONL telemetry ✅ DONE PR #1, commit f05a98d
5 — Risk class + connector audit log GATED on item 1 above + telemetry data Low-risk addition once item 1 lands and digest shows clean baseline.
6 — Memory distillation HOLD Only if Phase 2 retrieval is paying off after a few weeks of real use.
7 — Bounded connector chaining HOLD Only if connector-audit.jsonl (Phase 5) shows a recurring multi-call pattern.

How to pick up next session

If continuing this work in a future session, paste this prompt template:

Context: orchestrator modernization handoff at ~/Documents/repositories/orchestrator/HANDOFF-2026-05-15.md. Read that first. Phases 1, 2, 4 are merged. Open items #1 (retrieval test), #2 (atomic cache), #3 (log rotation) documented there. I want to <pick one: tackle item #1 / schedule the weekly digest / start Phase 5 / something else>.

Memory entries the next session will load automatically (already in MEMORY.md):

  • project_orchestration_migration.md
  • feedback_executor_mode_workflow.md — relevant if executing an approved plan
  • feedback_pre_action_smoke_test.md — relevant for any prompt-builder work
  • feedback_composio_direct_tools.md — relevant if anyone proposes meta-tool routing

References

  • Assessment: ~/Desktop/orchestrator-modernization-assessment.md — original analysis and phase plan (kept outside the repo as the originating artifact)
  • This handoff: ~/Documents/repositories/orchestrator/HANDOFF-2026-05-15.md (you're reading it)
  • Repo: https://github.com/amoussa1229/orchestrator
  • PRs: #1, #2, #3
  • Telemetry logs: ~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl (UTC date)
  • Embedding cache: ~/Documents/repositories/orchestrator/.cache/embeddings.json
  • Memory store: ~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/