187 lines
9.9 KiB
Markdown
187 lines
9.9 KiB
Markdown
# Sea Haven Orchestrator — Session Handoff
|
||
|
||
_Generated 2026-05-15 at end of modernization session. Companion to the original assessment at `~/Desktop/orchestrator-modernization-assessment.md`._
|
||
|
||
## TL;DR
|
||
|
||
Phases 1, 2, and 4 of the modernization assessment are **merged to `main`** (`db34d7e` is the latest commit). The orchestrator now has a memory-aware structured router, JSONL telemetry, CI, and is bumped to Sonnet 4.6. Three FIX-level review items were caught and fixed in a follow-up PR. One known follow-up to do **before Phase 5** (retrieval-augmented routing test) and two low-priority items can ride along on the next time anyone touches the relevant files.
|
||
|
||
---
|
||
|
||
## Current state
|
||
|
||
**Repo:** `amoussa1229/orchestrator` (personal, private) at `~/Documents/repositories/orchestrator/`
|
||
**Branch:** all work merged into `main` via PRs #1, #2, #3. No open branches.
|
||
**Architecture:** `START → retriever → router → agent → END`. Six agent nodes (implementer, reviewer, researcher, cross_reviewer, scanner, fast_coder) via an `AGENTS` registry + factory. Connector node (one-tool-call rule preserved). Unknown route returns an explicit error string.
|
||
**Memory:** `~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/` (87 files) embedded with `text-embedding-3-small`, cached to `.cache/embeddings.json`. Top-3 injected into router + agent prompts. Read-only.
|
||
**Telemetry:** every full run writes one JSONL line to `~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl` (UTC). `scripts/weekly_summary.py` produces a markdown digest.
|
||
**CI:** `.github/workflows/ci.yaml` runs ruff check + ruff format check + `pytest --collect-only`. Two jobs, both green.
|
||
**Models:** Sonnet 4.6 (`claude-sonnet-4-6`), Haiku 4.5 (`claude-haiku-4-5-20251001`), GPT-4.1, Gemini 2.5 Pro, DeepSeek Coder. All centralized as constants in `models.py`.
|
||
|
||
## What shipped this session
|
||
|
||
| PR | Commits | What |
|
||
|---|---|---|
|
||
| [#1](https://github.com/amoussa1229/orchestrator/pull/1) | `366d724`, `ac7101f`, `67180e1`, `f05a98d` | Phases 1, 2, 4 + README update |
|
||
| [#2](https://github.com/amoussa1229/orchestrator/pull/2) | `5378bac` | Code-review FIX items: Gemini retry coverage, batched embeddings, UTC log filenames, cost-includes-failures |
|
||
| [#3](https://github.com/amoussa1229/orchestrator/pull/3) | `fb5dbb2`, `3638d6b`, `2bb8365` | CI workflow, Sonnet 4.6 bump, golden-set CI collection fix |
|
||
|
||
Memory updates:
|
||
- `project_orchestration_migration.md` — rewritten for current architecture
|
||
- `feedback_executor_mode_workflow.md` — new (executor-mode workflow rules)
|
||
- `feedback_pre_action_smoke_test.md` — new (smoke-test prompt builders)
|
||
- `MEMORY.md` — indexed both new feedback entries
|
||
|
||
System updates:
|
||
- `~/.config/git/ignore` — `.DS_Store` added globally
|
||
|
||
---
|
||
|
||
## Open items
|
||
|
||
### 1. Retrieval-augmented routing test — **DO BEFORE PHASE 5**
|
||
|
||
**Status:** DONE (2026-05-26). Added 6 tests: 3 with live retrieval, 3 with off-topic fake memories. Also fixed Composio lazy loading so tests don't crash on expired keys.
|
||
|
||
**What's missing.** The golden-set test (`tests/test_routing_golden.py`) calls `router_node` with no `retrieved` key, so all 21 cases exercise the router on a bare prompt without memory injection. We currently have zero test signal on whether retrieval flips a routing decision.
|
||
|
||
**Why it blocks Phase 5.** Phase 5 modifies `RouteDecision` (adds `risk_class`, `confidence`). That changes the router prompt path that retrieval also touches. Without a test that validates memory-augmented routing, we can't tell whether Phase 5 breaks routing only in the with-memory case.
|
||
|
||
**Fix.** Add 2–3 cases with populated `retrieved`:
|
||
|
||
```python
|
||
from retriever import retrieve
|
||
|
||
@requires_keys
|
||
def test_router_unchanged_with_relevant_memory():
|
||
retrieved = retrieve("Send a Slack message to ops", k=3)
|
||
out = router_node({
|
||
"task": "Send a Slack message to ops about the deploy",
|
||
"messages": [],
|
||
"retrieved": retrieved,
|
||
})
|
||
assert out["route"] == "connector"
|
||
|
||
|
||
@requires_keys
|
||
def test_router_unchanged_with_offtopic_memory():
|
||
# Hand-built off-topic context — router should still pick the obvious route.
|
||
fake_retrieved = [
|
||
{"name": "feedback_emulator_coordinates", "score": 0.1,
|
||
"content": "ADB input uses landscape; getevent uses portrait"},
|
||
]
|
||
out = router_node({
|
||
"task": "Review this PR diff for security issues",
|
||
"messages": [],
|
||
"retrieved": fake_retrieved,
|
||
})
|
||
assert out["route"] == "reviewer"
|
||
```
|
||
|
||
**Priority: medium.** Not urgent for current usage, but it's the only gate on confidently merging Phase 5.
|
||
|
||
---
|
||
|
||
### 2. Non-atomic cache write — **LOW**
|
||
|
||
**Status:** DONE (2026-05-26). Atomic write via temp file in `retriever._save_cache`.
|
||
|
||
**What's wrong.** `retriever._save_cache` calls `CACHE_FILE.write_text(...)`, which truncates the file to zero before writing. A crash, Ctrl-C, or kill during the ~50ms write window leaves `.cache/embeddings.json` empty or half-written. Next run, `_load_cache` hits `JSONDecodeError`, returns `{}`, and the retriever re-embeds all 87 memories (~10s + ~$0.001).
|
||
|
||
**Impact severity.** Latency-only, not correctness or data-loss. Source memory files are untouched; re-embedded vectors are identical.
|
||
|
||
**Probability.** Low. Vulnerable window is short and only occurs when memories actually changed (cache-hit-only path doesn't write).
|
||
|
||
**Fix.** Atomic write via temp file:
|
||
|
||
```python
|
||
def _save_cache(cache: dict) -> None:
|
||
CACHE_DIR.mkdir(exist_ok=True)
|
||
tmp = CACHE_FILE.with_suffix(".json.tmp")
|
||
tmp.write_text(json.dumps(cache))
|
||
tmp.replace(CACHE_FILE)
|
||
```
|
||
|
||
**Priority: low.** Bundle with the next change that touches `retriever.py`. Not worth a dedicated PR.
|
||
|
||
---
|
||
|
||
### 3. Telemetry log rotation — **LOW**
|
||
|
||
**Status:** DONE (2026-05-26). `prune_old_logs()` added to `weekly_summary.py`; deletes JSONL files >90 days old, runs automatically after each digest.
|
||
|
||
**What's wrong.** `~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl` files accumulate forever. No deletion, no compression. At current volume (20 runs/day) the directory grows ~2 MB/year — disk is not the issue, but file count climbs and the digest only reads the last 7 days, so older logs are noise.
|
||
|
||
**Fix.** One-liner in a scheduled routine:
|
||
|
||
```bash
|
||
find ~/.claude/logs/orchestrator -name "*.jsonl" -mtime +90 -delete
|
||
```
|
||
|
||
90 days well exceeds the 7-day digest window.
|
||
|
||
**Priority: low.** Bundle with the weekly digest `/schedule` — same job can emit the digest and then prune.
|
||
|
||
---
|
||
|
||
## Scheduled-but-unscheduled
|
||
|
||
These were called out as "do it when ready" but never wired:
|
||
|
||
- **Weekly digest** → `python3 ~/Documents/repositories/orchestrator/scripts/weekly_summary.py` piped to Slack. Suggested cadence: Monday morning. Wire via `/schedule`.
|
||
- **Quarterly log prune** (see item 3 above) — bundle with the digest schedule.
|
||
|
||
---
|
||
|
||
## Telemetry collection period
|
||
|
||
The next decision point on the assessment roadmap (Phase 5 vs Phase 3 vs stop) was deliberately deferred until **~1 week of real-workload telemetry** is in the JSONL log. As of 2026-05-15 the log has 2 entries (both from smoke testing this session). Real usage starts now.
|
||
|
||
Re-run `python3 scripts/weekly_summary.py` after ~7 days of normal use to get the first meaningful digest. Decisions that depend on this:
|
||
|
||
- **Unknown rate** — target <2%. If higher, router prompt or AGENTS descriptions need work.
|
||
- **Top routes** — confirms model-diversity hypothesis is real (i.e., non-Sonnet routes get used).
|
||
- **Mean tokens per route** — flags runaway prompts.
|
||
- **Failed-run cost** — flags transient provider issues that retries aren't catching.
|
||
|
||
---
|
||
|
||
## Roadmap status
|
||
|
||
From `~/Desktop/orchestrator-modernization-assessment.md`:
|
||
|
||
| Phase | Status | Notes |
|
||
|---|---|---|
|
||
| 1 — Stabilize foundations | ✅ DONE | PR #1, commit `366d724` |
|
||
| 2 — Memory retriever | ✅ DONE | PR #1, commit `ac7101f` |
|
||
| 3 — Governance compiler | HOLD | Only if prompt drift pain shows up. AGENTS dict + handbook links currently cover ~80%. |
|
||
| 4 — JSONL telemetry | ✅ DONE | PR #1, commit `f05a98d` |
|
||
| 5 — Risk class + connector audit log | GATED on item 1 above + telemetry data | Low-risk addition once item 1 lands and digest shows clean baseline. |
|
||
| 6 — Memory distillation | HOLD | Only if Phase 2 retrieval is paying off after a few weeks of real use. |
|
||
| 7 — Bounded connector chaining | HOLD | Only if `connector-audit.jsonl` (Phase 5) shows a recurring multi-call pattern. |
|
||
|
||
---
|
||
|
||
## How to pick up next session
|
||
|
||
If continuing this work in a future session, paste this prompt template:
|
||
|
||
> Context: orchestrator modernization handoff at `~/Documents/repositories/orchestrator/HANDOFF-2026-05-15.md`. Read that first. Phases 1, 2, 4 are merged. Open items #1 (retrieval test), #2 (atomic cache), #3 (log rotation) documented there. I want to <pick one: tackle item #1 / schedule the weekly digest / start Phase 5 / something else>.
|
||
|
||
Memory entries the next session will load automatically (already in MEMORY.md):
|
||
- `project_orchestration_migration.md`
|
||
- `feedback_executor_mode_workflow.md` — relevant if executing an approved plan
|
||
- `feedback_pre_action_smoke_test.md` — relevant for any prompt-builder work
|
||
- `feedback_composio_direct_tools.md` — relevant if anyone proposes meta-tool routing
|
||
|
||
---
|
||
|
||
## References
|
||
|
||
- **Assessment**: `~/Desktop/orchestrator-modernization-assessment.md` — original analysis and phase plan (kept outside the repo as the originating artifact)
|
||
- **This handoff**: `~/Documents/repositories/orchestrator/HANDOFF-2026-05-15.md` (you're reading it)
|
||
- **Repo**: https://github.com/amoussa1229/orchestrator
|
||
- **PRs**: [#1](https://github.com/amoussa1229/orchestrator/pull/1), [#2](https://github.com/amoussa1229/orchestrator/pull/2), [#3](https://github.com/amoussa1229/orchestrator/pull/3)
|
||
- **Telemetry logs**: `~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl` (UTC date)
|
||
- **Embedding cache**: `~/Documents/repositories/orchestrator/.cache/embeddings.json`
|
||
- **Memory store**: `~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/`
|