Follow-up to the Phases 1-2-4 modernization PR.
- models.py: extend retry coverage to Gemini transient errors
(google.api_core: ResourceExhausted, InternalServerError,
ServiceUnavailable, DeadlineExceeded, GatewayTimeout). Before this,
a 429 or 5xx from Google would crash the scanner route. Retriable
count went from 6 to 11.
- retriever.py: batch cache-miss embeddings into a single
embed_documents() call instead of one embed_query() per memory.
Cold rebuild of 87 memories went from ~30s to ~10s; one HTTPS
round-trip instead of 87.
- telemetry.py + scripts/weekly_summary.py: use UTC date for log
filenames so they align with the UTC `timestamp` field inside each
record. Eliminates the off-by-one near local midnight and matches
Sea Haven's UTC-for-logs convention.
- scripts/weekly_summary.py: include failed runs in the total cost
estimate (they consumed tokens too) and surface a `(incl. $X on
failed runs)` breakdown so outages are visible in the digest.
Validated: ruff clean; weekly_summary.py still prints; retriever cold
rebuild + cache hit both verified end-to-end.
Plugs the orchestrator into Adam's existing memory store at
~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/. Every
run starts with a top-3 retrieval pass that is then surfaced in the CLI
output and injected as system context into the router and downstream agent.
- retriever.py: load *.md memories (skipping the MEMORY.md index), embed
with text-embedding-3-small, cache to .cache/embeddings.json keyed on
file mtime. Cosine similarity, top-k=3 default. Reads only — never
writes back to the memory store.
- state.py: add `retrieved: list[dict]` to OrchestratorState; relax to
total=False to match LangGraph's partial-update semantics.
- graph.py: new retriever_node wired as START -> retriever -> router.
router_node and connector_node now inject retrieved memories into their
SystemMessage. Retrieval failures are caught and the run continues with
empty memory context (logged).
- agents.py: make_agent_node injects retrieved memories into each agent's
system prompt.
- run.py: prints `[retrieved: name1, name2, name3]` (or `[retrieved: none]`)
before route/result for both --route-only and full-run modes, so bad
retrieval is visible at a glance.
- .gitignore: add .cache/, .pytest_cache/, .ruff_cache/.
Validated: golden-set still 21/21 passing; smoke tests retrieve plausible
memories ("Send a Slack message to ops about the new exec-aide deploy" ->
project_exec_aide, feedback_exec_aide_vip_management, project_seahaven_slack_bot).