- Add 6 retrieval-augmented routing tests (3 live retrieval, 3 off-topic
fake memories) to unblock Phase 5
- Defer Composio tool loading and graph construction to first use so
expired or missing keys don't crash imports
- Atomic cache write in retriever via temp file (open item #2)
- Log rotation in weekly_summary.py, pruning JSONL >90 days (open item #3)
Follow-up to the Phases 1-2-4 modernization PR.
- models.py: extend retry coverage to Gemini transient errors
(google.api_core: ResourceExhausted, InternalServerError,
ServiceUnavailable, DeadlineExceeded, GatewayTimeout). Before this,
a 429 or 5xx from Google would crash the scanner route. Retriable
count went from 6 to 11.
- retriever.py: batch cache-miss embeddings into a single
embed_documents() call instead of one embed_query() per memory.
Cold rebuild of 87 memories went from ~30s to ~10s; one HTTPS
round-trip instead of 87.
- telemetry.py + scripts/weekly_summary.py: use UTC date for log
filenames so they align with the UTC `timestamp` field inside each
record. Eliminates the off-by-one near local midnight and matches
Sea Haven's UTC-for-logs convention.
- scripts/weekly_summary.py: include failed runs in the total cost
estimate (they consumed tokens too) and surface a `(incl. $X on
failed runs)` breakdown so outages are visible in the digest.
Validated: ruff clean; weekly_summary.py still prints; retriever cold
rebuild + cache hit both verified end-to-end.
Observability slim layer. Every full run appends one JSON line to
~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl with timestamp, sha256-prefix
task hash (raw task is never logged), retrieved memory names, router
choice, runtime, tokens in/out, success/error. risk_class and confidence
fields are reserved nulls for Phase 5.
- telemetry.py: log_run(), build_record(), task_hash(), token-usage
extraction from AIMessage.usage_metadata. log_run swallows all
exceptions — telemetry never kills a run.
- run.py: wraps app.invoke in try/except with a monotonic-clock window;
logs on both success and failure. --route-only path is left unlogged
(no agent work, doesn't represent a "run").
- scripts/weekly_summary.py: scans the last 7 days of JSONL and prints
a markdown digest (routes, unknown rate, cross-review rate, success
rate, total spend, mean tokens/route). Schedule via /schedule and
pipe stdout to Slack from the scheduler.
Cost rates per route are rough Sonnet/Haiku/GPT/Gemini/DeepSeek defaults
suitable for spotting runaway prompts, not finance. Router tokens for
structured-output calls aren't captured (they don't surface through the
message trail); agent tokens are the dominant component anyway.
Validated: golden-set 21/21 still passing; full-run smoke writes
expected fields; weekly_summary.py prints clean markdown from a 2-run
log.