Commit graph

3 commits

Author SHA1 Message Date
Adam Moussa
c23ea679e4 Add retrieval-augmented routing tests, fix lazy loading and open items #2-3
- Add 6 retrieval-augmented routing tests (3 live retrieval, 3 off-topic
  fake memories) to unblock Phase 5
- Defer Composio tool loading and graph construction to first use so
  expired or missing keys don't crash imports
- Atomic cache write in retriever via temp file (open item #2)
- Log rotation in weekly_summary.py, pruning JSONL >90 days (open item #3)
2026-05-26 18:26:59 -04:00
Adam Moussa
5378bac69b Address code review FIX items from PR #1
Follow-up to the Phases 1-2-4 modernization PR.

- models.py: extend retry coverage to Gemini transient errors
  (google.api_core: ResourceExhausted, InternalServerError,
  ServiceUnavailable, DeadlineExceeded, GatewayTimeout). Before this,
  a 429 or 5xx from Google would crash the scanner route. Retriable
  count went from 6 to 11.
- retriever.py: batch cache-miss embeddings into a single
  embed_documents() call instead of one embed_query() per memory.
  Cold rebuild of 87 memories went from ~30s to ~10s; one HTTPS
  round-trip instead of 87.
- telemetry.py + scripts/weekly_summary.py: use UTC date for log
  filenames so they align with the UTC `timestamp` field inside each
  record. Eliminates the off-by-one near local midnight and matches
  Sea Haven's UTC-for-logs convention.
- scripts/weekly_summary.py: include failed runs in the total cost
  estimate (they consumed tokens too) and surface a `(incl. $X on
  failed runs)` breakdown so outages are visible in the digest.

Validated: ruff clean; weekly_summary.py still prints; retriever cold
rebuild + cache hit both verified end-to-end.
2026-05-15 13:01:14 -04:00
Adam Moussa
f05a98dba1 Add JSONL telemetry and weekly summary (Phase 4)
Observability slim layer. Every full run appends one JSON line to
~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl with timestamp, sha256-prefix
task hash (raw task is never logged), retrieved memory names, router
choice, runtime, tokens in/out, success/error. risk_class and confidence
fields are reserved nulls for Phase 5.

- telemetry.py: log_run(), build_record(), task_hash(), token-usage
  extraction from AIMessage.usage_metadata. log_run swallows all
  exceptions — telemetry never kills a run.
- run.py: wraps app.invoke in try/except with a monotonic-clock window;
  logs on both success and failure. --route-only path is left unlogged
  (no agent work, doesn't represent a "run").
- scripts/weekly_summary.py: scans the last 7 days of JSONL and prints
  a markdown digest (routes, unknown rate, cross-review rate, success
  rate, total spend, mean tokens/route). Schedule via /schedule and
  pipe stdout to Slack from the scheduler.

Cost rates per route are rough Sonnet/Haiku/GPT/Gemini/DeepSeek defaults
suitable for spotting runaway prompts, not finance. Router tokens for
structured-output calls aren't captured (they don't surface through the
message trail); agent tokens are the dominant component anyway.

Validated: golden-set 21/21 still passing; full-run smoke writes
expected fields; weekly_summary.py prints clean markdown from a 2-run
log.
2026-05-15 11:49:05 -04:00