Commit graph

10 commits

Author SHA1 Message Date
Adam Moussa
5378bac69b Address code review FIX items from PR #1
Follow-up to the Phases 1-2-4 modernization PR.

- models.py: extend retry coverage to Gemini transient errors
  (google.api_core: ResourceExhausted, InternalServerError,
  ServiceUnavailable, DeadlineExceeded, GatewayTimeout). Before this,
  a 429 or 5xx from Google would crash the scanner route. Retriable
  count went from 6 to 11.
- retriever.py: batch cache-miss embeddings into a single
  embed_documents() call instead of one embed_query() per memory.
  Cold rebuild of 87 memories went from ~30s to ~10s; one HTTPS
  round-trip instead of 87.
- telemetry.py + scripts/weekly_summary.py: use UTC date for log
  filenames so they align with the UTC `timestamp` field inside each
  record. Eliminates the off-by-one near local midnight and matches
  Sea Haven's UTC-for-logs convention.
- scripts/weekly_summary.py: include failed runs in the total cost
  estimate (they consumed tokens too) and surface a `(incl. $X on
  failed runs)` breakdown so outages are visible in the digest.

Validated: ruff clean; weekly_summary.py still prints; retriever cold
rebuild + cache hit both verified end-to-end.
2026-05-15 13:01:14 -04:00
Adam Moussa
f6a4071464
Merge pull request #1 from amoussa1229/phase-1-stabilize-foundations
Orchestrator modernization — Phases 1, 2, 4
2026-05-15 12:14:51 -04:00
Adam Moussa
f05a98dba1 Add JSONL telemetry and weekly summary (Phase 4)
Observability slim layer. Every full run appends one JSON line to
~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl with timestamp, sha256-prefix
task hash (raw task is never logged), retrieved memory names, router
choice, runtime, tokens in/out, success/error. risk_class and confidence
fields are reserved nulls for Phase 5.

- telemetry.py: log_run(), build_record(), task_hash(), token-usage
  extraction from AIMessage.usage_metadata. log_run swallows all
  exceptions — telemetry never kills a run.
- run.py: wraps app.invoke in try/except with a monotonic-clock window;
  logs on both success and failure. --route-only path is left unlogged
  (no agent work, doesn't represent a "run").
- scripts/weekly_summary.py: scans the last 7 days of JSONL and prints
  a markdown digest (routes, unknown rate, cross-review rate, success
  rate, total spend, mean tokens/route). Schedule via /schedule and
  pipe stdout to Slack from the scheduler.

Cost rates per route are rough Sonnet/Haiku/GPT/Gemini/DeepSeek defaults
suitable for spotting runaway prompts, not finance. Router tokens for
structured-output calls aren't captured (they don't surface through the
message trail); agent tokens are the dominant component anyway.

Validated: golden-set 21/21 still passing; full-run smoke writes
expected fields; weekly_summary.py prints clean markdown from a 2-run
log.
2026-05-15 11:49:05 -04:00
Adam Moussa
67180e124e Update README for Phases 1-2 architecture
Reflect retriever node, structured router with explicit unknown route,
AGENTS registry, model-ID constants, memory-aware CLI output, and the
golden-set test.
2026-05-15 11:37:11 -04:00
Adam Moussa
ac7101f3df Add memory retriever node (Phase 2)
Plugs the orchestrator into Adam's existing memory store at
~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/. Every
run starts with a top-3 retrieval pass that is then surfaced in the CLI
output and injected as system context into the router and downstream agent.

- retriever.py: load *.md memories (skipping the MEMORY.md index), embed
  with text-embedding-3-small, cache to .cache/embeddings.json keyed on
  file mtime. Cosine similarity, top-k=3 default. Reads only — never
  writes back to the memory store.
- state.py: add `retrieved: list[dict]` to OrchestratorState; relax to
  total=False to match LangGraph's partial-update semantics.
- graph.py: new retriever_node wired as START -> retriever -> router.
  router_node and connector_node now inject retrieved memories into their
  SystemMessage. Retrieval failures are caught and the run continues with
  empty memory context (logged).
- agents.py: make_agent_node injects retrieved memories into each agent's
  system prompt.
- run.py: prints `[retrieved: name1, name2, name3]` (or `[retrieved: none]`)
  before route/result for both --route-only and full-run modes, so bad
  retrieval is visible at a glance.
- .gitignore: add .cache/, .pytest_cache/, .ruff_cache/.

Validated: golden-set still 21/21 passing; smoke tests retrieve plausible
memories ("Send a Slack message to ops about the new exec-aide deploy" ->
project_exec_aide, feedback_exec_aide_vip_management, project_seahaven_slack_bot).
2026-05-15 11:36:03 -04:00
Adam Moussa
366d7247da Stabilize router and consolidate agent registry
Phase 1 stabilization. Removes the four-copy prompt/agent-description drift
surface and the silent router fallback.

- models.py: hoist model IDs to module-level constants; add with_retries()
  helper (2 retries on Anthropic+OpenAI transient errors via with_retry).
- agents.py: single AGENTS dict (model_fn, prompt, description) and a
  make_agent_node() factory that collapses six near-identical node functions.
- graph.py: router prompt is generated from AGENTS; router_node uses
  with_structured_output(RouteDecision) and returns an explicit "unknown"
  route instead of the silent "researcher" fallback. New unknown_node wires
  to END. All LLM invocations go through with_retries.
- state.py: add "unknown" to the route Literal.
- run.py: --route-only now imports router_node from graph.py, killing the
  fourth prompt copy.
- tests/: pytest golden-set (20 labelled tasks + size guard). Skips cleanly
  without ANTHROPIC_API_KEY or COMPOSIO_API_KEY. Validated 21/21 passing.
2026-05-15 11:20:28 -04:00
Adam Moussa
acfeb543d9 Document hybrid delegation model for Claude Code
README now explains when Claude Code delegates to the orchestrator
vs. handles tasks natively, matching the updated CLAUDE.md rules.
2026-05-08 16:41:46 -04:00
Adam Moussa
ca96b488e6 Add README and fix .env.example project name 2026-05-08 13:12:02 -04:00
Adam Moussa
b125956ee5 Add CLI entry point and fix Composio user_id
- Add run.py with --route-only flag for Claude Code integration
- Fix Composio user_id from "default" to "amoussa" to match connected accounts
2026-05-08 13:10:39 -04:00
Adam Moussa
02e29495c0 Initial LangGraph + Composio orchestration graph
Multi-model agent routing with 7 agent nodes (Sonnet, Haiku,
GPT-4.1, Gemini, DeepSeek) and 15 pre-loaded Composio tools
for Slack, Notion, GitHub, and Google Drive integration.
2026-05-08 12:54:00 -04:00