Commit graph

5 commits

Author SHA1 Message Date
8fb7d188b3 fix(ws5): allowlist save_memory names + symlink-safe write
Security follow-up from the per-PR review (non-blocking, defense-in-depth):
- Replace the save_memory name blocklist with an allowlist regex
  (^[A-Za-z0-9][A-Za-z0-9._-]*$, max 128) so dot-only/hidden/backslash/NUL/
  over-long names are rejected outright, not written as malformed-but-contained
  files.
- Write via os.open(..., O_NOFOLLOW): the open fails (ELOOP) if the final
  path component is a pre-planted symlink, closing the TOCTOU where a symlink
  in _box-drafts/ could redirect the write outside the dir. O_CREAT|O_TRUNC
  keeps overwrite-on-resave for regular files.

Tests: adds allowlist-rejection + symlink-refusal cases (21 pass).
2026-06-23 12:29:25 -04:00
Claude
a12de32a34
feat(ws5): add memory/handbook injection seams for context-provider pattern
- retriever: add save_memory() writing to _box-drafts/ review queue, add
  memory_dir param to retrieve() for isolated test routing
- handbook: new load_handbook_conventions() with safe no-op contract (returns
  "" when dir absent/empty/unreadable, never raises)
- clarifier_llm: add context_provider=None seam to ClaudeClarifier and
  build_claude_clarifier_callables(); failure in provider is silent
- planner: add context_provider=None seam to build_plan_prompt() and plan_node()
- coordinator: thread context_provider through default_plan_node_factory(),
  only pass kwarg when non-None to preserve stub-monkeypatching in tests
- tests: 19 new WS5 tests covering all seams (1063 total, all passing)
2026-06-23 01:08:34 +00:00
Adam Moussa
c23ea679e4 Add retrieval-augmented routing tests, fix lazy loading and open items #2-3
- Add 6 retrieval-augmented routing tests (3 live retrieval, 3 off-topic
  fake memories) to unblock Phase 5
- Defer Composio tool loading and graph construction to first use so
  expired or missing keys don't crash imports
- Atomic cache write in retriever via temp file (open item #2)
- Log rotation in weekly_summary.py, pruning JSONL >90 days (open item #3)
2026-05-26 18:26:59 -04:00
Adam Moussa
5378bac69b Address code review FIX items from PR #1
Follow-up to the Phases 1-2-4 modernization PR.

- models.py: extend retry coverage to Gemini transient errors
  (google.api_core: ResourceExhausted, InternalServerError,
  ServiceUnavailable, DeadlineExceeded, GatewayTimeout). Before this,
  a 429 or 5xx from Google would crash the scanner route. Retriable
  count went from 6 to 11.
- retriever.py: batch cache-miss embeddings into a single
  embed_documents() call instead of one embed_query() per memory.
  Cold rebuild of 87 memories went from ~30s to ~10s; one HTTPS
  round-trip instead of 87.
- telemetry.py + scripts/weekly_summary.py: use UTC date for log
  filenames so they align with the UTC `timestamp` field inside each
  record. Eliminates the off-by-one near local midnight and matches
  Sea Haven's UTC-for-logs convention.
- scripts/weekly_summary.py: include failed runs in the total cost
  estimate (they consumed tokens too) and surface a `(incl. $X on
  failed runs)` breakdown so outages are visible in the digest.

Validated: ruff clean; weekly_summary.py still prints; retriever cold
rebuild + cache hit both verified end-to-end.
2026-05-15 13:01:14 -04:00
Adam Moussa
ac7101f3df Add memory retriever node (Phase 2)
Plugs the orchestrator into Adam's existing memory store at
~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/. Every
run starts with a top-3 retrieval pass that is then surfaced in the CLI
output and injected as system context into the router and downstream agent.

- retriever.py: load *.md memories (skipping the MEMORY.md index), embed
  with text-embedding-3-small, cache to .cache/embeddings.json keyed on
  file mtime. Cosine similarity, top-k=3 default. Reads only — never
  writes back to the memory store.
- state.py: add `retrieved: list[dict]` to OrchestratorState; relax to
  total=False to match LangGraph's partial-update semantics.
- graph.py: new retriever_node wired as START -> retriever -> router.
  router_node and connector_node now inject retrieved memories into their
  SystemMessage. Retrieval failures are caught and the run continues with
  empty memory context (logged).
- agents.py: make_agent_node injects retrieved memories into each agent's
  system prompt.
- run.py: prints `[retrieved: name1, name2, name3]` (or `[retrieved: none]`)
  before route/result for both --route-only and full-run modes, so bad
  retrieval is visible at a glance.
- .gitignore: add .cache/, .pytest_cache/, .ruff_cache/.

Validated: golden-set still 21/21 passing; smoke tests retrieve plausible
memories ("Send a Slack message to ops about the new exec-aide deploy" ->
project_exec_aide, feedback_exec_aide_vip_management, project_seahaven_slack_bot).
2026-05-15 11:36:03 -04:00