Commit graph

97 commits

Author SHA1 Message Date
b0d8b842e5 security-review: surgically exclude only cdk.out asset bundles from checkov (A2)
The prior exclusion (--skip-path cdk.out) stopped the CDK-repo stall but also
silenced checkov on cdk.out/<stack>.template.json — the actual deploy artifact —
losing real S3/IAM IaC coverage (CKV_AWS_53-56, CKV_AWS_111). Switch to skipping
only the cdk.out asset.<hash>/ dependency bundles (the stall cause) so the
synthesized templates are still scanned.

- --skip-path 'cdk\.out/asset\.' anchors to cdk.out so a source file literally
  named asset.* is not also excluded; keeps cdk.out/*.template.json scanned.
- venv/dist/build kept as bare names (match anywhere); .venv/.aws-sam escaped.

cfn-lint + semgrep already prune these trees (prior commit). gitleaks runs in
git-mode and respects .gitignore, so cdk.out is already skipped there.

Verified: shellcheck clean; synthetic cdk.out test confirms the stack template is
scanned while asset.* is skipped; the orchestrator's own --scanners-only gate
still exits 0 with suppressions (inert on non-CDK repos: A1==A2 findings here).
2026-06-17 17:06:05 -04:00
Adam Moussa
af6dd98783
Merge pull request #8 from amoussa1229/feature/r720-plane2-scaffold
R720 agent-team: Plane-2 SDLC pipeline scaffold (pre-deployment)
2026-06-17 15:21:27 -04:00
502828f75c Fix CI collection: isolate agent-team tests; add agent-team CI job
Repo-root 'pytest --collect-only' failed with ImportPathMismatchError because
agent-team/tests/ and the root tests/ are both the 'tests' package. Add a root
conftest.py that excludes agent-team from the root collection, and a dedicated
agent-team-tests CI job that runs the (API-key-free) agent-team suite in its own
working dir.
2026-06-17 15:19:51 -04:00
721cec5315 Resolve security-review BLOCK: CI-guard bypasses, denylist parity, force-resume
Addresses the confirmed findings from /sh-security-review + the GPT-4.1
cross-review of the Plane-2 scaffold. Full suite: 589 passed; ruff clean.

FIXED (proven-exploitable):
- CI-guard denylist bypass (HIGH): Python fnmatch '**/' is non-recursive, so
  root-level template.yaml/*.tf/cdk.json/*.pem/*.key/*-stack.* evaded the
  trust-control surface. Replaced fnmatch with a recursive, case-insensitive
  glob->regex matcher. (verified: fnmatch('template.yaml','**/template.yaml')==False)
- CI-guard scope bypass (HIGH): a '**' declared_scope made every path in-scope.
  Scope is now concrete-prefix confinement (reduces a glob to its leading
  metacharacter-free segments; '**' -> empty -> dropped -> unscoped reject).
- Box-side vs CI denylist divergence (MED): builders.py _DENY_PATTERNS now covers
  Terraform, *.pem/*.key, CDK stack files, .github/actions, *iam*, bare policy*.json
  (case-insensitive), matching the CI surface.
- force-resume was backwards (MED): it superseded the answered row recovery
  resumes from, making a stuck task permanently un-resumable while printing
  success. Now re-opens an EXPIRED (parked) question via a new reopen_question
  CAS helper; never supersedes an answered row; honest exit codes.
- operator attribution (MED): run-team.py --operator defaulted to "" -> now the
  OS login, so destructive actions are always attributable.
- audit-log append race (MED): replaced read-modify-rewrite (lost records under
  concurrent operators) with an O_APPEND single-line write, mode 600 enforced.
- lstrip("ab/") path-mangling in the symlink error path -> regex prefix strip.

Regression tests added across test_ci_gate_workflow / test_builders / test_run_team
/ test_schema. Design-level findings (resume-worker durability, egress breadth,
answered_at ordering, DB-swap TOCTOU, diff-hash threat-model) are pre-deployment
/ P1-build-proper and recorded with written justification in
agent-team/.security-review/suppressions.json; CI README diff-hash wording made
honest.
2026-06-17 15:16:12 -04:00
0eae5dbfc3 Prove P1 exit criteria against the real LangGraph graph; fix question_id stability
Reworks the P1 sim so the four §7.1 exit criteria are demonstrated against the
ACTUAL mechanic, not a model (resolves the verifier's "sim models the ledger,
not the LangGraph integration" finding).

- New tests/sim/test_p1_graph_integration.py drives the real agent_team.graph
  StateGraph (interrupt/Command(resume)) + the real langgraph SqliteSaver
  checkpointer + the committed pending_questions compare-and-set, proving:
  (a) suspend survives a simulated restart (drop saver/conn, rebuild over the
  same checkpoint DB) and resumes; (b) duplicate answer loses the CAS and the
  graph never double-advances; (c) a post-deadline answer loses to expire and
  the task is not resumed; (d) two concurrent tasks resume to the correct
  thread, with a turn-guarded no-double-apply check.
- graph.py: derive a STABLE question_id from uuid5(thread_id, turn). The
  clarifier node replays on resume, so the prior fresh-uuid id changed between
  the delivered/ledgered question and the qa_history entry — breaking the
  §3.3.1 identity contract. Now the delivered id == ledger key == history entry
  (unit-tested in test_graph.py).
- harness._connect() now uses the committed schema.connect() (WAL + busy_timeout)
  instead of a raw sqlite3.connect, so concurrent responders genuinely serialize;
  the criterion-(d) concurrency test no longer swallows OperationalError (it
  asserts zero errors + exactly one CAS winner).
- requirements.txt: pin langgraph-checkpoint-sqlite==3.1.0 (design D9 durable
  checkpointer), now exercised by the integration test.

Full suite: 564 passed; ruff + format clean.
2026-06-17 15:16:12 -04:00
8f28fbe6a5 Harden CI guard: symlink-escape reject, strict decode, scope canon; back with tests
Resolves the GPT-4.1 cross-review FIX items on the §3.3.2 CI apply/verify guard:
- Symlink-escape (Medium-High): reject any candidate diff that introduces a
  symlink (git mode 120000). A symlink can redirect a later in-diff write into a
  denied path that textual canonicalization cannot see; auto-built diffs have no
  legitimate symlinks, so this fails closed (exit 7).
- Diff-parse robustness (Medium): decode the diff as strict UTF-8 and fail closed
  (exit 8) instead of errors='replace', closing homoglyph/encoding evasion.
- Declared-scope canonicalization (Medium): drop parent-escaping scope globs so a
  malformed scope can only shrink coverage, never widen it past repo root.
- Egress allowlist (Low-Med): explicit DEPLOY marker to parameterize the
  build-test registries per target repo before enabling.

Backs the gate's correctness claim with a committed, runnable suite
(tests/test_ci_gate_workflow.py) that extracts the inline guard from the YAML and
exercises good + adversarial diffs (clean, hash mismatch, workflow delete,
copy-into-denied, symlink, non-UTF-8, out-of-scope, unscoped, escaping scope).
Corrects the README "Tests" section that claimed coverage that did not exist.
Full suite: 557 passed, 1 skipped; ruff clean.
2026-06-17 15:16:12 -04:00
3c29ce3fdb Fix verified P1 findings: denylist bypasses, CAS concurrency, operator CLI
Resolves three execution-proven verifier findings from the scaffold review.
Full suite: 548 passed, 1 skipped (stable across repeated runs); ruff clean.

builders denylist (§3.3.2 #2): scan was +++-only and missed header-only
sections. Now section-driven off `diff --git a/<src> b/<dest>`, catching the 4
proven bypasses — delete of a denied path, mode-change-only, `copy to` a denied
path, out-of-scope delete (regression tests for each).

§3.3.1 compare-and-set concurrency: BEGIN IMMEDIATE moved inside guarded retry;
each CAS now runs on its own connection (shared sqlite3.Connection cannot hold
two transactions, and is unsafe for concurrent use even for reads). connect()
stashes the db path on a Connection subclass so the path is derived by a
thread-safe attribute read, not a PRAGMA on the shared conn; busy_timeout set
before the WAL pragma. Added shared-connection concurrent regression tests
(distinct + same question) — previously raised "transaction within a
transaction".

operator CLI (run-team.py): added the design-named re-deliver and force-resume
verbs (were missing); audit now records the attempt BEFORE the mutation and the
outcome after, so a ledger mutation can never land without a trail; main()
catches OSError instead of leaving an uncaught traceback on audit-write failure.
2026-06-17 15:16:12 -04:00
15a416d31a Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.

Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled

KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven

Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 15:16:12 -04:00
dc2da449fd Plane 2 foundation: interfaces, SQLite schemas, state-store, billing seam 2026-06-17 15:16:12 -04:00
Adam Moussa
c7661460e9
Merge pull request #6 from amoussa1229/feature/security-review-two-tier-auto-discovery
Finalize security-review: repo-sourced hooks, two-tier nightly auto-discovery
2026-06-17 15:03:15 -04:00
c6ecc621be Prune generated/vendored trees from the scanners
cfn-lint, semgrep, and checkov were scanning synthesized/vendored output
(cdk.out, node_modules, .venv/venv, .aws-sam, dist, build). On CDK repos this
explodes the find/xargs arg list and stalls the scan, and flagging synthesized
templates is wrong. Prune those trees in the cfn-lint find, and pass
--exclude / --skip-path to semgrep / checkov.
2026-06-17 14:55:41 -04:00
dbe2f8c186 Raise nightly agentic budget $20 -> $120 for full per-night coverage
First-run data (2026-06-17) showed the $20 ceiling covered only the
canary + 5 of ~22 scannable repos before pausing the rotation, leaving
16 repos un-deep-scanned that night. Raise TOTAL_BUDGET_USD default to
$120 so every repo gets a deep agentic pass each night (~22 x ~$5 +
canary, with headroom). Spend draws on the Max subscription pool; the
per-target cap ($12) and round-robin rotation are unchanged, so this is
a ceiling raise, not a per-repo cost change. Updates README, DEPLOY, and
the systemd Environment example to match.
2026-06-17 13:18:46 -04:00
15a16f8b88 Tune nightly defaults from first-run data
First full VM run: canary recall 18 with the expanded Node/.NET corpus, and a real
agentic repo cost ~$3.50 (vs the $1.4 testbed). Raise CANARY_FLOOR 10->14 (catches a
language-blindness recall collapse to ~9 while keeping margin under 18) and MAX_CYCLE_NIGHTS
4->6 (at ~4-6 agentic repos/night a full 22-repo rotation takes ~4-5 nights; 4 would
false-fire the coverage alarm). Both stay env-overridable and re-tunable as data accrues.
2026-06-16 18:02:57 -04:00
f4bb2bce8a Make scan_scanners return 0 so a passing repo doesn't trip set -e
scan_scanners communicates results via globals (T1_BLOCK/T1_CRIT/T1_HIGH); its last
statement was a bare [ $rc -eq 1 ] && T1_BLOCK=1. On a PASS (rc=0) that test is false,
so the function returned non-zero and set -e killed the whole sweep at the first passing
repo in tier 1 (right after the unbound-variable fix let it get that far). Add an explicit
return 0. Audited the rest of the tier1/tier2/summary path; scan_agentic and the others
already end on a zero-status command.
2026-06-16 16:43:35 -04:00
6d65c54b58 Fix unbound-variable crash in the nightly sweep tier-1 loop
scan_scanners referenced ${slug} inside the SAME local statement that defines it
(local target=... slug=$2 result_json=...${slug}...). bash expands the local's
arguments before the builtin assigns them, so under set -u ${slug} is unbound and
the sweep died right after the canary, before tier 1 ever ran. Split the local so
slug exists first. Fix the same latent self-reference in mirror_repo (dir=...$name),
which only worked by accident because the discovery loop left a global $name.
2026-06-16 16:00:21 -04:00
86ad29dcc2 Pin direct dependencies to exact versions
requirements.txt used loose >= ranges; pin the 7 direct deps to the known-good installed
versions (Python 3.12.13) for a reproducible venv.
2026-06-16 15:55:01 -04:00
028c844c86 Move the May handoff doc under docs/
De-clutter the repo root; HANDOFF-2026-05-15.md now lives in docs/.
2026-06-16 15:55:01 -04:00
572ae3ca41 Move orchestrator gitleaks suppression out of repo history
Per the machine-level-suppressions policy, the .env.example FP suppression now lives at
~/.config/sea-haven/security-review/orchestrator/suppressions.json (resolved first by the
hooks), not a committed .security-review/suppressions.json. Gate still passes via the
machine-level file.
2026-06-16 15:55:01 -04:00
80d5391785 Document the security-review subsystem in the root README
The root README covered only the router and omitted the security-review/ subsystem
entirely (Path A skill, review.sh gate, Path B nightly sweep, global hooks). Add a
Security Review section summarizing both paths and linking the subsystem README +
DEPLOY-R720 runbook.
2026-06-16 15:55:01 -04:00
c217c5656d Add machine-level suppressions to the repo-sourced hooks
The live global pre-push hook was hand-edited to resolve suppressions from a
machine-level file (${SH_SECURITY_SUPPRESSIONS_DIR:-~/.config/sea-haven/security-review}/
<repo-basename>/suppressions.json) kept out of repo history, falling back to a
repo-local .security-review/suppressions.json. The repo-sourced hooks lacked it, so
install-hooks.sh --global would overwrite the live hook and lose the feature.

Port the prefer-machine/fallback-repo-local block into hooks/pre-push, align
hooks/pre-commit to the same (else a machine-suppressed finding passes at push but
blocks at commit), and document the path + SH_SECURITY_SUPPRESSIONS_DIR override +
basename-collision caveat in the README.
2026-06-16 15:33:43 -04:00
f4dd72ced9 Add headless detector fan-out + proof-or-kill verifier runner
run_headless.py is the Path B engine the nightly sweep calls: the 6 fresh-context
detectors + proof-or-kill verifier from /sh-security-review, run unattended via the
Claude Agent SDK on subscription OAuth (pops ANTHROPIC_API_KEY so the API key can't
silently win). Read-only tools, hermetic, per-call + total budget caps, fails toward
over-reporting. Emits the finding schema that review.sh --agent-findings consumes.
2026-06-16 15:00:10 -04:00
a05afb5ebc Update security-review docs for two-tier auto-discovery
Rewrite README + DEPLOY-R720 for the clean-clone mirror model, the read-only GH_TOKEN
PAT recipe, the global hook installer, and the no-CI-by-design decision.
2026-06-16 15:00:00 -04:00
af5ce7362e Suppress orchestrator's own .env.example gitleaks false positive
gitleaks git-mode scans committed history, so the .env.example placeholder flags from
history and would block the orchestrator's own pre-push gate. Suppress it with a written
justification. Also gitignore stray .adf_final*.json left by an unrelated tool.
2026-06-16 15:00:00 -04:00
537e83975b Rewrite nightly sweep as two-tier clean-clone auto-discovery
Replace the opt-in sweep-targets allowlist with zero-wiring discovery: enumerate org
repos via the GitHub REST API (curl + read-only GH_TOKEN, no gh dependency) and mirror
each as a shallow clean clone (git clone --depth=1, default branch from the API) into
~/repo-mirrors. Scanning server-side clones keeps local .env secrets out of scope.

Tier 1 runs deterministic scanners over every repo nightly ($0 Claude); tier 2 runs the
agentic pass over a budget-bounded round-robin rotation with a persistent cycle pointer,
so the draw on the shared Max limits stays bounded and coverage never goes silently
incomplete (COVERAGE ALARM if the rotation falls behind). Skip = committed marker or
central list (marker-skips logged). Redact secrets from Slack; reports mode 600. Raise
the systemd timeout to 6h for the longer two-tier run.
2026-06-16 15:00:00 -04:00
a3ab3f5f40 Make security-review hooks and skill installable from the repo
The global pre-push hook, the /sh-security-review prompt, and finding.schema.json
previously lived only in ~/.config/git and ~/.claude (untracked) — unreproducible.
Source them here: add hooks/pre-push, rewrite install-hooks.sh with a --global mode
(lays down both hooks, sets core.hooksPath, links skill+schema into ~/.claude) and a
per-repo mode. Align pre-commit with pre-push (honor skip marker + suppressions). Add
semgrep p/javascript so the scanners cover the org's Node/.NET repos.
2026-06-16 15:00:00 -04:00
18412c7482 Remove parked security-review CI drafts
CI was swapped for the global git hooks + nightly VM sweep (solo dev), so
the parked ci/*.yml and CI-BACKSTOP-NOTES.md were dead weight — a defective
workflow in-tree is a foot-gun. Recover from history if the team grows.
2026-06-16 14:59:42 -04:00
Adam Moussa
d066690b4a
Merge pull request #5 from amoussa1229/security-review-gate
Add security-review gate (review.sh + scanners + CI backstop)
2026-06-15 15:59:00 -04:00
f90f759e12 Tune checkov severity, wire npm audit, add CI backstop + R720 runbook
checkov: high-signal exposure/access checks -> high, best-practice noise -> low
(was 51 undifferentiated mediums). npm audit wired for Node dep CVEs. Report
collapses the low/info tail to a count. Adds ci/security-review.yml (PR backstop)
and DEPLOY-R720.md (Phase 3 host runbook).
2026-06-15 15:57:34 -04:00
7e5ce1f5b2 Add security-review gate (review.sh + scanners + pre-commit hook)
Trigger-agnostic pure-code gate that merges deterministic-scanner findings
(semgrep/gitleaks/checkov/cfn-lint/pip-audit) with agent findings from
/sh-security-review, dedups, applies justification-required suppressions, and
makes the block decision (exit 1 on confirmed critical/high). Phase 2 of the
Sea Haven security-review agent; Path B (CI/headless) wiring lands in Phase 3.
2026-06-15 15:52:15 -04:00
Adam Moussa
c23ea679e4 Add retrieval-augmented routing tests, fix lazy loading and open items #2-3
- Add 6 retrieval-augmented routing tests (3 live retrieval, 3 off-topic
  fake memories) to unblock Phase 5
- Defer Composio tool loading and graph construction to first use so
  expired or missing keys don't crash imports
- Atomic cache write in retriever via temp file (open item #2)
- Log rotation in weekly_summary.py, pruning JSONL >90 days (open item #3)
2026-05-26 18:26:59 -04:00
Adam Moussa
4d4381b5eb
Merge pull request #4 from amoussa1229/docs/handoff-2026-05-15
Add session handoff doc (2026-05-15)
2026-05-15 13:34:09 -04:00
Adam Moussa
2705ee38c5 Add session handoff doc for 2026-05-15 modernization
Captures end-of-session state after Phases 1-2-4 merge: current
architecture, shipped commits, three documented open items
(retrieval-augmented routing test before Phase 5; non-atomic cache
write; log rotation), unscheduled-but-ready items (weekly digest + log
prune), telemetry collection window, roadmap status, and a pickup
prompt template for the next session.

Dated filename so future handoffs can coexist without overwriting.
2026-05-15 13:32:25 -04:00
Adam Moussa
db34d7e0ba
Merge pull request #3 from amoussa1229/chore/ci-workflow-sonnet-bump
Add CI workflow and bump Sonnet to 4.6
2026-05-15 13:21:43 -04:00
Adam Moussa
2bb8365972 Fix golden-set collection when provider keys are absent
CI's `pytest --collect-only` was exiting 5 (no tests collected) because
the test module used `pytest.skip(..., allow_module_level=True)` at
import time — skipped modules never enter the collection phase.

Switch to a per-test `skipif` marker driven by ANTHROPIC_API_KEY +
COMPOSIO_API_KEY presence, and lazy-import `graph.router_node` inside
the test body so module import works without COMPOSIO_API_KEY.

Result: 21 tests collect in both environments; live tests run only
when both keys are present. Verified locally with env -i.
2026-05-15 13:18:04 -04:00
Adam Moussa
3638d6ba96 Bump Sonnet model ID to claude-sonnet-4-6
claude-sonnet-4-20250514 reaches EOL 2026-06-15. Bumping ahead of the
deadline while the centralized models.py constant makes it a one-line
change. Smoke-tested via run.py --route-only — no deprecation warning,
router behavior unchanged.
2026-05-15 13:14:18 -04:00
Adam Moussa
fb5dbb20ee Add CI workflow and ignore .DS_Store
Closes the no-CI gap surfaced in the post-merge retrospective. The org
reusable workflows live under Sea-Haven-Industries and assume SAM/CDK
projects — orchestrator is a personal CLI tool with neither, so this is
a standalone workflow on the same conventions (actions/checkout@v6,
Python 3.12, ruff).

- lint job: ruff check + ruff format --check.
- test-collect job: installs requirements + runs `pytest --collect-only`.
  Catches import errors and golden-set test discovery regressions without
  needing live ANTHROPIC/COMPOSIO secrets — full pytest stays a local
  pre-push responsibility.

Also adds .DS_Store to the local .gitignore (also covered by the user
global gitignore, but belt-and-suspenders).
2026-05-15 13:13:41 -04:00
Adam Moussa
652eaaae17
Merge pull request #2 from amoussa1229/fix/code-review-followups
Address FIX items from PR #1 code review
2026-05-15 13:04:18 -04:00
Adam Moussa
5378bac69b Address code review FIX items from PR #1
Follow-up to the Phases 1-2-4 modernization PR.

- models.py: extend retry coverage to Gemini transient errors
  (google.api_core: ResourceExhausted, InternalServerError,
  ServiceUnavailable, DeadlineExceeded, GatewayTimeout). Before this,
  a 429 or 5xx from Google would crash the scanner route. Retriable
  count went from 6 to 11.
- retriever.py: batch cache-miss embeddings into a single
  embed_documents() call instead of one embed_query() per memory.
  Cold rebuild of 87 memories went from ~30s to ~10s; one HTTPS
  round-trip instead of 87.
- telemetry.py + scripts/weekly_summary.py: use UTC date for log
  filenames so they align with the UTC `timestamp` field inside each
  record. Eliminates the off-by-one near local midnight and matches
  Sea Haven's UTC-for-logs convention.
- scripts/weekly_summary.py: include failed runs in the total cost
  estimate (they consumed tokens too) and surface a `(incl. $X on
  failed runs)` breakdown so outages are visible in the digest.

Validated: ruff clean; weekly_summary.py still prints; retriever cold
rebuild + cache hit both verified end-to-end.
2026-05-15 13:01:14 -04:00
Adam Moussa
f6a4071464
Merge pull request #1 from amoussa1229/phase-1-stabilize-foundations
Orchestrator modernization — Phases 1, 2, 4
2026-05-15 12:14:51 -04:00
Adam Moussa
f05a98dba1 Add JSONL telemetry and weekly summary (Phase 4)
Observability slim layer. Every full run appends one JSON line to
~/.claude/logs/orchestrator/YYYY-MM-DD.jsonl with timestamp, sha256-prefix
task hash (raw task is never logged), retrieved memory names, router
choice, runtime, tokens in/out, success/error. risk_class and confidence
fields are reserved nulls for Phase 5.

- telemetry.py: log_run(), build_record(), task_hash(), token-usage
  extraction from AIMessage.usage_metadata. log_run swallows all
  exceptions — telemetry never kills a run.
- run.py: wraps app.invoke in try/except with a monotonic-clock window;
  logs on both success and failure. --route-only path is left unlogged
  (no agent work, doesn't represent a "run").
- scripts/weekly_summary.py: scans the last 7 days of JSONL and prints
  a markdown digest (routes, unknown rate, cross-review rate, success
  rate, total spend, mean tokens/route). Schedule via /schedule and
  pipe stdout to Slack from the scheduler.

Cost rates per route are rough Sonnet/Haiku/GPT/Gemini/DeepSeek defaults
suitable for spotting runaway prompts, not finance. Router tokens for
structured-output calls aren't captured (they don't surface through the
message trail); agent tokens are the dominant component anyway.

Validated: golden-set 21/21 still passing; full-run smoke writes
expected fields; weekly_summary.py prints clean markdown from a 2-run
log.
2026-05-15 11:49:05 -04:00
Adam Moussa
67180e124e Update README for Phases 1-2 architecture
Reflect retriever node, structured router with explicit unknown route,
AGENTS registry, model-ID constants, memory-aware CLI output, and the
golden-set test.
2026-05-15 11:37:11 -04:00
Adam Moussa
ac7101f3df Add memory retriever node (Phase 2)
Plugs the orchestrator into Adam's existing memory store at
~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/. Every
run starts with a top-3 retrieval pass that is then surfaced in the CLI
output and injected as system context into the router and downstream agent.

- retriever.py: load *.md memories (skipping the MEMORY.md index), embed
  with text-embedding-3-small, cache to .cache/embeddings.json keyed on
  file mtime. Cosine similarity, top-k=3 default. Reads only — never
  writes back to the memory store.
- state.py: add `retrieved: list[dict]` to OrchestratorState; relax to
  total=False to match LangGraph's partial-update semantics.
- graph.py: new retriever_node wired as START -> retriever -> router.
  router_node and connector_node now inject retrieved memories into their
  SystemMessage. Retrieval failures are caught and the run continues with
  empty memory context (logged).
- agents.py: make_agent_node injects retrieved memories into each agent's
  system prompt.
- run.py: prints `[retrieved: name1, name2, name3]` (or `[retrieved: none]`)
  before route/result for both --route-only and full-run modes, so bad
  retrieval is visible at a glance.
- .gitignore: add .cache/, .pytest_cache/, .ruff_cache/.

Validated: golden-set still 21/21 passing; smoke tests retrieve plausible
memories ("Send a Slack message to ops about the new exec-aide deploy" ->
project_exec_aide, feedback_exec_aide_vip_management, project_seahaven_slack_bot).
2026-05-15 11:36:03 -04:00
Adam Moussa
366d7247da Stabilize router and consolidate agent registry
Phase 1 stabilization. Removes the four-copy prompt/agent-description drift
surface and the silent router fallback.

- models.py: hoist model IDs to module-level constants; add with_retries()
  helper (2 retries on Anthropic+OpenAI transient errors via with_retry).
- agents.py: single AGENTS dict (model_fn, prompt, description) and a
  make_agent_node() factory that collapses six near-identical node functions.
- graph.py: router prompt is generated from AGENTS; router_node uses
  with_structured_output(RouteDecision) and returns an explicit "unknown"
  route instead of the silent "researcher" fallback. New unknown_node wires
  to END. All LLM invocations go through with_retries.
- state.py: add "unknown" to the route Literal.
- run.py: --route-only now imports router_node from graph.py, killing the
  fourth prompt copy.
- tests/: pytest golden-set (20 labelled tasks + size guard). Skips cleanly
  without ANTHROPIC_API_KEY or COMPOSIO_API_KEY. Validated 21/21 passing.
2026-05-15 11:20:28 -04:00
Adam Moussa
acfeb543d9 Document hybrid delegation model for Claude Code
README now explains when Claude Code delegates to the orchestrator
vs. handles tasks natively, matching the updated CLAUDE.md rules.
2026-05-08 16:41:46 -04:00
Adam Moussa
ca96b488e6 Add README and fix .env.example project name 2026-05-08 13:12:02 -04:00
Adam Moussa
b125956ee5 Add CLI entry point and fix Composio user_id
- Add run.py with --route-only flag for Claude Code integration
- Fix Composio user_id from "default" to "amoussa" to match connected accounts
2026-05-08 13:10:39 -04:00
Adam Moussa
02e29495c0 Initial LangGraph + Composio orchestration graph
Multi-model agent routing with 7 agent nodes (Sonnet, Haiku,
GPT-4.1, Gemini, DeepSeek) and 15 pre-loaded Composio tools
for Slack, Notion, GitHub, and Google Drive integration.
2026-05-08 12:54:00 -04:00