* feat(secrev): doc-drift Plane-1 Tier-1 checker (UNGATED)
Third Plane-1 checker on the Phase-0 shared substrate, mirroring
compliance-drift.sh / dependency-cve.sh conventions verbatim (set -euo pipefail,
sourced substrate, --canary/--dry-run/--no-api/--refresh/--targets, mode-600
reports under $REPORT_ROOT/doc-drift/<UTC-date>/, ALARM-only, finding.schema
spirit JSON, exit 0/2/3, dotgit->.git fixture trick).
Detects documentation drift deterministically (design §4 doc-drift row):
- readme-omits-component: README omits an existing major component in the tree
(top-level service dir, SAM/CDK stack, Lambda handler dir, openapi/docs spec)
- readme-stale-vs-code: README last-touch far older than newest code commit
(two-factor: >=DOC_DRIFT_STALE_DAYS AND >=DOC_DRIFT_STALE_COMMITS)
A repo with NO README is SKIPPED (compliance-drift owns readme-present; no
double-flag). Future Gemini large-context judge (§4) is an inert stub (maybe_judge),
off in canary/dry-run/offline.
Planted-drift fixture corpus + EXPECTED_DRIFT_COUNT=4, canary-asserted (exit 3 on
miss). shellcheck -x clean (only accepted SC1091 source-line info).
Does NOT touch checker_coordinator.sh, requirements.txt, or aws-posture.
Wiring/systemd is gated (PROVISIONING footer). Design refs §4, §7 Phase 3.
* feat(secrev): Phase-3 IAM artifacts for cross-review (aws-posture gated)
Authored FILES (not applied to AWS — provisioning gated behind the mandatory
GPT-4.1 IAM cross-review + Adam, design §7 B3) for the aws-posture checker's
read-only AWS identity. Decision D5: box stays read-only, auths via IAM Roles
Anywhere short-lived leaf certs from a new internal step-ca; NO long-lived AWS key.
- aws-posture-readonly-policy.json least-privilege read-only (ce:Get*,
cloudwatch:GetMetric*/DescribeAlarms, ec2/elb/rds:Describe*, lambda list +
GetFunctionConfiguration, s3:ListAllMyBuckets/GetBucketLocation). No write,
no iam:* mutation, no s3:GetObject/secrets/kms/logs data reads, no wildcard
actions. Resource:* only where AWS has no resource-level support.
- aws-posture-readonly-policy.rationale.md per-statement least-privilege rationale.
- aws-posture-trust-policy.json pins Roles Anywhere principal + leaf subject CN +
issuer CN + trust-anchor SourceArn (three conditions, all required).
- roles-anywhere-config.json trust anchor (pins step-ca root) + profile (1h session).
- step-ca-config-sketch.md internal CA config + systemd-timer leaf auto-renewal.
- CROSS-REVIEW-PACKET.md end-to-end trust model, blast radius, EXERCISED rollback,
reviewer scrutiny list.
Does NOT build aws-posture.sh, touch checker_coordinator.sh, or requirements.txt.
* fix(secrev): apply IAM cross-review FIXes
GPT-4.1 IAM cross-review 2026-06-18: APPROVE, no BLOCKs. Applied FIXes:
- trust policy: add aws:SourceAccount=328440206208 (confused-deputy guard)
alongside the existing aws:SourceArn trust-anchor pin
- readonly policy: remove ec2:DescribeImages (data minimization — AMIs are
not an idle-spend signal)
- aws:RequestedRegion NIT: deliberately SKIPPED — ce:* and s3:ListAllMyBuckets
are global-endpoint services a blanket region condition could DENY; rationale
recorded in aws-posture-readonly-policy.rationale.md
- rationale.md + CROSS-REVIEW-PACKET.md: record APPROVE + FIXes + NIT answers
(snapshots=account-owned idle signal; s3 list=names-only; no logs:* needed)
* feat(secrev): aws-posture checker (Tier-2, provisioning-gated)
Read-only Tier-2 idle/anomalous-spend + idle-resource posture checker for the
R720 agent-team (design D5 / §4 / §6.3 / §7 Phase 3). Mirrors the Tier-1 checker
conventions verbatim (flags --canary/--dry-run/--no-api/--targets, mode-600
report under $REPORT_ROOT/aws-posture/<date>/, ALARM-only, finding.schema.json
spirit, exit 0/2/3, shared substrate redact/post_slack_alarm).
Detectors (complement GuardDuty/SecurityHub/Config, do not replace):
- anomalous Cost Explorer deltas (ce get-anomalies, $-impact threshold)
- stopped EC2 still paying for attached EBS
- unattached EBS volumes
- unassociated Elastic IPs
- idle NAT gateways (≈0 bytes out)
- idle load balancers (0 healthy targets)
- idle RDS (0 connections over window)
Live AWS calls are PROVISIONING-GATED: they run ONLY when Roles Anywhere creds
are available (STS identity probe) AND not --no-api/--canary. With no creds or
--no-api/--canary the checker SKIPS live calls and notes them — NEVER alarms on
missing data (memory feedback_cloudwatch_alarms). Roles Anywhere/step-ca are not
stood up (IAM cross-review PASSED 2026-06-18; see security-review/iam/).
Offline canary: fixtures of mocked AWS responses (cost/describe-* JSON) under
fixtures/aws-posture/ + EXPECTED_FINDING_COUNT=7, asserted fully offline (no aws,
no network). Identical detector code runs online and offline. shellcheck-clean
(only accepted SC1091), chmod +x.
* fix(secrev): doc-drift fixture py ruff-clean (root CI runs check + format --check)
The repo-root CI lint runs both 'ruff check .' and 'ruff format --check .' over
all fixtures. Fixed E701 one-liners and ruff-formatted the sample-service .py
files (handlers/*, feature_*.py). Fixture content is irrelevant to doc-drift
(keys on file/dir presence + git staleness).
|
||
|---|---|---|
| .github | ||
| agent-team | ||
| docs | ||
| scripts | ||
| security-review | ||
| tests | ||
| .env.example | ||
| .gitignore | ||
| agents.py | ||
| conftest.py | ||
| graph.py | ||
| models.py | ||
| README.md | ||
| requirements.txt | ||
| retriever.py | ||
| run.py | ||
| state.py | ||
| telemetry.py | ||
| tools.py | ||
orchestrator
Multi-model AI agent orchestration via LangGraph + Composio. Routes tasks to the best-fit model and connects to external services (Slack, Notion, GitHub, Google Drive). Memory-aware — each run is enriched with the top-3 most relevant notes from Adam's project/feedback/reference memory store.
Architecture
Claude Code ──► run.py ──► LangGraph StateGraph
│
▼
retriever ──► top-3 memories from
│ ~/.claude/projects/.../memory/
▼
router (Sonnet, structured output)
│
┌─────────────┼─────────────────────┐
▼ ▼ ▼
┌──────────────┐ ┌───────────┐ ┌──────────┐
│ implementer │ │ connector │ │ unknown │
│ reviewer │ │ (Composio)│ │ (no fit) │
│ researcher │ └───────────┘ └──────────┘
│ cross_reviewer│ │
│ scanner │ ▼
│ fast_coder │ tool_executor ──► summarizer
└──────────────┘
The retriever embeds Adam's memory files once and caches vectors to .cache/embeddings.json (mtime-keyed; only changed files re-embed). Each run picks the top-3 most relevant memories and surfaces them in the CLI output before the route line.
The router uses Pydantic structured output (RouteDecision) and returns an explicit "unknown" route when no agent fits — no silent fallback. All LLM invocations are wrapped with retry-on-transient-error.
Files
| File | Purpose |
|---|---|
run.py |
CLI entry point — python3 run.py "<task>" |
graph.py |
LangGraph graph: retriever, router, connector, summarizer, unknown nodes |
agents.py |
AGENTS registry (label → model_fn, prompt, description) + make_agent_node factory |
models.py |
LLM factories, model-ID constants, with_retries() helper |
state.py |
OrchestratorState TypedDict |
retriever.py |
Memory loader, embedder, cache, top-k retrieval |
tools.py |
Composio tool loading (Slack, Notion, GitHub, Google Drive) |
tests/test_routing_golden.py |
20-case golden-set regression test for the router |
Usage
# Full execution — retrieves memory, routes, and runs the task
python3 run.py "What is the LangGraph checkpoint API?"
# Route-only — retrieves memory and prints the agent that would handle the task
python3 run.py --route-only "Review this code for security issues"
Output shape:
[retrieved: project_seahaven_slack_bot, feedback_secrets_manager, reference_sea_haven_aws]
[reviewer]
<agent output>
From Claude Code (via CLAUDE.md hybrid delegation):
python3 ~/Documents/repositories/orchestrator/run.py "<task description>"
python3 ~/Documents/repositories/orchestrator/run.py --route-only "<task description>"
When Claude Code delegates vs. handles natively
Claude Code uses a hybrid model — it delegates to the orchestrator when a different model has a genuine advantage, and handles everything else natively:
| Delegate to orchestrator | Handle natively in Claude Code |
|---|---|
| Cross-family code review (GPT-4.1) | File editing, refactoring, bug fixes |
| Large codebase scanning (Gemini) | Git operations, PRs, merges |
| Quick bounded coding (DeepSeek) | AWS/SAM/CDK deployments |
| External service actions (Composio) | Shell commands, system admin |
| Interactive planning and conversation |
Agents
| Agent | Model | Use Case |
|---|---|---|
| implementer | Claude Sonnet | Write code with a clear spec |
| reviewer | Claude Sonnet | Code review (BLOCK/FIX/NIT/QUESTION) |
| researcher | Claude Haiku | Doc lookups, API research |
| cross_reviewer | GPT-4.1 | Independent second-opinion review |
| scanner | Gemini 2.5 Pro | Large codebase analysis |
| fast_coder | DeepSeek Coder | Quick, bounded coding tasks |
| connector | Sonnet + Composio | Slack, Notion, GitHub, Google Drive |
The router can also return done (no agent needed) or unknown (no clear fit). Model IDs are centralized as constants in models.py.
Memory retrieval
The retriever reads ~/.claude/projects/-Users-adammoussa-Documents-repositories/memory/*.md (skipping the MEMORY.md index), embeds each file once with text-embedding-3-small, and caches the vectors to .cache/embeddings.json. On subsequent runs:
- Only files whose mtime changed are re-embedded.
- Top-3 memories by cosine similarity are injected as system context into both the router and the agent.
- Retrieved names are printed as the first line of every run so bad retrieval is visible.
- Retrieval is read-only. The orchestrator never writes back to the memory store.
If retrieval fails (network, missing key), the run continues with no memory context and logs the failure into the message trail.
Connectors (via Composio)
All connections authenticated under Composio user amoussa:
- Slack: send messages, read channels/threads, find users, add reactions
- Notion: search/read/create/update pages, add content
- GitHub: create issues, list issues, get repo info
- Google Drive: find files, get metadata
The connector node is restricted to one tool call per run — a load-bearing rule learned from a 1.9M-token incident with meta-tool routing.
Security Review
The security-review/ subsystem is a high-recall, anti-complacency security gate. It is separate from the router — it does not route through run.py or LangGraph. One pure-code script, review.sh, owns the block decision (confirmed critical/high → block); no agent decides.
- Path A — interactive: the
/sh-security-reviewClaude Code skill (Max-covered). Narrow fresh-context detector fan-out + a proof-or-kill verifier; emits the structured finding schema forreview.shto gate. - Path B — unattended: a nightly two-tier sweep on the
sh-secrevR720 VM. Tier 1 runs deterministic scanners (review.sh --scanners-only) over every Sea-Haven-Industries org repo; Tier 2 is a budget-bounded agentic pass (run_headless.py) on a round-robin rotation. Clean-clone auto-discovery via a read-only GitHub PAT; ALARM-only Slack (a clean night posts nothing). - Git hooks: global pre-commit / pre-push hooks (
install-hooks.sh --global) gate every local repo viareview.sh --scanners-only.
See security-review/README.md for full detail and security-review/DEPLOY-R720.md for the VM runbook.
Setup
- Install dependencies:
pip install -r requirements.txt - Copy
.env.exampleto.envand fill in API keys - Authenticate Composio integrations at app.composio.dev
Configuration
All API keys are stored in .env (gitignored):
ANTHROPIC_API_KEY— Claude models + routerOPENAI_API_KEY— GPT-4.1 cross-reviewer + text-embedding-3-smallGOOGLE_API_KEY— Gemini scannerDEEPSEEK_API_KEY— DeepSeek fast-coderCOMPOSIO_API_KEY— Composio connectorsLANGSMITH_API_KEY— LangSmith tracing
Tracing is enabled via LangSmith (project: orchestration).
Testing
pytest tests/test_routing_golden.py -v
20 labelled tasks → expected agent. Skipped cleanly if ANTHROPIC_API_KEY or COMPOSIO_API_KEY are unset.