commit 3e2d99a1eb7a633106f7d17cbfe0a4e8780b3731 Author: Adam Moussa Date: Tue Jun 9 19:25:24 2026 -0400 Initial commit: sh-mcp design and plan Design doc for the Sea Haven MCP platform that replaces seahaven-slack-bot and exec-aide: trust-tiered MCP servers, Google-SSO via a Cognito broker, TypeScript monorepo. Cross-review (cross_reviewer/GPT-4.1) folded in. Status: planning only, nothing built. diff --git a/.gitignore b/.gitignore new file mode 100644 index 0000000..61740ab --- /dev/null +++ b/.gitignore @@ -0,0 +1,23 @@ +# Node / TypeScript +node_modules/ +dist/ +build/ +*.tsbuildinfo + +# CDK +cdk.out/ +.cdk.staging/ + +# Env / secrets +.env +.env.* +*.local + +# Logs / OS +*.log +npm-debug.log* +.DS_Store + +# Test / coverage +coverage/ +.nyc_output/ diff --git a/README.md b/README.md new file mode 100644 index 0000000..30512b8 --- /dev/null +++ b/README.md @@ -0,0 +1,40 @@ +# sh-mcp + +Sea Haven MCP platform. A TypeScript monorepo of trust-tiered **MCP servers** that expose +Sea Haven's proprietary integrations as tools, plus the **Cognito/Google auth broker** and the +**rebuilt scheduled jobs**. This service replaces `seahaven-slack-bot` and `exec-aide`, which are +deprecated completely; the conversational surface becomes a configurable Slack task agent. + +> **Status: DESIGN / PLANNING. Not built. No stack deployed.** +> The full design, auth architecture, scope matrix, and build plan live in +> [`docs/design.md`](docs/design.md). Read it before writing any code. + +## Shape (planned) + +- **MCP servers** (trust-tiered, remote HTTP, per-server IAM): + - `sh-mcp-ops` — read-mostly, agent-facing (WO/PO/site lookups, KB search, Google Maps, + Gmail/Calendar, tasks, reminders). + - `sh-mcp-finance` — sensitive, read-only, audited (QBO vendor search, payment lookups). + - `sh-mcp-physical` — DEFERRED, admin/out-of-band only (Lenel/Yealink/3CX control). +- **Auth** — Google Workspace is the single IdP; an Amazon Cognito user pool federated to Google + issues scoped, audience-bound JWTs; group → scope mapping via a pre-token Lambda. See design §2. +- **Jobs** — rebuilt proactive Lambdas (email classify/digest, KB syncs). +- **Language** — TypeScript everywhere (servers, packages, CDK, jobs). + +## Open decisions + +- Task-agent surface: Agentforce (recommended) vs marketplace Claude app vs custom Bolt assistant + (design §12). Drives the model + guardrail story. +- Endpoint exposure specifics (Slack egress ranges / WAF) — design §9 / §11. + +## Layout (target) + +``` +packages/ shared + one package per integration +servers/ sh-mcp-ops, sh-mcp-finance (CDK stacks) +auth/ cognito, pre-token-lambda, group-sync-lambda +jobs/ rebuilt scheduled Lambdas +docs/ design.md (the canonical plan) +``` + +See [`docs/design.md`](docs/design.md) for the authoritative spec. diff --git a/docs/design.md b/docs/design.md new file mode 100644 index 0000000..8aa978e --- /dev/null +++ b/docs/design.md @@ -0,0 +1,519 @@ +# Sea Haven MCP Platform — Auth Architecture, Scope Matrix & Build Plan + +Status: DRAFT for review (cross-review gate not yet run) +Date: 2026-06-09 +Owner: Adam Moussa + +Supersedes the ad-hoc design discussion. This is the consolidated plan for deprecating +`seahaven-slack-bot` and `exec-aide` in favor of Slack AI agents backed by a set of +Sea Haven MCP servers, with authentication centralized on Google Workspace SSO. + +--- + +## 1. Decisions locked in + +This is a **greenfield service**. `seahaven-slack-bot` and `exec-aide` are deprecated +**completely** (repos archived, stacks decommissioned, Bedrock agents + guardrails torn down). +Nothing is migrated as a running component. Where proactive functionality is still needed +(Lauren's email classifier/digest, the KB syncs) it is **rebuilt fresh** in this new service's +`jobs/` (§5), not carried over. + +1. **Conversational surface = a configurable Slack task agent** that users @mention (in channels, + threads, DMs) and that calls our MCP tools to do work end to end. Slack hosts the agent, the + LLM, and the orchestration; we run no Bolt backend of our own (unless we pick the custom-assistant + option). Our system is: MCP servers (the tools) + the Cognito/Google auth broker + the rebuilt jobs. + IMPORTANT (corrected, §12): this is NOT "native Slack AI." Native Slack AI is the summarize / + recap / search feature suite over Slack's own content; it does not act as a configurable agent + that calls our tools, so it cannot replace the bots. The real surface candidates are **Agentforce** + (recommended), a **marketplace agent** (e.g. the Claude app), or a **custom Bolt assistant**. + Surface choice is still OPEN (§12) and it drives the model + guardrail story. + Guardrail consequence (investigated, §11): Slack's native guardrails cover prompt-injection and + content safety, but do NOT redact PII inside our tool responses, so we rebuild PII masking at the + MCP layer (§2.5) regardless of surface. (If we pick Agentforce + BYOLLM-on-Bedrock, we could also + keep our own Bedrock guardrail on the model path — defense in depth, not a substitute.) +2. **Tools = Sea Haven MCP servers**, grouped by trust tier (`ops`, `finance`; `physical` deferred). +3. **Identity = Google Workspace**, the same IdP already used for Slack SSO. Google is the + single source of truth for authentication AND group membership. +4. **Authorization = Amazon Cognito** user pool federated to Google, acting as the OAuth 2.1 + authorization server that issues scoped, audience-bound tokens the MCP servers validate. + Slack's built-in agent is the MCP client; each user does per-user OAuth consent against the + broker (Slack's MCP client supports per-user OAuth), so tool calls carry the real user identity. +5. **Language = TypeScript** for the entire monorepo (servers, packages, CDK, rebuilt jobs). + Packages are shared by both the MCP servers and the rebuilt scheduled jobs (one client per + service everywhere); a single language is what makes that sharing real. + +--- + +## 2. Auth architecture + +### 2.1 Planes + +Three planes kept separate (the credential that reaches the broker is never the credential +that reaches a downstream API): + +- **Inbound auth (who):** Cognito-issued JWT, identity federated from Google. Carries `sub` + (the user's Google identity), `aud` (the target MCP server), and `scope` claims. +- **Access control (what):** each MCP server validates issuer + audience and enforces a + required scope per tool. Tools whose scope the token lacks are hidden from the user. +- **Outbound auth (how we reach the API):** + - **Google-API tools** (Gmail, Calendar): act *as the user* via a direct per-user Google + OAuth grant (see 2.4). + - **Non-Google tools** (QBO, payments, Lenel, 3CX, Afi, Front, Maps): broker token + authorizes; the server uses its own service credentials from Secrets Manager. + +### 2.2 Flow + +``` +User (already Google-SSO'd into Slack) + │ OAuth 2.1 auth-code (PKCE), federated login → Google (silent if logged in) + ▼ +Cognito user pool (Google = external OIDC IdP) + │ pre-token-generation Lambda: group → scope mapping + ▼ +Access token (JWT): + sub = lauren@seahavenind.com + aud = sh-mcp-ops + scope = "ops:read ops:tasks gmail:self calendar:self" + │ + ├──► sh-mcp-ops validates aud+scope, runs permitted tools + ├──► sh-mcp-finance token lacks finance:read → finance tools hidden + └──► sh-mcp-physical token lacks physical:* → hidden +``` + +### 2.3 Group → scope mapping + +Google Groups are the single control point. Membership is synced to Cognito groups by a +scheduled sync Lambda (every 5 min, tightened from hourly per cross-review — group changes are +infrequent but revocation latency on sensitive groups must be low). The sync Lambda has a +CloudWatch ALARM on failure (a silent sync failure freezes group state). The Cognito pre-token +Lambda maps Cognito group → scope claims at token mint, so the hot path makes no Directory API +call. (Alternative: pre-token Lambda calls the Google Directory API with a short cache. Chose +the sync approach to keep token minting fast and resilient to Directory API throttling.) + +Revocation: a 5-min sync + 60-min token TTL means a removed user could retain access for ~65 min +worst case. For `finance:*`, TTL is shortened to 15 min, and a Cognito-backed deny-list (checked +by the servers) gives immediate hard revocation when a user is pulled from a sensitive group. + +| Google Group | Members | Scopes granted | +|------------------------|--------------------------------|-------------------------------------------------| +| `sh-mcp-ops@` | all staff | `ops:read` | +| `sh-mcp-assistant@` | Lauren, Adam | `ops:read`, `ops:tasks`, `gmail:self`, `calendar:self` | +| `sh-mcp-finance@` | Adam, Lauren, accounting | `ops:read`, `finance:read` | +| `sh-mcp-admin@` | Adam | all of the above + `finance:admin` | + +A user in multiple groups gets the union of scopes (Lauren is in `-assistant@` and `-finance@`, +so she gets `ops:read ops:tasks gmail:self calendar:self finance:read`). + +`sh-mcp-physical@` / `physical:*` scopes are NOT issued to any agent surface at launch — the +physical tier is admin/out-of-band only (see §3 and §6). + +### 2.4 Google-API tools (Gmail / Calendar) — outbound detail + +Pulling the upstream Google API access token *through* Cognito federation is limited (Cognito +requests `openid profile email` and does not cleanly expose/refresh Google API scopes like +`gmail.readonly`). So the clean implementation is: + +- Cognito-federated-Google handles **authentication + authorization scopes** (single SSO login). +- The Gmail/Calendar MCP tools do a **separate direct per-user Google OAuth grant** at first + use (incremental consent for `gmail.readonly` + `calendar`), storing a per-user refresh + token encrypted (Secrets Manager or a KMS-encrypted DynamoDB item, key = user `sub`). +- Both consents resolve to the same Google account, so identity stays consistent; it is one + extra one-time consent, not a second identity. + +Result: `search_inbox` / `get_calendar_events` act as the signed-in user. Lauren physically +cannot read anyone else's mailbox; segregation is enforced by Google, not by our code. + +Token-handling rules (cross-review BLOCK items): +- The Google client is instantiated with the **minimal scope for the requested tool only** + (`gmail:self` → `gmail.readonly`, `calendar:self` → `calendar`), never the union of a user's + scopes. A `gmail:self`-only user cannot mint a Calendar token and vice versa. +- The per-user refresh token is **never** returned to the Slack agent or any client, and never + appears in logs or error messages. Access tokens are minted per request, not cached or reused. +- Per-user refresh tokens are stored in a KMS-CMK-encrypted DynamoDB item keyed by user `sub`, + with the key policy scoped to the owning server role only. IAM on that table is partitioned by + `sub` (ABAC `dynamodb:LeadingKeys`) so a server can only read the calling user's token, not all + users' tokens. + +### 2.5 Token & endpoint hardening + +- Short-lived access tokens (60 min general; **15 min for `finance:*`**); refresh by the client. +- **Audience-bound**: each server validates `aud` and rejects tokens not minted for its resource + server. An `ops` token presented to `finance` is rejected. +- **Server-side enforcement is authoritative**: every server independently validates issuer + + `aud` + required scope per tool on every call. Tool-hiding in the agent UI is a convenience, + never the access-control boundary. The client is never trusted. +- **No broker-JWT passthrough**: the Cognito JWT is never forwarded to a downstream API. Servers + reach external APIs only via their own service credentials or the per-user Google token (§2.4). +- **Endpoints are PUBLIC HTTPS** (decided — see §11). Slack's agent runtime is cloud-hosted and + initiates outbound Streamable-HTTP to a public MCP URL with an OAuth bearer, so VPC-only / + PrivateLink is not an option. The OAuth bearer (Cognito) is the real access boundary; in front + of it: AWS WAF, plus IP-allowlisting to Slack's egress ranges IF Slack publishes them (to + verify). Do not depend on mTLS — Slack's MCP client is not documented to present client certs. + Never an unauthenticated path. PKCE required if any public client is added. +- **MCP-layer PII redaction** (decided — see §11): Slack's native guardrails do NOT redact PII + inside our tool responses, so the Bedrock guardrail's PII-anonymization is reimplemented in the + servers' output path. Every tool response is inspected and sensitive fields (bank account/ + routing, card, SSN) masked before it leaves the server, heaviest on `finance`. +- **Physical tier step-up**: `physical:*` tools require recent re-auth (`auth_time` / `max_age`) + AND an out-of-band human approval before execution. Scope grants the right to *request*, never + to execute unattended. +- **Audit**: every `finance:*` and `physical:*` tool call logged to CloudWatch (user `sub`, + tool, args hash, decision, result). CloudWatch ALARM on any `physical:*` invocation + (ALARM-state only, per standing alarm preference). +- Connection-level governance: which MCP connectors a user can even add is also gated by group, + so `physical` is unreachable for non-members regardless of token contents (defense in depth). +- **IAM least-privilege per server** (cross-review BLOCK): each server role gets + `secretsmanager:GetSecretValue` on only its own secrets (no wildcard) and KMS decrypt on only + the relevant CMK. `sh-mcp-finance` can read the QBO + payments secrets but not Maps or per-user + Google tokens; `sh-mcp-ops` cannot read finance secrets. Access to QBO and per-user Google + tokens is logged and alarmed on anomalous patterns (mass/unexpected-role access). +- **Prompt-injection containment** (no Bedrock guardrail to fall back on): tool-call parameters + are validated server-side against strict schemas/allow-lists, never trusted from the agent. + Tool-returned content is treated as data, never instructions. Side-effecting `ops` tools + (`create_calendar_event`, `create_task`, `create_reminder`) get sane bounds; the trust-tier + split keeps any injection blast radius inside `ops` (no finance/physical reachable in-session). + +--- + +## 3. Scope matrix (every tool → required scope, outbound auth, risk) + +### sh-mcp-ops (agent-facing, read-mostly) + +| Tool | Source pkg | Scope | Outbound auth | Risk | +|--------------------------|-----------------|----------------|------------------------|------| +| lookup_work_order | internal-data | `ops:read` | service creds (DDB ro) | low | +| lookup_purchase_order | internal-data | `ops:read` | service creds (DDB ro) | low | +| lookup_site | internal-data | `ops:read` | service creds (DDB ro) | low | +| search_knowledge_base | knowledge-base | `ops:read` | service (Bedrock Retrieve) | low | +| search_nearby_vendors | google-maps | `ops:read` | service (Maps API key) | low (external $, rate-limit) | +| search_inbox | gmail | `gmail:self` | forwarded user Google token | med (private data) | +| get_email_thread_detail | gmail | `gmail:self` | forwarded user Google token | med | +| get_calendar_events | calendar | `calendar:self`| forwarded user Google token | low | +| create_calendar_event | calendar | `calendar:self`| forwarded user Google token | med (writes, external attendees) | +| check_availability | calendar | `calendar:self`| forwarded user Google token | low | +| create_task | tasks | `ops:tasks` | service creds (DDB, partitioned by sub) | low | +| list_tasks | tasks | `ops:tasks` | service creds | low | +| complete_task | tasks | `ops:tasks` | service creds | low | +| delete_task | tasks | `ops:tasks` | service creds | low | +| create_reminder | reminders | `ops:tasks` | service creds (EventBridge Scheduler) | low | + +### sh-mcp-finance (sensitive, read-only, fully audited) + +| Tool | Source pkg | Scope | Outbound auth | Risk | +|---------------------------|------------|----------------|------------------------------|------| +| search_vendors | qbo | `finance:read` | service creds (QBO OAuth, server-held) | med (token can read all of QBO) | +| lookup_payment_by_vendor | payments | `finance:read` | service creds (DDB PaymentsDashboard) | med | +| lookup_payment_by_invoice | payments | `finance:read` | service creds (DDB) | med | +| lookup_payment_by_check | payments | `finance:read` | service creds (DDB) | med | + +QBO OAuth maintenance (`/qbo/connect`, `/qbo/callback`, `/qbo/disconnect`) stays as +**admin web endpoints** (API Gateway), NOT exposed as agent tools. Gated by `finance:admin`. + +### sh-mcp-physical (DECIDED: admin / out-of-band only — NOT in the launch build) + +Neither bot does physical actions today, so this tier is documented to keep the trust model +complete but is DEFERRED. No `physical:*` scope is issued to any agent. The tools below stay as +admin/out-of-band capability (existing door-unlock-api path) until a future decision to expose them. + +| Tool | Source pkg | Scope | Outbound auth | Risk | +|-----------------------|------------|----------------------|-----------------------|------| +| request_door_unlock | lenel | `physical:unlock` | service creds (Elements API) | HIGH — human approval required | +| initiate_lockdown | lenel | `physical:lockdown` | service creds | HIGH — Adam only + approval | +| release_lockdown | lenel | `physical:lockdown` | service creds | HIGH — Adam only + approval | +| push_xml | yealink | `physical:telephony` | service creds (Push XML) | med + approval | +| reroute_call | threecx | `physical:telephony` | service creds (3CX xapi) | med + approval | +| set_dnd | threecx | `physical:telephony` | service creds (3CX xapi) | med + approval | + +### Trifecta rationale + +The dangerous combination (read-from-untrusted-source + sensitive-action in one session) is +prevented by tier separation: `finance` and `physical` are never reachable by a session that +also holds the Gmail/web read tools. Within `ops`, the actions (`create_task`, `create_reminder`, +`create_calendar_event`, `search_nearby_vendors`) are low-consequence, so injected content in an +email can at worst create a spurious task or run a Maps query — not move money or unlock a door. +Outbound items in `ops` (calendar invites to external attendees) are flagged for monitoring. + +--- + +## 4. Monorepo layout + +Packages are consumed by BOTH the MCP servers and the surviving scheduled Lambdas — one client +per external service, used everywhere. This is the core payoff over per-service repos. + +``` +sh-mcp/ (one monorepo, org: Sea-Haven-Industries) + packages/ + shared/ MCP scaffold, JWT validation, requires_scope guard, secrets, audit log + qbo/ google-maps/ internal-data/ payments/ knowledge-base/ + gmail/ calendar/ tasks/ reminders/ notion/ + # DEFERRED (physical tier, not in launch build): lenel/ yealink/ threecx/ + servers/ + sh-mcp-ops/ → remote HTTP MCP stack (CDK) + sh-mcp-finance/ → remote HTTP MCP stack (CDK) + # DEFERRED: sh-mcp-physical/ (admin/out-of-band only) + auth/ + cognito/ user pool, Google federation, resource servers, app clients + pre-token-lambda/ group → scope mapping + group-sync-lambda/ Google Groups → Cognito groups (hourly) + jobs/ (surviving scheduled Lambdas, import packages/) + notion-sync/ po-sync/ workorder-sync/ (feed Bedrock KB) + fetch-classify/ daily-digest/ reminder/ (exec-aide proactive, Lauren) +``` + +--- + +## 5. Proactive jobs — rebuilt fresh in `sh-mcp/jobs` (NOT MCP, NOT migrated) + +Proactive/event-driven work has no conversational equivalent and Slack's reactive agent cannot +replace it, so it is **rebuilt** in the new service as TypeScript Lambdas that import the shared +packages and use stored offline credentials (not interactive SSO). The old exec-aide/slack-bot +Lambdas are decommissioned, not refactored in place. + +| Job | Trigger | Notes | +|--------------------|----------------|-------------------------------------------------------------| +| fetch-classify | every 15 min | Gmail history → Haiku classify → HIGH DM. Offline Google refresh token (`exec-aide/gmail-oauth`). | +| daily-digest | 5pm ET M-F | Query DDB → digest → DM. | +| reminder (one-shot)| Scheduler | Fired by the `create_reminder` tool. | +| notion-sync | 02:00 UTC | Notion → S3 → Bedrock KB ingestion. | +| po-sync | 02:00 UTC | DDB purchase-orders → S3 → KB. | +| workorder-sync | 02:00 UTC | DDB work orders/comments → S3 → KB. | + +The Bedrock **agents** (`seahaven-alex`, exec-aide Sonnet loop) are what get deprecated. The +Bedrock **guardrail** (PII anonymization, prompt-attack filtering) must be explicitly replaced +on the Slack surface before deletion — do not assume the new surface provides parity. + +--- + +## 6. Phasing + +1. **Foundation:** monorepo + `shared` + CI/CD (reusable workflows, deploy role first, OIDC, + Dependabot, naming) + the test harness and coverage gate (see §7). Stand up Cognito + Google + federation + pre-token + group-sync. Prove SSO login end to end. +2. **Pilot — `sh-mcp-ops`** wired to the `seahaven-slack-bot` replacement agent: vendor + + WO/PO/site + KB. Validate per-user auth, tool-hiding, and parity vs the old Bedrock agent. +3. **`sh-mcp-finance`** with full audit logging once ops is proven. +4. **Convert Lauren → Slack AI agent:** add `gmail:self`/`calendar:self`/`ops:tasks` tools to + ops, do the per-user Google grant, keep her agent DM/personal-scoped. Refactor her scheduled + Lambdas to import the packages. +5. **`sh-mcp-physical`** is DEFERRED — admin/out-of-band only, not part of this build. +6. **Deprecate each bot only after its replacement is proven at parity.** Lauren's workflow + needs sign-off before `exec-aide` is retired. + +--- + +## 7. CI/CD & test suite + +### 7.0 Language: TypeScript (DECIDED) + +TypeScript for everything: servers, packages, CDK, and the rebuilt jobs. Toolchain is +`tsc --noEmit` + eslint/prettier + vitest. Since the old stacks are fully deprecated (not +migrated), exec-aide's former Python logic (Gmail/Calendar/Haiku classify) is rewritten in TS +as part of the greenfield build, not ported line-for-line. No polyglot matrix needed. + +### 7.1 Reusable-workflow pattern (Sea Haven standard) + +Per the handbook `cicd.md`, repos call org reusable workflows rather than inlining steps. Two +caller workflows: + +- `.github/workflows/ci.yaml` — on `pull_request`. Calls the org reusable `ci` workflow. + Aggregated under a single required check named `ci / ci` for branch protection (so adding + matrix legs never breaks the required-context list). +- `.github/workflows/deploy.yaml` — on push to `main` (deploy-then-merge per git-workflow). + Calls the org reusable `cd-cdk` workflow per server, OIDC into the deploy role. + +Gotchas baked in from prior repos: +- `permissions:` declared in the CALLER (reusable-workflow permissions don't inherit — missing + them yields `startup_failure`). +- `concurrency` keyed to caller context, not the reusable workflow, to avoid cross-PR collision. +- ARM64 Lambda containers: `enable-qemu: true` in BOTH ci.yaml and deploy.yaml, AND + `platform: LINUX_ARM64` on every Docker image asset (Lambda) or "exec format error". +- `cd-cdk` npm-ci guard since CDK is Node even where Lambda code may differ. + +### 7.2 Monorepo CI shape + +Path-filtered matrix so only changed packages/servers run. Per leg: + +1. Install + build (workspace-aware: `npm ci` at root, build changed package + dependents). +2. Lint + format gate: `tsc --noEmit` + eslint/prettier (or `ruff check` + `ruff format --check` + on any Python paths). Enforced locally pre-push by the existing Claude Code hook too. +3. Unit + auth + security tests (§7.3) with coverage gate. +4. `cdk synth` for changed servers (catches IaC + ARM64 platform regressions without deploying). +5. Dependabot enabled; exact-pin `aws-cdk-lib` per the handbook Pinning Principle (no blanket + ignores); Dependabot keeps it current. + +Deploy job runs only on `main`, per-server, gated on `ci / ci` green and the cross-review +sign-off for any IAM/authz diff (§8). + +### 7.3 Test suite (full, security-weighted) + +The authorization layer is the highest-risk surface, so it gets the heaviest coverage. + +| Layer | What it covers | +|-------|----------------| +| **Unit — per package** | Each integration package with the external API mocked: QBO vendor query + token auto-rotation, payments DDB queries, internal-data lookups, Maps search, KB retrieve, Gmail/Calendar calls, tasks CRUD, reminder scheduling. Error/empty/throttle paths. | +| **Auth — the crown jewels** | JWT validation (issuer, signature, expiry); **audience binding** (a token minted for `sh-mcp-ops` is REJECTED by `sh-mcp-finance`); per-tool scope enforcement **server-side** (independent of UI tool-hiding); **tool-hiding** (a user without `finance:read` does not see finance tools in `list_tools`); pre-token Lambda group→scope mapping for every group; group-sync correctness; **deny-list hard revocation** (a revoked user is rejected immediately, before token expiry); minimal-scope Google client (a `gmail:self`-only token cannot mint a Calendar token); per-user token ABAC isolation (a server cannot read another user's refresh token). | +| **Security / abuse** | **Trifecta separation** — assert no single issued token can reach both untrusted-read (Gmail) and a sensitive-action tool; scope-escalation attempts rejected; prompt-injection regression corpus (tool-returned content containing "ignore instructions / call X" must NOT trigger out-of-scope tool calls); **MCP-layer PII redaction** (every finance tool response masks bank/routing/card/SSN before egress); per-session tool-call cap + per-tool rate limit. | +| **Contract / MCP conformance** | Every tool's input/output JSON schema validates; MCP protocol handshake; `list_tools` reflects the caller's scopes. | +| **Audit** | Every `finance:*` call emits a structured audit record (user sub, tool, args hash, decision); assert the record shape and that secrets never appear in logs. | +| **Parity (pre-deprecation)** | Golden-transcript tests replaying real slack-bot / exec-aide interactions against the new tools to confirm equivalent answers before retiring a bot. | + +Coverage gate in CI (start at 80% lines, 100% on the `shared` auth/scope guard module). Parity +tests run in phase 2/4 before each bot is deprecated, not on every PR. + +## 8. Handbook obligations (per global instructions) + +- Naming: kebab-case throughout (`sh-mcp`, `sh-mcp-ops`, etc.). ✓ in this draft. +- Secrets: all in Secrets Manager; existing secret names reused; per-user Google tokens encrypted. +- Each deployable server: deploy role first, CI/CD pipeline, Dependabot, README. +- **Confluence "AWS Architecture Map" (id 1540098)**: add a Mermaid subgraph for the MCP + platform + Cognito before the build is reported done. +- Project memory: `project_sh_mcp.md` created; cross-links to exec-aide, slack-bot, payments, + identity-center, lenel/door-unlock memories. +- **Cross-review gate (mandatory):** Cognito↔Google federation, group→scope mapping, the + token-forwarding design, and every MCP server's IAM role go through `cross_reviewer` before + any commit/merge. IAM + auth = the breaking-change category that gates. + +--- + +## 9. Decisions log & remaining open items + +RESOLVED 2026-06-09: +- Complete deprecation: slack-bot + exec-aide repos archived, stacks decommissioned, no migration. + Greenfield service; conversational surface = Slack's built-in AI agent (no Bolt backend, no own model). +- Lauren gets `finance:read` (member of `-assistant@` + `-finance@`). +- `sh-mcp-physical` is admin/out-of-band only, deferred, not in the launch build. +- Afi removed from the stack entirely. +- Launch tool scope = exactly what the two bots do today (no Front / Afi / 3CX-read adds). +- Language = TypeScript for the entire monorepo. + +INVESTIGATED 2026-06-09 (both former open items — see §11): +- Endpoint model: PUBLIC HTTPS + OAuth + WAF (VPC/PrivateLink not viable; mTLS not relied on). +- Guardrail parity: Slack covers prompt-injection/content safety; PII redaction is NOT covered and + is rebuilt at the MCP layer. + +TOP OPEN DECISION (drives model + guardrail + whether Salesforce enters the stack — see §12): +- Task-agent surface: **Agentforce** (recommended), a **marketplace agent** (Claude app), or a + **custom Bolt assistant**. Native Slack AI is a complement, not a candidate. MCP + auth design is + identical across all three, so this does not block foundation work. + +RESIDUAL VERIFY (at build, not blocking design): +1. Whether Slack publishes egress IP ranges for custom MCP connectors (for WAF IP-allowlisting) + and whether its MCP client supports client certs — confirm with Slack docs/support. +2. Exact Slack-AI-guardrail config knobs available to us as the tool provider. + +## 10. Cross-review (cross_reviewer / GPT-4.1, 2026-06-09) + +Verdict: design fundamentally sound. BLOCK/FIX items below were folded into §2; remaining are +operational and tracked for the build. Full output in the conversation log. + +BLOCK (resolved in §2.4 / §2.5): +- Minimal-scope Google client per tool, no scope union, refresh token never exposed/logged, + access tokens not cached. → §2.4 +- Audience + scope enforced server-side at every boundary; tool-hiding is not the boundary. → §2.5 +- No broker-JWT passthrough to downstream APIs. → §2.5 +- IAM least-privilege per server role; per-user Google tokens partitioned by `sub` (ABAC), KMS + key policy scoped to role. → §2.4 / §2.5 + +FIX (resolved / tracked): +- Group-revocation latency: sync tightened to 5 min, `finance:*` TTL → 15 min, deny-list for + immediate hard revocation, ALARM on sync failure. → §2.3 +- Prompt-injection: server-side param validation + allow-lists, tool output treated as data. → §2.5 +- Audit access to sensitive secrets. → §2.5 +- Google refresh-token rotation/revocation on group removal. → build task. + +NIT: PKCE if a public client is added (§2.5); keep pre-token Lambda fast/no external calls (§2.3). + +These auth/IAM specifics are signed off by the cross-review gate. Re-run `cross_reviewer` on the +actual IAM policy JSON and pre-token Lambda code once written (the reviewer asked for the concrete +artifacts for a deeper pass). + +## 11. Investigation: endpoint connectivity & guardrail parity (2026-06-09) + +### 11.1 How does Slack's built-in agent reach our MCP servers? → PUBLIC HTTPS + +Slack's agent platform has two MCP directions, and ours is the outbound one: +- **Inbound** (`https://mcp.slack.com/mcp`): external clients connect IN to act on Slack. Not us. +- **Outbound** (our case): Slack's built-in agent / Slackbot acts as an MCP client and reaches + OUT to external MCP servers, configured by URL + `Authorization` bearer, with an OAuth callback + like `https://mcp./upstream-auth/callback`. Transport is JSON-RPC 2.0 over + Streamable HTTP (Slack does not support SSE or Dynamic Client Registration). + +Implication: the connection originates from Slack's cloud to a **public, internet-reachable HTTPS +endpoint**. VPC-only / PrivateLink is therefore not viable for the Slack surface. Decision: + +- API Gateway (HTTP API) or ALB, public, behind **AWS WAF**. +- The **Cognito OAuth bearer is the access boundary** (we already validate issuer/aud/scope). +- **IP-allowlist** Slack's egress ranges in WAF IF Slack publishes them (unconfirmed in docs; + it is a known open question in the custom-MCP-connector community — verify with Slack). +- **Do not rely on mTLS**: Slack's MCP client is not documented to present client certificates. + +(If we later add non-Slack MCP clients that run inside our network, e.g. internal tools, those +specific servers could be VPC-only — but anything Slack consumes must be public.) + +### 11.2 Do Slack's safeguards replace the Bedrock guardrail? → PARTIALLY + +Slack provides an enterprise "Slack AI guardrails" framework that DOES cover: +- **Prompt injection / jailbreak**: context engineering to mitigate injection, real-time + jailbreak/prompt-attack detection, URL filtering, output validation. +- **Access governance**: AI only accesses what the user is authorized to see; admin control over + which data sources and tools each assistant can reach; per-user/feature toggles. +- **Audit**: comprehensive logging of what each assistant accessed; real-time monitoring; + incident-investigation ("who asked what"). +- **Zero-training guarantee**; models run without outbound network access. + +What it does NOT cover, and is the gap we must own: +- **PII redaction of OUR tool responses.** Slack's DLP/tombstoning applies to Slack messages and + AI summaries derived from them, not to payloads our MCP tools return. Industry guidance is + explicit that regulated data needs **MCP-layer DLP that inspects and redacts every tool response + before it reaches the model**. So the Bedrock guardrail's PII-anonymization is reimplemented in + our servers' response path (§2.5), heaviest on `finance` (bank/routing/card/SSN). + +Net: prompt-attack protection transfers to Slack acceptably (and was arguably the surface's job +anyway); PII protection does not transfer and stays our responsibility at the MCP layer. + +Sources: docs.slack.dev/ai/slack-mcp-server; slack.com/blog/transformation/securing-the-agentic- +enterprise; slack.com/blog/news/how-we-built-slack-ai-to-be-secure-and-private; custom-MCP- +connector egress-range discussion (OpenAI dev community); strac.io MCP-layer DLP guidance. + +## 12. Conversational surface options (corrected 2026-06-09) + +Earlier drafts loosely said "native Slack AI" was the surface. That is wrong: native Slack AI and +a task agent are different categories. + +- **Native Slack AI is NOT our surface.** It is the built-in comprehension suite: Summarize + (channel/thread/DM), daily recaps, natural-language search, Slackbot catch-up. You invoke it via + menu buttons and the search bar, not by @mentioning a named agent. It reads/retrieves Slack's own + content and does not call our MCP tools or run multi-step tasks. It is a useful COMPLEMENT, not a + replacement for seahaven-slack-bot / exec-aide. +- **The task-agent surface** (what actually replaces the bots) is a named agent users @mention or + DM that reasons and takes action via tools. Three real candidates: + +| Surface | UX | Calls our MCP servers | Model control | Notes | +|---------|-----|-----------------------|---------------|-------| +| **Agentforce** (recommended) | Named AI teammates in an Agents tab; @mention in channels/DMs/threads; per-agent DM history | Yes | Salesforce-default (GPT-4o mix) OR BYOLLM (Bedrock/Azure/OpenAI/Vertex) → can keep our Claude + guardrail | Configured in Agentforce Studio (Topics/Actions); pulls Salesforce/Agentforce into the stack (licensing) | +| **Marketplace agent** (e.g. Claude app) | @mention/DM a vendor agent | Via Slackbot-MCP-client outward tools | Vendor's models (Claude app = Claude) | Least config; least control over tool wiring/scoping | +| **Custom Bolt assistant** | Assistant pane / @mention; we build the UX | Yes, we wire directly | Whatever we call (incl. our Bedrock Claude + guardrail) | Most control, most build/maintenance; contradicts the "no Bolt backend" goal | + +### Why Agentforce fits our design + +Agentforce supports **multiple specialized named agents**, which maps almost 1:1 onto our +trust-tiered servers + Google-group scoping: + +- **Seahaven-Ops agent** → `sh-mcp-ops`, granted to `sh-mcp-ops@` / `sh-mcp-assistant@`. +- **Seahaven-Finance agent** → `sh-mcp-finance`, granted to `sh-mcp-finance@`. +- **Lauren's Exec agent** (DM-scoped) → ops `*:self` tools + finance read. + +Each agent is a separate teammate with its own published audience and tool access, which reinforces +the trust-tier separation at the UX layer, not just in the token. + +### Decision needed + +Pick the task-agent surface: **Agentforce** (recommended — native multi-agent, MCP support, and +BYOLLM-on-Bedrock to retain Claude + our guardrail), a **marketplace agent** (fastest, least +control), or a **custom Bolt assistant** (most control, most build). This is now the top open +decision because it determines the model, the guardrail story, and whether Salesforce enters the +stack. The MCP servers + auth design below are identical across all three. + +Sources: slack.com/help (Guide to AI features in Slack; Use Agentforce in Slack); +slack.com/blog/news/turn-agents-into-teammates-with-slack; slack.dev illustrated Agentforce guide; +developer.salesforce.com Agentforce supported-models + BYOLLM.