Initial commit: sh-mcp design and plan

Design doc for the Sea Haven MCP platform that replaces seahaven-slack-bot
and exec-aide: trust-tiered MCP servers, Google-SSO via a Cognito broker,
TypeScript monorepo. Cross-review (cross_reviewer/GPT-4.1) folded in.

Status: planning only, nothing built.
This commit is contained in:
Adam Moussa 2026-06-09 19:25:24 -04:00
commit 3e2d99a1eb
3 changed files with 582 additions and 0 deletions

23
.gitignore vendored Normal file
View file

@ -0,0 +1,23 @@
# Node / TypeScript
node_modules/
dist/
build/
*.tsbuildinfo
# CDK
cdk.out/
.cdk.staging/
# Env / secrets
.env
.env.*
*.local
# Logs / OS
*.log
npm-debug.log*
.DS_Store
# Test / coverage
coverage/
.nyc_output/

40
README.md Normal file
View file

@ -0,0 +1,40 @@
# sh-mcp
Sea Haven MCP platform. A TypeScript monorepo of trust-tiered **MCP servers** that expose
Sea Haven's proprietary integrations as tools, plus the **Cognito/Google auth broker** and the
**rebuilt scheduled jobs**. This service replaces `seahaven-slack-bot` and `exec-aide`, which are
deprecated completely; the conversational surface becomes a configurable Slack task agent.
> **Status: DESIGN / PLANNING. Not built. No stack deployed.**
> The full design, auth architecture, scope matrix, and build plan live in
> [`docs/design.md`](docs/design.md). Read it before writing any code.
## Shape (planned)
- **MCP servers** (trust-tiered, remote HTTP, per-server IAM):
- `sh-mcp-ops` — read-mostly, agent-facing (WO/PO/site lookups, KB search, Google Maps,
Gmail/Calendar, tasks, reminders).
- `sh-mcp-finance` — sensitive, read-only, audited (QBO vendor search, payment lookups).
- `sh-mcp-physical` — DEFERRED, admin/out-of-band only (Lenel/Yealink/3CX control).
- **Auth** — Google Workspace is the single IdP; an Amazon Cognito user pool federated to Google
issues scoped, audience-bound JWTs; group → scope mapping via a pre-token Lambda. See design §2.
- **Jobs** — rebuilt proactive Lambdas (email classify/digest, KB syncs).
- **Language** — TypeScript everywhere (servers, packages, CDK, jobs).
## Open decisions
- Task-agent surface: Agentforce (recommended) vs marketplace Claude app vs custom Bolt assistant
(design §12). Drives the model + guardrail story.
- Endpoint exposure specifics (Slack egress ranges / WAF) — design §9 / §11.
## Layout (target)
```
packages/ shared + one package per integration
servers/ sh-mcp-ops, sh-mcp-finance (CDK stacks)
auth/ cognito, pre-token-lambda, group-sync-lambda
jobs/ rebuilt scheduled Lambdas
docs/ design.md (the canonical plan)
```
See [`docs/design.md`](docs/design.md) for the authoritative spec.

519
docs/design.md Normal file
View file

@ -0,0 +1,519 @@
# Sea Haven MCP Platform — Auth Architecture, Scope Matrix & Build Plan
Status: DRAFT for review (cross-review gate not yet run)
Date: 2026-06-09
Owner: Adam Moussa
Supersedes the ad-hoc design discussion. This is the consolidated plan for deprecating
`seahaven-slack-bot` and `exec-aide` in favor of Slack AI agents backed by a set of
Sea Haven MCP servers, with authentication centralized on Google Workspace SSO.
---
## 1. Decisions locked in
This is a **greenfield service**. `seahaven-slack-bot` and `exec-aide` are deprecated
**completely** (repos archived, stacks decommissioned, Bedrock agents + guardrails torn down).
Nothing is migrated as a running component. Where proactive functionality is still needed
(Lauren's email classifier/digest, the KB syncs) it is **rebuilt fresh** in this new service's
`jobs/` (§5), not carried over.
1. **Conversational surface = a configurable Slack task agent** that users @mention (in channels,
threads, DMs) and that calls our MCP tools to do work end to end. Slack hosts the agent, the
LLM, and the orchestration; we run no Bolt backend of our own (unless we pick the custom-assistant
option). Our system is: MCP servers (the tools) + the Cognito/Google auth broker + the rebuilt jobs.
IMPORTANT (corrected, §12): this is NOT "native Slack AI." Native Slack AI is the summarize /
recap / search feature suite over Slack's own content; it does not act as a configurable agent
that calls our tools, so it cannot replace the bots. The real surface candidates are **Agentforce**
(recommended), a **marketplace agent** (e.g. the Claude app), or a **custom Bolt assistant**.
Surface choice is still OPEN (§12) and it drives the model + guardrail story.
Guardrail consequence (investigated, §11): Slack's native guardrails cover prompt-injection and
content safety, but do NOT redact PII inside our tool responses, so we rebuild PII masking at the
MCP layer (§2.5) regardless of surface. (If we pick Agentforce + BYOLLM-on-Bedrock, we could also
keep our own Bedrock guardrail on the model path — defense in depth, not a substitute.)
2. **Tools = Sea Haven MCP servers**, grouped by trust tier (`ops`, `finance`; `physical` deferred).
3. **Identity = Google Workspace**, the same IdP already used for Slack SSO. Google is the
single source of truth for authentication AND group membership.
4. **Authorization = Amazon Cognito** user pool federated to Google, acting as the OAuth 2.1
authorization server that issues scoped, audience-bound tokens the MCP servers validate.
Slack's built-in agent is the MCP client; each user does per-user OAuth consent against the
broker (Slack's MCP client supports per-user OAuth), so tool calls carry the real user identity.
5. **Language = TypeScript** for the entire monorepo (servers, packages, CDK, rebuilt jobs).
Packages are shared by both the MCP servers and the rebuilt scheduled jobs (one client per
service everywhere); a single language is what makes that sharing real.
---
## 2. Auth architecture
### 2.1 Planes
Three planes kept separate (the credential that reaches the broker is never the credential
that reaches a downstream API):
- **Inbound auth (who):** Cognito-issued JWT, identity federated from Google. Carries `sub`
(the user's Google identity), `aud` (the target MCP server), and `scope` claims.
- **Access control (what):** each MCP server validates issuer + audience and enforces a
required scope per tool. Tools whose scope the token lacks are hidden from the user.
- **Outbound auth (how we reach the API):**
- **Google-API tools** (Gmail, Calendar): act *as the user* via a direct per-user Google
OAuth grant (see 2.4).
- **Non-Google tools** (QBO, payments, Lenel, 3CX, Afi, Front, Maps): broker token
authorizes; the server uses its own service credentials from Secrets Manager.
### 2.2 Flow
```
User (already Google-SSO'd into Slack)
│ OAuth 2.1 auth-code (PKCE), federated login → Google (silent if logged in)
▼
Cognito user pool (Google = external OIDC IdP)
│ pre-token-generation Lambda: group → scope mapping
▼
Access token (JWT):
sub = lauren@seahavenind.com
aud = sh-mcp-ops
scope = "ops:read ops:tasks gmail:self calendar:self"
│
├──► sh-mcp-ops validates aud+scope, runs permitted tools
├──► sh-mcp-finance token lacks finance:read → finance tools hidden
└──► sh-mcp-physical token lacks physical:* → hidden
```
### 2.3 Group → scope mapping
Google Groups are the single control point. Membership is synced to Cognito groups by a
scheduled sync Lambda (every 5 min, tightened from hourly per cross-review — group changes are
infrequent but revocation latency on sensitive groups must be low). The sync Lambda has a
CloudWatch ALARM on failure (a silent sync failure freezes group state). The Cognito pre-token
Lambda maps Cognito group → scope claims at token mint, so the hot path makes no Directory API
call. (Alternative: pre-token Lambda calls the Google Directory API with a short cache. Chose
the sync approach to keep token minting fast and resilient to Directory API throttling.)
Revocation: a 5-min sync + 60-min token TTL means a removed user could retain access for ~65 min
worst case. For `finance:*`, TTL is shortened to 15 min, and a Cognito-backed deny-list (checked
by the servers) gives immediate hard revocation when a user is pulled from a sensitive group.
| Google Group | Members | Scopes granted |
|------------------------|--------------------------------|-------------------------------------------------|
| `sh-mcp-ops@` | all staff | `ops:read` |
| `sh-mcp-assistant@` | Lauren, Adam | `ops:read`, `ops:tasks`, `gmail:self`, `calendar:self` |
| `sh-mcp-finance@` | Adam, Lauren, accounting | `ops:read`, `finance:read` |
| `sh-mcp-admin@` | Adam | all of the above + `finance:admin` |
A user in multiple groups gets the union of scopes (Lauren is in `-assistant@` and `-finance@`,
so she gets `ops:read ops:tasks gmail:self calendar:self finance:read`).
`sh-mcp-physical@` / `physical:*` scopes are NOT issued to any agent surface at launch — the
physical tier is admin/out-of-band only (see §3 and §6).
### 2.4 Google-API tools (Gmail / Calendar) — outbound detail
Pulling the upstream Google API access token *through* Cognito federation is limited (Cognito
requests `openid profile email` and does not cleanly expose/refresh Google API scopes like
`gmail.readonly`). So the clean implementation is:
- Cognito-federated-Google handles **authentication + authorization scopes** (single SSO login).
- The Gmail/Calendar MCP tools do a **separate direct per-user Google OAuth grant** at first
use (incremental consent for `gmail.readonly` + `calendar`), storing a per-user refresh
token encrypted (Secrets Manager or a KMS-encrypted DynamoDB item, key = user `sub`).
- Both consents resolve to the same Google account, so identity stays consistent; it is one
extra one-time consent, not a second identity.
Result: `search_inbox` / `get_calendar_events` act as the signed-in user. Lauren physically
cannot read anyone else's mailbox; segregation is enforced by Google, not by our code.
Token-handling rules (cross-review BLOCK items):
- The Google client is instantiated with the **minimal scope for the requested tool only**
(`gmail:self` → `gmail.readonly`, `calendar:self` → `calendar`), never the union of a user's
scopes. A `gmail:self`-only user cannot mint a Calendar token and vice versa.
- The per-user refresh token is **never** returned to the Slack agent or any client, and never
appears in logs or error messages. Access tokens are minted per request, not cached or reused.
- Per-user refresh tokens are stored in a KMS-CMK-encrypted DynamoDB item keyed by user `sub`,
with the key policy scoped to the owning server role only. IAM on that table is partitioned by
`sub` (ABAC `dynamodb:LeadingKeys`) so a server can only read the calling user's token, not all
users' tokens.
### 2.5 Token & endpoint hardening
- Short-lived access tokens (60 min general; **15 min for `finance:*`**); refresh by the client.
- **Audience-bound**: each server validates `aud` and rejects tokens not minted for its resource
server. An `ops` token presented to `finance` is rejected.
- **Server-side enforcement is authoritative**: every server independently validates issuer +
`aud` + required scope per tool on every call. Tool-hiding in the agent UI is a convenience,
never the access-control boundary. The client is never trusted.
- **No broker-JWT passthrough**: the Cognito JWT is never forwarded to a downstream API. Servers
reach external APIs only via their own service credentials or the per-user Google token (§2.4).
- **Endpoints are PUBLIC HTTPS** (decided — see §11). Slack's agent runtime is cloud-hosted and
initiates outbound Streamable-HTTP to a public MCP URL with an OAuth bearer, so VPC-only /
PrivateLink is not an option. The OAuth bearer (Cognito) is the real access boundary; in front
of it: AWS WAF, plus IP-allowlisting to Slack's egress ranges IF Slack publishes them (to
verify). Do not depend on mTLS — Slack's MCP client is not documented to present client certs.
Never an unauthenticated path. PKCE required if any public client is added.
- **MCP-layer PII redaction** (decided — see §11): Slack's native guardrails do NOT redact PII
inside our tool responses, so the Bedrock guardrail's PII-anonymization is reimplemented in the
servers' output path. Every tool response is inspected and sensitive fields (bank account/
routing, card, SSN) masked before it leaves the server, heaviest on `finance`.
- **Physical tier step-up**: `physical:*` tools require recent re-auth (`auth_time` / `max_age`)
AND an out-of-band human approval before execution. Scope grants the right to *request*, never
to execute unattended.
- **Audit**: every `finance:*` and `physical:*` tool call logged to CloudWatch (user `sub`,
tool, args hash, decision, result). CloudWatch ALARM on any `physical:*` invocation
(ALARM-state only, per standing alarm preference).
- Connection-level governance: which MCP connectors a user can even add is also gated by group,
so `physical` is unreachable for non-members regardless of token contents (defense in depth).
- **IAM least-privilege per server** (cross-review BLOCK): each server role gets
`secretsmanager:GetSecretValue` on only its own secrets (no wildcard) and KMS decrypt on only
the relevant CMK. `sh-mcp-finance` can read the QBO + payments secrets but not Maps or per-user
Google tokens; `sh-mcp-ops` cannot read finance secrets. Access to QBO and per-user Google
tokens is logged and alarmed on anomalous patterns (mass/unexpected-role access).
- **Prompt-injection containment** (no Bedrock guardrail to fall back on): tool-call parameters
are validated server-side against strict schemas/allow-lists, never trusted from the agent.
Tool-returned content is treated as data, never instructions. Side-effecting `ops` tools
(`create_calendar_event`, `create_task`, `create_reminder`) get sane bounds; the trust-tier
split keeps any injection blast radius inside `ops` (no finance/physical reachable in-session).
---
## 3. Scope matrix (every tool → required scope, outbound auth, risk)
### sh-mcp-ops (agent-facing, read-mostly)
| Tool | Source pkg | Scope | Outbound auth | Risk |
|--------------------------|-----------------|----------------|------------------------|------|
| lookup_work_order | internal-data | `ops:read` | service creds (DDB ro) | low |
| lookup_purchase_order | internal-data | `ops:read` | service creds (DDB ro) | low |
| lookup_site | internal-data | `ops:read` | service creds (DDB ro) | low |
| search_knowledge_base | knowledge-base | `ops:read` | service (Bedrock Retrieve) | low |
| search_nearby_vendors | google-maps | `ops:read` | service (Maps API key) | low (external $, rate-limit) |
| search_inbox | gmail | `gmail:self` | forwarded user Google token | med (private data) |
| get_email_thread_detail | gmail | `gmail:self` | forwarded user Google token | med |
| get_calendar_events | calendar | `calendar:self`| forwarded user Google token | low |
| create_calendar_event | calendar | `calendar:self`| forwarded user Google token | med (writes, external attendees) |
| check_availability | calendar | `calendar:self`| forwarded user Google token | low |
| create_task | tasks | `ops:tasks` | service creds (DDB, partitioned by sub) | low |
| list_tasks | tasks | `ops:tasks` | service creds | low |
| complete_task | tasks | `ops:tasks` | service creds | low |
| delete_task | tasks | `ops:tasks` | service creds | low |
| create_reminder | reminders | `ops:tasks` | service creds (EventBridge Scheduler) | low |
### sh-mcp-finance (sensitive, read-only, fully audited)
| Tool | Source pkg | Scope | Outbound auth | Risk |
|---------------------------|------------|----------------|------------------------------|------|
| search_vendors | qbo | `finance:read` | service creds (QBO OAuth, server-held) | med (token can read all of QBO) |
| lookup_payment_by_vendor | payments | `finance:read` | service creds (DDB PaymentsDashboard) | med |
| lookup_payment_by_invoice | payments | `finance:read` | service creds (DDB) | med |
| lookup_payment_by_check | payments | `finance:read` | service creds (DDB) | med |
QBO OAuth maintenance (`/qbo/connect`, `/qbo/callback`, `/qbo/disconnect`) stays as
**admin web endpoints** (API Gateway), NOT exposed as agent tools. Gated by `finance:admin`.
### sh-mcp-physical (DECIDED: admin / out-of-band only — NOT in the launch build)
Neither bot does physical actions today, so this tier is documented to keep the trust model
complete but is DEFERRED. No `physical:*` scope is issued to any agent. The tools below stay as
admin/out-of-band capability (existing door-unlock-api path) until a future decision to expose them.
| Tool | Source pkg | Scope | Outbound auth | Risk |
|-----------------------|------------|----------------------|-----------------------|------|
| request_door_unlock | lenel | `physical:unlock` | service creds (Elements API) | HIGH — human approval required |
| initiate_lockdown | lenel | `physical:lockdown` | service creds | HIGH — Adam only + approval |
| release_lockdown | lenel | `physical:lockdown` | service creds | HIGH — Adam only + approval |
| push_xml | yealink | `physical:telephony` | service creds (Push XML) | med + approval |
| reroute_call | threecx | `physical:telephony` | service creds (3CX xapi) | med + approval |
| set_dnd | threecx | `physical:telephony` | service creds (3CX xapi) | med + approval |
### Trifecta rationale
The dangerous combination (read-from-untrusted-source + sensitive-action in one session) is
prevented by tier separation: `finance` and `physical` are never reachable by a session that
also holds the Gmail/web read tools. Within `ops`, the actions (`create_task`, `create_reminder`,
`create_calendar_event`, `search_nearby_vendors`) are low-consequence, so injected content in an
email can at worst create a spurious task or run a Maps query — not move money or unlock a door.
Outbound items in `ops` (calendar invites to external attendees) are flagged for monitoring.
---
## 4. Monorepo layout
Packages are consumed by BOTH the MCP servers and the surviving scheduled Lambdas — one client
per external service, used everywhere. This is the core payoff over per-service repos.
```
sh-mcp/ (one monorepo, org: Sea-Haven-Industries)
packages/
shared/ MCP scaffold, JWT validation, requires_scope guard, secrets, audit log
qbo/ google-maps/ internal-data/ payments/ knowledge-base/
gmail/ calendar/ tasks/ reminders/ notion/
# DEFERRED (physical tier, not in launch build): lenel/ yealink/ threecx/
servers/
sh-mcp-ops/ → remote HTTP MCP stack (CDK)
sh-mcp-finance/ → remote HTTP MCP stack (CDK)
# DEFERRED: sh-mcp-physical/ (admin/out-of-band only)
auth/
cognito/ user pool, Google federation, resource servers, app clients
pre-token-lambda/ group → scope mapping
group-sync-lambda/ Google Groups → Cognito groups (hourly)
jobs/ (surviving scheduled Lambdas, import packages/)
notion-sync/ po-sync/ workorder-sync/ (feed Bedrock KB)
fetch-classify/ daily-digest/ reminder/ (exec-aide proactive, Lauren)
```
---
## 5. Proactive jobs — rebuilt fresh in `sh-mcp/jobs` (NOT MCP, NOT migrated)
Proactive/event-driven work has no conversational equivalent and Slack's reactive agent cannot
replace it, so it is **rebuilt** in the new service as TypeScript Lambdas that import the shared
packages and use stored offline credentials (not interactive SSO). The old exec-aide/slack-bot
Lambdas are decommissioned, not refactored in place.
| Job | Trigger | Notes |
|--------------------|----------------|-------------------------------------------------------------|
| fetch-classify | every 15 min | Gmail history → Haiku classify → HIGH DM. Offline Google refresh token (`exec-aide/gmail-oauth`). |
| daily-digest | 5pm ET M-F | Query DDB → digest → DM. |
| reminder (one-shot)| Scheduler | Fired by the `create_reminder` tool. |
| notion-sync | 02:00 UTC | Notion → S3 → Bedrock KB ingestion. |
| po-sync | 02:00 UTC | DDB purchase-orders → S3 → KB. |
| workorder-sync | 02:00 UTC | DDB work orders/comments → S3 → KB. |
The Bedrock **agents** (`seahaven-alex`, exec-aide Sonnet loop) are what get deprecated. The
Bedrock **guardrail** (PII anonymization, prompt-attack filtering) must be explicitly replaced
on the Slack surface before deletion — do not assume the new surface provides parity.
---
## 6. Phasing
1. **Foundation:** monorepo + `shared` + CI/CD (reusable workflows, deploy role first, OIDC,
Dependabot, naming) + the test harness and coverage gate (see §7). Stand up Cognito + Google
federation + pre-token + group-sync. Prove SSO login end to end.
2. **Pilot — `sh-mcp-ops`** wired to the `seahaven-slack-bot` replacement agent: vendor +
WO/PO/site + KB. Validate per-user auth, tool-hiding, and parity vs the old Bedrock agent.
3. **`sh-mcp-finance`** with full audit logging once ops is proven.
4. **Convert Lauren → Slack AI agent:** add `gmail:self`/`calendar:self`/`ops:tasks` tools to
ops, do the per-user Google grant, keep her agent DM/personal-scoped. Refactor her scheduled
Lambdas to import the packages.
5. **`sh-mcp-physical`** is DEFERRED — admin/out-of-band only, not part of this build.
6. **Deprecate each bot only after its replacement is proven at parity.** Lauren's workflow
needs sign-off before `exec-aide` is retired.
---
## 7. CI/CD & test suite
### 7.0 Language: TypeScript (DECIDED)
TypeScript for everything: servers, packages, CDK, and the rebuilt jobs. Toolchain is
`tsc --noEmit` + eslint/prettier + vitest. Since the old stacks are fully deprecated (not
migrated), exec-aide's former Python logic (Gmail/Calendar/Haiku classify) is rewritten in TS
as part of the greenfield build, not ported line-for-line. No polyglot matrix needed.
### 7.1 Reusable-workflow pattern (Sea Haven standard)
Per the handbook `cicd.md`, repos call org reusable workflows rather than inlining steps. Two
caller workflows:
- `.github/workflows/ci.yaml` — on `pull_request`. Calls the org reusable `ci` workflow.
Aggregated under a single required check named `ci / ci` for branch protection (so adding
matrix legs never breaks the required-context list).
- `.github/workflows/deploy.yaml` — on push to `main` (deploy-then-merge per git-workflow).
Calls the org reusable `cd-cdk` workflow per server, OIDC into the deploy role.
Gotchas baked in from prior repos:
- `permissions:` declared in the CALLER (reusable-workflow permissions don't inherit — missing
them yields `startup_failure`).
- `concurrency` keyed to caller context, not the reusable workflow, to avoid cross-PR collision.
- ARM64 Lambda containers: `enable-qemu: true` in BOTH ci.yaml and deploy.yaml, AND
`platform: LINUX_ARM64` on every Docker image asset (Lambda) or "exec format error".
- `cd-cdk` npm-ci guard since CDK is Node even where Lambda code may differ.
### 7.2 Monorepo CI shape
Path-filtered matrix so only changed packages/servers run. Per leg:
1. Install + build (workspace-aware: `npm ci` at root, build changed package + dependents).
2. Lint + format gate: `tsc --noEmit` + eslint/prettier (or `ruff check` + `ruff format --check`
on any Python paths). Enforced locally pre-push by the existing Claude Code hook too.
3. Unit + auth + security tests (§7.3) with coverage gate.
4. `cdk synth` for changed servers (catches IaC + ARM64 platform regressions without deploying).
5. Dependabot enabled; exact-pin `aws-cdk-lib` per the handbook Pinning Principle (no blanket
ignores); Dependabot keeps it current.
Deploy job runs only on `main`, per-server, gated on `ci / ci` green and the cross-review
sign-off for any IAM/authz diff (§8).
### 7.3 Test suite (full, security-weighted)
The authorization layer is the highest-risk surface, so it gets the heaviest coverage.
| Layer | What it covers |
|-------|----------------|
| **Unit — per package** | Each integration package with the external API mocked: QBO vendor query + token auto-rotation, payments DDB queries, internal-data lookups, Maps search, KB retrieve, Gmail/Calendar calls, tasks CRUD, reminder scheduling. Error/empty/throttle paths. |
| **Auth — the crown jewels** | JWT validation (issuer, signature, expiry); **audience binding** (a token minted for `sh-mcp-ops` is REJECTED by `sh-mcp-finance`); per-tool scope enforcement **server-side** (independent of UI tool-hiding); **tool-hiding** (a user without `finance:read` does not see finance tools in `list_tools`); pre-token Lambda group→scope mapping for every group; group-sync correctness; **deny-list hard revocation** (a revoked user is rejected immediately, before token expiry); minimal-scope Google client (a `gmail:self`-only token cannot mint a Calendar token); per-user token ABAC isolation (a server cannot read another user's refresh token). |
| **Security / abuse** | **Trifecta separation** — assert no single issued token can reach both untrusted-read (Gmail) and a sensitive-action tool; scope-escalation attempts rejected; prompt-injection regression corpus (tool-returned content containing "ignore instructions / call X" must NOT trigger out-of-scope tool calls); **MCP-layer PII redaction** (every finance tool response masks bank/routing/card/SSN before egress); per-session tool-call cap + per-tool rate limit. |
| **Contract / MCP conformance** | Every tool's input/output JSON schema validates; MCP protocol handshake; `list_tools` reflects the caller's scopes. |
| **Audit** | Every `finance:*` call emits a structured audit record (user sub, tool, args hash, decision); assert the record shape and that secrets never appear in logs. |
| **Parity (pre-deprecation)** | Golden-transcript tests replaying real slack-bot / exec-aide interactions against the new tools to confirm equivalent answers before retiring a bot. |
Coverage gate in CI (start at 80% lines, 100% on the `shared` auth/scope guard module). Parity
tests run in phase 2/4 before each bot is deprecated, not on every PR.
## 8. Handbook obligations (per global instructions)
- Naming: kebab-case throughout (`sh-mcp`, `sh-mcp-ops`, etc.). ✓ in this draft.
- Secrets: all in Secrets Manager; existing secret names reused; per-user Google tokens encrypted.
- Each deployable server: deploy role first, CI/CD pipeline, Dependabot, README.
- **Confluence "AWS Architecture Map" (id 1540098)**: add a Mermaid subgraph for the MCP
platform + Cognito before the build is reported done.
- Project memory: `project_sh_mcp.md` created; cross-links to exec-aide, slack-bot, payments,
identity-center, lenel/door-unlock memories.
- **Cross-review gate (mandatory):** Cognito↔Google federation, group→scope mapping, the
token-forwarding design, and every MCP server's IAM role go through `cross_reviewer` before
any commit/merge. IAM + auth = the breaking-change category that gates.
---
## 9. Decisions log & remaining open items
RESOLVED 2026-06-09:
- Complete deprecation: slack-bot + exec-aide repos archived, stacks decommissioned, no migration.
Greenfield service; conversational surface = Slack's built-in AI agent (no Bolt backend, no own model).
- Lauren gets `finance:read` (member of `-assistant@` + `-finance@`).
- `sh-mcp-physical` is admin/out-of-band only, deferred, not in the launch build.
- Afi removed from the stack entirely.
- Launch tool scope = exactly what the two bots do today (no Front / Afi / 3CX-read adds).
- Language = TypeScript for the entire monorepo.
INVESTIGATED 2026-06-09 (both former open items — see §11):
- Endpoint model: PUBLIC HTTPS + OAuth + WAF (VPC/PrivateLink not viable; mTLS not relied on).
- Guardrail parity: Slack covers prompt-injection/content safety; PII redaction is NOT covered and
is rebuilt at the MCP layer.
TOP OPEN DECISION (drives model + guardrail + whether Salesforce enters the stack — see §12):
- Task-agent surface: **Agentforce** (recommended), a **marketplace agent** (Claude app), or a
**custom Bolt assistant**. Native Slack AI is a complement, not a candidate. MCP + auth design is
identical across all three, so this does not block foundation work.
RESIDUAL VERIFY (at build, not blocking design):
1. Whether Slack publishes egress IP ranges for custom MCP connectors (for WAF IP-allowlisting)
and whether its MCP client supports client certs — confirm with Slack docs/support.
2. Exact Slack-AI-guardrail config knobs available to us as the tool provider.
## 10. Cross-review (cross_reviewer / GPT-4.1, 2026-06-09)
Verdict: design fundamentally sound. BLOCK/FIX items below were folded into §2; remaining are
operational and tracked for the build. Full output in the conversation log.
BLOCK (resolved in §2.4 / §2.5):
- Minimal-scope Google client per tool, no scope union, refresh token never exposed/logged,
access tokens not cached. → §2.4
- Audience + scope enforced server-side at every boundary; tool-hiding is not the boundary. → §2.5
- No broker-JWT passthrough to downstream APIs. → §2.5
- IAM least-privilege per server role; per-user Google tokens partitioned by `sub` (ABAC), KMS
key policy scoped to role. → §2.4 / §2.5
FIX (resolved / tracked):
- Group-revocation latency: sync tightened to 5 min, `finance:*` TTL → 15 min, deny-list for
immediate hard revocation, ALARM on sync failure. → §2.3
- Prompt-injection: server-side param validation + allow-lists, tool output treated as data. → §2.5
- Audit access to sensitive secrets. → §2.5
- Google refresh-token rotation/revocation on group removal. → build task.
NIT: PKCE if a public client is added (§2.5); keep pre-token Lambda fast/no external calls (§2.3).
These auth/IAM specifics are signed off by the cross-review gate. Re-run `cross_reviewer` on the
actual IAM policy JSON and pre-token Lambda code once written (the reviewer asked for the concrete
artifacts for a deeper pass).
## 11. Investigation: endpoint connectivity & guardrail parity (2026-06-09)
### 11.1 How does Slack's built-in agent reach our MCP servers? → PUBLIC HTTPS
Slack's agent platform has two MCP directions, and ours is the outbound one:
- **Inbound** (`https://mcp.slack.com/mcp`): external clients connect IN to act on Slack. Not us.
- **Outbound** (our case): Slack's built-in agent / Slackbot acts as an MCP client and reaches
OUT to external MCP servers, configured by URL + `Authorization` bearer, with an OAuth callback
like `https://mcp.<our-domain>/upstream-auth/callback`. Transport is JSON-RPC 2.0 over
Streamable HTTP (Slack does not support SSE or Dynamic Client Registration).
Implication: the connection originates from Slack's cloud to a **public, internet-reachable HTTPS
endpoint**. VPC-only / PrivateLink is therefore not viable for the Slack surface. Decision:
- API Gateway (HTTP API) or ALB, public, behind **AWS WAF**.
- The **Cognito OAuth bearer is the access boundary** (we already validate issuer/aud/scope).
- **IP-allowlist** Slack's egress ranges in WAF IF Slack publishes them (unconfirmed in docs;
it is a known open question in the custom-MCP-connector community — verify with Slack).
- **Do not rely on mTLS**: Slack's MCP client is not documented to present client certificates.
(If we later add non-Slack MCP clients that run inside our network, e.g. internal tools, those
specific servers could be VPC-only — but anything Slack consumes must be public.)
### 11.2 Do Slack's safeguards replace the Bedrock guardrail? → PARTIALLY
Slack provides an enterprise "Slack AI guardrails" framework that DOES cover:
- **Prompt injection / jailbreak**: context engineering to mitigate injection, real-time
jailbreak/prompt-attack detection, URL filtering, output validation.
- **Access governance**: AI only accesses what the user is authorized to see; admin control over
which data sources and tools each assistant can reach; per-user/feature toggles.
- **Audit**: comprehensive logging of what each assistant accessed; real-time monitoring;
incident-investigation ("who asked what").
- **Zero-training guarantee**; models run without outbound network access.
What it does NOT cover, and is the gap we must own:
- **PII redaction of OUR tool responses.** Slack's DLP/tombstoning applies to Slack messages and
AI summaries derived from them, not to payloads our MCP tools return. Industry guidance is
explicit that regulated data needs **MCP-layer DLP that inspects and redacts every tool response
before it reaches the model**. So the Bedrock guardrail's PII-anonymization is reimplemented in
our servers' response path (§2.5), heaviest on `finance` (bank/routing/card/SSN).
Net: prompt-attack protection transfers to Slack acceptably (and was arguably the surface's job
anyway); PII protection does not transfer and stays our responsibility at the MCP layer.
Sources: docs.slack.dev/ai/slack-mcp-server; slack.com/blog/transformation/securing-the-agentic-
enterprise; slack.com/blog/news/how-we-built-slack-ai-to-be-secure-and-private; custom-MCP-
connector egress-range discussion (OpenAI dev community); strac.io MCP-layer DLP guidance.
## 12. Conversational surface options (corrected 2026-06-09)
Earlier drafts loosely said "native Slack AI" was the surface. That is wrong: native Slack AI and
a task agent are different categories.
- **Native Slack AI is NOT our surface.** It is the built-in comprehension suite: Summarize
(channel/thread/DM), daily recaps, natural-language search, Slackbot catch-up. You invoke it via
menu buttons and the search bar, not by @mentioning a named agent. It reads/retrieves Slack's own
content and does not call our MCP tools or run multi-step tasks. It is a useful COMPLEMENT, not a
replacement for seahaven-slack-bot / exec-aide.
- **The task-agent surface** (what actually replaces the bots) is a named agent users @mention or
DM that reasons and takes action via tools. Three real candidates:
| Surface | UX | Calls our MCP servers | Model control | Notes |
|---------|-----|-----------------------|---------------|-------|
| **Agentforce** (recommended) | Named AI teammates in an Agents tab; @mention in channels/DMs/threads; per-agent DM history | Yes | Salesforce-default (GPT-4o mix) OR BYOLLM (Bedrock/Azure/OpenAI/Vertex) → can keep our Claude + guardrail | Configured in Agentforce Studio (Topics/Actions); pulls Salesforce/Agentforce into the stack (licensing) |
| **Marketplace agent** (e.g. Claude app) | @mention/DM a vendor agent | Via Slackbot-MCP-client outward tools | Vendor's models (Claude app = Claude) | Least config; least control over tool wiring/scoping |
| **Custom Bolt assistant** | Assistant pane / @mention; we build the UX | Yes, we wire directly | Whatever we call (incl. our Bedrock Claude + guardrail) | Most control, most build/maintenance; contradicts the "no Bolt backend" goal |
### Why Agentforce fits our design
Agentforce supports **multiple specialized named agents**, which maps almost 1:1 onto our
trust-tiered servers + Google-group scoping:
- **Seahaven-Ops agent** → `sh-mcp-ops`, granted to `sh-mcp-ops@` / `sh-mcp-assistant@`.
- **Seahaven-Finance agent** → `sh-mcp-finance`, granted to `sh-mcp-finance@`.
- **Lauren's Exec agent** (DM-scoped) → ops `*:self` tools + finance read.
Each agent is a separate teammate with its own published audience and tool access, which reinforces
the trust-tier separation at the UX layer, not just in the token.
### Decision needed
Pick the task-agent surface: **Agentforce** (recommended — native multi-agent, MCP support, and
BYOLLM-on-Bedrock to retain Claude + our guardrail), a **marketplace agent** (fastest, least
control), or a **custom Bolt assistant** (most control, most build). This is now the top open
decision because it determines the model, the guardrail story, and whether Salesforce enters the
stack. The MCP servers + auth design below are identical across all three.
Sources: slack.com/help (Guide to AI features in Slack; Use Agentforce in Slack);
slack.com/blog/news/turn-agents-into-teammates-with-slack; slack.dev illustrated Agentforce guide;
developer.salesforce.com Agentforce supported-models + BYOLLM.