mirror of
https://github.com/Sea-Haven-Industries/sh-mcp.git
synced 2026-09-30 08:53:18 +00:00
Design doc for the Sea Haven MCP platform that replaces seahaven-slack-bot and exec-aide: trust-tiered MCP servers, Google-SSO via a Cognito broker, TypeScript monorepo. Cross-review (cross_reviewer/GPT-4.1) folded in. Status: planning only, nothing built.
519 lines
34 KiB
Markdown
519 lines
34 KiB
Markdown
# Sea Haven MCP Platform — Auth Architecture, Scope Matrix & Build Plan
|
|
|
|
Status: DRAFT for review (cross-review gate not yet run)
|
|
Date: 2026-06-09
|
|
Owner: Adam Moussa
|
|
|
|
Supersedes the ad-hoc design discussion. This is the consolidated plan for deprecating
|
|
`seahaven-slack-bot` and `exec-aide` in favor of Slack AI agents backed by a set of
|
|
Sea Haven MCP servers, with authentication centralized on Google Workspace SSO.
|
|
|
|
---
|
|
|
|
## 1. Decisions locked in
|
|
|
|
This is a **greenfield service**. `seahaven-slack-bot` and `exec-aide` are deprecated
|
|
**completely** (repos archived, stacks decommissioned, Bedrock agents + guardrails torn down).
|
|
Nothing is migrated as a running component. Where proactive functionality is still needed
|
|
(Lauren's email classifier/digest, the KB syncs) it is **rebuilt fresh** in this new service's
|
|
`jobs/` (§5), not carried over.
|
|
|
|
1. **Conversational surface = a configurable Slack task agent** that users @mention (in channels,
|
|
threads, DMs) and that calls our MCP tools to do work end to end. Slack hosts the agent, the
|
|
LLM, and the orchestration; we run no Bolt backend of our own (unless we pick the custom-assistant
|
|
option). Our system is: MCP servers (the tools) + the Cognito/Google auth broker + the rebuilt jobs.
|
|
IMPORTANT (corrected, §12): this is NOT "native Slack AI." Native Slack AI is the summarize /
|
|
recap / search feature suite over Slack's own content; it does not act as a configurable agent
|
|
that calls our tools, so it cannot replace the bots. The real surface candidates are **Agentforce**
|
|
(recommended), a **marketplace agent** (e.g. the Claude app), or a **custom Bolt assistant**.
|
|
Surface choice is still OPEN (§12) and it drives the model + guardrail story.
|
|
Guardrail consequence (investigated, §11): Slack's native guardrails cover prompt-injection and
|
|
content safety, but do NOT redact PII inside our tool responses, so we rebuild PII masking at the
|
|
MCP layer (§2.5) regardless of surface. (If we pick Agentforce + BYOLLM-on-Bedrock, we could also
|
|
keep our own Bedrock guardrail on the model path — defense in depth, not a substitute.)
|
|
2. **Tools = Sea Haven MCP servers**, grouped by trust tier (`ops`, `finance`; `physical` deferred).
|
|
3. **Identity = Google Workspace**, the same IdP already used for Slack SSO. Google is the
|
|
single source of truth for authentication AND group membership.
|
|
4. **Authorization = Amazon Cognito** user pool federated to Google, acting as the OAuth 2.1
|
|
authorization server that issues scoped, audience-bound tokens the MCP servers validate.
|
|
Slack's built-in agent is the MCP client; each user does per-user OAuth consent against the
|
|
broker (Slack's MCP client supports per-user OAuth), so tool calls carry the real user identity.
|
|
5. **Language = TypeScript** for the entire monorepo (servers, packages, CDK, rebuilt jobs).
|
|
Packages are shared by both the MCP servers and the rebuilt scheduled jobs (one client per
|
|
service everywhere); a single language is what makes that sharing real.
|
|
|
|
---
|
|
|
|
## 2. Auth architecture
|
|
|
|
### 2.1 Planes
|
|
|
|
Three planes kept separate (the credential that reaches the broker is never the credential
|
|
that reaches a downstream API):
|
|
|
|
- **Inbound auth (who):** Cognito-issued JWT, identity federated from Google. Carries `sub`
|
|
(the user's Google identity), `aud` (the target MCP server), and `scope` claims.
|
|
- **Access control (what):** each MCP server validates issuer + audience and enforces a
|
|
required scope per tool. Tools whose scope the token lacks are hidden from the user.
|
|
- **Outbound auth (how we reach the API):**
|
|
- **Google-API tools** (Gmail, Calendar): act *as the user* via a direct per-user Google
|
|
OAuth grant (see 2.4).
|
|
- **Non-Google tools** (QBO, payments, Lenel, 3CX, Afi, Front, Maps): broker token
|
|
authorizes; the server uses its own service credentials from Secrets Manager.
|
|
|
|
### 2.2 Flow
|
|
|
|
```
|
|
User (already Google-SSO'd into Slack)
|
|
│ OAuth 2.1 auth-code (PKCE), federated login → Google (silent if logged in)
|
|
▼
|
|
Cognito user pool (Google = external OIDC IdP)
|
|
│ pre-token-generation Lambda: group → scope mapping
|
|
▼
|
|
Access token (JWT):
|
|
sub = lauren@seahavenind.com
|
|
aud = sh-mcp-ops
|
|
scope = "ops:read ops:tasks gmail:self calendar:self"
|
|
│
|
|
├──► sh-mcp-ops validates aud+scope, runs permitted tools
|
|
├──► sh-mcp-finance token lacks finance:read → finance tools hidden
|
|
└──► sh-mcp-physical token lacks physical:* → hidden
|
|
```
|
|
|
|
### 2.3 Group → scope mapping
|
|
|
|
Google Groups are the single control point. Membership is synced to Cognito groups by a
|
|
scheduled sync Lambda (every 5 min, tightened from hourly per cross-review — group changes are
|
|
infrequent but revocation latency on sensitive groups must be low). The sync Lambda has a
|
|
CloudWatch ALARM on failure (a silent sync failure freezes group state). The Cognito pre-token
|
|
Lambda maps Cognito group → scope claims at token mint, so the hot path makes no Directory API
|
|
call. (Alternative: pre-token Lambda calls the Google Directory API with a short cache. Chose
|
|
the sync approach to keep token minting fast and resilient to Directory API throttling.)
|
|
|
|
Revocation: a 5-min sync + 60-min token TTL means a removed user could retain access for ~65 min
|
|
worst case. For `finance:*`, TTL is shortened to 15 min, and a Cognito-backed deny-list (checked
|
|
by the servers) gives immediate hard revocation when a user is pulled from a sensitive group.
|
|
|
|
| Google Group | Members | Scopes granted |
|
|
|------------------------|--------------------------------|-------------------------------------------------|
|
|
| `sh-mcp-ops@` | all staff | `ops:read` |
|
|
| `sh-mcp-assistant@` | Lauren, Adam | `ops:read`, `ops:tasks`, `gmail:self`, `calendar:self` |
|
|
| `sh-mcp-finance@` | Adam, Lauren, accounting | `ops:read`, `finance:read` |
|
|
| `sh-mcp-admin@` | Adam | all of the above + `finance:admin` |
|
|
|
|
A user in multiple groups gets the union of scopes (Lauren is in `-assistant@` and `-finance@`,
|
|
so she gets `ops:read ops:tasks gmail:self calendar:self finance:read`).
|
|
|
|
`sh-mcp-physical@` / `physical:*` scopes are NOT issued to any agent surface at launch — the
|
|
physical tier is admin/out-of-band only (see §3 and §6).
|
|
|
|
### 2.4 Google-API tools (Gmail / Calendar) — outbound detail
|
|
|
|
Pulling the upstream Google API access token *through* Cognito federation is limited (Cognito
|
|
requests `openid profile email` and does not cleanly expose/refresh Google API scopes like
|
|
`gmail.readonly`). So the clean implementation is:
|
|
|
|
- Cognito-federated-Google handles **authentication + authorization scopes** (single SSO login).
|
|
- The Gmail/Calendar MCP tools do a **separate direct per-user Google OAuth grant** at first
|
|
use (incremental consent for `gmail.readonly` + `calendar`), storing a per-user refresh
|
|
token encrypted (Secrets Manager or a KMS-encrypted DynamoDB item, key = user `sub`).
|
|
- Both consents resolve to the same Google account, so identity stays consistent; it is one
|
|
extra one-time consent, not a second identity.
|
|
|
|
Result: `search_inbox` / `get_calendar_events` act as the signed-in user. Lauren physically
|
|
cannot read anyone else's mailbox; segregation is enforced by Google, not by our code.
|
|
|
|
Token-handling rules (cross-review BLOCK items):
|
|
- The Google client is instantiated with the **minimal scope for the requested tool only**
|
|
(`gmail:self` → `gmail.readonly`, `calendar:self` → `calendar`), never the union of a user's
|
|
scopes. A `gmail:self`-only user cannot mint a Calendar token and vice versa.
|
|
- The per-user refresh token is **never** returned to the Slack agent or any client, and never
|
|
appears in logs or error messages. Access tokens are minted per request, not cached or reused.
|
|
- Per-user refresh tokens are stored in a KMS-CMK-encrypted DynamoDB item keyed by user `sub`,
|
|
with the key policy scoped to the owning server role only. IAM on that table is partitioned by
|
|
`sub` (ABAC `dynamodb:LeadingKeys`) so a server can only read the calling user's token, not all
|
|
users' tokens.
|
|
|
|
### 2.5 Token & endpoint hardening
|
|
|
|
- Short-lived access tokens (60 min general; **15 min for `finance:*`**); refresh by the client.
|
|
- **Audience-bound**: each server validates `aud` and rejects tokens not minted for its resource
|
|
server. An `ops` token presented to `finance` is rejected.
|
|
- **Server-side enforcement is authoritative**: every server independently validates issuer +
|
|
`aud` + required scope per tool on every call. Tool-hiding in the agent UI is a convenience,
|
|
never the access-control boundary. The client is never trusted.
|
|
- **No broker-JWT passthrough**: the Cognito JWT is never forwarded to a downstream API. Servers
|
|
reach external APIs only via their own service credentials or the per-user Google token (§2.4).
|
|
- **Endpoints are PUBLIC HTTPS** (decided — see §11). Slack's agent runtime is cloud-hosted and
|
|
initiates outbound Streamable-HTTP to a public MCP URL with an OAuth bearer, so VPC-only /
|
|
PrivateLink is not an option. The OAuth bearer (Cognito) is the real access boundary; in front
|
|
of it: AWS WAF, plus IP-allowlisting to Slack's egress ranges IF Slack publishes them (to
|
|
verify). Do not depend on mTLS — Slack's MCP client is not documented to present client certs.
|
|
Never an unauthenticated path. PKCE required if any public client is added.
|
|
- **MCP-layer PII redaction** (decided — see §11): Slack's native guardrails do NOT redact PII
|
|
inside our tool responses, so the Bedrock guardrail's PII-anonymization is reimplemented in the
|
|
servers' output path. Every tool response is inspected and sensitive fields (bank account/
|
|
routing, card, SSN) masked before it leaves the server, heaviest on `finance`.
|
|
- **Physical tier step-up**: `physical:*` tools require recent re-auth (`auth_time` / `max_age`)
|
|
AND an out-of-band human approval before execution. Scope grants the right to *request*, never
|
|
to execute unattended.
|
|
- **Audit**: every `finance:*` and `physical:*` tool call logged to CloudWatch (user `sub`,
|
|
tool, args hash, decision, result). CloudWatch ALARM on any `physical:*` invocation
|
|
(ALARM-state only, per standing alarm preference).
|
|
- Connection-level governance: which MCP connectors a user can even add is also gated by group,
|
|
so `physical` is unreachable for non-members regardless of token contents (defense in depth).
|
|
- **IAM least-privilege per server** (cross-review BLOCK): each server role gets
|
|
`secretsmanager:GetSecretValue` on only its own secrets (no wildcard) and KMS decrypt on only
|
|
the relevant CMK. `sh-mcp-finance` can read the QBO + payments secrets but not Maps or per-user
|
|
Google tokens; `sh-mcp-ops` cannot read finance secrets. Access to QBO and per-user Google
|
|
tokens is logged and alarmed on anomalous patterns (mass/unexpected-role access).
|
|
- **Prompt-injection containment** (no Bedrock guardrail to fall back on): tool-call parameters
|
|
are validated server-side against strict schemas/allow-lists, never trusted from the agent.
|
|
Tool-returned content is treated as data, never instructions. Side-effecting `ops` tools
|
|
(`create_calendar_event`, `create_task`, `create_reminder`) get sane bounds; the trust-tier
|
|
split keeps any injection blast radius inside `ops` (no finance/physical reachable in-session).
|
|
|
|
---
|
|
|
|
## 3. Scope matrix (every tool → required scope, outbound auth, risk)
|
|
|
|
### sh-mcp-ops (agent-facing, read-mostly)
|
|
|
|
| Tool | Source pkg | Scope | Outbound auth | Risk |
|
|
|--------------------------|-----------------|----------------|------------------------|------|
|
|
| lookup_work_order | internal-data | `ops:read` | service creds (DDB ro) | low |
|
|
| lookup_purchase_order | internal-data | `ops:read` | service creds (DDB ro) | low |
|
|
| lookup_site | internal-data | `ops:read` | service creds (DDB ro) | low |
|
|
| search_knowledge_base | knowledge-base | `ops:read` | service (Bedrock Retrieve) | low |
|
|
| search_nearby_vendors | google-maps | `ops:read` | service (Maps API key) | low (external $, rate-limit) |
|
|
| search_inbox | gmail | `gmail:self` | forwarded user Google token | med (private data) |
|
|
| get_email_thread_detail | gmail | `gmail:self` | forwarded user Google token | med |
|
|
| get_calendar_events | calendar | `calendar:self`| forwarded user Google token | low |
|
|
| create_calendar_event | calendar | `calendar:self`| forwarded user Google token | med (writes, external attendees) |
|
|
| check_availability | calendar | `calendar:self`| forwarded user Google token | low |
|
|
| create_task | tasks | `ops:tasks` | service creds (DDB, partitioned by sub) | low |
|
|
| list_tasks | tasks | `ops:tasks` | service creds | low |
|
|
| complete_task | tasks | `ops:tasks` | service creds | low |
|
|
| delete_task | tasks | `ops:tasks` | service creds | low |
|
|
| create_reminder | reminders | `ops:tasks` | service creds (EventBridge Scheduler) | low |
|
|
|
|
### sh-mcp-finance (sensitive, read-only, fully audited)
|
|
|
|
| Tool | Source pkg | Scope | Outbound auth | Risk |
|
|
|---------------------------|------------|----------------|------------------------------|------|
|
|
| search_vendors | qbo | `finance:read` | service creds (QBO OAuth, server-held) | med (token can read all of QBO) |
|
|
| lookup_payment_by_vendor | payments | `finance:read` | service creds (DDB PaymentsDashboard) | med |
|
|
| lookup_payment_by_invoice | payments | `finance:read` | service creds (DDB) | med |
|
|
| lookup_payment_by_check | payments | `finance:read` | service creds (DDB) | med |
|
|
|
|
QBO OAuth maintenance (`/qbo/connect`, `/qbo/callback`, `/qbo/disconnect`) stays as
|
|
**admin web endpoints** (API Gateway), NOT exposed as agent tools. Gated by `finance:admin`.
|
|
|
|
### sh-mcp-physical (DECIDED: admin / out-of-band only — NOT in the launch build)
|
|
|
|
Neither bot does physical actions today, so this tier is documented to keep the trust model
|
|
complete but is DEFERRED. No `physical:*` scope is issued to any agent. The tools below stay as
|
|
admin/out-of-band capability (existing door-unlock-api path) until a future decision to expose them.
|
|
|
|
| Tool | Source pkg | Scope | Outbound auth | Risk |
|
|
|-----------------------|------------|----------------------|-----------------------|------|
|
|
| request_door_unlock | lenel | `physical:unlock` | service creds (Elements API) | HIGH — human approval required |
|
|
| initiate_lockdown | lenel | `physical:lockdown` | service creds | HIGH — Adam only + approval |
|
|
| release_lockdown | lenel | `physical:lockdown` | service creds | HIGH — Adam only + approval |
|
|
| push_xml | yealink | `physical:telephony` | service creds (Push XML) | med + approval |
|
|
| reroute_call | threecx | `physical:telephony` | service creds (3CX xapi) | med + approval |
|
|
| set_dnd | threecx | `physical:telephony` | service creds (3CX xapi) | med + approval |
|
|
|
|
### Trifecta rationale
|
|
|
|
The dangerous combination (read-from-untrusted-source + sensitive-action in one session) is
|
|
prevented by tier separation: `finance` and `physical` are never reachable by a session that
|
|
also holds the Gmail/web read tools. Within `ops`, the actions (`create_task`, `create_reminder`,
|
|
`create_calendar_event`, `search_nearby_vendors`) are low-consequence, so injected content in an
|
|
email can at worst create a spurious task or run a Maps query — not move money or unlock a door.
|
|
Outbound items in `ops` (calendar invites to external attendees) are flagged for monitoring.
|
|
|
|
---
|
|
|
|
## 4. Monorepo layout
|
|
|
|
Packages are consumed by BOTH the MCP servers and the surviving scheduled Lambdas — one client
|
|
per external service, used everywhere. This is the core payoff over per-service repos.
|
|
|
|
```
|
|
sh-mcp/ (one monorepo, org: Sea-Haven-Industries)
|
|
packages/
|
|
shared/ MCP scaffold, JWT validation, requires_scope guard, secrets, audit log
|
|
qbo/ google-maps/ internal-data/ payments/ knowledge-base/
|
|
gmail/ calendar/ tasks/ reminders/ notion/
|
|
# DEFERRED (physical tier, not in launch build): lenel/ yealink/ threecx/
|
|
servers/
|
|
sh-mcp-ops/ → remote HTTP MCP stack (CDK)
|
|
sh-mcp-finance/ → remote HTTP MCP stack (CDK)
|
|
# DEFERRED: sh-mcp-physical/ (admin/out-of-band only)
|
|
auth/
|
|
cognito/ user pool, Google federation, resource servers, app clients
|
|
pre-token-lambda/ group → scope mapping
|
|
group-sync-lambda/ Google Groups → Cognito groups (hourly)
|
|
jobs/ (surviving scheduled Lambdas, import packages/)
|
|
notion-sync/ po-sync/ workorder-sync/ (feed Bedrock KB)
|
|
fetch-classify/ daily-digest/ reminder/ (exec-aide proactive, Lauren)
|
|
```
|
|
|
|
---
|
|
|
|
## 5. Proactive jobs — rebuilt fresh in `sh-mcp/jobs` (NOT MCP, NOT migrated)
|
|
|
|
Proactive/event-driven work has no conversational equivalent and Slack's reactive agent cannot
|
|
replace it, so it is **rebuilt** in the new service as TypeScript Lambdas that import the shared
|
|
packages and use stored offline credentials (not interactive SSO). The old exec-aide/slack-bot
|
|
Lambdas are decommissioned, not refactored in place.
|
|
|
|
| Job | Trigger | Notes |
|
|
|--------------------|----------------|-------------------------------------------------------------|
|
|
| fetch-classify | every 15 min | Gmail history → Haiku classify → HIGH DM. Offline Google refresh token (`exec-aide/gmail-oauth`). |
|
|
| daily-digest | 5pm ET M-F | Query DDB → digest → DM. |
|
|
| reminder (one-shot)| Scheduler | Fired by the `create_reminder` tool. |
|
|
| notion-sync | 02:00 UTC | Notion → S3 → Bedrock KB ingestion. |
|
|
| po-sync | 02:00 UTC | DDB purchase-orders → S3 → KB. |
|
|
| workorder-sync | 02:00 UTC | DDB work orders/comments → S3 → KB. |
|
|
|
|
The Bedrock **agents** (`seahaven-alex`, exec-aide Sonnet loop) are what get deprecated. The
|
|
Bedrock **guardrail** (PII anonymization, prompt-attack filtering) must be explicitly replaced
|
|
on the Slack surface before deletion — do not assume the new surface provides parity.
|
|
|
|
---
|
|
|
|
## 6. Phasing
|
|
|
|
1. **Foundation:** monorepo + `shared` + CI/CD (reusable workflows, deploy role first, OIDC,
|
|
Dependabot, naming) + the test harness and coverage gate (see §7). Stand up Cognito + Google
|
|
federation + pre-token + group-sync. Prove SSO login end to end.
|
|
2. **Pilot — `sh-mcp-ops`** wired to the `seahaven-slack-bot` replacement agent: vendor +
|
|
WO/PO/site + KB. Validate per-user auth, tool-hiding, and parity vs the old Bedrock agent.
|
|
3. **`sh-mcp-finance`** with full audit logging once ops is proven.
|
|
4. **Convert Lauren → Slack AI agent:** add `gmail:self`/`calendar:self`/`ops:tasks` tools to
|
|
ops, do the per-user Google grant, keep her agent DM/personal-scoped. Refactor her scheduled
|
|
Lambdas to import the packages.
|
|
5. **`sh-mcp-physical`** is DEFERRED — admin/out-of-band only, not part of this build.
|
|
6. **Deprecate each bot only after its replacement is proven at parity.** Lauren's workflow
|
|
needs sign-off before `exec-aide` is retired.
|
|
|
|
---
|
|
|
|
## 7. CI/CD & test suite
|
|
|
|
### 7.0 Language: TypeScript (DECIDED)
|
|
|
|
TypeScript for everything: servers, packages, CDK, and the rebuilt jobs. Toolchain is
|
|
`tsc --noEmit` + eslint/prettier + vitest. Since the old stacks are fully deprecated (not
|
|
migrated), exec-aide's former Python logic (Gmail/Calendar/Haiku classify) is rewritten in TS
|
|
as part of the greenfield build, not ported line-for-line. No polyglot matrix needed.
|
|
|
|
### 7.1 Reusable-workflow pattern (Sea Haven standard)
|
|
|
|
Per the handbook `cicd.md`, repos call org reusable workflows rather than inlining steps. Two
|
|
caller workflows:
|
|
|
|
- `.github/workflows/ci.yaml` — on `pull_request`. Calls the org reusable `ci` workflow.
|
|
Aggregated under a single required check named `ci / ci` for branch protection (so adding
|
|
matrix legs never breaks the required-context list).
|
|
- `.github/workflows/deploy.yaml` — on push to `main` (deploy-then-merge per git-workflow).
|
|
Calls the org reusable `cd-cdk` workflow per server, OIDC into the deploy role.
|
|
|
|
Gotchas baked in from prior repos:
|
|
- `permissions:` declared in the CALLER (reusable-workflow permissions don't inherit — missing
|
|
them yields `startup_failure`).
|
|
- `concurrency` keyed to caller context, not the reusable workflow, to avoid cross-PR collision.
|
|
- ARM64 Lambda containers: `enable-qemu: true` in BOTH ci.yaml and deploy.yaml, AND
|
|
`platform: LINUX_ARM64` on every Docker image asset (Lambda) or "exec format error".
|
|
- `cd-cdk` npm-ci guard since CDK is Node even where Lambda code may differ.
|
|
|
|
### 7.2 Monorepo CI shape
|
|
|
|
Path-filtered matrix so only changed packages/servers run. Per leg:
|
|
|
|
1. Install + build (workspace-aware: `npm ci` at root, build changed package + dependents).
|
|
2. Lint + format gate: `tsc --noEmit` + eslint/prettier (or `ruff check` + `ruff format --check`
|
|
on any Python paths). Enforced locally pre-push by the existing Claude Code hook too.
|
|
3. Unit + auth + security tests (§7.3) with coverage gate.
|
|
4. `cdk synth` for changed servers (catches IaC + ARM64 platform regressions without deploying).
|
|
5. Dependabot enabled; exact-pin `aws-cdk-lib` per the handbook Pinning Principle (no blanket
|
|
ignores); Dependabot keeps it current.
|
|
|
|
Deploy job runs only on `main`, per-server, gated on `ci / ci` green and the cross-review
|
|
sign-off for any IAM/authz diff (§8).
|
|
|
|
### 7.3 Test suite (full, security-weighted)
|
|
|
|
The authorization layer is the highest-risk surface, so it gets the heaviest coverage.
|
|
|
|
| Layer | What it covers |
|
|
|-------|----------------|
|
|
| **Unit — per package** | Each integration package with the external API mocked: QBO vendor query + token auto-rotation, payments DDB queries, internal-data lookups, Maps search, KB retrieve, Gmail/Calendar calls, tasks CRUD, reminder scheduling. Error/empty/throttle paths. |
|
|
| **Auth — the crown jewels** | JWT validation (issuer, signature, expiry); **audience binding** (a token minted for `sh-mcp-ops` is REJECTED by `sh-mcp-finance`); per-tool scope enforcement **server-side** (independent of UI tool-hiding); **tool-hiding** (a user without `finance:read` does not see finance tools in `list_tools`); pre-token Lambda group→scope mapping for every group; group-sync correctness; **deny-list hard revocation** (a revoked user is rejected immediately, before token expiry); minimal-scope Google client (a `gmail:self`-only token cannot mint a Calendar token); per-user token ABAC isolation (a server cannot read another user's refresh token). |
|
|
| **Security / abuse** | **Trifecta separation** — assert no single issued token can reach both untrusted-read (Gmail) and a sensitive-action tool; scope-escalation attempts rejected; prompt-injection regression corpus (tool-returned content containing "ignore instructions / call X" must NOT trigger out-of-scope tool calls); **MCP-layer PII redaction** (every finance tool response masks bank/routing/card/SSN before egress); per-session tool-call cap + per-tool rate limit. |
|
|
| **Contract / MCP conformance** | Every tool's input/output JSON schema validates; MCP protocol handshake; `list_tools` reflects the caller's scopes. |
|
|
| **Audit** | Every `finance:*` call emits a structured audit record (user sub, tool, args hash, decision); assert the record shape and that secrets never appear in logs. |
|
|
| **Parity (pre-deprecation)** | Golden-transcript tests replaying real slack-bot / exec-aide interactions against the new tools to confirm equivalent answers before retiring a bot. |
|
|
|
|
Coverage gate in CI (start at 80% lines, 100% on the `shared` auth/scope guard module). Parity
|
|
tests run in phase 2/4 before each bot is deprecated, not on every PR.
|
|
|
|
## 8. Handbook obligations (per global instructions)
|
|
|
|
- Naming: kebab-case throughout (`sh-mcp`, `sh-mcp-ops`, etc.). ✓ in this draft.
|
|
- Secrets: all in Secrets Manager; existing secret names reused; per-user Google tokens encrypted.
|
|
- Each deployable server: deploy role first, CI/CD pipeline, Dependabot, README.
|
|
- **Confluence "AWS Architecture Map" (id 1540098)**: add a Mermaid subgraph for the MCP
|
|
platform + Cognito before the build is reported done.
|
|
- Project memory: `project_sh_mcp.md` created; cross-links to exec-aide, slack-bot, payments,
|
|
identity-center, lenel/door-unlock memories.
|
|
- **Cross-review gate (mandatory):** Cognito↔Google federation, group→scope mapping, the
|
|
token-forwarding design, and every MCP server's IAM role go through `cross_reviewer` before
|
|
any commit/merge. IAM + auth = the breaking-change category that gates.
|
|
|
|
---
|
|
|
|
## 9. Decisions log & remaining open items
|
|
|
|
RESOLVED 2026-06-09:
|
|
- Complete deprecation: slack-bot + exec-aide repos archived, stacks decommissioned, no migration.
|
|
Greenfield service; conversational surface = Slack's built-in AI agent (no Bolt backend, no own model).
|
|
- Lauren gets `finance:read` (member of `-assistant@` + `-finance@`).
|
|
- `sh-mcp-physical` is admin/out-of-band only, deferred, not in the launch build.
|
|
- Afi removed from the stack entirely.
|
|
- Launch tool scope = exactly what the two bots do today (no Front / Afi / 3CX-read adds).
|
|
- Language = TypeScript for the entire monorepo.
|
|
|
|
INVESTIGATED 2026-06-09 (both former open items — see §11):
|
|
- Endpoint model: PUBLIC HTTPS + OAuth + WAF (VPC/PrivateLink not viable; mTLS not relied on).
|
|
- Guardrail parity: Slack covers prompt-injection/content safety; PII redaction is NOT covered and
|
|
is rebuilt at the MCP layer.
|
|
|
|
TOP OPEN DECISION (drives model + guardrail + whether Salesforce enters the stack — see §12):
|
|
- Task-agent surface: **Agentforce** (recommended), a **marketplace agent** (Claude app), or a
|
|
**custom Bolt assistant**. Native Slack AI is a complement, not a candidate. MCP + auth design is
|
|
identical across all three, so this does not block foundation work.
|
|
|
|
RESIDUAL VERIFY (at build, not blocking design):
|
|
1. Whether Slack publishes egress IP ranges for custom MCP connectors (for WAF IP-allowlisting)
|
|
and whether its MCP client supports client certs — confirm with Slack docs/support.
|
|
2. Exact Slack-AI-guardrail config knobs available to us as the tool provider.
|
|
|
|
## 10. Cross-review (cross_reviewer / GPT-4.1, 2026-06-09)
|
|
|
|
Verdict: design fundamentally sound. BLOCK/FIX items below were folded into §2; remaining are
|
|
operational and tracked for the build. Full output in the conversation log.
|
|
|
|
BLOCK (resolved in §2.4 / §2.5):
|
|
- Minimal-scope Google client per tool, no scope union, refresh token never exposed/logged,
|
|
access tokens not cached. → §2.4
|
|
- Audience + scope enforced server-side at every boundary; tool-hiding is not the boundary. → §2.5
|
|
- No broker-JWT passthrough to downstream APIs. → §2.5
|
|
- IAM least-privilege per server role; per-user Google tokens partitioned by `sub` (ABAC), KMS
|
|
key policy scoped to role. → §2.4 / §2.5
|
|
|
|
FIX (resolved / tracked):
|
|
- Group-revocation latency: sync tightened to 5 min, `finance:*` TTL → 15 min, deny-list for
|
|
immediate hard revocation, ALARM on sync failure. → §2.3
|
|
- Prompt-injection: server-side param validation + allow-lists, tool output treated as data. → §2.5
|
|
- Audit access to sensitive secrets. → §2.5
|
|
- Google refresh-token rotation/revocation on group removal. → build task.
|
|
|
|
NIT: PKCE if a public client is added (§2.5); keep pre-token Lambda fast/no external calls (§2.3).
|
|
|
|
These auth/IAM specifics are signed off by the cross-review gate. Re-run `cross_reviewer` on the
|
|
actual IAM policy JSON and pre-token Lambda code once written (the reviewer asked for the concrete
|
|
artifacts for a deeper pass).
|
|
|
|
## 11. Investigation: endpoint connectivity & guardrail parity (2026-06-09)
|
|
|
|
### 11.1 How does Slack's built-in agent reach our MCP servers? → PUBLIC HTTPS
|
|
|
|
Slack's agent platform has two MCP directions, and ours is the outbound one:
|
|
- **Inbound** (`https://mcp.slack.com/mcp`): external clients connect IN to act on Slack. Not us.
|
|
- **Outbound** (our case): Slack's built-in agent / Slackbot acts as an MCP client and reaches
|
|
OUT to external MCP servers, configured by URL + `Authorization` bearer, with an OAuth callback
|
|
like `https://mcp.<our-domain>/upstream-auth/callback`. Transport is JSON-RPC 2.0 over
|
|
Streamable HTTP (Slack does not support SSE or Dynamic Client Registration).
|
|
|
|
Implication: the connection originates from Slack's cloud to a **public, internet-reachable HTTPS
|
|
endpoint**. VPC-only / PrivateLink is therefore not viable for the Slack surface. Decision:
|
|
|
|
- API Gateway (HTTP API) or ALB, public, behind **AWS WAF**.
|
|
- The **Cognito OAuth bearer is the access boundary** (we already validate issuer/aud/scope).
|
|
- **IP-allowlist** Slack's egress ranges in WAF IF Slack publishes them (unconfirmed in docs;
|
|
it is a known open question in the custom-MCP-connector community — verify with Slack).
|
|
- **Do not rely on mTLS**: Slack's MCP client is not documented to present client certificates.
|
|
|
|
(If we later add non-Slack MCP clients that run inside our network, e.g. internal tools, those
|
|
specific servers could be VPC-only — but anything Slack consumes must be public.)
|
|
|
|
### 11.2 Do Slack's safeguards replace the Bedrock guardrail? → PARTIALLY
|
|
|
|
Slack provides an enterprise "Slack AI guardrails" framework that DOES cover:
|
|
- **Prompt injection / jailbreak**: context engineering to mitigate injection, real-time
|
|
jailbreak/prompt-attack detection, URL filtering, output validation.
|
|
- **Access governance**: AI only accesses what the user is authorized to see; admin control over
|
|
which data sources and tools each assistant can reach; per-user/feature toggles.
|
|
- **Audit**: comprehensive logging of what each assistant accessed; real-time monitoring;
|
|
incident-investigation ("who asked what").
|
|
- **Zero-training guarantee**; models run without outbound network access.
|
|
|
|
What it does NOT cover, and is the gap we must own:
|
|
- **PII redaction of OUR tool responses.** Slack's DLP/tombstoning applies to Slack messages and
|
|
AI summaries derived from them, not to payloads our MCP tools return. Industry guidance is
|
|
explicit that regulated data needs **MCP-layer DLP that inspects and redacts every tool response
|
|
before it reaches the model**. So the Bedrock guardrail's PII-anonymization is reimplemented in
|
|
our servers' response path (§2.5), heaviest on `finance` (bank/routing/card/SSN).
|
|
|
|
Net: prompt-attack protection transfers to Slack acceptably (and was arguably the surface's job
|
|
anyway); PII protection does not transfer and stays our responsibility at the MCP layer.
|
|
|
|
Sources: docs.slack.dev/ai/slack-mcp-server; slack.com/blog/transformation/securing-the-agentic-
|
|
enterprise; slack.com/blog/news/how-we-built-slack-ai-to-be-secure-and-private; custom-MCP-
|
|
connector egress-range discussion (OpenAI dev community); strac.io MCP-layer DLP guidance.
|
|
|
|
## 12. Conversational surface options (corrected 2026-06-09)
|
|
|
|
Earlier drafts loosely said "native Slack AI" was the surface. That is wrong: native Slack AI and
|
|
a task agent are different categories.
|
|
|
|
- **Native Slack AI is NOT our surface.** It is the built-in comprehension suite: Summarize
|
|
(channel/thread/DM), daily recaps, natural-language search, Slackbot catch-up. You invoke it via
|
|
menu buttons and the search bar, not by @mentioning a named agent. It reads/retrieves Slack's own
|
|
content and does not call our MCP tools or run multi-step tasks. It is a useful COMPLEMENT, not a
|
|
replacement for seahaven-slack-bot / exec-aide.
|
|
- **The task-agent surface** (what actually replaces the bots) is a named agent users @mention or
|
|
DM that reasons and takes action via tools. Three real candidates:
|
|
|
|
| Surface | UX | Calls our MCP servers | Model control | Notes |
|
|
|---------|-----|-----------------------|---------------|-------|
|
|
| **Agentforce** (recommended) | Named AI teammates in an Agents tab; @mention in channels/DMs/threads; per-agent DM history | Yes | Salesforce-default (GPT-4o mix) OR BYOLLM (Bedrock/Azure/OpenAI/Vertex) → can keep our Claude + guardrail | Configured in Agentforce Studio (Topics/Actions); pulls Salesforce/Agentforce into the stack (licensing) |
|
|
| **Marketplace agent** (e.g. Claude app) | @mention/DM a vendor agent | Via Slackbot-MCP-client outward tools | Vendor's models (Claude app = Claude) | Least config; least control over tool wiring/scoping |
|
|
| **Custom Bolt assistant** | Assistant pane / @mention; we build the UX | Yes, we wire directly | Whatever we call (incl. our Bedrock Claude + guardrail) | Most control, most build/maintenance; contradicts the "no Bolt backend" goal |
|
|
|
|
### Why Agentforce fits our design
|
|
|
|
Agentforce supports **multiple specialized named agents**, which maps almost 1:1 onto our
|
|
trust-tiered servers + Google-group scoping:
|
|
|
|
- **Seahaven-Ops agent** → `sh-mcp-ops`, granted to `sh-mcp-ops@` / `sh-mcp-assistant@`.
|
|
- **Seahaven-Finance agent** → `sh-mcp-finance`, granted to `sh-mcp-finance@`.
|
|
- **Lauren's Exec agent** (DM-scoped) → ops `*:self` tools + finance read.
|
|
|
|
Each agent is a separate teammate with its own published audience and tool access, which reinforces
|
|
the trust-tier separation at the UX layer, not just in the token.
|
|
|
|
### Decision needed
|
|
|
|
Pick the task-agent surface: **Agentforce** (recommended — native multi-agent, MCP support, and
|
|
BYOLLM-on-Bedrock to retain Claude + our guardrail), a **marketplace agent** (fastest, least
|
|
control), or a **custom Bolt assistant** (most control, most build). This is now the top open
|
|
decision because it determines the model, the guardrail story, and whether Salesforce enters the
|
|
stack. The MCP servers + auth design below are identical across all three.
|
|
|
|
Sources: slack.com/help (Guide to AI features in Slack; Use Agentforce in Slack);
|
|
slack.com/blog/news/turn-agents-into-teammates-with-slack; slack.dev illustrated Agentforce guide;
|
|
developer.salesforce.com Agentforce supported-models + BYOLLM.
|