sh-mcp/docs/design.md
Adam Moussa faa497de73
Some checks failed
deploy / deploy (push) Has been cancelled
docs: repoint pending cross_reviewer follow-ups to security-review cross_review.py (#30)
2026-07-14 19:24:12 -04:00

34 KiB

Sea Haven MCP Platform — Auth Architecture, Scope Matrix & Build Plan

Status: DRAFT for review (cross-review gate not yet run) Date: 2026-06-09 Owner: Adam Moussa

Supersedes the ad-hoc design discussion. This is the consolidated plan for deprecating seahaven-slack-bot and exec-aide in favor of Slack AI agents backed by a set of Sea Haven MCP servers, with authentication centralized on Google Workspace SSO.


1. Decisions locked in

This is a greenfield service. seahaven-slack-bot and exec-aide are deprecated completely (repos archived, stacks decommissioned, Bedrock agents + guardrails torn down). Nothing is migrated as a running component. Where proactive functionality is still needed (Lauren's email classifier/digest, the KB syncs) it is rebuilt fresh in this new service's jobs/ (§5), not carried over.

  1. Conversational surface = a configurable Slack task agent that users @mention (in channels, threads, DMs) and that calls our MCP tools to do work end to end. Slack hosts the agent, the LLM, and the orchestration; we run no Bolt backend of our own (unless we pick the custom-assistant option). Our system is: MCP servers (the tools) + the Cognito/Google auth broker + the rebuilt jobs. IMPORTANT (corrected, §12): this is NOT "native Slack AI." Native Slack AI is the summarize / recap / search feature suite over Slack's own content; it does not act as a configurable agent that calls our tools, so it cannot replace the bots. The real surface candidates are Agentforce (recommended), a marketplace agent (e.g. the Claude app), or a custom Bolt assistant. Surface choice is still OPEN (§12) and it drives the model + guardrail story. Guardrail consequence (investigated, §11): Slack's native guardrails cover prompt-injection and content safety, but do NOT redact PII inside our tool responses, so we rebuild PII masking at the MCP layer (§2.5) regardless of surface. (If we pick Agentforce + BYOLLM-on-Bedrock, we could also keep our own Bedrock guardrail on the model path — defense in depth, not a substitute.)
  2. Tools = Sea Haven MCP servers, grouped by trust tier (ops, finance; physical deferred).
  3. Identity = Google Workspace, the same IdP already used for Slack SSO. Google is the single source of truth for authentication AND group membership.
  4. Authorization = Amazon Cognito user pool federated to Google, acting as the OAuth 2.1 authorization server that issues scoped, audience-bound tokens the MCP servers validate. Slack's built-in agent is the MCP client; each user does per-user OAuth consent against the broker (Slack's MCP client supports per-user OAuth), so tool calls carry the real user identity.
  5. Language = TypeScript for the entire monorepo (servers, packages, CDK, rebuilt jobs). Packages are shared by both the MCP servers and the rebuilt scheduled jobs (one client per service everywhere); a single language is what makes that sharing real.

2. Auth architecture

2.1 Planes

Three planes kept separate (the credential that reaches the broker is never the credential that reaches a downstream API):

  • Inbound auth (who): Cognito-issued JWT, identity federated from Google. Carries sub (the user's Google identity), aud (the target MCP server), and scope claims.
  • Access control (what): each MCP server validates issuer + audience and enforces a required scope per tool. Tools whose scope the token lacks are hidden from the user.
  • Outbound auth (how we reach the API):
    • Google-API tools (Gmail, Calendar): act as the user via a direct per-user Google OAuth grant (see 2.4).
    • Non-Google tools (QBO, payments, Lenel, 3CX, Afi, Front, Maps): broker token authorizes; the server uses its own service credentials from Secrets Manager.

2.2 Flow

User (already Google-SSO'd into Slack)
      │  OAuth 2.1 auth-code (PKCE), federated login → Google (silent if logged in)
      ▼
Cognito user pool  (Google = external OIDC IdP)
      │  pre-token-generation Lambda: group → scope mapping
      ▼
Access token (JWT):
   sub = lauren@seahavenind.com
   aud = sh-mcp-ops
   scope = "ops:read ops:tasks gmail:self calendar:self"
      │
      ├──► sh-mcp-ops      validates aud+scope, runs permitted tools
      ├──► sh-mcp-finance  token lacks finance:read → finance tools hidden
      └──► sh-mcp-physical token lacks physical:* → hidden

2.3 Group → scope mapping

Google Groups are the single control point. Membership is synced to Cognito groups by a scheduled sync Lambda (every 5 min, tightened from hourly per cross-review — group changes are infrequent but revocation latency on sensitive groups must be low). The sync Lambda has a CloudWatch ALARM on failure (a silent sync failure freezes group state). The Cognito pre-token Lambda maps Cognito group → scope claims at token mint, so the hot path makes no Directory API call. (Alternative: pre-token Lambda calls the Google Directory API with a short cache. Chose the sync approach to keep token minting fast and resilient to Directory API throttling.)

Revocation: a 5-min sync + 60-min token TTL means a removed user could retain access for ~65 min worst case. For finance:*, TTL is shortened to 15 min, and a Cognito-backed deny-list (checked by the servers) gives immediate hard revocation when a user is pulled from a sensitive group.

Google Group Members Scopes granted
sh-mcp-ops@ all staff ops:read
sh-mcp-assistant@ Lauren, Adam ops:read, ops:tasks, gmail:self, calendar:self
sh-mcp-finance@ Adam, Lauren, accounting ops:read, finance:read
sh-mcp-admin@ Adam all of the above + finance:admin

A user in multiple groups gets the union of scopes (Lauren is in -assistant@ and -finance@, so she gets ops:read ops:tasks gmail:self calendar:self finance:read).

sh-mcp-physical@ / physical:* scopes are NOT issued to any agent surface at launch — the physical tier is admin/out-of-band only (see §3 and §6).

2.4 Google-API tools (Gmail / Calendar) — outbound detail

Pulling the upstream Google API access token through Cognito federation is limited (Cognito requests openid profile email and does not cleanly expose/refresh Google API scopes like gmail.readonly). So the clean implementation is:

  • Cognito-federated-Google handles authentication + authorization scopes (single SSO login).
  • The Gmail/Calendar MCP tools do a separate direct per-user Google OAuth grant at first use (incremental consent for gmail.readonly + calendar), storing a per-user refresh token encrypted (Secrets Manager or a KMS-encrypted DynamoDB item, key = user sub).
  • Both consents resolve to the same Google account, so identity stays consistent; it is one extra one-time consent, not a second identity.

Result: search_inbox / get_calendar_events act as the signed-in user. Lauren physically cannot read anyone else's mailbox; segregation is enforced by Google, not by our code.

Token-handling rules (cross-review BLOCK items):

  • The Google client is instantiated with the minimal scope for the requested tool only (gmail:self → gmail.readonly, calendar:self → calendar), never the union of a user's scopes. A gmail:self-only user cannot mint a Calendar token and vice versa.
  • The per-user refresh token is never returned to the Slack agent or any client, and never appears in logs or error messages. Access tokens are minted per request, not cached or reused.
  • Per-user refresh tokens are stored in a KMS-CMK-encrypted DynamoDB item keyed by user sub, with the key policy scoped to the owning server role only. IAM on that table is partitioned by sub (ABAC dynamodb:LeadingKeys) so a server can only read the calling user's token, not all users' tokens.

2.5 Token & endpoint hardening

  • Short-lived access tokens (60 min general; 15 min for finance:*); refresh by the client.
  • Audience-bound: each server validates aud and rejects tokens not minted for its resource server. An ops token presented to finance is rejected.
  • Server-side enforcement is authoritative: every server independently validates issuer + aud + required scope per tool on every call. Tool-hiding in the agent UI is a convenience, never the access-control boundary. The client is never trusted.
  • No broker-JWT passthrough: the Cognito JWT is never forwarded to a downstream API. Servers reach external APIs only via their own service credentials or the per-user Google token (§2.4).
  • Endpoints are PUBLIC HTTPS (decided — see §11). Slack's agent runtime is cloud-hosted and initiates outbound Streamable-HTTP to a public MCP URL with an OAuth bearer, so VPC-only / PrivateLink is not an option. The OAuth bearer (Cognito) is the real access boundary; in front of it: AWS WAF, plus IP-allowlisting to Slack's egress ranges IF Slack publishes them (to verify). Do not depend on mTLS — Slack's MCP client is not documented to present client certs. Never an unauthenticated path. PKCE required if any public client is added.
  • MCP-layer PII redaction (decided — see §11): Slack's native guardrails do NOT redact PII inside our tool responses, so the Bedrock guardrail's PII-anonymization is reimplemented in the servers' output path. Every tool response is inspected and sensitive fields (bank account/ routing, card, SSN) masked before it leaves the server, heaviest on finance.
  • Physical tier step-up: physical:* tools require recent re-auth (auth_time / max_age) AND an out-of-band human approval before execution. Scope grants the right to request, never to execute unattended.
  • Audit: every finance:* and physical:* tool call logged to CloudWatch (user sub, tool, args hash, decision, result). CloudWatch ALARM on any physical:* invocation (ALARM-state only, per standing alarm preference).
  • Connection-level governance: which MCP connectors a user can even add is also gated by group, so physical is unreachable for non-members regardless of token contents (defense in depth).
  • IAM least-privilege per server (cross-review BLOCK): each server role gets secretsmanager:GetSecretValue on only its own secrets (no wildcard) and KMS decrypt on only the relevant CMK. sh-mcp-finance can read the QBO + payments secrets but not Maps or per-user Google tokens; sh-mcp-ops cannot read finance secrets. Access to QBO and per-user Google tokens is logged and alarmed on anomalous patterns (mass/unexpected-role access).
  • Prompt-injection containment (no Bedrock guardrail to fall back on): tool-call parameters are validated server-side against strict schemas/allow-lists, never trusted from the agent. Tool-returned content is treated as data, never instructions. Side-effecting ops tools (create_calendar_event, create_task, create_reminder) get sane bounds; the trust-tier split keeps any injection blast radius inside ops (no finance/physical reachable in-session).

3. Scope matrix (every tool → required scope, outbound auth, risk)

sh-mcp-ops (agent-facing, read-mostly)

Tool Source pkg Scope Outbound auth Risk
lookup_work_order internal-data ops:read service creds (DDB ro) low
lookup_purchase_order internal-data ops:read service creds (DDB ro) low
lookup_site internal-data ops:read service creds (DDB ro) low
search_knowledge_base knowledge-base ops:read service (Bedrock Retrieve) low
search_nearby_vendors google-maps ops:read service (Maps API key) low (external $, rate-limit)
search_inbox gmail gmail:self forwarded user Google token med (private data)
get_email_thread_detail gmail gmail:self forwarded user Google token med
get_calendar_events calendar calendar:self forwarded user Google token low
create_calendar_event calendar calendar:self forwarded user Google token med (writes, external attendees)
check_availability calendar calendar:self forwarded user Google token low
create_task tasks ops:tasks service creds (DDB, partitioned by sub) low
list_tasks tasks ops:tasks service creds low
complete_task tasks ops:tasks service creds low
delete_task tasks ops:tasks service creds low
create_reminder reminders ops:tasks service creds (EventBridge Scheduler) low

sh-mcp-finance (sensitive, read-only, fully audited)

Tool Source pkg Scope Outbound auth Risk
search_vendors qbo finance:read service creds (QBO OAuth, server-held) med (token can read all of QBO)
lookup_payment_by_vendor payments finance:read service creds (DDB PaymentsDashboard) med
lookup_payment_by_invoice payments finance:read service creds (DDB) med
lookup_payment_by_check payments finance:read service creds (DDB) med

QBO OAuth maintenance (/qbo/connect, /qbo/callback, /qbo/disconnect) stays as admin web endpoints (API Gateway), NOT exposed as agent tools. Gated by finance:admin.

sh-mcp-physical (DECIDED: admin / out-of-band only — NOT in the launch build)

Neither bot does physical actions today, so this tier is documented to keep the trust model complete but is DEFERRED. No physical:* scope is issued to any agent. The tools below stay as admin/out-of-band capability (existing door-unlock-api path) until a future decision to expose them.

Tool Source pkg Scope Outbound auth Risk
request_door_unlock lenel physical:unlock service creds (Elements API) HIGH — human approval required
initiate_lockdown lenel physical:lockdown service creds HIGH — Adam only + approval
release_lockdown lenel physical:lockdown service creds HIGH — Adam only + approval
push_xml yealink physical:telephony service creds (Push XML) med + approval
reroute_call threecx physical:telephony service creds (3CX xapi) med + approval
set_dnd threecx physical:telephony service creds (3CX xapi) med + approval

Trifecta rationale

The dangerous combination (read-from-untrusted-source + sensitive-action in one session) is prevented by tier separation: finance and physical are never reachable by a session that also holds the Gmail/web read tools. Within ops, the actions (create_task, create_reminder, create_calendar_event, search_nearby_vendors) are low-consequence, so injected content in an email can at worst create a spurious task or run a Maps query — not move money or unlock a door. Outbound items in ops (calendar invites to external attendees) are flagged for monitoring.


4. Monorepo layout

Packages are consumed by BOTH the MCP servers and the surviving scheduled Lambdas — one client per external service, used everywhere. This is the core payoff over per-service repos.

sh-mcp/                              (one monorepo, org: Sea-Haven-Industries)
  packages/
    shared/        MCP scaffold, JWT validation, requires_scope guard, secrets, audit log
    qbo/  google-maps/  internal-data/  payments/  knowledge-base/
    gmail/  calendar/  tasks/  reminders/  notion/
    # DEFERRED (physical tier, not in launch build): lenel/  yealink/  threecx/
  servers/
    sh-mcp-ops/      → remote HTTP MCP stack (CDK)
    sh-mcp-finance/  → remote HTTP MCP stack (CDK)
    # DEFERRED: sh-mcp-physical/ (admin/out-of-band only)
  auth/
    cognito/         user pool, Google federation, resource servers, app clients
    pre-token-lambda/  group → scope mapping
    group-sync-lambda/ Google Groups → Cognito groups (hourly)
  jobs/                              (surviving scheduled Lambdas, import packages/)
    notion-sync/  po-sync/  workorder-sync/          (feed Bedrock KB)
    fetch-classify/  daily-digest/  reminder/         (exec-aide proactive, Lauren)

5. Proactive jobs — rebuilt fresh in sh-mcp/jobs (NOT MCP, NOT migrated)

Proactive/event-driven work has no conversational equivalent and Slack's reactive agent cannot replace it, so it is rebuilt in the new service as TypeScript Lambdas that import the shared packages and use stored offline credentials (not interactive SSO). The old exec-aide/slack-bot Lambdas are decommissioned, not refactored in place.

Job Trigger Notes
fetch-classify every 15 min Gmail history → Haiku classify → HIGH DM. Offline Google refresh token (exec-aide/gmail-oauth).
daily-digest 5pm ET M-F Query DDB → digest → DM.
reminder (one-shot) Scheduler Fired by the create_reminder tool.
notion-sync 02:00 UTC Notion → S3 → Bedrock KB ingestion.
po-sync 02:00 UTC DDB purchase-orders → S3 → KB.
workorder-sync 02:00 UTC DDB work orders/comments → S3 → KB.

The Bedrock agents (seahaven-alex, exec-aide Sonnet loop) are what get deprecated. The Bedrock guardrail (PII anonymization, prompt-attack filtering) must be explicitly replaced on the Slack surface before deletion — do not assume the new surface provides parity.


6. Phasing

  1. Foundation: monorepo + shared + CI/CD (reusable workflows, deploy role first, OIDC, Dependabot, naming) + the test harness and coverage gate (see §7). Stand up Cognito + Google federation + pre-token + group-sync. Prove SSO login end to end.
  2. Pilot — sh-mcp-ops wired to the seahaven-slack-bot replacement agent: vendor + WO/PO/site + KB. Validate per-user auth, tool-hiding, and parity vs the old Bedrock agent.
  3. sh-mcp-finance with full audit logging once ops is proven.
  4. Convert Lauren → Slack AI agent: add gmail:self/calendar:self/ops:tasks tools to ops, do the per-user Google grant, keep her agent DM/personal-scoped. Refactor her scheduled Lambdas to import the packages.
  5. sh-mcp-physical is DEFERRED — admin/out-of-band only, not part of this build.
  6. Deprecate each bot only after its replacement is proven at parity. Lauren's workflow needs sign-off before exec-aide is retired.

7. CI/CD & test suite

7.0 Language: TypeScript (DECIDED)

TypeScript for everything: servers, packages, CDK, and the rebuilt jobs. Toolchain is tsc --noEmit + eslint/prettier + vitest. Since the old stacks are fully deprecated (not migrated), exec-aide's former Python logic (Gmail/Calendar/Haiku classify) is rewritten in TS as part of the greenfield build, not ported line-for-line. No polyglot matrix needed.

7.1 Reusable-workflow pattern (Sea Haven standard)

Per the handbook cicd.md, repos call org reusable workflows rather than inlining steps. Two caller workflows:

  • .github/workflows/ci.yaml — on pull_request. Calls the org reusable ci workflow. Aggregated under a single required check named ci / ci for branch protection (so adding matrix legs never breaks the required-context list).
  • .github/workflows/deploy.yaml — on push to main (deploy-then-merge per git-workflow). Calls the org reusable cd-cdk workflow per server, OIDC into the deploy role.

Gotchas baked in from prior repos:

  • permissions: declared in the CALLER (reusable-workflow permissions don't inherit — missing them yields startup_failure).
  • concurrency keyed to caller context, not the reusable workflow, to avoid cross-PR collision.
  • ARM64 Lambda containers: enable-qemu: true in BOTH ci.yaml and deploy.yaml, AND platform: LINUX_ARM64 on every Docker image asset (Lambda) or "exec format error".
  • cd-cdk npm-ci guard since CDK is Node even where Lambda code may differ.

7.2 Monorepo CI shape

Path-filtered matrix so only changed packages/servers run. Per leg:

  1. Install + build (workspace-aware: npm ci at root, build changed package + dependents).
  2. Lint + format gate: tsc --noEmit + eslint/prettier (or ruff check + ruff format --check on any Python paths). Enforced locally pre-push by the existing Claude Code hook too.
  3. Unit + auth + security tests (§7.3) with coverage gate.
  4. cdk synth for changed servers (catches IaC + ARM64 platform regressions without deploying).
  5. Dependabot enabled; exact-pin aws-cdk-lib per the handbook Pinning Principle (no blanket ignores); Dependabot keeps it current.

Deploy job runs only on main, per-server, gated on ci / ci green and the cross-review sign-off for any IAM/authz diff (§8).

7.3 Test suite (full, security-weighted)

The authorization layer is the highest-risk surface, so it gets the heaviest coverage.

Layer What it covers
Unit — per package Each integration package with the external API mocked: QBO vendor query + token auto-rotation, payments DDB queries, internal-data lookups, Maps search, KB retrieve, Gmail/Calendar calls, tasks CRUD, reminder scheduling. Error/empty/throttle paths.
Auth — the crown jewels JWT validation (issuer, signature, expiry); audience binding (a token minted for sh-mcp-ops is REJECTED by sh-mcp-finance); per-tool scope enforcement server-side (independent of UI tool-hiding); tool-hiding (a user without finance:read does not see finance tools in list_tools); pre-token Lambda group→scope mapping for every group; group-sync correctness; deny-list hard revocation (a revoked user is rejected immediately, before token expiry); minimal-scope Google client (a gmail:self-only token cannot mint a Calendar token); per-user token ABAC isolation (a server cannot read another user's refresh token).
Security / abuse Trifecta separation — assert no single issued token can reach both untrusted-read (Gmail) and a sensitive-action tool; scope-escalation attempts rejected; prompt-injection regression corpus (tool-returned content containing "ignore instructions / call X" must NOT trigger out-of-scope tool calls); MCP-layer PII redaction (every finance tool response masks bank/routing/card/SSN before egress); per-session tool-call cap + per-tool rate limit.
Contract / MCP conformance Every tool's input/output JSON schema validates; MCP protocol handshake; list_tools reflects the caller's scopes.
Audit Every finance:* call emits a structured audit record (user sub, tool, args hash, decision); assert the record shape and that secrets never appear in logs.
Parity (pre-deprecation) Golden-transcript tests replaying real slack-bot / exec-aide interactions against the new tools to confirm equivalent answers before retiring a bot.

Coverage gate in CI (start at 80% lines, 100% on the shared auth/scope guard module). Parity tests run in phase 2/4 before each bot is deprecated, not on every PR.

8. Handbook obligations (per global instructions)

  • Naming: kebab-case throughout (sh-mcp, sh-mcp-ops, etc.). ✓ in this draft.
  • Secrets: all in Secrets Manager; existing secret names reused; per-user Google tokens encrypted.
  • Each deployable server: deploy role first, CI/CD pipeline, Dependabot, README.
  • Confluence "AWS Architecture Map" (id 1540098): add a Mermaid subgraph for the MCP platform + Cognito before the build is reported done.
  • Project memory: project_sh_mcp.md created; cross-links to exec-aide, slack-bot, payments, identity-center, lenel/door-unlock memories.
  • Cross-review gate (mandatory): Cognito↔Google federation, group→scope mapping, the token-forwarding design, and every MCP server's IAM role go through cross-family review before any commit/merge. IAM + auth = the breaking-change category that gates.

    Note: the orchestrator repo was archived 2026-07-14; cross-family review now runs via python3 ~/Documents/repositories/seahaven/security-review/cross_review.py (repo Sea-Haven-Industries/security-review). Mentions of cross_reviewer elsewhere in this doc are historical, naming the reviewer as it was invoked through the now-retired orchestrator.


9. Decisions log & remaining open items

RESOLVED 2026-06-09:

  • Complete deprecation: slack-bot + exec-aide repos archived, stacks decommissioned, no migration. Greenfield service; conversational surface = Slack's built-in AI agent (no Bolt backend, no own model).
  • Lauren gets finance:read (member of -assistant@ + -finance@).
  • sh-mcp-physical is admin/out-of-band only, deferred, not in the launch build.
  • Afi removed from the stack entirely.
  • Launch tool scope = exactly what the two bots do today (no Front / Afi / 3CX-read adds).
  • Language = TypeScript for the entire monorepo.

INVESTIGATED 2026-06-09 (both former open items — see §11):

  • Endpoint model: PUBLIC HTTPS + OAuth + WAF (VPC/PrivateLink not viable; mTLS not relied on).
  • Guardrail parity: Slack covers prompt-injection/content safety; PII redaction is NOT covered and is rebuilt at the MCP layer.

TOP OPEN DECISION (drives model + guardrail + whether Salesforce enters the stack — see §12):

  • Task-agent surface: Agentforce (recommended), a marketplace agent (Claude app), or a custom Bolt assistant. Native Slack AI is a complement, not a candidate. MCP + auth design is identical across all three, so this does not block foundation work.

RESIDUAL VERIFY (at build, not blocking design):

  1. Whether Slack publishes egress IP ranges for custom MCP connectors (for WAF IP-allowlisting) and whether its MCP client supports client certs — confirm with Slack docs/support.
  2. Exact Slack-AI-guardrail config knobs available to us as the tool provider.

10. Cross-review (cross_reviewer / GPT-4.1, 2026-06-09)

Verdict: design fundamentally sound. BLOCK/FIX items below were folded into §2; remaining are operational and tracked for the build. Full output in the conversation log.

BLOCK (resolved in §2.4 / §2.5):

  • Minimal-scope Google client per tool, no scope union, refresh token never exposed/logged, access tokens not cached. → §2.4
  • Audience + scope enforced server-side at every boundary; tool-hiding is not the boundary. → §2.5
  • No broker-JWT passthrough to downstream APIs. → §2.5
  • IAM least-privilege per server role; per-user Google tokens partitioned by sub (ABAC), KMS key policy scoped to role. → §2.4 / §2.5

FIX (resolved / tracked):

  • Group-revocation latency: sync tightened to 5 min, finance:* TTL → 15 min, deny-list for immediate hard revocation, ALARM on sync failure. → §2.3
  • Prompt-injection: server-side param validation + allow-lists, tool output treated as data. → §2.5
  • Audit access to sensitive secrets. → §2.5
  • Google refresh-token rotation/revocation on group removal. → build task.

NIT: PKCE if a public client is added (§2.5); keep pre-token Lambda fast/no external calls (§2.3).

These auth/IAM specifics are signed off by the cross-review gate. Re-run cross-family review via python3 ~/Documents/repositories/seahaven/security-review/cross_review.py (the orchestrator's cross_reviewer is retired as of 2026-07-14) on the actual IAM policy JSON and pre-token Lambda code once written (the reviewer asked for the concrete artifacts for a deeper pass).

11. Investigation: endpoint connectivity & guardrail parity (2026-06-09)

11.1 How does Slack's built-in agent reach our MCP servers? → PUBLIC HTTPS

Slack's agent platform has two MCP directions, and ours is the outbound one:

  • Inbound (https://mcp.slack.com/mcp): external clients connect IN to act on Slack. Not us.
  • Outbound (our case): Slack's built-in agent / Slackbot acts as an MCP client and reaches OUT to external MCP servers, configured by URL + Authorization bearer, with an OAuth callback like https://mcp.<our-domain>/upstream-auth/callback. Transport is JSON-RPC 2.0 over Streamable HTTP (Slack does not support SSE or Dynamic Client Registration).

Implication: the connection originates from Slack's cloud to a public, internet-reachable HTTPS endpoint. VPC-only / PrivateLink is therefore not viable for the Slack surface. Decision:

  • API Gateway (HTTP API) or ALB, public, behind AWS WAF.
  • The Cognito OAuth bearer is the access boundary (we already validate issuer/aud/scope).
  • IP-allowlist Slack's egress ranges in WAF IF Slack publishes them (unconfirmed in docs; it is a known open question in the custom-MCP-connector community — verify with Slack).
  • Do not rely on mTLS: Slack's MCP client is not documented to present client certificates.

(If we later add non-Slack MCP clients that run inside our network, e.g. internal tools, those specific servers could be VPC-only — but anything Slack consumes must be public.)

11.2 Do Slack's safeguards replace the Bedrock guardrail? → PARTIALLY

Slack provides an enterprise "Slack AI guardrails" framework that DOES cover:

  • Prompt injection / jailbreak: context engineering to mitigate injection, real-time jailbreak/prompt-attack detection, URL filtering, output validation.
  • Access governance: AI only accesses what the user is authorized to see; admin control over which data sources and tools each assistant can reach; per-user/feature toggles.
  • Audit: comprehensive logging of what each assistant accessed; real-time monitoring; incident-investigation ("who asked what").
  • Zero-training guarantee; models run without outbound network access.

What it does NOT cover, and is the gap we must own:

  • PII redaction of OUR tool responses. Slack's DLP/tombstoning applies to Slack messages and AI summaries derived from them, not to payloads our MCP tools return. Industry guidance is explicit that regulated data needs MCP-layer DLP that inspects and redacts every tool response before it reaches the model. So the Bedrock guardrail's PII-anonymization is reimplemented in our servers' response path (§2.5), heaviest on finance (bank/routing/card/SSN).

Net: prompt-attack protection transfers to Slack acceptably (and was arguably the surface's job anyway); PII protection does not transfer and stays our responsibility at the MCP layer.

Sources: docs.slack.dev/ai/slack-mcp-server; slack.com/blog/transformation/securing-the-agentic- enterprise; slack.com/blog/news/how-we-built-slack-ai-to-be-secure-and-private; custom-MCP- connector egress-range discussion (OpenAI dev community); strac.io MCP-layer DLP guidance.

12. Conversational surface options (corrected 2026-06-09)

Earlier drafts loosely said "native Slack AI" was the surface. That is wrong: native Slack AI and a task agent are different categories.

  • Native Slack AI is NOT our surface. It is the built-in comprehension suite: Summarize (channel/thread/DM), daily recaps, natural-language search, Slackbot catch-up. You invoke it via menu buttons and the search bar, not by @mentioning a named agent. It reads/retrieves Slack's own content and does not call our MCP tools or run multi-step tasks. It is a useful COMPLEMENT, not a replacement for seahaven-slack-bot / exec-aide.
  • The task-agent surface (what actually replaces the bots) is a named agent users @mention or DM that reasons and takes action via tools. Three real candidates:
Surface UX Calls our MCP servers Model control Notes
Agentforce (recommended) Named AI teammates in an Agents tab; @mention in channels/DMs/threads; per-agent DM history Yes Salesforce-default (GPT-4o mix) OR BYOLLM (Bedrock/Azure/OpenAI/Vertex) → can keep our Claude + guardrail Configured in Agentforce Studio (Topics/Actions); pulls Salesforce/Agentforce into the stack (licensing)
Marketplace agent (e.g. Claude app) @mention/DM a vendor agent Via Slackbot-MCP-client outward tools Vendor's models (Claude app = Claude) Least config; least control over tool wiring/scoping
Custom Bolt assistant Assistant pane / @mention; we build the UX Yes, we wire directly Whatever we call (incl. our Bedrock Claude + guardrail) Most control, most build/maintenance; contradicts the "no Bolt backend" goal

Why Agentforce fits our design

Agentforce supports multiple specialized named agents, which maps almost 1:1 onto our trust-tiered servers + Google-group scoping:

  • Seahaven-Ops agent → sh-mcp-ops, granted to sh-mcp-ops@ / sh-mcp-assistant@.
  • Seahaven-Finance agent → sh-mcp-finance, granted to sh-mcp-finance@.
  • Lauren's Exec agent (DM-scoped) → ops *:self tools + finance read.

Each agent is a separate teammate with its own published audience and tool access, which reinforces the trust-tier separation at the UX layer, not just in the token.

Decision needed

Pick the task-agent surface: Agentforce (recommended — native multi-agent, MCP support, and BYOLLM-on-Bedrock to retain Claude + our guardrail), a marketplace agent (fastest, least control), or a custom Bolt assistant (most control, most build). This is now the top open decision because it determines the model, the guardrail story, and whether Salesforce enters the stack. The MCP servers + auth design below are identical across all three.

Sources: slack.com/help (Guide to AI features in Slack; Use Agentforce in Slack); slack.com/blog/news/turn-agents-into-teammates-with-slack; slack.dev illustrated Agentforce guide; developer.salesforce.com Agentforce supported-models + BYOLLM.