Design doc for the Sea Haven MCP platform that replaces seahaven-slack-bot and exec-aide: trust-tiered MCP servers, Google-SSO via a Cognito broker, TypeScript monorepo. Cross-review (cross_reviewer/GPT-4.1) folded in. Status: planning only, nothing built.
34 KiB
Sea Haven MCP Platform — Auth Architecture, Scope Matrix & Build Plan
Status: DRAFT for review (cross-review gate not yet run) Date: 2026-06-09 Owner: Adam Moussa
Supersedes the ad-hoc design discussion. This is the consolidated plan for deprecating
seahaven-slack-bot and exec-aide in favor of Slack AI agents backed by a set of
Sea Haven MCP servers, with authentication centralized on Google Workspace SSO.
1. Decisions locked in
This is a greenfield service. seahaven-slack-bot and exec-aide are deprecated
completely (repos archived, stacks decommissioned, Bedrock agents + guardrails torn down).
Nothing is migrated as a running component. Where proactive functionality is still needed
(Lauren's email classifier/digest, the KB syncs) it is rebuilt fresh in this new service's
jobs/ (§5), not carried over.
- Conversational surface = a configurable Slack task agent that users @mention (in channels, threads, DMs) and that calls our MCP tools to do work end to end. Slack hosts the agent, the LLM, and the orchestration; we run no Bolt backend of our own (unless we pick the custom-assistant option). Our system is: MCP servers (the tools) + the Cognito/Google auth broker + the rebuilt jobs. IMPORTANT (corrected, §12): this is NOT "native Slack AI." Native Slack AI is the summarize / recap / search feature suite over Slack's own content; it does not act as a configurable agent that calls our tools, so it cannot replace the bots. The real surface candidates are Agentforce (recommended), a marketplace agent (e.g. the Claude app), or a custom Bolt assistant. Surface choice is still OPEN (§12) and it drives the model + guardrail story. Guardrail consequence (investigated, §11): Slack's native guardrails cover prompt-injection and content safety, but do NOT redact PII inside our tool responses, so we rebuild PII masking at the MCP layer (§2.5) regardless of surface. (If we pick Agentforce + BYOLLM-on-Bedrock, we could also keep our own Bedrock guardrail on the model path — defense in depth, not a substitute.)
- Tools = Sea Haven MCP servers, grouped by trust tier (
ops,finance;physicaldeferred). - Identity = Google Workspace, the same IdP already used for Slack SSO. Google is the single source of truth for authentication AND group membership.
- Authorization = Amazon Cognito user pool federated to Google, acting as the OAuth 2.1 authorization server that issues scoped, audience-bound tokens the MCP servers validate. Slack's built-in agent is the MCP client; each user does per-user OAuth consent against the broker (Slack's MCP client supports per-user OAuth), so tool calls carry the real user identity.
- Language = TypeScript for the entire monorepo (servers, packages, CDK, rebuilt jobs). Packages are shared by both the MCP servers and the rebuilt scheduled jobs (one client per service everywhere); a single language is what makes that sharing real.
2. Auth architecture
2.1 Planes
Three planes kept separate (the credential that reaches the broker is never the credential that reaches a downstream API):
- Inbound auth (who): Cognito-issued JWT, identity federated from Google. Carries
sub(the user's Google identity),aud(the target MCP server), andscopeclaims. - Access control (what): each MCP server validates issuer + audience and enforces a required scope per tool. Tools whose scope the token lacks are hidden from the user.
- Outbound auth (how we reach the API):
- Google-API tools (Gmail, Calendar): act as the user via a direct per-user Google OAuth grant (see 2.4).
- Non-Google tools (QBO, payments, Lenel, 3CX, Afi, Front, Maps): broker token authorizes; the server uses its own service credentials from Secrets Manager.
2.2 Flow
User (already Google-SSO'd into Slack)
│ OAuth 2.1 auth-code (PKCE), federated login → Google (silent if logged in)
▼
Cognito user pool (Google = external OIDC IdP)
│ pre-token-generation Lambda: group → scope mapping
▼
Access token (JWT):
sub = lauren@seahavenind.com
aud = sh-mcp-ops
scope = "ops:read ops:tasks gmail:self calendar:self"
│
├──► sh-mcp-ops validates aud+scope, runs permitted tools
├──► sh-mcp-finance token lacks finance:read → finance tools hidden
└──► sh-mcp-physical token lacks physical:* → hidden
2.3 Group → scope mapping
Google Groups are the single control point. Membership is synced to Cognito groups by a scheduled sync Lambda (every 5 min, tightened from hourly per cross-review — group changes are infrequent but revocation latency on sensitive groups must be low). The sync Lambda has a CloudWatch ALARM on failure (a silent sync failure freezes group state). The Cognito pre-token Lambda maps Cognito group → scope claims at token mint, so the hot path makes no Directory API call. (Alternative: pre-token Lambda calls the Google Directory API with a short cache. Chose the sync approach to keep token minting fast and resilient to Directory API throttling.)
Revocation: a 5-min sync + 60-min token TTL means a removed user could retain access for ~65 min
worst case. For finance:*, TTL is shortened to 15 min, and a Cognito-backed deny-list (checked
by the servers) gives immediate hard revocation when a user is pulled from a sensitive group.
| Google Group | Members | Scopes granted |
|---|---|---|
sh-mcp-ops@ |
all staff | ops:read |
sh-mcp-assistant@ |
Lauren, Adam | ops:read, ops:tasks, gmail:self, calendar:self |
sh-mcp-finance@ |
Adam, Lauren, accounting | ops:read, finance:read |
sh-mcp-admin@ |
Adam | all of the above + finance:admin |
A user in multiple groups gets the union of scopes (Lauren is in -assistant@ and -finance@,
so she gets ops:read ops:tasks gmail:self calendar:self finance:read).
sh-mcp-physical@ / physical:* scopes are NOT issued to any agent surface at launch — the
physical tier is admin/out-of-band only (see §3 and §6).
2.4 Google-API tools (Gmail / Calendar) — outbound detail
Pulling the upstream Google API access token through Cognito federation is limited (Cognito
requests openid profile email and does not cleanly expose/refresh Google API scopes like
gmail.readonly). So the clean implementation is:
- Cognito-federated-Google handles authentication + authorization scopes (single SSO login).
- The Gmail/Calendar MCP tools do a separate direct per-user Google OAuth grant at first
use (incremental consent for
gmail.readonly+calendar), storing a per-user refresh token encrypted (Secrets Manager or a KMS-encrypted DynamoDB item, key = usersub). - Both consents resolve to the same Google account, so identity stays consistent; it is one extra one-time consent, not a second identity.
Result: search_inbox / get_calendar_events act as the signed-in user. Lauren physically
cannot read anyone else's mailbox; segregation is enforced by Google, not by our code.
Token-handling rules (cross-review BLOCK items):
- The Google client is instantiated with the minimal scope for the requested tool only
(
gmail:self→gmail.readonly,calendar:self→calendar), never the union of a user's scopes. Agmail:self-only user cannot mint a Calendar token and vice versa. - The per-user refresh token is never returned to the Slack agent or any client, and never appears in logs or error messages. Access tokens are minted per request, not cached or reused.
- Per-user refresh tokens are stored in a KMS-CMK-encrypted DynamoDB item keyed by user
sub, with the key policy scoped to the owning server role only. IAM on that table is partitioned bysub(ABACdynamodb:LeadingKeys) so a server can only read the calling user's token, not all users' tokens.
2.5 Token & endpoint hardening
- Short-lived access tokens (60 min general; 15 min for
finance:*); refresh by the client. - Audience-bound: each server validates
audand rejects tokens not minted for its resource server. Anopstoken presented tofinanceis rejected. - Server-side enforcement is authoritative: every server independently validates issuer +
aud+ required scope per tool on every call. Tool-hiding in the agent UI is a convenience, never the access-control boundary. The client is never trusted. - No broker-JWT passthrough: the Cognito JWT is never forwarded to a downstream API. Servers reach external APIs only via their own service credentials or the per-user Google token (§2.4).
- Endpoints are PUBLIC HTTPS (decided — see §11). Slack's agent runtime is cloud-hosted and initiates outbound Streamable-HTTP to a public MCP URL with an OAuth bearer, so VPC-only / PrivateLink is not an option. The OAuth bearer (Cognito) is the real access boundary; in front of it: AWS WAF, plus IP-allowlisting to Slack's egress ranges IF Slack publishes them (to verify). Do not depend on mTLS — Slack's MCP client is not documented to present client certs. Never an unauthenticated path. PKCE required if any public client is added.
- MCP-layer PII redaction (decided — see §11): Slack's native guardrails do NOT redact PII
inside our tool responses, so the Bedrock guardrail's PII-anonymization is reimplemented in the
servers' output path. Every tool response is inspected and sensitive fields (bank account/
routing, card, SSN) masked before it leaves the server, heaviest on
finance. - Physical tier step-up:
physical:*tools require recent re-auth (auth_time/max_age) AND an out-of-band human approval before execution. Scope grants the right to request, never to execute unattended. - Audit: every
finance:*andphysical:*tool call logged to CloudWatch (usersub, tool, args hash, decision, result). CloudWatch ALARM on anyphysical:*invocation (ALARM-state only, per standing alarm preference). - Connection-level governance: which MCP connectors a user can even add is also gated by group,
so
physicalis unreachable for non-members regardless of token contents (defense in depth). - IAM least-privilege per server (cross-review BLOCK): each server role gets
secretsmanager:GetSecretValueon only its own secrets (no wildcard) and KMS decrypt on only the relevant CMK.sh-mcp-financecan read the QBO + payments secrets but not Maps or per-user Google tokens;sh-mcp-opscannot read finance secrets. Access to QBO and per-user Google tokens is logged and alarmed on anomalous patterns (mass/unexpected-role access). - Prompt-injection containment (no Bedrock guardrail to fall back on): tool-call parameters
are validated server-side against strict schemas/allow-lists, never trusted from the agent.
Tool-returned content is treated as data, never instructions. Side-effecting
opstools (create_calendar_event,create_task,create_reminder) get sane bounds; the trust-tier split keeps any injection blast radius insideops(no finance/physical reachable in-session).
3. Scope matrix (every tool → required scope, outbound auth, risk)
sh-mcp-ops (agent-facing, read-mostly)
| Tool | Source pkg | Scope | Outbound auth | Risk |
|---|---|---|---|---|
| lookup_work_order | internal-data | ops:read |
service creds (DDB ro) | low |
| lookup_purchase_order | internal-data | ops:read |
service creds (DDB ro) | low |
| lookup_site | internal-data | ops:read |
service creds (DDB ro) | low |
| search_knowledge_base | knowledge-base | ops:read |
service (Bedrock Retrieve) | low |
| search_nearby_vendors | google-maps | ops:read |
service (Maps API key) | low (external $, rate-limit) |
| search_inbox | gmail | gmail:self |
forwarded user Google token | med (private data) |
| get_email_thread_detail | gmail | gmail:self |
forwarded user Google token | med |
| get_calendar_events | calendar | calendar:self |
forwarded user Google token | low |
| create_calendar_event | calendar | calendar:self |
forwarded user Google token | med (writes, external attendees) |
| check_availability | calendar | calendar:self |
forwarded user Google token | low |
| create_task | tasks | ops:tasks |
service creds (DDB, partitioned by sub) | low |
| list_tasks | tasks | ops:tasks |
service creds | low |
| complete_task | tasks | ops:tasks |
service creds | low |
| delete_task | tasks | ops:tasks |
service creds | low |
| create_reminder | reminders | ops:tasks |
service creds (EventBridge Scheduler) | low |
sh-mcp-finance (sensitive, read-only, fully audited)
| Tool | Source pkg | Scope | Outbound auth | Risk |
|---|---|---|---|---|
| search_vendors | qbo | finance:read |
service creds (QBO OAuth, server-held) | med (token can read all of QBO) |
| lookup_payment_by_vendor | payments | finance:read |
service creds (DDB PaymentsDashboard) | med |
| lookup_payment_by_invoice | payments | finance:read |
service creds (DDB) | med |
| lookup_payment_by_check | payments | finance:read |
service creds (DDB) | med |
QBO OAuth maintenance (/qbo/connect, /qbo/callback, /qbo/disconnect) stays as
admin web endpoints (API Gateway), NOT exposed as agent tools. Gated by finance:admin.
sh-mcp-physical (DECIDED: admin / out-of-band only — NOT in the launch build)
Neither bot does physical actions today, so this tier is documented to keep the trust model
complete but is DEFERRED. No physical:* scope is issued to any agent. The tools below stay as
admin/out-of-band capability (existing door-unlock-api path) until a future decision to expose them.
| Tool | Source pkg | Scope | Outbound auth | Risk |
|---|---|---|---|---|
| request_door_unlock | lenel | physical:unlock |
service creds (Elements API) | HIGH — human approval required |
| initiate_lockdown | lenel | physical:lockdown |
service creds | HIGH — Adam only + approval |
| release_lockdown | lenel | physical:lockdown |
service creds | HIGH — Adam only + approval |
| push_xml | yealink | physical:telephony |
service creds (Push XML) | med + approval |
| reroute_call | threecx | physical:telephony |
service creds (3CX xapi) | med + approval |
| set_dnd | threecx | physical:telephony |
service creds (3CX xapi) | med + approval |
Trifecta rationale
The dangerous combination (read-from-untrusted-source + sensitive-action in one session) is
prevented by tier separation: finance and physical are never reachable by a session that
also holds the Gmail/web read tools. Within ops, the actions (create_task, create_reminder,
create_calendar_event, search_nearby_vendors) are low-consequence, so injected content in an
email can at worst create a spurious task or run a Maps query — not move money or unlock a door.
Outbound items in ops (calendar invites to external attendees) are flagged for monitoring.
4. Monorepo layout
Packages are consumed by BOTH the MCP servers and the surviving scheduled Lambdas — one client per external service, used everywhere. This is the core payoff over per-service repos.
sh-mcp/ (one monorepo, org: Sea-Haven-Industries)
packages/
shared/ MCP scaffold, JWT validation, requires_scope guard, secrets, audit log
qbo/ google-maps/ internal-data/ payments/ knowledge-base/
gmail/ calendar/ tasks/ reminders/ notion/
# DEFERRED (physical tier, not in launch build): lenel/ yealink/ threecx/
servers/
sh-mcp-ops/ → remote HTTP MCP stack (CDK)
sh-mcp-finance/ → remote HTTP MCP stack (CDK)
# DEFERRED: sh-mcp-physical/ (admin/out-of-band only)
auth/
cognito/ user pool, Google federation, resource servers, app clients
pre-token-lambda/ group → scope mapping
group-sync-lambda/ Google Groups → Cognito groups (hourly)
jobs/ (surviving scheduled Lambdas, import packages/)
notion-sync/ po-sync/ workorder-sync/ (feed Bedrock KB)
fetch-classify/ daily-digest/ reminder/ (exec-aide proactive, Lauren)
5. Proactive jobs — rebuilt fresh in sh-mcp/jobs (NOT MCP, NOT migrated)
Proactive/event-driven work has no conversational equivalent and Slack's reactive agent cannot replace it, so it is rebuilt in the new service as TypeScript Lambdas that import the shared packages and use stored offline credentials (not interactive SSO). The old exec-aide/slack-bot Lambdas are decommissioned, not refactored in place.
| Job | Trigger | Notes |
|---|---|---|
| fetch-classify | every 15 min | Gmail history → Haiku classify → HIGH DM. Offline Google refresh token (exec-aide/gmail-oauth). |
| daily-digest | 5pm ET M-F | Query DDB → digest → DM. |
| reminder (one-shot) | Scheduler | Fired by the create_reminder tool. |
| notion-sync | 02:00 UTC | Notion → S3 → Bedrock KB ingestion. |
| po-sync | 02:00 UTC | DDB purchase-orders → S3 → KB. |
| workorder-sync | 02:00 UTC | DDB work orders/comments → S3 → KB. |
The Bedrock agents (seahaven-alex, exec-aide Sonnet loop) are what get deprecated. The
Bedrock guardrail (PII anonymization, prompt-attack filtering) must be explicitly replaced
on the Slack surface before deletion — do not assume the new surface provides parity.
6. Phasing
- Foundation: monorepo +
shared+ CI/CD (reusable workflows, deploy role first, OIDC, Dependabot, naming) + the test harness and coverage gate (see §7). Stand up Cognito + Google federation + pre-token + group-sync. Prove SSO login end to end. - Pilot —
sh-mcp-opswired to theseahaven-slack-botreplacement agent: vendor + WO/PO/site + KB. Validate per-user auth, tool-hiding, and parity vs the old Bedrock agent. sh-mcp-financewith full audit logging once ops is proven.- Convert Lauren → Slack AI agent: add
gmail:self/calendar:self/ops:taskstools to ops, do the per-user Google grant, keep her agent DM/personal-scoped. Refactor her scheduled Lambdas to import the packages. sh-mcp-physicalis DEFERRED — admin/out-of-band only, not part of this build.- Deprecate each bot only after its replacement is proven at parity. Lauren's workflow
needs sign-off before
exec-aideis retired.
7. CI/CD & test suite
7.0 Language: TypeScript (DECIDED)
TypeScript for everything: servers, packages, CDK, and the rebuilt jobs. Toolchain is
tsc --noEmit + eslint/prettier + vitest. Since the old stacks are fully deprecated (not
migrated), exec-aide's former Python logic (Gmail/Calendar/Haiku classify) is rewritten in TS
as part of the greenfield build, not ported line-for-line. No polyglot matrix needed.
7.1 Reusable-workflow pattern (Sea Haven standard)
Per the handbook cicd.md, repos call org reusable workflows rather than inlining steps. Two
caller workflows:
.github/workflows/ci.yaml— onpull_request. Calls the org reusableciworkflow. Aggregated under a single required check namedci / cifor branch protection (so adding matrix legs never breaks the required-context list)..github/workflows/deploy.yaml— on push tomain(deploy-then-merge per git-workflow). Calls the org reusablecd-cdkworkflow per server, OIDC into the deploy role.
Gotchas baked in from prior repos:
permissions:declared in the CALLER (reusable-workflow permissions don't inherit — missing them yieldsstartup_failure).concurrencykeyed to caller context, not the reusable workflow, to avoid cross-PR collision.- ARM64 Lambda containers:
enable-qemu: truein BOTH ci.yaml and deploy.yaml, ANDplatform: LINUX_ARM64on every Docker image asset (Lambda) or "exec format error". cd-cdknpm-ci guard since CDK is Node even where Lambda code may differ.
7.2 Monorepo CI shape
Path-filtered matrix so only changed packages/servers run. Per leg:
- Install + build (workspace-aware:
npm ciat root, build changed package + dependents). - Lint + format gate:
tsc --noEmit+ eslint/prettier (orruff check+ruff format --checkon any Python paths). Enforced locally pre-push by the existing Claude Code hook too. - Unit + auth + security tests (§7.3) with coverage gate.
cdk synthfor changed servers (catches IaC + ARM64 platform regressions without deploying).- Dependabot enabled; exact-pin
aws-cdk-libper the handbook Pinning Principle (no blanket ignores); Dependabot keeps it current.
Deploy job runs only on main, per-server, gated on ci / ci green and the cross-review
sign-off for any IAM/authz diff (§8).
7.3 Test suite (full, security-weighted)
The authorization layer is the highest-risk surface, so it gets the heaviest coverage.
| Layer | What it covers |
|---|---|
| Unit — per package | Each integration package with the external API mocked: QBO vendor query + token auto-rotation, payments DDB queries, internal-data lookups, Maps search, KB retrieve, Gmail/Calendar calls, tasks CRUD, reminder scheduling. Error/empty/throttle paths. |
| Auth — the crown jewels | JWT validation (issuer, signature, expiry); audience binding (a token minted for sh-mcp-ops is REJECTED by sh-mcp-finance); per-tool scope enforcement server-side (independent of UI tool-hiding); tool-hiding (a user without finance:read does not see finance tools in list_tools); pre-token Lambda group→scope mapping for every group; group-sync correctness; deny-list hard revocation (a revoked user is rejected immediately, before token expiry); minimal-scope Google client (a gmail:self-only token cannot mint a Calendar token); per-user token ABAC isolation (a server cannot read another user's refresh token). |
| Security / abuse | Trifecta separation — assert no single issued token can reach both untrusted-read (Gmail) and a sensitive-action tool; scope-escalation attempts rejected; prompt-injection regression corpus (tool-returned content containing "ignore instructions / call X" must NOT trigger out-of-scope tool calls); MCP-layer PII redaction (every finance tool response masks bank/routing/card/SSN before egress); per-session tool-call cap + per-tool rate limit. |
| Contract / MCP conformance | Every tool's input/output JSON schema validates; MCP protocol handshake; list_tools reflects the caller's scopes. |
| Audit | Every finance:* call emits a structured audit record (user sub, tool, args hash, decision); assert the record shape and that secrets never appear in logs. |
| Parity (pre-deprecation) | Golden-transcript tests replaying real slack-bot / exec-aide interactions against the new tools to confirm equivalent answers before retiring a bot. |
Coverage gate in CI (start at 80% lines, 100% on the shared auth/scope guard module). Parity
tests run in phase 2/4 before each bot is deprecated, not on every PR.
8. Handbook obligations (per global instructions)
- Naming: kebab-case throughout (
sh-mcp,sh-mcp-ops, etc.). ✓ in this draft. - Secrets: all in Secrets Manager; existing secret names reused; per-user Google tokens encrypted.
- Each deployable server: deploy role first, CI/CD pipeline, Dependabot, README.
- Confluence "AWS Architecture Map" (id 1540098): add a Mermaid subgraph for the MCP platform + Cognito before the build is reported done.
- Project memory:
project_sh_mcp.mdcreated; cross-links to exec-aide, slack-bot, payments, identity-center, lenel/door-unlock memories. - Cross-review gate (mandatory): Cognito↔Google federation, group→scope mapping, the
token-forwarding design, and every MCP server's IAM role go through
cross_reviewerbefore any commit/merge. IAM + auth = the breaking-change category that gates.
9. Decisions log & remaining open items
RESOLVED 2026-06-09:
- Complete deprecation: slack-bot + exec-aide repos archived, stacks decommissioned, no migration. Greenfield service; conversational surface = Slack's built-in AI agent (no Bolt backend, no own model).
- Lauren gets
finance:read(member of-assistant@+-finance@). sh-mcp-physicalis admin/out-of-band only, deferred, not in the launch build.- Afi removed from the stack entirely.
- Launch tool scope = exactly what the two bots do today (no Front / Afi / 3CX-read adds).
- Language = TypeScript for the entire monorepo.
INVESTIGATED 2026-06-09 (both former open items — see §11):
- Endpoint model: PUBLIC HTTPS + OAuth + WAF (VPC/PrivateLink not viable; mTLS not relied on).
- Guardrail parity: Slack covers prompt-injection/content safety; PII redaction is NOT covered and is rebuilt at the MCP layer.
TOP OPEN DECISION (drives model + guardrail + whether Salesforce enters the stack — see §12):
- Task-agent surface: Agentforce (recommended), a marketplace agent (Claude app), or a custom Bolt assistant. Native Slack AI is a complement, not a candidate. MCP + auth design is identical across all three, so this does not block foundation work.
RESIDUAL VERIFY (at build, not blocking design):
- Whether Slack publishes egress IP ranges for custom MCP connectors (for WAF IP-allowlisting) and whether its MCP client supports client certs — confirm with Slack docs/support.
- Exact Slack-AI-guardrail config knobs available to us as the tool provider.
10. Cross-review (cross_reviewer / GPT-4.1, 2026-06-09)
Verdict: design fundamentally sound. BLOCK/FIX items below were folded into §2; remaining are operational and tracked for the build. Full output in the conversation log.
BLOCK (resolved in §2.4 / §2.5):
- Minimal-scope Google client per tool, no scope union, refresh token never exposed/logged, access tokens not cached. → §2.4
- Audience + scope enforced server-side at every boundary; tool-hiding is not the boundary. → §2.5
- No broker-JWT passthrough to downstream APIs. → §2.5
- IAM least-privilege per server role; per-user Google tokens partitioned by
sub(ABAC), KMS key policy scoped to role. → §2.4 / §2.5
FIX (resolved / tracked):
- Group-revocation latency: sync tightened to 5 min,
finance:*TTL → 15 min, deny-list for immediate hard revocation, ALARM on sync failure. → §2.3 - Prompt-injection: server-side param validation + allow-lists, tool output treated as data. → §2.5
- Audit access to sensitive secrets. → §2.5
- Google refresh-token rotation/revocation on group removal. → build task.
NIT: PKCE if a public client is added (§2.5); keep pre-token Lambda fast/no external calls (§2.3).
These auth/IAM specifics are signed off by the cross-review gate. Re-run cross_reviewer on the
actual IAM policy JSON and pre-token Lambda code once written (the reviewer asked for the concrete
artifacts for a deeper pass).
11. Investigation: endpoint connectivity & guardrail parity (2026-06-09)
11.1 How does Slack's built-in agent reach our MCP servers? → PUBLIC HTTPS
Slack's agent platform has two MCP directions, and ours is the outbound one:
- Inbound (
https://mcp.slack.com/mcp): external clients connect IN to act on Slack. Not us. - Outbound (our case): Slack's built-in agent / Slackbot acts as an MCP client and reaches
OUT to external MCP servers, configured by URL +
Authorizationbearer, with an OAuth callback likehttps://mcp.<our-domain>/upstream-auth/callback. Transport is JSON-RPC 2.0 over Streamable HTTP (Slack does not support SSE or Dynamic Client Registration).
Implication: the connection originates from Slack's cloud to a public, internet-reachable HTTPS endpoint. VPC-only / PrivateLink is therefore not viable for the Slack surface. Decision:
- API Gateway (HTTP API) or ALB, public, behind AWS WAF.
- The Cognito OAuth bearer is the access boundary (we already validate issuer/aud/scope).
- IP-allowlist Slack's egress ranges in WAF IF Slack publishes them (unconfirmed in docs; it is a known open question in the custom-MCP-connector community — verify with Slack).
- Do not rely on mTLS: Slack's MCP client is not documented to present client certificates.
(If we later add non-Slack MCP clients that run inside our network, e.g. internal tools, those specific servers could be VPC-only — but anything Slack consumes must be public.)
11.2 Do Slack's safeguards replace the Bedrock guardrail? → PARTIALLY
Slack provides an enterprise "Slack AI guardrails" framework that DOES cover:
- Prompt injection / jailbreak: context engineering to mitigate injection, real-time jailbreak/prompt-attack detection, URL filtering, output validation.
- Access governance: AI only accesses what the user is authorized to see; admin control over which data sources and tools each assistant can reach; per-user/feature toggles.
- Audit: comprehensive logging of what each assistant accessed; real-time monitoring; incident-investigation ("who asked what").
- Zero-training guarantee; models run without outbound network access.
What it does NOT cover, and is the gap we must own:
- PII redaction of OUR tool responses. Slack's DLP/tombstoning applies to Slack messages and
AI summaries derived from them, not to payloads our MCP tools return. Industry guidance is
explicit that regulated data needs MCP-layer DLP that inspects and redacts every tool response
before it reaches the model. So the Bedrock guardrail's PII-anonymization is reimplemented in
our servers' response path (§2.5), heaviest on
finance(bank/routing/card/SSN).
Net: prompt-attack protection transfers to Slack acceptably (and was arguably the surface's job anyway); PII protection does not transfer and stays our responsibility at the MCP layer.
Sources: docs.slack.dev/ai/slack-mcp-server; slack.com/blog/transformation/securing-the-agentic- enterprise; slack.com/blog/news/how-we-built-slack-ai-to-be-secure-and-private; custom-MCP- connector egress-range discussion (OpenAI dev community); strac.io MCP-layer DLP guidance.
12. Conversational surface options (corrected 2026-06-09)
Earlier drafts loosely said "native Slack AI" was the surface. That is wrong: native Slack AI and a task agent are different categories.
- Native Slack AI is NOT our surface. It is the built-in comprehension suite: Summarize (channel/thread/DM), daily recaps, natural-language search, Slackbot catch-up. You invoke it via menu buttons and the search bar, not by @mentioning a named agent. It reads/retrieves Slack's own content and does not call our MCP tools or run multi-step tasks. It is a useful COMPLEMENT, not a replacement for seahaven-slack-bot / exec-aide.
- The task-agent surface (what actually replaces the bots) is a named agent users @mention or DM that reasons and takes action via tools. Three real candidates:
| Surface | UX | Calls our MCP servers | Model control | Notes |
|---|---|---|---|---|
| Agentforce (recommended) | Named AI teammates in an Agents tab; @mention in channels/DMs/threads; per-agent DM history | Yes | Salesforce-default (GPT-4o mix) OR BYOLLM (Bedrock/Azure/OpenAI/Vertex) → can keep our Claude + guardrail | Configured in Agentforce Studio (Topics/Actions); pulls Salesforce/Agentforce into the stack (licensing) |
| Marketplace agent (e.g. Claude app) | @mention/DM a vendor agent | Via Slackbot-MCP-client outward tools | Vendor's models (Claude app = Claude) | Least config; least control over tool wiring/scoping |
| Custom Bolt assistant | Assistant pane / @mention; we build the UX | Yes, we wire directly | Whatever we call (incl. our Bedrock Claude + guardrail) | Most control, most build/maintenance; contradicts the "no Bolt backend" goal |
Why Agentforce fits our design
Agentforce supports multiple specialized named agents, which maps almost 1:1 onto our trust-tiered servers + Google-group scoping:
- Seahaven-Ops agent →
sh-mcp-ops, granted tosh-mcp-ops@/sh-mcp-assistant@. - Seahaven-Finance agent →
sh-mcp-finance, granted tosh-mcp-finance@. - Lauren's Exec agent (DM-scoped) → ops
*:selftools + finance read.
Each agent is a separate teammate with its own published audience and tool access, which reinforces the trust-tier separation at the UX layer, not just in the token.
Decision needed
Pick the task-agent surface: Agentforce (recommended — native multi-agent, MCP support, and BYOLLM-on-Bedrock to retain Claude + our guardrail), a marketplace agent (fastest, least control), or a custom Bolt assistant (most control, most build). This is now the top open decision because it determines the model, the guardrail story, and whether Salesforce enters the stack. The MCP servers + auth design below are identical across all three.
Sources: slack.com/help (Guide to AI features in Slack; Use Agentforce in Slack); slack.com/blog/news/turn-agents-into-teammates-with-slack; slack.dev illustrated Agentforce guide; developer.salesforce.com Agentforce supported-models + BYOLLM.