| D2 | Deploy as **Agentforce Employee Agents in Slack** | Slack is our surface; only Employee Agents deploy to Slack | `FACT` slack.com/help 36218109305875 |
| D3 | **Refine design.md §12**: remove `finance:read` from Lauren Exec; finance lookups go through the Finance agent | Gmail-read + finance-read + a write/egress tool in one session is the exfiltration trifecta | design.md §2.5, §3 (trifecta); §5 |
| D4 | **RESOLVED (2026-06-11, §0.1):** reach MCP tools via **GA External Service (OpenAPI) actions** (or Apex `@InvocableMethod`) authenticated by a **Per-User OAuth 2.1 Browser Flow External Credential → Cognito**. Native remote-MCP connector is NOT used for the per-user hop. | Per-user OAuth is GA; native custom-MCP per-user binding is unconfirmed and documented per-user identity exists only for SF-**hosted** MCP. Trade-off: lose the MCP transport on the Agentforce→tool hop (tools re-exposed as OpenAPI/Apex). | `FACT` research wf_1bf9e142; Named Credentials OAuth dev guide; G1/G2 §5 |
| D5 | **RESOLVED (2026-06-11, §0.1):** agent reasoning/planner model = **AWS-Hosted Claude (Salesforce-managed, on Bedrock)**. **BYOLLM is NOT used for the planner.** | BYOLLM is documented only for custom actions, never as a planner option, and routes back through Salesforce's Models API/Trust Layer anyway. AWS-Hosted is the only documented way to keep Claude as planner. | `FACT` research wf_1bf9e142; developer.salesforce.com supported-models |
| D6 | KB → **Agentforce Data Library (unstructured) on Data Cloud**, replacing Bedrock KB `LSDCNHTH6O` + AOSS | Data Libraries auto-create a Data Cloud search index + retriever; managed RAG | `FACT` Trailhead/Atrium Data Libraries; design.md §3 |
| D7 | **Retire**`po-sync` / `workorder-sync` KB feeds; structured WO/PO/payments stay **live MCP lookup tools** | Data Libraries are unstructured-only; structured data belongs in DDB-backed MCP tools | `FACT` Data Libraries unstructured-only; design.md §3 |
| D8 | Proactive jobs (`fetch-classify`, `daily-digest`, `reminder`, `notion-sync`) **stay as our scheduled Lambdas**, not Agentforce; HIGH-priority alerts continue as **Slack DMs** | No conversational equivalent; keeps Haiku classifier in our Bedrock account; cheapest reliable path | design.md §5; §1f below |
| D9 | New separate **`sh-agentforce` SFDX repo** for agent metadata (Bot/GenAiPlannerBundle/GenAiPlugin/GenAiFunction) under Agentforce DX | Different toolchain (sf CLI, scratch orgs, deploy-to-org) than the CDK monorepo; one deploy target per repo | `FACT` developer.salesforce.com Agent DX metadata; handbook |
| D10 | Build **Ops + Finance agents now**; HR/Gusto, SA8000 Q&A, IT/onboarding **later, corpus-gated**; add near-term capabilities as **Topics**, not new agents | Front/ops corpus is rich and current; HR/Amazon/SA8000 corpus is thin/stale/missing | Notion (below); design.md corpus notes |
| D11 | **REVISED (2026-06-11, §0.1):** guardrail (prompt-attack + PII redaction) lives **entirely at the MCP/action layer**, not on any model path; the Einstein Trust Layer covers planner-path moderation | The planner runs AWS-Hosted Claude inside Salesforce's boundary (D5) — there is no BYOLLM model path to attach our guardrail to; Trust Layer does not redact OUR tool-response payloads, so MCP-layer redaction is authoritative | design.md §2.5, §11.2; §0.1 |
| D12 | **RESOLVED (2026-06-11, §0.1):****OpenAPI-first, MCP-ready transport-agnostic core.** Tool handlers + the shared auth/scope/PII/audit guard are built independent of transport; ship the **OpenAPI adapter now** (Agentforce External Service actions; Apex only where an action needs request/response shaping); **defer the MCP adapter** to a named trigger. | Agentforce's per-user hop needs OpenAPI regardless (D4); no committed MCP consumer today, so a second public interface + its conformance/hardening cost isn't justified yet; a transport-agnostic core makes the MCP adapter a cheap later add, not a re-platform; deferring it also drops the IDE-session trifecta nuance until then. | Adam 2026-06-11; research wf_1bf9e142 (D4) |
**Agent roster (the answer to "how many"):** **3** — Seahaven Ops, Seahaven Finance, Lauren Exec (§1a, §2).
**Top 5 decisions for Adam — ALL RESOLVED 2026-06-11 (§0.1):** (1) Agentforce-in-Slack + licensing **accepted**;
(2) reasoning model = **AWS-Hosted Claude** (BYOLLM ruled out for the planner); (3) **D3 accepted** (strip finance
from Lauren Exec); (4) **D4** per-user auth = ES/Apex actions + Per-User Browser Flow → Cognito; (5) corpus
| 2. Reasoning model | **AWS-Hosted Claude (SF-managed)** as the planner. BYOLLM ruled out for the planner (only Salesforce Default / AWS-Hosted drive the Atlas reasoning engine; BYOLLM is custom-action-only and still routes through SF's Models API/Trust Layer). | research wf_1bf9e142 (Q2) |
| 4. Per-user auth path (D4) | **External Service (OpenAPI) / Apex actions + Per-User OAuth Browser Flow External Credential → Cognito.** Native remote-MCP connector is Beta and its per-user binding is unconfirmed — do NOT depend on it for the per-user hop. | research wf_1bf9e142 (Q1) |
| 5. Fund SA8000/handbook corpus (G8) | **DEFERRED — fund when needed** (not now). SA8000/compliance + HR agents stay blocked until content is authored; the agents must decline rather than hallucinate in the interim. | Adam, 2026-06-11 |
**Two consequences that ripple into the design (must be honored downstream):**
1.**The Agentforce→tool hop is NOT MCP.** Because per-user identity is only achievable on the GA External
Service/Apex action path (not the Beta MCP connector), Agentforce reaches our tools over **OpenAPI/Apex
actions** authenticated per-user to Cognito. **Day-one architecture (B3-resolved):** there is NO MCP wire
protocol at launch — the trust-tier servers (`sh-mcp-ops`, `sh-mcp-finance`) expose their tools as **OpenAPI
endpoints** via the transport-agnostic core (D12); the MCP adapter is deferred (D12 trigger). The servers'
JWT validation, per-tool scope enforcement, audience-binding, PII redaction, and audit are **unchanged** —
they sit below the adapter and are transport-agnostic. Browser Flow requires an interactive first-auth per
user (fine for Slack users; the proactive Lambdas in §1f are unaffected — they use offline creds).
2.**Our Bedrock guardrail cannot sit on the planner path.** Both AWS-Hosted and BYOLLM keep Salesforce in the
inference path, so keeping inference + our own guardrail in account `328440206208` is **unsatisfiable on the
planner path today**. The Einstein **Trust Layer** covers planner-path moderation; **our guardrail + PII
redaction move entirely to the MCP/action layer** (already the plan for PII — design.md §2.5). This **revises
D11**: the guardrail is applied at the MCP/action layer, not "on the BYOLLM model path."
**Transport architecture (D12) — OpenAPI-first, MCP-ready core.** design.md §4 framed each server as a remote
**MCP** server because the assumed consumer was Slack's built-in MCP client. That consumer is gone, and the chosen
Agentforce per-user path (D4) is OpenAPI/Apex, not MCP. Decision:
- **Build a transport-agnostic core.** The tool handlers and the shared auth/scope/PII-redaction/audit guard
(design.md §4 `shared` package) are written independent of wire protocol. The OpenAPI spec **and** any future MCP
tool list are generated from one tool registry, so the two can never drift.
- **Ship the OpenAPI adapter now** — Agentforce **External Service (OpenAPI) actions** are the primary transport;
use **Apex `@InvocableMethod`** only for actions needing request/response shaping (e.g. extra redaction,
pagination).
- **One facade per trust tier (B3-resolved — audience rejected EARLY at the edge; the server remains the
boundary, FIX-3/CR-1).** Deploy a **separate API Gateway + authorizer per server**: the `sh-mcp-ops` facade
rejects non-ops tokens, the `sh-mcp-finance` facade rejects non-finance tokens (mechanism per BLOCK-1 above —
scope-prefix / `client_id` allow-list / custom `aud`). The edge is an early-reject convenience; it does **not**
move the boundary off the server (next bullet). A **shared AWS WAF web ACL** fronts both (WAF is rate-limit/IP
defense-in-depth ONLY, never an authz boundary — CR-7).
- **Audience validated at the edge AND the server (defense in depth — CR-1, design.md §2.5).** The per-tier
authorizer is NOT the only check: every server **independently re-validates issuer + `aud` + required scope on
every call** (design.md §2.5 makes server-side enforcement authoritative — the edge does not replace it). The
per-tier gateway adds an early-reject layer; it does not move the boundary off the server. **Alarm on any token
presented to the wrong audience** (a finance-aud token at the ops endpoint, or vice-versa) — that signature
means either a misconfig or an attack.
- **Defer the MCP adapter** until a **named trigger**: (a) Claude Code / IDE consumption becomes a real recurring
The first commit of this plan was audited by the Fable adversarial gate (verdict: REQUEST CHANGES). Every BLOCK
and FIX is addressed below; this revision supersedes the pre-audit text wherever they differ.
| Finding | Resolution (where) |
|---------|--------------------|
| **B1** Teardown self-contradiction (Alex deleted Phase 1 *and* the Phase 2 rollback target; shared AOSS deleted before exec-aide retires; KB starved during parallel run) | §3 rewritten: **all** teardown moved to Phase 4 with a shared-consumer audit; legacy bots + KB + AOSS + all sync jobs run untouched through Phase 3; `notion-sync`**dual-feeds** (Bedrock KB + Data Cloud) until Phase 4; rollback never targets a deleted resource. |
| **B2** Plan body still stated pre-resolution decisions (BYOLLM preferred; ~30%; guardrail-on-BYOLLM-path; "decide D4") | §0.1 propagated into **D5, D11, §1d, §2.1–2.3 model lines, §3 guardrail-teardown condition, G6, G9, §6, Phase 0**. |
| **B3** D12 launch architecture ambiguous; one-gateway vs per-server audience | §0.1: **no MCP at launch**; **one facade (API GW + Cognito authorizer) per trust tier**, audience rejection at the edge; §1h list_tools test moved to deferred set. |
| **B4** Token-layer trifecta asserted, not designed | §0.1 per-agent credential block above; §2 Connections name each per-agent app client/credential; new §1h test. |
| **F1** Parity gate runs on Beta tooling; Testing-Center identity under per-user creds | §6 verify-#5 added; **G12** registers the Beta/identity dependency. |
| **F2** design.md "annotation, not reopening" inaccurate | §4 reframed as **amendments to locked decisions** (design.md §1 item 4 now false; D3 reverses §12). |
| **F3**`sh-agentforce` new-repo obligations + `cd-sfdx` Salesforce auth missing | §1g + §4 expanded: full new-repo checklist; **JWT-bearer connected-app cert/key in Secrets Manager**, sandbox/prod-scoped, rotation; `cd-sfdx` gets its own design/verify story. |
| **F4** No phase builds the servers | §3 **Phase 0b "Platform build"** added with deploy-role-first, CI/CD, coverage-gate exit criteria. |
| **F5** New trust boundary: Salesforce now processes tool payloads | **G13** registered (Trust Layer / Data Cloud as new data processor; verify retention + zero-training). |
The doc-answerable 0a unknowns were resolved by web research *before* the live spike. Impacts:
| # | Finding | Confidence | Plan impact |
|---|---------|-----------|-------------|
| 1 | **Cognito `AllowedOAuthScopes` does NOT cap a pre-token V2 Lambda's `scopesToAdd`** (only a no-blank-space rule). | med-high | **Design changed:** the pre-token Lambda is now **suppress-only** (never `scopesToAdd` for a tier scope), which keeps `AllowedOAuthScopes` as the genuine per-tier ceiling. Hard rule + test added (§0.1). |
| 2 | Pre-token **V2 requires the Essentials/Plus feature plan**; new pools default to Essentials; **Lite silently ignores V2**. ~2.7× MAU cost vs Lite (negligible at our scale). | high | Provision the pool on Essentials; confirm not Lite. Cost negligible. (§0.1 hard rule.) |
| 3 | Cognito access tokens carry `client_id` + scopes, **no `aud`** by default (only via managed-login resource binding). API GW authorizer checks `client_id` when `aud` absent. | high | **Mechanism chosen:** per-gateway `client_id` allow-list + resource-server scope-prefix as the audience proxy; optional resource-binding for a real `aud` (test in 0a). (§0.1.) |
| 4 | **Employee-Agent callouts may carry SERVICE identity, not the user's** — "user identity is lost by default" for agent→external client-credentials callouts (primary SF Architects doc). No primary example of per-user Browser Flow + Cognito + `sub` end-to-end. `offline_access` needed or tokens die at 1h. | med (doc) | **NEW CRITICAL gap G16 — now THE top risk** (above the Cognito mechanism). 0a fatal #1 = prove the Slack Employee-Agent callout carries the user's token. Fallbacks: MuleSoft Trusted Agent Identity (RFC 8693), signed user-header, or custom-Bolt. |
| 5 | **Per-user actions are UNTESTABLE via batch Testing Center** — batch/Testing-API run as the client-credentials "Run As" service account (no `UserExternalCredential`). | high | G12 resolved: per-user parity runs as **scripted interactive sessions** (Agent Builder preview / Agent API `bypassUser=false`), not batch. |
| 6 | All-staff agent is **not free** — Grid not required, Identity license waives the *seat* only; **Employee agents consume Flex Credits (~$0.10/action) or the $125/user/mo add-on**. | med | G9 updated with the real cost model; verify the exact unmetered PSL name in-org. |
| 7 | Org-wide planner = 2 options (Default GPT-4o / AWS-Hosted Claude Sonnet 4 → routed to 4.6). **BUT Agent Script `model_config` (Summer '26) can override per-agent to ANY supported model** (50+, incl. Gemini 3.1 Pro, Claude Haiku 4.5); whether a **BYOLLM endpoint** can be a per-agent planner is **unresolved/possibly favorable**. | med | D5 (org-level AWS-Hosted Claude) **stands**, but **reopen as a verify**: can `model_config` place our BYO Bedrock Claude as a per-agent planner? Added to §6. Does NOT block; potential upside. |
**Net:** the auth design is now research-grounded (suppress-only Lambda; client_id/scope-prefix audience; Essentials plan), and the **single biggest risk shifted** from "does Cognito bound scopes" to **"does an Employee-Agent-in-Slack callout even carry the user's identity" (G16)** — which only the live 0a spike can settle, with named fallbacks if it doesn't.
---
## 1. Platform-level design
### 1a. Agent roster — how many, and why
**Decision: three Agentforce agents**, mapped 1:1 to trust tier + Google group, preserving the lethal-trifecta
separation in the UX layer as well as the token layer (design.md §3, §12). More agents would proliferate config;
| **Dispatch / ops helpdesk** (Front workflow, tags/statuses, scheduling) | **NOW** | Topic in Seahaven Ops | RICH, CURRENT — Notion Front subtree: *Inboxes & How Email Flows*`33d2ecdd…8174`, *Dispatcher Workflow*`33d2ecdd…81f1`, *Scheduling Manager Workflow*`33d2ecdd…81d1`, *Tags & Statuses*`33d2ecdd…8169`, *Getting Started with Front*`33d2ecdd…810f`, *Tips & FAQ*`33d2ecdd…8134` |
| **Procurement / intake** (Customer Proposal Request, Invoice Payment Submission) | **NOW (read), Later (write)** | Topic in Seahaven Ops; intake *answers* now, intake *actions* via Flow later | Intake SOPs exist in Notion; the write paths are today Slack workflows |
| **SA8000 / labor-compliance Q&A** | **LATER — blocked pending content** | Topic in Seahaven Ops once sourced | SA8000 docs **not found in Notion** (design.md corpus gap; Gap G8) — author + ingest first |
| **HR / IT onboarding-offboarding** | **LATER — blocked pending content** | Topic in Seahaven Ops, or its own agent if PII-heavy | Notion *Departments & Roles* + HR onboarding pages are near-empty stubs (design.md) |
| **Gusto-backed HR / payroll self-service** | **LATER — new agent + new tier** | New **`sh-mcp-hr`** server + `hr:self`/`hr:read` tier + **Seahaven HR** agent | Employment-of-record is **Nacre Ventures Inc.** (W2), operating brand is Sea Haven — the agent must state this correctly; PII-heavy, warrants its own audited tier |
| **Amazon AMOC / Site-Lead ops** | **LATER — content stale + access-restricted** | Topic, audience-restricted | Amazon subtree is THIN/STALE ("migrated from BookStack"): *Operations (Amazon)*`33a2ecdd…8118`, *AMOC*`33a2ecdd…8162` (comms restricted to Adam & Robert), *Site-Lead*`33a2ecdd…8166` |
`ASSUMPTION`: the highest near-term ROI is the dispatch/ops helpdesk Topic, because the Front corpus is the
single richest, most current body of SOPs we have. SA8000/HR agents are *demand-real but supply-blocked* on
content — calling them out now lets us fund authoring in parallel (Gap G8).
### 1c. MCP functionality to ADD beyond the legacy agents
Each new capability tagged trust tier + scope + outbound-auth, consistent with design.md §2–§3:
| New tool / server | Tier | Scope | Outbound auth | Build window |
| **`sh-mcp-hr`** (Gusto): `get_my_paystub`, `get_pto_balance`, `list_benefits` (self-service) | new `hr` | `hr:self` | service creds (Gusto API token, Secrets Manager); ABAC-partitioned by `sub` like Gmail | Later (corpus + Gusto API) |
| **Front read tools** in `sh-mcp-ops`: `lookup_front_conversation`, `get_sla_status` | ops | `ops:read` | service creds (Front API key) | Optional add — design.md §9 deliberately excluded Front at launch; add only if dispatch Topic needs live conversation state |
| **Procurement intake writes**: `submit_proposal_request`, `submit_invoice_payment` | ops | `ops:tasks` | service creds (DDB/Front) **or** Agentforce **Flow** action | Later; today these are Slack workflows |
| **WO comment free-text search** (replaces a lost KB feature, see D7) | ops | `ops:read` | service creds (DDB / a small text retriever) | Optional — see Gap G4 |
Physical tier (`physical:*`) remains **deferred / admin-out-of-band**, no scope issued to any agent
(design.md §3, §6). Unchanged.
### 1d. AI models — which, and where
`FACT` (research wf_1bf9e142, 25/25 verified; developer.salesforce.com supported-models; "Agentforce 360 for
AWS"; Salesforce×Anthropic Oct-2025): Agentforce's reasoning engine/planner is driven by the **org/agent model
selection**, whose **only two documented options are "Salesforce Default"** (managed mix, currently GPT-4o) **and
"AWS-Hosted"** (Anthropic Claude Sonnet 4 on Bedrock, **inside the Salesforce Trust Boundary**, not our account).
**BYOLLM (Models API) is documented ONLY for custom actions** (prompt templates/Apex/Models-API calls), is **not a
reasoning-engine option**, and even on its action path inference still routes **through Salesforce's Models API /
Trust Layer** — so BYOLLM cannot keep inference in our boundary either.
**Decision (D5) — RESOLVED, §0.1:**
- **Agent reasoning/planner model = AWS-Hosted Claude (Salesforce-managed on Bedrock).** This is the **only**
documented way to keep Claude as the planner. **BYOLLM is ruled out for the planner** (G6, resolved). Consequence:
inference and our own Bedrock guardrail **cannot sit on the planner path** (it runs in Salesforce's boundary) —
our guardrail + PII redaction move **entirely to the MCP/action layer** (D11) and the **Einstein Trust Layer**
covers planner-path moderation. We still retain the Claude lineage of the legacy bots (Sonnet 4.5/4.6).
`ASSUMPTION`: AWS-Hosted's model name is drifting (Sonnet 4 → 4.6 / Haiku 4.5 as of May 2026) — pin the current
name in the Phase-0 spike (§6 verify #4); the structural fact (Claude as planner, SF-managed) holds.
- **Classification stays out of Agentforce.** The 15-min `fetch-classify` job keeps using **Bedrock Haiku 4.5**
in our account (design.md §5, D8) — proactive/event-driven, no conversational surface, no Einstein-Request spend.
- **Per-agent model selection** is set in Setup → Agentforce Agents / `model_config` in Agent Script (`FACT`).
All three agents use the AWS-Hosted Claude reasoning model.
- Licensing/capability flag: Agentforce + Data Cloud carry consumption/licensing cost (Gap G9). The earlier
"~30% fewer Einstein Requests" figure is **dropped — uncorroborated** (§0.1).
content to **Data Cloud**, which **auto-creates a search index** (chunked + vectorized) **and a retriever**
(the link between prompt and index). **Data Libraries support UNSTRUCTURED data only.**
**Design:**
| Corpus | → Data Library | → Retriever | Notes |
|--------|----------------|-------------|-------|
| Notion How-To/Front SOPs + intake SOPs (unstructured) | **Seahaven Ops Knowledge** | `Seahaven_Ops_Knowledge_Retriever` | Highest-value, current. **Ingestion = S3 → Data Cloud (N1, pinned):**`notion-sync` keeps writing `s3://seahaven-kb-docs-328440206208` and Data Cloud ingests from S3 — reuses the existing sync write path, avoids a new Notion-connector dependency, and lets the **same S3 dual-feed** the legacy Bedrock KB during parallel run (B1). |
| Amazon/AMOC subtree (unstructured, **stale**) | same library, separate index segment or tagged | same | Audience-restrict AMOC content; flag staleness (Gap G7) |
| SA8000 docs + employee handbook + company policies | **must be authored**, then ingested | same | **Not in Notion** (Gap G8) — stage in `s3://seahaven-kb-docs-328440206208` or Drive, then ingest |
| WorkOrders / purchase-orders / SiteAssignments / payments (**structured, live**) | **NOT a Data Library** | n/a | Stay **MCP lookup tools** over DDB (design.md §3); Data Libraries can't hold structured data (D7) |
**Relationship to the legacy Bedrock KB + 3 sync jobs (D6/D7):**
- The **Bedrock KB `LSDCNHTH6O`** + **OpenSearch Serverless `gv1540frh1crb79gtr4b`** + **Titan Embed V2** are
**replaced** by the Data Cloud search index + Salesforce-managed embeddings. (We lose control of the embedding
model — Gap G5, low.)
- **`notion-sync`** is **kept and DUAL-FEEDS via S3 until Phase 4** (FIX-1/N1/B1): it keeps writing
`s3://seahaven-kb-docs-328440206208`; the **legacy Bedrock KB** ingests from that S3 (unchanged) AND **Data
Cloud** ingests from the same S3. It does NOT stop feeding the Bedrock KB until teardown (§3) — repointing it
away early would starve the rollback target. Still a scheduled Lambda in `sh-mcp/jobs` (design.md §5).
- **`po-sync` / `workorder-sync` keep feeding the legacy KB until Phase 4** (B1); they are retired as KB feeds at
teardown because the NEW platform serves POs/WOs live via MCP lookup tools (D7). The only thing lost post-cutover
is free-text search over WO *comments* that Alex's KB allowed — Gap G4 (low; workaround: a small dedicated
retriever or `lookup`-by-id only).
**SA8000 / handbook sourcing gap (explicit):** these are referenced as KB inputs but were **not found as Notion
pages** (design.md corpus notes). Resolution: **author them** (Jira stories §4), stage in S3/Drive, ingest into
the Seahaven Ops Knowledge Data Library. **Until authored, SA8000/handbook Q&A is BLOCKED** (Gap G8) — the agent
must say it cannot answer rather than hallucinate, and SA8000-misconduct questions must not be suppressed (legacy
guardrail set MISCONDUCT output to MEDIUM precisely so they aren't — design.md/legacy Alex guardrail).
### 1g. Agentforce DX
`FACT` (developer.salesforce.com Agent DX metadata; "New Agentforce Metadata and Development Lifecycle", May
2026): agents are metadata — **Bot + BotVersion** + a single **GenAiPlannerBundle** per agent (container for
subagents/actions) + **GenAiPlugin** per Topic/subagent + **GenAiFunction** per custom action +
**GenAiPromptTemplate**. Agentforce DX = sf CLI + VS Code extension + Agentforce Vibes IDE; supports scratch
orgs, sandboxes, and VCS as source of truth.
**Decision (D9):** create a **separate `sh-agentforce` SFDX repo** under the GitHub org, NOT a folder in the
CDK monorepo — the toolchains are disjoint (sf CLI / metadata deploy-to-org vs `cdk deploy` to AWS), and the
handbook is one-deploy-target-per-repo. Coexistence:
- *(Proactive digest + 15-min HIGH-priority classification are NOT topics here — they remain scheduled
Lambdas that DM the user; D8, design.md §5.)*
---
## 3. Migration & cutover plan
Follows design.md §6 phasing; each legacy bot is deprecated **only at proven parity** (golden-transcript gate,
§1h).
| Phase | Work | Parity / exit gate | Rollback |
|-------|------|--------------------|----------|
| **0a — Auth spike (GATING, B5)** | **Before any other Phase-0 spend.** Stand up a **minimal throwaway kit** (one Cognito user pool + one app client + one ES action + one Cognito-fronted smoke endpoint; licensing already accepted) and prove the **Per-User OAuth Browser Flow action in Slack** carries the real Google `sub` (§6 verify #1, #3, #8). Confirm the model selector + AWS-Hosted name (#4) and the Testing-Center identity question (#5). **0a also implicitly tests verify #6 — if the Slack plan does not support Employee Agents, 0a cannot start (escalate immediately).** | **FATAL gate (fail → STOP, switch to fallback, do NOT proceed to 0b):** real Google `sub` reaches the server; first-call Slack consent works; token refresh survives; the `aud`/authorizer mechanism (#8) works. **Decision-input (non-fatal):** Testing-Center per-user invocation (#5) — if it can't, parity runs as scripted per-user sessions (G12), not a fallback trigger. | N/A (legacy untouched) |
| **0b — Platform build (F4)** | Only after 0a passes. **Build the servers:** monorepo + `shared` auth/scope/PII/audit guard + per-service packages + **transport-agnostic core + OpenAPI adapter** + **per-tier API Gateway + Cognito authorizer + shared WAF** (B3); Cognito + Google federation + pre-token (per-audience scoping, B4) + 5-min group-sync (design.md §6.1); **deploy-role-first + 1:1 repo↔role isolation (NIT):**`githubdeploy-sh-mcp` exists; provision a **NEW isolated `githubdeploy-sh-agentforce`** — never reuse/expand the sh-mcp role — scoped minimally to read `sh-agentforce/sfdx-jwt-key` (the Salesforce deploy itself uses the JWT key, not an AWS role). Thin `ci.yaml`/`deploy.yaml` callers on **Node 24**: sh-mcp → AWS/CDK reusable workflows; sh-agentforce → the **new `cd-sfdx`** (NOT the AWS templates — they don't fit the `sf` toolchain). **ARM64 markers (FIX, per `reference_cicd_arm64_qemu` — bit slack-bot + exec-aide twice):** any Docker-bundled tier task needs `enable-qemu: true` in BOTH ci.yaml AND deploy.yaml **and**`platform: LINUX_ARM64` on every image asset. Coverage gate. Create `sh-agentforce` SFDX repo + **design & verify `cd-sfdx`** (G10, F3). Stand up Data Cloud + **Seahaven Ops Knowledge** Data Library; **`notion-sync` DUAL-FEEDS** Bedrock KB **and** Data Cloud (B1 — legacy KB stays fresh for rollback). | SSO end-to-end; per-user JWT reaches a smoke-test tool with real `sub`; **core + OpenAPI adapter green at the design.md §7.3 coverage gate** (80% lines, 100% shared auth guard); **`githubdeploy-sh-agentforce` provisioned (isolated); CI/CD green with ARM64 markers — no `exec format error`**; Data Library retriever returns Front SOP answers; legacy Bedrock KB still fed. | N/A (legacy untouched) |
| **1 — Seahaven Ops** | Wire Ops agent → `sh-mcp-ops` facade; vendors (KB/Maps) + WO/PO/site + knowledge + tasks. | Behavioral suite + golden-transcripts vs **Alex** green; per-user auth + per-tool scope verified at the server. | **Flip Slack default back to Alex** (Alex still running — NOT torn down until Phase 4). |
| **2 — Seahaven Finance** | Wire Finance agent → `sh-mcp-finance` facade; full audit logging; 15-min TTL + deny-list. | Audit records emitted; PII-redaction tests green; **per-tier audience-binding rejection verified** (B3). | **Flip back to Alex** (whose QBO action group is still live — Alex runs through Phase 4). |
| **4 — Teardown** | Only after ALL three replacements are signed off at parity, and after the shared-consumer audit below. | Consumer audit passes (no live reader of the KB/AOSS remains). | — |
**Fallback if the Phase-0a auth spike fails (B5).** The plan does not discard design.md §12's alternatives without
a Plan B. If Per-User Browser Flow cannot carry the real `sub` to our server (or Slack consent / refresh /
Testing-Center identity is unworkable), fall back to a **custom Bolt assistant** as the surface (design.md §12) —
it keeps the **MCP transport + per-user OAuth** design intact end-to-end and therefore **un-defers the MCP adapter
(D12)**. This is a surface change, not an auth/scope/trust-tier redesign — the `sh-mcp` core, Cognito, scopes, and
trust tiers are unchanged. Decision point escalates to Adam; it does NOT mean restarting.
**What gets torn down — ALL IN PHASE 4 (B1-corrected; nothing shared is deleted while a consumer or rollback path
still needs it):**
- **Pre-req: shared-consumer audit.** The AOSS collection `gv1540frh1crb79gtr4b` / KB `LSDCNHTH6O` is a shared
dependency of **BOTH**`seahaven-slack-bot`**and**`exec-aide` (INFRA-92). Before any KB/AOSS delete, confirm
**zero** live consumers remain (both legacy stacks retired; no other reader). Run the INFRA-92 consumer check.
- **Bedrock agents** `seahaven-alex` (`QVL5GEJN9B`) and the exec-aide Sonnet loop — torn down in **Phase 4**, after
their parity sign-off AND after they are no longer any phase's rollback target. (Alex is the Phase 1 *and* Phase
2 rollback target, so it must survive to Phase 4 — this was the B1 contradiction.)
- **Bedrock KB `LSDCNHTH6O`** + **AOSS `gv1540frh1crb79gtr4b`** — Phase 4, after the shared-consumer audit.
- **Bedrock guardrail `seahaven-alex-guardrail`** — **not** deleted until the **MCP/action-layer PII redaction +
prompt-attack replacement are live and tested** (D11, revised — there is no BYOLLM model path; design.md §5
requires the guardrail be explicitly replaced before deletion, no parity assumed).
- **Sync Lambdas:** `notion-sync`**dual-feeds through Phase 3**, then drops the Bedrock-KB feed at Phase 4 (keeps
the Data Cloud feed). `po-sync` + `workorder-sync`**keep feeding the legacy KB until Phase 4** (so a rollback to
Alex is never stale — corrects the original "decommission at Phase 1"); they are retired as KB feeds at teardown
(D7 — the NEW platform never used them; structured data is live MCP tools). `fetch-classify`/`daily-digest`/
`reminder` rebuilt in Phase 3; old exec-aide versions decommissioned at Phase 4.
| **G1** | **MITIGATION CHOSEN — UNVERIFIED, SPIKE-GATED (B5-relabel).** Per-user OAuth Browser Flow on the native custom remote-MCP connector is unconfirmed (per-user identity is documented only for SF-**hosted** MCP). The chosen GA path (D4) is *plausible* but **not yet proven** to carry our Cognito `sub` end-to-end in Slack. | **C (open until Phase-0a)** | If the mitigation also fails to propagate the real `sub`: per-user identity, ABAC Gmail isolation, and the trifecta all break — and there is no Agentforce-native surface. | **Decided (D4):** ES (OpenAPI)/Apex actions + Per-User OAuth Browser Flow External Credential → Cognito. **Gated by the Phase-0a spike; if it fails → custom-Bolt fallback (§3).** Status is *mitigation chosen, not verified* — do not call this resolved until 0a passes. | Phase-0a: real `sub` reaches the server, first-call consent in Slack, refresh, Testing-Center identity (§6 verify #1/#3/#5) |
| **G2** | **RESOLVED 2026-06-11.** Custom remote MCP client is **Beta** (Pilot Jul 2025 → Beta Jan 2026); SF-*hosted* MCP is GA (Apr 2026) but that's not our remote servers. | **C→ avoided** | Schedule/stability risk on the native-MCP path. | **Avoided by D4** — the GA ES/Apex action path is the per-user transport; the Beta MCP connector is off the critical path. Revisit native MCP once remote-MCP per-user binding reaches GA. | SF release notes for remote-MCP GA date |
| **G3** | **Trust Layer masking may over-mask.** Agentforce Trust Layer can mask PII; legacy Alex **deliberately left vendor names/contacts unmasked** (lookup is the bot's job). | M | Vendor lookups could be degraded if Trust Layer masks names. | Configure Trust Layer masking to exclude names/contacts; keep MCP-layer redaction authoritative for bank/routing/card/SSN (design.md §2.5). | Trust Layer data-masking config |
| **G4** | **Loss of free-text search over WO comments** (old KB feature) — Data Libraries are unstructured-only, WO/PO are structured MCP lookups. | L | Can't fuzzy-search WO comment text. | Lookup-by-id via MCP tools; or a dedicated retriever over a comment text export if demand appears. | — |
| **G5** | **Embedding model not selectable** — Data Cloud uses Salesforce-managed embeddings (vs legacy Titan Embed V2 1024-dim). | L | Less control over retrieval tuning. | Accept managed embeddings; tune chunking. | Data Cloud index config options |
| **G6** | **RESOLVED 2026-06-11 (research wf_1bf9e142): BYOLLM cannot drive the planner.** Only Salesforce Default / AWS-Hosted options drive the Atlas reasoning engine; BYOLLM is custom-action-only and still routes through SF's Models API/Trust Layer. | M→ resolved | Our-account inference + our guardrail are **unsatisfiable on the planner path**. | **Decided (D5):** use **AWS-Hosted Claude** as the planner; move our guardrail + PII redaction to the **MCP/action layer** (revises D11); rely on the Trust Layer for planner-path moderation. Note AWS-Hosted model is drifting Sonnet 4 → 4.6/Haiku 4.5 (May 2026). | In-org: confirm the reasoning-model selector offers only Default/AWS-Hosted; confirm current AWS-Hosted model name |
| **G7** | **Amazon/AMOC corpus is stale** ("migrated from BookStack, may need updating") and access-restricted (AMOC comms = Adam & Robert only). | M | Agent could give outdated Amazon-site guidance. | Audience-restrict the AMOC topic; flag content as stale; re-author before exposing widely. | Notion *AMOC*`33a2ecdd…8162` currency |
| **G8** | **SA8000 docs + employee handbook + company policies missing** from Notion (referenced as KB inputs, not found). | M | SA8000/compliance + HR Q&A **blocked**. | **Blocked pending content authoring** — author, stage in S3/Drive, ingest into Ops Knowledge Data Library; until then the agent must decline (without suppressing legitimate misconduct questions). | design.md corpus notes; locate any S3/Drive originals |
| **G9** | **Cost — all-staff Ops agent is NOT free (research Finding 6) + parallel-run double-spend.** Agentforce in Slack works on **all paid Slack plans (Grid not required)**, and a **no-cost Salesforce Identity license** covers non-CRM users — BUT that only waives the *CRM seat*; **Employee agents consume Flex Credits (~$0.10/action, 20 credits/action)** OR need the **$125/user/mo Agentforce add-on**. "Free for all staff" is marketing framing about seats, not consumption. Plus dual-run: Phases 1–3 run legacy (2 Bedrock agents + KB + AOSS) AND Agentforce/Data Cloud at once. | M | Real per-action or per-user cost at all-staff scale; the unmetered PSL name is uncertain (docs vs blog differ) — wrong name silently meters everything. | Size a Flex-Credit pool or budget the add-on; **verify the exact unmetered PSL name in-org**; budget a dual-run AWS line; **time-box Phases 1–3 (~6–8 wks)**. | slack.com pricing; Trailhead Agentforce-for-Employees (Flex Credits); Salesforce account team |
| **G10** | **No SFDX reusable CI/CD workflow + no OIDC for Salesforce deploy** — org reusable workflows are AWS/CDK/OIDC-shaped (design.md §7.1). | M | `sh-agentforce` can't deploy via the existing pattern; the JWT-bearer connected-app key is a long-lived prod-deploy secret. | Author reusable **`cd-sfdx`** (deploy-then-merge, scratch-org validate); store the connected-app **JWT key in Secrets Manager** (`sh-agentforce/sfdx-jwt-key`), sandbox/prod-scoped, rotated (F3, §1g). Its own design/verify story. | handbook `cicd.md`; SFDX JWT-bearer flow |
| **G11** | **Two access-control planes.** Agentforce-in-Slack assigns member access via **Salesforce permissions**; our authz is **Google Groups → Cognito → scopes**. | M | Drift: a user could see an agent but lack the MCP scope, or vice-versa. | Provision Agentforce/Salesforce user access **from the same Google Groups** (SCIM/identity sync) so Google Groups stays the single source of truth (design.md §2.3). | Salesforce SCIM/Google provisioning |
| **G12** | **RESOLVED by research (Finding 5): per-user actions are UNTESTABLE via batch Testing Center.** Testing Center batch/AI-generated runs and the Testing API execute under the External Client App's **client-credentials "Run As" SERVICE account** — which has no `UserExternalCredential`, so any Per-User Browser Flow action fails at callout. | M | The behavioral/parity suite (which authorizes teardown) cannot exercise per-user (Lauren Exec / Gmail) tools via batch. | **Decided:** drive per-user parity via **scripted interactive sessions** (Agent Builder preview, or Agent API with `bypassUser=false` + a pre-authorized real user token) — NOT batch. Non-per-user (Ops/Finance read) tools can still use batch. | developer.salesforce.com testing-api-connect; agent-api-get-started (`bypassUser`) |
| **G16** | **CRITICAL — Employee-Agent callout may carry SERVICE identity, not the user's (research Finding 4, the new top risk).** Salesforce Architects docs state verbatim: "User identity is lost by default. When an agent calls an MCP server or downstream API using client credentials, the request carries the agent's service identity and not the end user's." Per-user External Credentials forward the user's token ONLY in a **user-session** context; an autonomously-invoked agent uses the service identity. Whether an **Employee Agent in Slack** runs each tool callout in the end-user's session (so the per-user Cognito token + real `sub` flow) or autonomously (service identity) is **undocumented and unproven** — there is NO primary Salesforce example of per-user Browser Flow + Cognito + `sub` propagation end-to-end. | **C** | If callouts use service identity, **per-user `sub`, ABAC Gmail isolation, and the lethal-trifecta all collapse** — every user looks like one service account to our MCP servers. This is now THE controlling risk, above the Cognito mechanism. | **0a fatal #1:** prove an Employee-Agent-in-Slack tool callout carries the signed-in user's Cognito token (real Google `sub`) to our server. If it does NOT: fallbacks are (a) MuleSoft Flex Gateway "Trusted Agent Identity" (RFC 8693 token exchange), (b) a signed user-identity header our server trusts, or (c) the custom-Bolt surface (design.md §12, keeps per-user OAuth intact). Also: grant **`offline_access`** on the Cognito client or per-user tokens die at 1h with no documented auto-refresh (Finding 4). | architect.salesforce.com end-user-identity-propagation; nc-use-oauth-cred-in-callout; in-org spike |
| **G13** | **New data processor: Salesforce now processes our tool payloads (F5).** Finance/Gmail tool responses transit the Agentforce planner / Einstein Trust Layer; grounding flows through Data Cloud. design.md's data boundary was AWS + Slack. | M | Vendor/payment/inbox content (PII is masked, but business content is not) is processed by Salesforce — a new processor not in the original threat model. | Explicit acceptance; verify Trust Layer **retention + zero-training** guarantees and Data Cloud data-residency; keep MCP-layer PII redaction authoritative. | Einstein Trust Layer retention/zero-training docs; DPA |
| **G14** | **Per-user tool visibility degraded (F7).** MCP `list_tools` hid tools per-*user* by scope; per-agent action assignment is per-*agent*, so a user lacking a scope still sees the action and fails server-side (403). | L | UX degradation (confusing failures), not a security hole — the per-tool scope check is the boundary. | Graceful-refusal copy in agent instructions; behavioral test for a clean "no access" message (§1h). | — |
| **G15** | **Three unverified-load-bearing platform assumptions (Q1/Q2/Q3).** (a) Sea Haven's Slack plan actually supports **Agentforce Employee Agents**; (b) per-employee Agentforce **user+license** required for the all-staff Ops agent (drives G9 cost + G11 SCIM); (c) Agentforce exposes **`$User.GoogleGroups`** and **`$Session.Channel`** to agent instructions (Lauren Exec's DM-only + scope gating depend on these). | M | If (a) wrong, the surface decision (D2) collapses; if (c) wrong, privacy/DM controls need a different enforcement point (e.g. a context Apex action). | Verify all three in Phase-0 (§6); for (c), design a fallback enforcement now so it's not on the critical path. | slack.com/help Agentforce-in-Slack plan reqs; Salesforce licensing; Agentforce agent variables doc |
`FACT` for G1/G2: Agentforce MCP support (Pilot Jul 2025, Beta Jan 2026; SF-hosted MCP GA Apr 2026; OAuth 2.0,
JSON-RPC over Streamable HTTP) and Named/External Credential **Per User** identity type — verified via Salesforce
sources (§Sources). The **specific** per-user binding for MCP connectors is the unverified piece, hence the gap.
---
## 6. Open decisions for Adam (with recommendation)
> **STATUS — all resolved 2026-06-11. See §0.1 for the committed resolutions and basis.** Summary: (1) surface +
> licensing **accepted**; (2) reasoning model = **AWS-Hosted Claude** (BYOLLM ruled out for the planner);