From 8c768f04958919d238ed29895b2ab99b10aebbf1 Mon Sep 17 00:00:00 2001 From: Adam Moussa Date: Thu, 11 Jun 2026 12:44:00 -0400 Subject: [PATCH] Remediate sh-plan-review findings (B1-B5, F1-F7, NITs, Qs) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Address the Fable plan-review gate (REQUEST CHANGES): - B1: move ALL teardown to Phase 4 + shared-consumer audit; notion-sync dual-feeds; rollback never targets a deleted/starved resource. - B2: propagate the §0.1 resolutions through D5/D11/§1d/§2.x/§3/G6/G9/§6 (model, guardrail, dropped ~30% claim, Phase-0 'verify' not 'decide'). - B3: pin D12 to one facade per trust tier (per-server Cognito audience at the edge); reconcile the §1h list_tools test (deferred with MCP adapter). - B4: specify per-agent External Credential + Cognito app client minting only its tier's scopes; add the token-layer trifecta test. - B5: split Phase 0 (0a auth spike gates 0b); add custom-Bolt fallback; relabel G1 as mitigation-chosen/unverified. - F1-F7: Testing-Center identity (G12); design.md amendments reframed (F2); sh-agentforce new-repo checklist + cd-sfdx JWT-key auth (F3); platform-build phase (F4); Salesforce data-processor gap (G13/F5); memory+README obligations (F6); per-user visibility degradation (G14/F7). - NITs: S3->Data Cloud ingestion pinned; facade+notion-sync ALARM monitoring; DevName naming boundary. Qs Q1-Q3 registered as G15 + verify items. - §0.3 remediation log + gap count 11->15. --- docs/agentforce-plan.md | 368 ++++++++++++++++++++++++++++------------ 1 file changed, 259 insertions(+), 109 deletions(-) diff --git a/docs/agentforce-plan.md b/docs/agentforce-plan.md index 4dde647..59a7ff9 100644 --- a/docs/agentforce-plan.md +++ b/docs/agentforce-plan.md @@ -46,18 +46,26 @@ off the per-user hop, so the risk is mitigated rather than load-bearing. | D8 | Proactive jobs (`fetch-classify`, `daily-digest`, `reminder`, `notion-sync`) **stay as our scheduled Lambdas**, not Agentforce; HIGH-priority alerts continue as **Slack DMs** | No conversational equivalent; keeps Haiku classifier in our Bedrock account; cheapest reliable path | design.md §5; §1f below | | D9 | New separate **`sh-agentforce` SFDX repo** for agent metadata (Bot/GenAiPlannerBundle/GenAiPlugin/GenAiFunction) under Agentforce DX | Different toolchain (sf CLI, scratch orgs, deploy-to-org) than the CDK monorepo; one deploy target per repo | `FACT` developer.salesforce.com Agent DX metadata; handbook | | D10 | Build **Ops + Finance agents now**; HR/Gusto, SA8000 Q&A, IT/onboarding **later, corpus-gated**; add near-term capabilities as **Topics**, not new agents | Front/ops corpus is rich and current; HR/Amazon/SA8000 corpus is thin/stale/missing | Notion (below); design.md corpus notes | -| D11 | Keep the **Bedrock guardrail on the BYOLLM model path** + MCP-layer PII redaction (defense in depth) | Agentforce Trust Layer does not redact OUR tool-response payloads | design.md §2.5, §11.2 | +| D11 | **REVISED (2026-06-11, §0.1):** guardrail (prompt-attack + PII redaction) lives **entirely at the MCP/action layer**, not on any model path; the Einstein Trust Layer covers planner-path moderation | The planner runs AWS-Hosted Claude inside Salesforce's boundary (D5) — there is no BYOLLM model path to attach our guardrail to; Trust Layer does not redact OUR tool-response payloads, so MCP-layer redaction is authoritative | design.md §2.5, §11.2; §0.1 | | D12 | **RESOLVED (2026-06-11, §0.1):** **OpenAPI-first, MCP-ready transport-agnostic core.** Tool handlers + the shared auth/scope/PII/audit guard are built independent of transport; ship the **OpenAPI adapter now** (Agentforce External Service actions; Apex only where an action needs request/response shaping); **defer the MCP adapter** to a named trigger. | Agentforce's per-user hop needs OpenAPI regardless (D4); no committed MCP consumer today, so a second public interface + its conformance/hardening cost isn't justified yet; a transport-agnostic core makes the MCP adapter a cheap later add, not a re-platform; deferring it also drops the IDE-session trifecta nuance until then. | Adam 2026-06-11; research wf_1bf9e142 (D4) | **Agent roster (the answer to "how many"):** **3** — Seahaven Ops, Seahaven Finance, Lauren Exec (§1a, §2). -**Top 5 decisions for Adam (§6):** (1) confirm Agentforce-in-Slack as the surface + accept Agentforce/Data -Cloud licensing; (2) BYOLLM-on-our-Bedrock vs AWS-Hosted Claude Sonnet 4; (3) accept D3 (strip finance from -Lauren's exec agent); (4) accept the Apex/External-Service per-user wrapper (D4) rather than waiting for native -MCP per-user GA; (5) fund authoring the missing SA8000 + employee-handbook corpus. +**Top 5 decisions for Adam — ALL RESOLVED 2026-06-11 (§0.1):** (1) Agentforce-in-Slack + licensing **accepted**; +(2) reasoning model = **AWS-Hosted Claude** (BYOLLM ruled out for the planner); (3) **D3 accepted** (strip finance +from Lauren Exec); (4) **D4** per-user auth = ES/Apex actions + Per-User Browser Flow → Cognito; (5) corpus +authoring **deferred** (fund when needed). -**Capability-gap count: 11** (§5). Severity: **2 critical** (G1 per-user MCP auth, G2 remote-MCP Beta), **6 -medium**, **3 low**. +**Capability-gap count: 15** (§5). Severity: **1 critical** (G1 per-user auth — mitigation chosen, spike-gated, +**unverified**), **9 medium**, **2 low**, **3 resolved/avoided** (G2, G6, and the BYOLLM cost claim). + +> **REVISION — plan-review remediation (2026-06-11, §0.3).** This plan was audited by the `sh-plan-review` Fable +> gate; §0.3 logs every BLOCK/FIX addressed. Key structural changes since the first commit: teardown moved wholly +> to Phase 4 with a shared-consumer audit (was contradictory); §0.1 propagated into D5/D11/§1d/§2.x/§3/G6/G9/§6 +> (model/guardrail consistency); D12 transport pinned to **one facade per trust tier** (per-server audience +> preserved); per-agent External Credential / Cognito app-client scoping specified (token-layer trifecta); +> Phase 0 split so the auth spike gates the rest, with a **custom-Bolt fallback** if it fails; a platform-build +> phase added; new gaps G12–G15 registered. --- @@ -78,12 +86,13 @@ provisional D4/D5 wording above and the open items in §6. **Two consequences that ripple into the design (must be honored downstream):** 1. **The Agentforce→tool hop is NOT MCP.** Because per-user identity is only achievable on the GA External - Service/Apex action path (not the Beta MCP connector), Agentforce reaches our tools through an **OpenAPI/Apex - facade in front of each MCP server**, authenticated per-user to Cognito. The `sh-mcp-ops`/`-finance` servers, - their JWT validation, scopes, audience-binding, and PII redaction are **unchanged** — only the Agentforce-side - transport changes from MCP to OpenAPI/Apex actions. (Other MCP clients can still speak MCP to the same servers.) - Browser Flow requires an interactive first-auth per user (fine for Slack users; the proactive Lambdas in §1f - are unaffected — they use offline creds). + Service/Apex action path (not the Beta MCP connector), Agentforce reaches our tools over **OpenAPI/Apex + actions** authenticated per-user to Cognito. **Day-one architecture (B3-resolved):** there is NO MCP wire + protocol at launch — the trust-tier servers (`sh-mcp-ops`, `sh-mcp-finance`) expose their tools as **OpenAPI + endpoints** via the transport-agnostic core (D12); the MCP adapter is deferred (D12 trigger). The servers' + JWT validation, per-tool scope enforcement, audience-binding, PII redaction, and audit are **unchanged** — + they sit below the adapter and are transport-agnostic. Browser Flow requires an interactive first-auth per + user (fine for Slack users; the proactive Lambdas in §1f are unaffected — they use offline creds). 2. **Our Bedrock guardrail cannot sit on the planner path.** Both AWS-Hosted and BYOLLM keep Salesforce in the inference path, so keeping inference + our own guardrail in account `328440206208` is **unsatisfiable on the planner path today**. The Einstein **Trust Layer** covers planner-path moderation; **our guardrail + PII @@ -99,8 +108,14 @@ Agentforce per-user path (D4) is OpenAPI/Apex, not MCP. Decision: tool list are generated from one tool registry, so the two can never drift. - **Ship the OpenAPI adapter now** — Agentforce **External Service (OpenAPI) actions** are the primary transport; use **Apex `@InvocableMethod`** only for actions needing request/response shaping (e.g. extra redaction, - pagination). This is the one interface to build, secure (one API Gateway + Cognito authorizer + WAF web ACL), - and test on the critical path. + pagination). +- **One facade per trust tier (B3-resolved — the audience boundary stays in infrastructure, not in shared + handlers).** Deploy a **separate API Gateway + Cognito authorizer per server**: `sh-mcp-ops` facade validates + `aud=sh-mcp-ops`, `sh-mcp-finance` facade validates `aud=sh-mcp-finance`. A token minted for ops is rejected at + the finance gateway's authorizer (and vice-versa) — the audience-binding rejection design.md §2.5 requires, + enforced at the edge. A **shared AWS WAF web ACL** fronts both. This is deliberately NOT one shared gateway with + audience checks pushed into application code: keeping per-tier audiences at separate authorizers preserves the + trust-tier firewall even if a handler bug slips through. - **Defer the MCP adapter** until a **named trigger**: (a) Claude Code / IDE consumption becomes a real recurring workflow (not nice-to-have), **or** (b) Salesforce native remote-MCP per-user binding reaches GA (then MCP-native could also collapse the Agentforce-side OpenAPI facade). When triggered, the MCP adapter is a thin add over the @@ -113,11 +128,55 @@ Agentforce per-user path (D4) is OpenAPI/Apex, not MCP. Decision: "don't co-connect finance with Gmail-read in one IDE session." The `sh-mcp` name stays accurate — MCP remains the strategic protocol, just not the day-one transport. +**Per-agent credential & token-layer trifecta (B4-resolved — the guarantee is designed, not asserted).** The +claim "separate audiences → no single token spans tiers" is only true if each agent's tokens are minted from a +credential that requests **only its tier's scopes**. Mechanism (specified, testable): + +- **One Cognito app client + one External Credential per agent**, each registered with **only its tier's + resource-server scopes**: Seahaven Ops client → `ops:read [ops:tasks]` (`aud=sh-mcp-ops`); Seahaven Finance + client → `finance:read` (`aud=sh-mcp-finance`); Lauren Exec client → `ops:read ops:tasks gmail:self + calendar:self` (`aud=sh-mcp-ops`, **never finance**). +- design.md §2.3 mints the **union** of a user's group-scopes into a token *only for the audience requested*. + Because Lauren Exec's app client requests only the ops audience and ops scopes, **Lauren-the-person's + `finance:read` (from `-finance@`) is never present in any token the Exec agent obtains** — even though she + holds it. She reaches finance only through the Finance agent's separate client/audience. The boundary is the + per-agent credential, not per-agent action assignment (which the plan concedes is "a convenience, never the + boundary," §1h). +- The Cognito pre-token Lambda must therefore scope the minted scopes to the requested `aud` (resource server), + not blanket-union across all audiences. **This is a per-server IAM/authz behavior → mandatory cross-review + (GPT-4.1) before commit** (design.md §8/§10). +- Tested by: "per-agent credential mints only its tier's scopes" (a token obtained via the Exec agent's + credential never contains `finance:read`) — added to the §1h security suite. + **Dropped claim:** the "BYOLLM = ~30% fewer Einstein Requests" figure was **not corroborated** by any verified source — removed from cost modeling. --- +## 0.3 Plan-review remediation log — 2026-06-11 (Fable `sh-plan-review` gate) + +The first commit of this plan was audited by the Fable adversarial gate (verdict: REQUEST CHANGES). Every BLOCK +and FIX is addressed below; this revision supersedes the pre-audit text wherever they differ. + +| Finding | Resolution (where) | +|---------|--------------------| +| **B1** Teardown self-contradiction (Alex deleted Phase 1 *and* the Phase 2 rollback target; shared AOSS deleted before exec-aide retires; KB starved during parallel run) | §3 rewritten: **all** teardown moved to Phase 4 with a shared-consumer audit; legacy bots + KB + AOSS + all sync jobs run untouched through Phase 3; `notion-sync` **dual-feeds** (Bedrock KB + Data Cloud) until Phase 4; rollback never targets a deleted resource. | +| **B2** Plan body still stated pre-resolution decisions (BYOLLM preferred; ~30%; guardrail-on-BYOLLM-path; "decide D4") | §0.1 propagated into **D5, D11, §1d, §2.1–2.3 model lines, §3 guardrail-teardown condition, G6, G9, §6, Phase 0**. | +| **B3** D12 launch architecture ambiguous; one-gateway vs per-server audience | §0.1: **no MCP at launch**; **one facade (API GW + Cognito authorizer) per trust tier**, audience rejection at the edge; §1h list_tools test moved to deferred set. | +| **B4** Token-layer trifecta asserted, not designed | §0.1 per-agent credential block above; §2 Connections name each per-agent app client/credential; new §1h test. | +| **B5** Phase-0 spike "gates everything" but no fallback; spend precedes it; G1 overstated | §3 **Phase 0 split (0a auth spike → 0b platform/Data-Cloud/SFDX)**; explicit **custom-Bolt fallback** if the spike fails (design.md §12); G1 relabelled "mitigation chosen — UNVERIFIED, spike-gated." | +| **F1** Parity gate runs on Beta tooling; Testing-Center identity under per-user creds | §6 verify-#5 added; **G12** registers the Beta/identity dependency. | +| **F2** design.md "annotation, not reopening" inaccurate | §4 reframed as **amendments to locked decisions** (design.md §1 item 4 now false; D3 reverses §12). | +| **F3** `sh-agentforce` new-repo obligations + `cd-sfdx` Salesforce auth missing | §1g + §4 expanded: full new-repo checklist; **JWT-bearer connected-app cert/key in Secrets Manager**, sandbox/prod-scoped, rotation; `cd-sfdx` gets its own design/verify story. | +| **F4** No phase builds the servers | §3 **Phase 0b "Platform build"** added with deploy-role-first, CI/CD, coverage-gate exit criteria. | +| **F5** New trust boundary: Salesforce now processes tool payloads | **G13** registered (Trust Layer / Data Cloud as new data processor; verify retention + zero-training). | +| **F6** Memory/README obligations omitted | §4 **Repo docs** expanded: `sh-agentforce` project memory, `project_sh_mcp.md` update, cross-project refs, both READMEs. | +| **F7** Per-user tool visibility degraded with no replacement | §1h states it; behavioral test + graceful-refusal UX added; **G14**. | +| **N1/N2/N3** | §1f ingestion pinned to **S3 → Data Cloud**; ALARM-only monitoring for the facades + repointed `notion-sync`; Salesforce DevName/kebab boundary noted. | +| **Q1/Q2/Q3** | **G15** (Slack plan supports Employee Agents; per-employee Agentforce licensing; `$User.GoogleGroups`/`$Session.Channel` existence) — all unverified-load-bearing, on the §6 verify gate. | + +--- + ## 1. Platform-level design ### 1a. Agent roster — how many, and why @@ -182,26 +241,28 @@ Physical tier (`physical:*`) remains **deferred / admin-out-of-band**, no scope ### 1d. AI models — which, and where -`FACT` (developer.salesforce.com supported-models; Salesforce "Agentforce 360 for AWS"; Salesforce×Anthropic -Oct-2025 partnership): Agentforce's **Atlas reasoning engine is model-agnostic**; the default is a -Salesforce-managed mix (incl. GPT-4o); an **AWS-Hosted option runs Anthropic Claude Sonnet 4 on Amazon Bedrock** -and can power Atlas; **BYOLLM** (Models API) supports **Amazon Bedrock, Azure OpenAI, OpenAI, Vertex**, runs on -your own credentials/instance, keeps the **Trust Layer**, and consumes **~30% fewer Einstein Requests**. +`FACT` (research wf_1bf9e142, 25/25 verified; developer.salesforce.com supported-models; "Agentforce 360 for +AWS"; Salesforce×Anthropic Oct-2025): Agentforce's reasoning engine/planner is driven by the **org/agent model +selection**, whose **only two documented options are "Salesforce Default"** (managed mix, currently GPT-4o) **and +"AWS-Hosted"** (Anthropic Claude Sonnet 4 on Bedrock, **inside the Salesforce Trust Boundary**, not our account). +**BYOLLM (Models API) is documented ONLY for custom actions** (prompt templates/Apex/Models-API calls), is **not a +reasoning-engine option**, and even on its action path inference still routes **through Salesforce's Models API / +Trust Layer** — so BYOLLM cannot keep inference in our boundary either. -**Decision (D5):** -- **Agent reasoning / planner model = Anthropic Claude on Bedrock.** Preferred path: **BYOLLM pointed at our - Bedrock** (account `328440206208`, `us-east-1`) so inference stays in our trust boundary, we keep our own - guardrail on the model path (D11), and we cut Einstein-Request spend. **Gap G6:** confirm a BYOLLM endpoint can - be the *agent reasoning model* (Atlas planner), not only a prompt-template/Models-API call. If not, fall back - to the **AWS-Hosted Claude Sonnet 4** managed option (confirmed to power Atlas) — same model family, less - control. Either way we retain the Claude lineage of the legacy bots (Sonnet 4.5 for Alex, Sonnet 4.6 for - Lauren's conversation loop). +**Decision (D5) — RESOLVED, §0.1:** +- **Agent reasoning/planner model = AWS-Hosted Claude (Salesforce-managed on Bedrock).** This is the **only** + documented way to keep Claude as the planner. **BYOLLM is ruled out for the planner** (G6, resolved). Consequence: + inference and our own Bedrock guardrail **cannot sit on the planner path** (it runs in Salesforce's boundary) — + our guardrail + PII redaction move **entirely to the MCP/action layer** (D11) and the **Einstein Trust Layer** + covers planner-path moderation. We still retain the Claude lineage of the legacy bots (Sonnet 4.5/4.6). + `ASSUMPTION`: AWS-Hosted's model name is drifting (Sonnet 4 → 4.6 / Haiku 4.5 as of May 2026) — pin the current + name in the Phase-0 spike (§6 verify #4); the structural fact (Claude as planner, SF-managed) holds. - **Classification stays out of Agentforce.** The 15-min `fetch-classify` job keeps using **Bedrock Haiku 4.5** - in our account (design.md §5, D8). It is proactive/event-driven, has no conversational surface, and shouldn't - consume Einstein Requests. -- **Per-agent model selection** is set in Setup → Agentforce Agents (`FACT`). All three agents use the same - Claude reasoning model; Finance's lower latency tolerance is fine. -- Licensing/capability flag: BYOLLM and Data Cloud both carry consumption/licensing cost (Gap G9, O5). + in our account (design.md §5, D8) — proactive/event-driven, no conversational surface, no Einstein-Request spend. +- **Per-agent model selection** is set in Setup → Agentforce Agents / `model_config` in Agent Script (`FACT`). + All three agents use the AWS-Hosted Claude reasoning model. +- Licensing/capability flag: Agentforce + Data Cloud carry consumption/licensing cost (Gap G9). The earlier + "~30% fewer Einstein Requests" figure is **dropped — uncorroborated** (§0.1). ### 1e. Prompt Builder / Template Library structure @@ -232,7 +293,7 @@ content to **Data Cloud**, which **auto-creates a search index** (chunked + vect | Corpus | → Data Library | → Retriever | Notes | |--------|----------------|-------------|-------| -| Notion How-To/Front SOPs + intake SOPs (unstructured) | **Seahaven Ops Knowledge** | `Seahaven_Ops_Knowledge_Retriever` | Highest-value, current. Source via `notion-sync` repointed to Data Cloud ingestion (S3 → Data Cloud, or Notion connector) | +| Notion How-To/Front SOPs + intake SOPs (unstructured) | **Seahaven Ops Knowledge** | `Seahaven_Ops_Knowledge_Retriever` | Highest-value, current. **Ingestion = S3 → Data Cloud (N1, pinned):** `notion-sync` keeps writing `s3://seahaven-kb-docs-328440206208` and Data Cloud ingests from S3 — reuses the existing sync write path, avoids a new Notion-connector dependency, and lets the **same S3 dual-feed** the legacy Bedrock KB during parallel run (B1). | | Amazon/AMOC subtree (unstructured, **stale**) | same library, separate index segment or tagged | same | Audience-restrict AMOC content; flag staleness (Gap G7) | | SA8000 docs + employee handbook + company policies | **must be authored**, then ingested | same | **Not in Notion** (Gap G8) — stage in `s3://seahaven-kb-docs-328440206208` or Drive, then ingest | | WorkOrders / purchase-orders / SiteAssignments / payments (**structured, live**) | **NOT a Data Library** | n/a | Stay **MCP lookup tools** over DDB (design.md §3); Data Libraries can't hold structured data (D7) | @@ -267,12 +328,25 @@ handbook is one-deploy-target-per-repo. Coexistence: - `sh-mcp` (existing): MCP servers, Cognito/auth, jobs — TypeScript/CDK, OIDC-into-AWS, `ci / ci` required check (design.md §7). Unchanged. - `sh-agentforce` (new): agent metadata. CI runs `sf` validate-deploy against a scratch org; CD does - **deploy-then-merge** to sandbox → prod org. **Gap G10:** the org's reusable workflows are AWS/CDK-shaped; - we need a **new reusable `cd-sfdx` workflow** (Jira story). Naming kebab-case; Dependabot N/A (no npm), but pin - `@salesforce/cli` version. + **deploy-then-merge** to sandbox → prod org. +- **`sh-agentforce` new-repo checklist (F3 — was missing; per `feedback_new_repo_checklist`):** private repo in + org; **branch protection + a required check** (the SFDX analog of `ci / ci`); **enable vulnerability-alerts + + `dependabot_security_updates` + secret_scanning** — Dependabot is **NOT N/A** (the `github-actions` ecosystem + applies even without npm; also watch `@salesforce/cli`); pin `@salesforce/cli`; kebab-case repo/workflow names. +- **`cd-sfdx` deploy auth (F3 — the prod-deploy key for the entire agent surface; design/verify story of its + own, NOT one Jira bullet).** Salesforce has **no OIDC**; CD authenticates via the **JWT-bearer flow of a + connected app** using a private key + cert. That key is a **long-lived secret** — store it in **AWS Secrets + Manager** (`sh-agentforce/sfdx-jwt-key`, per secrets-and-config.md; NOT a GitHub secret blob, NOT an env var), + pull it in the workflow, **scope separate connected apps for sandbox vs prod**, and define a **rotation cadence**. + `cd-sfdx` is a brand-new org-wide reusable workflow (Gap G10) implementing deploy-then-merge against + sandbox→prod orgs — it gets its **own design + verification step** before any agent metadata depends on it. - **Cross-review gate extends to Agentforce metadata** that changes tool exposure, audience, or scope binding — - those are security-relevant just like an IAM diff (design.md §8). `ASSUMPTION`: GenAiPlannerBundle/connection - changes go through `cross_reviewer` the same as IAM. + security-relevant just like an IAM diff (design.md §8). `ASSUMPTION`: GenAiPlannerBundle/connection changes and + the `cd-sfdx` connected-app trust go through `cross_reviewer` (GPT-4.1) the same as IAM. +- **Naming boundary (N3):** Salesforce Developer/API names (`Seahaven_Ops`, Flex template names) **cannot contain + hyphens** and conventionally use PascalCase/snake. The handbook kebab-case rule governs **AWS resources + repo + names** (the repo `sh-agentforce` IS kebab); it does NOT apply to Salesforce-side metadata API names. Recorded so + a future compliance pass doesn't "fix" a non-violation. ### 1h. Test suite — Agentforce Testing Center + the MCP-layer security tests @@ -300,21 +374,35 @@ real home: 8. **Parity golden-transcripts** — replay real Alex/Lauren interactions; assert equivalent answers **before** deprecating each bot (design.md §6, §7.3 parity gate). -**MCP-layer security tests (authoritative, in `sh-mcp`, design.md §7.3 — unchanged):** -audience-binding rejection (an `ops` token rejected by `finance`); per-tool **server-side** scope enforcement; -`list_tools` tool-hiding reflects caller scopes; deny-list **hard revocation**; minimal-scope Google client -(a `gmail:self` token can't mint a Calendar token); **per-user refresh-token ABAC isolation**; -**finance PII redaction** (bank/routing/card/SSN masked before egress) **while leaving vendor names/contacts -UNMASKED** (legacy Alex deliberately left names unmasked — Trust Layer must not re-mask them, Gap G3); per-tool -rate limit + per-session cap; finance audit-record shape. +**Core security tests (authoritative, transport-agnostic, in `sh-mcp`, design.md §7.3):** +**audience-binding rejection at the per-tier facade** (an `aud=sh-mcp-ops` token rejected by the `sh-mcp-finance` +gateway authorizer, and vice-versa — B3); per-tool **server-side** scope enforcement; **per-agent credential +mints only its tier's scopes** (a token obtained via the Lauren Exec app client never contains `finance:read`, +even though the person holds it — B4); deny-list **hard revocation**; minimal-scope Google client (a `gmail:self` +token can't mint a Calendar token); **per-user refresh-token ABAC isolation**; **finance PII redaction** +(bank/routing/card/SSN masked before egress) **while leaving vendor names/contacts UNMASKED** (legacy Alex +deliberately left names unmasked — Trust Layer must not re-mask them, Gap G3); per-tool rate limit + per-session +cap; finance audit-record shape. -**Transport-test note (D12):** the active interface is the **OpenAPI adapter**, so the design.md §7.3 -"Contract / MCP conformance" layer is split — **OpenAPI contract tests** (schema validation, the generated spec -matches the tool registry) run now; **MCP protocol-conformance tests defer with the MCP adapter**. The security -tests above are transport-agnostic and unchanged: server-side per-tool scope enforcement is the authoritative -boundary on either transport, and `list_tools` tool-hiding (MCP-only) is replaced on the OpenAPI path by -**per-agent action assignment** (each Agentforce agent is granted only its tier's actions) — still a convenience, -never the boundary (design.md §2.5). +**Transport-test note (D12, B3-reconciled):** the active interface is the **OpenAPI adapter**, so the design.md +§7.3 "Contract / MCP conformance" layer is split — **OpenAPI contract tests** (schema validation, the generated +spec matches the tool registry) run now; **MCP `list_tools` tool-hiding and MCP protocol-conformance tests defer +WITH the MCP adapter** (not in the launch suite — they were MCP-only). On the OpenAPI path the per-*user* tool +visibility that `list_tools` gave is replaced by per-*agent* action assignment (each agent granted only its +tier's actions) — a *convenience, never the boundary*; the per-tool server-side scope check is the boundary +(design.md §2.5). + +**Per-user visibility degradation (F7 — stated, not silently dropped):** per-agent action assignment is +per-*agent*, not per-*user*. A user **without** `ops:tasks` talking to Seahaven Ops still *sees* the task actions +offered and they fail server-side (403). Boundary holds; UX degrades. Handle with **graceful-refusal copy** in +the agent instructions ("you may not have access to that — server will confirm") and a behavioral test: +a no-`ops:tasks` user gets a clean "you don't have access" rather than a raw error. + +**Testing-Center identity caveat (F1):** Testing Center batch/AI-generated runs invoke actions — **under whose +identity** when the action uses a Per-User Browser Flow credential? Headless batch may fail (no interactive +first-auth) or run as a different principal, which would make the behavioral suite unable to exercise the real +per-user tools (the suite authorizes teardown — design.md §6). Resolve in the Phase-0 spike (§6 verify #5); +registered as Gap **G12** (Beta dependency). Eval criteria: behavioral suite ≥ agreed pass rate before each cutover; MCP/core suite at design.md coverage gate (80% lines, 100% on the shared auth/scope guard) — both green are the parity gate for retiring a bot. @@ -343,13 +431,17 @@ Eval criteria: behavioral suite ≥ agreed pass rate before each cutover; MCP/co - **Error Message (≤255):** "Sorry — I hit a problem reaching that information. Please try again in a moment; if it keeps failing, post in #it-help and we'll take a look." - **Languages:** English (US). `ASSUMPTION`: no multilingual requirement today. -- **Variables:** `$User.Email`, `$User.GoogleGroups` (for scope context), `$Session.Channel` (public vs DM, for - privacy gating). -- **Connections:** `sh-mcp-ops` MCP server — trust tier **ops**, `aud=sh-mcp-ops`, scopes **`ops:read`** - (+ `ops:tasks` only when the caller's token carries it). Per-user OAuth 2.1 → Cognito (D4). +- **Variables:** `$User.Email`; `$User.GoogleGroups` (scope context); `$Session.Channel` (public vs DM, privacy + gating). `ASSUMPTION` (Q3, unverified-load-bearing): that Agentforce exposes Google-group membership and a + channel-visibility variable to agent instructions — **verify in Phase-0**; if absent, gate privacy/scope via a + different mechanism (e.g. a context Apex action). Registered as Gap **G15**. +- **Connections:** `sh-mcp-ops` **tier (OpenAPI facade, `aud=sh-mcp-ops`)** via **External Service actions** over + a **per-agent External Credential `ec-seahaven-ops` (Per-User OAuth Browser Flow → Cognito app client + `sh-agentforce-ops`, scopes `ops:read` [+`ops:tasks` when the caller's token carries it])** — D4/B4. The app + client requests **only ops-tier scopes**; no finance/gmail scope is reachable through this agent. - **Data:** Data Library **Seahaven Ops Knowledge** via `Seahaven_Ops_Knowledge_Retriever` (Front SOPs, intake SOPs, Amazon subtree [restricted], SA8000/handbook once authored). -- **Model:** Claude on Bedrock (BYOLLM preferred; AWS-Hosted Claude Sonnet 4 fallback) — D5. +- **Model:** AWS-Hosted Claude (Salesforce-managed) — D5. - **Topics / Subagents:** - **Work Orders & Sites** — *Description:* WO/PO/site-assignment lookups. *Reasoning:* identify the record id or natural-language key; call the lookup tool; summarize with the WorkOrderSummary Flex template. *Actions:* @@ -383,12 +475,15 @@ Eval criteria: behavioral suite ≥ agreed pass rate before each cutover; MCP/co - **Error Message (≤255):** "I couldn't complete that finance lookup. Please retry; if it persists, contact Adam or accounting. (All lookups are audited.)" - **Languages:** English (US). -- **Variables:** `$User.Email`, `$Session.Channel` (decline sensitive detail in public channels). -- **Connections:** `sh-mcp-finance` MCP server — trust tier **finance**, `aud=sh-mcp-finance`, scope - **`finance:read`**, **15-min token TTL + deny-list** hard revocation (design.md §2.3/§2.5). Per-user OAuth → - Cognito (D4). **No ops/gmail connection on this agent.** +- **Variables:** `$User.Email`, `$Session.Channel` (decline sensitive detail in public channels — Q3/G15 caveat + as in §2.1). +- **Connections:** `sh-mcp-finance` **tier (separate OpenAPI facade, `aud=sh-mcp-finance`)** via External Service + actions over a **per-agent External Credential `ec-seahaven-finance` (Per-User OAuth Browser Flow → Cognito app + client `sh-agentforce-finance`, scope `finance:read` only)** — D4/B4. **15-min token TTL + deny-list** hard + revocation (design.md §2.3/§2.5). **No ops/gmail connection, and the app client cannot request ops/gmail + scopes** — the finance audience is reached only here. - **Data:** none (structured lookups only; no Data Library grounding). -- **Model:** Claude on Bedrock (D5). +- **Model:** AWS-Hosted Claude (Salesforce-managed) — D5. - **Topics / Subagents:** - **Vendor Search** — *Description:* QBO vendor lookup. *Reasoning:* query QBO; return vetted vendor records. *Actions:* `search_vendors` (MCP, `finance:read`, QBO server-held OAuth). @@ -416,15 +511,18 @@ Eval criteria: behavioral suite ≥ agreed pass rate before each cutover; MCP/co - **Error Message (≤255):** "I couldn't complete that. Please try again in DM; if it keeps failing, let Adam know. I only operate in direct messages." - **Languages:** English (US). -- **Variables:** `$User.Email` (the Google identity to act as), `$Session.Channel` (must be DM), per-user Google - grant status. -- **Connections:** `sh-mcp-ops` MCP server — trust tier **ops**, `aud=sh-mcp-ops`, scopes - **`ops:read ops:tasks gmail:self calendar:self`**. Gmail/Calendar act as the user via a **separate per-user - Google OAuth grant** (design.md §2.4), refresh tokens KMS-encrypted, **ABAC-partitioned by `sub`**. Per-user - Cognito OAuth (D4) is what carries the real `sub` so the server selects the right Google token — **directly - dependent on Gap G1**. **No finance connection.** +- **Variables:** `$User.Email` (the Google identity to act as), `$Session.Channel` (must be DM — **the DM-only + control depends on this variable existing; Q3/G15 — verify in Phase-0, else enforce DM-scoping another way**), + per-user Google grant status. +- **Connections:** `sh-mcp-ops` **tier (OpenAPI facade, `aud=sh-mcp-ops`)** via External Service actions over a + **per-agent External Credential `ec-lauren-exec` (Per-User OAuth Browser Flow → Cognito app client + `sh-agentforce-exec`, scopes `ops:read ops:tasks gmail:self calendar:self`)** — D4/B4. **The app client requests + NO finance scope**, so even though Lauren-the-person holds `finance:read`, no token this agent obtains carries + it (§0.1 B4). Gmail/Calendar act as the user via a **separate per-user Google OAuth grant** (design.md §2.4), + refresh tokens KMS-encrypted, **ABAC-partitioned by `sub`**; the per-user Cognito token carries the real `sub` + so the server selects the right Google token — **dependent on the Phase-0 auth spike (G1, unverified)**. - **Data:** none (operates on the user's live Gmail/Calendar, not a Data Library). -- **Model:** Claude on Bedrock (D5). +- **Model:** AWS-Hosted Claude (Salesforce-managed) — D5. - **Topics / Subagents:** - **Inbox Triage** — *Description:* summaries, high-priority, unanswered threads, bypassed work orders, search-by-sender. *Reasoning:* call read tools as the user; compose with the InboxDigest Flex template; @@ -447,27 +545,42 @@ Follows design.md §6 phasing; each legacy bot is deprecated **only at proven pa | Phase | Work | Parity / exit gate | Rollback | |-------|------|--------------------|----------| -| **0 — Foundation** | Stand up Cognito + Google federation + pre-token + 5-min group-sync (design.md §6.1). Create `sh-agentforce` SFDX repo + `cd-sfdx` reusable workflow (Gap G10). Stand up the Data Cloud org + **Seahaven Ops Knowledge** Data Library; repoint `notion-sync`. Decide D4 wrapper (Apex/External Service per-user named credential vs native MCP connector) after verifying G1/G2. | SSO login end-to-end; per-user JWT reaches a smoke-test MCP tool carrying the real `sub`; Data Library retriever returns Front SOP answers. | N/A (legacy still running) | -| **1 — Seahaven Ops** | Wire Ops agent → `sh-mcp-ops`; vendors (KB/Maps) + WO/PO/site + knowledge + tasks. | Behavioral suite + golden-transcripts vs **Alex** green; per-user auth + tool-hiding verified at MCP layer. | Keep Alex running in parallel; flip Slack default back to Alex. | -| **2 — Seahaven Finance** | Wire Finance agent → `sh-mcp-finance`; full audit logging; 15-min TTL + deny-list. | Audit records emitted; PII-redaction tests green; audience-binding rejection verified. | Finance lookups revert to Alex's QBO action group temporarily. | -| **3 — Lauren Exec** | Wire Exec agent → `sh-mcp-ops` `*:self`; per-user Google grant; DM-scoped. Refactor `fetch-classify`/`daily-digest`/`reminder` to import shared packages (design.md §6.4). | Channel-aware-privacy + golden-transcripts vs **Lauren** green; ABAC Gmail isolation proven. | Keep exec-aide running; Lauren's workflow needs explicit sign-off before retiring exec-aide (design.md §6.6). | -| **4 — Teardown** | Only after each replacement is signed off at parity. | — | — | +| **0a — Auth spike (GATING, B5)** | **Before any other Phase-0 spend.** Prove the **Per-User OAuth Browser Flow External Service action in Slack** carries the real `sub` to a Cognito-fronted smoke-test endpoint (§6 verify #1, #3, #5). Confirm reasoning-model selector + AWS-Hosted model name (#4). | Real `sub` reaches the server; first-call consent works in Slack; token refresh survives; Testing-Center can invoke per-user actions. **If the spike FAILS → STOP and switch to the fallback (below); do not proceed to 0b.** | N/A (legacy untouched) | +| **0b — Platform build (F4)** | Only after 0a passes. **Build the servers:** monorepo + `shared` auth/scope/PII/audit guard + per-service packages + **transport-agnostic core + OpenAPI adapter** + **per-tier API Gateway + Cognito authorizer + shared WAF** (B3); Cognito + Google federation + pre-token (per-audience scoping, B4) + 5-min group-sync (design.md §6.1); **deploy-role-first** (`githubdeploy-sh-mcp` already exists; add for `sh-agentforce`), CI/CD reusable workflows, coverage gate. Create `sh-agentforce` SFDX repo + **design & verify `cd-sfdx`** (G10, F3). Stand up Data Cloud + **Seahaven Ops Knowledge** Data Library; **`notion-sync` DUAL-FEEDS** Bedrock KB **and** Data Cloud (B1 — legacy KB stays fresh for rollback). | SSO end-to-end; per-user JWT reaches a smoke-test tool with real `sub`; **core + OpenAPI adapter green at the design.md §7.3 coverage gate** (80% lines, 100% shared auth guard); Data Library retriever returns Front SOP answers; legacy Bedrock KB still fed. | N/A (legacy untouched) | +| **1 — Seahaven Ops** | Wire Ops agent → `sh-mcp-ops` facade; vendors (KB/Maps) + WO/PO/site + knowledge + tasks. | Behavioral suite + golden-transcripts vs **Alex** green; per-user auth + per-tool scope verified at the server. | **Flip Slack default back to Alex** (Alex still running — NOT torn down until Phase 4). | +| **2 — Seahaven Finance** | Wire Finance agent → `sh-mcp-finance` facade; full audit logging; 15-min TTL + deny-list. | Audit records emitted; PII-redaction tests green; **per-tier audience-binding rejection verified** (B3). | **Flip back to Alex** (whose QBO action group is still live — Alex runs through Phase 4). | +| **3 — Lauren Exec** | Wire Exec agent → `sh-mcp-ops` facade `*:self`; per-user Google grant; DM-scoped. Refactor `fetch-classify`/`daily-digest`/`reminder` to import shared packages (design.md §6.4). | Channel-aware-privacy + golden-transcripts vs **Lauren** green; ABAC Gmail isolation proven. | **Keep exec-aide running**; Lauren's workflow needs explicit sign-off before retiring exec-aide (design.md §6.6). | +| **4 — Teardown** | Only after ALL three replacements are signed off at parity, and after the shared-consumer audit below. | Consumer audit passes (no live reader of the KB/AOSS remains). | — | -**What gets torn down, and when (design.md §1, §5):** -- **Bedrock agents** `seahaven-alex` (`QVL5GEJN9B`) and the exec-aide Sonnet loop — after their respective - agent's parity sign-off (Alex after Phase 1; Lauren after Phase 3). -- **Bedrock KB `LSDCNHTH6O`** + **OpenSearch Serverless `gv1540frh1crb79gtr4b`** (AOSS, INFRA-92) — after the - Data Cloud Data Library is proven in Phase 1 (both Bedrock retrieval Lambdas are non-VPC consumers of this - collection; confirm no other consumer before delete). -- **Bedrock guardrail `seahaven-alex-guardrail`** — **not** deleted until the MCP-layer PII redaction + (D11) - BYOLLM-path guardrail are live and tested (design.md §5: the guardrail must be explicitly replaced before - deletion — no parity assumed). -- **Sync Lambdas:** `po-sync` + `workorder-sync` decommissioned at Phase 1 (D7); `notion-sync` repointed, not - deleted; `fetch-classify`/`daily-digest`/`reminder` rebuilt, old exec-aide versions decommissioned at Phase 3. -- Archive `seahaven-slack-bot` + `exec-aide` repos; decommission their CDK stacks (design.md §1). +**Fallback if the Phase-0a auth spike fails (B5).** The plan does not discard design.md §12's alternatives without +a Plan B. If Per-User Browser Flow cannot carry the real `sub` to our server (or Slack consent / refresh / +Testing-Center identity is unworkable), fall back to a **custom Bolt assistant** as the surface (design.md §12) — +it keeps the **MCP transport + per-user OAuth** design intact end-to-end and therefore **un-defers the MCP adapter +(D12)**. This is a surface change, not an auth/scope/trust-tier redesign — the `sh-mcp` core, Cognito, scopes, and +trust tiers are unchanged. Decision point escalates to Adam; it does NOT mean restarting. -**Rollback principle:** legacy and replacement run **in parallel** through each phase; the Slack default agent is -the single flip point; no legacy component is deleted until the corresponding parity gate is signed off. +**What gets torn down — ALL IN PHASE 4 (B1-corrected; nothing shared is deleted while a consumer or rollback path +still needs it):** +- **Pre-req: shared-consumer audit.** The AOSS collection `gv1540frh1crb79gtr4b` / KB `LSDCNHTH6O` is a shared + dependency of **BOTH** `seahaven-slack-bot` **and** `exec-aide` (INFRA-92). Before any KB/AOSS delete, confirm + **zero** live consumers remain (both legacy stacks retired; no other reader). Run the INFRA-92 consumer check. +- **Bedrock agents** `seahaven-alex` (`QVL5GEJN9B`) and the exec-aide Sonnet loop — torn down in **Phase 4**, after + their parity sign-off AND after they are no longer any phase's rollback target. (Alex is the Phase 1 *and* Phase + 2 rollback target, so it must survive to Phase 4 — this was the B1 contradiction.) +- **Bedrock KB `LSDCNHTH6O`** + **AOSS `gv1540frh1crb79gtr4b`** — Phase 4, after the shared-consumer audit. +- **Bedrock guardrail `seahaven-alex-guardrail`** — **not** deleted until the **MCP/action-layer PII redaction + + prompt-attack replacement are live and tested** (D11, revised — there is no BYOLLM model path; design.md §5 + requires the guardrail be explicitly replaced before deletion, no parity assumed). +- **Sync Lambdas:** `notion-sync` **dual-feeds through Phase 3**, then drops the Bedrock-KB feed at Phase 4 (keeps + the Data Cloud feed). `po-sync` + `workorder-sync` **keep feeding the legacy KB until Phase 4** (so a rollback to + Alex is never stale — corrects the original "decommission at Phase 1"); they are retired as KB feeds at teardown + (D7 — the NEW platform never used them; structured data is live MCP tools). `fetch-classify`/`daily-digest`/ + `reminder` rebuilt in Phase 3; old exec-aide versions decommissioned at Phase 4. +- Archive `seahaven-slack-bot` + `exec-aide` repos; decommission their CDK stacks (design.md §1) — Phase 4. + +**Rollback principle:** legacy and replacement run **in parallel** through Phases 1–3; the Slack default agent is +the single flip point; **no legacy component — especially the shared KB/AOSS and any rollback-target bot — is +deleted, repointed-away, or starved of its sync feed until Phase 4.** --- @@ -484,24 +597,45 @@ the single flip point; no legacy component is deleted until the corresponding pa per-user auth model + Gap G1, Data Library/retriever design, model choice, Testing Center plan); "Agentforce DX & Release Process" (the `sh-agentforce` repo, `cd-sfdx`, scratch-org flow). -**Repo docs:** -- **`sh-mcp/docs/design.md` §4 (annotation, D12):** design.md §4 frames each server as a remote *MCP* server - (assumed Slack MCP-client consumer). Annotate it to record the **transport-agnostic core** decision — OpenAPI - adapter shipped now for Agentforce, MCP adapter deferred to its named trigger — so the build doesn't couple - tool logic to either protocol. (Annotation, not a reopening — the locked auth/scope/trust-tier decisions are - unchanged.) +**design.md amendments (F2 — these AMEND locked decisions; record them as amendments, not "annotations"):** +Two locked design.md decisions are genuinely changed by this plan (Adam-approved) and must be written into +design.md as dated amendments so the next builder isn't misled by stale locked text: +- **design.md §1 item 4 / §11 / §12** ("Slack's built-in agent is the MCP client; per-user OAuth via Slack's MCP + client; endpoint reasoning based on Slack egress") is **now false** — the consumer is Agentforce via OpenAPI + per-user actions; the MCP transport is deferred (D12); WAF egress reasoning re-scopes to **Salesforce/Hyperforce** + egress, not Slack. Amend §4 (transport-agnostic core, OpenAPI-first) and §11 (endpoint/egress). +- **design.md §9/§12** ("Lauren gets `finance:read`" *as an agent mapping*) is **reversed** by **D3** — the person + keeps the scope, the Exec agent does not. Amend the §12 agent mapping. +The locked auth/scope/trust-tier *primitives* (Cognito federation, audience-binding, per-tool scope, PII at the +MCP layer) are unchanged — those stay locked. + +**Repo READMEs & memory (F6 — global-instruction obligations, were omitted):** +- `sh-mcp` README updated to **OpenAPI-first** (servers expose OpenAPI; MCP adapter deferred). +- `sh-agentforce` README created (agent metadata repo, `cd-sfdx`, scratch-org flow). +- **Project memory:** create `project_sh_agentforce` before the repo-creating conversation ends; update + `project_sh_mcp` for the transport/auth changes; add **cross-project references** `sh-mcp` ↔ `sh-agentforce` + ↔ `project_seahaven_slack_bot` ↔ `project_exec_aide` (each side links the other). + +**Monitoring (N2 — ALARM-state-only convention, design.md alarm preferences):** +- Each per-tier **API Gateway facade**: 5xx-rate + auth-failure-spike alarms → site-alerts SNS (ALARM action + only, no OK/no-data). +- Repointed **`notion-sync`** carries the design.md ALARM-on-sync-failure through to the dual-feed (two-alarm + pattern: Errors≥1 missing=notBreaching; Invocations<1 over cadence missing=breaching). **Jira (INFRA project — no migration epic exists yet; related done: INFRA-92 AOSS lockdown, INFRA-37 reminder removal):** - **New epic:** "Agentforce migration & MCP platform." -- Stories (representative): foundation/Cognito+Google federation; `sh-agentforce` repo + `cd-sfdx` reusable - workflow (G10); **Phase-0 per-user auth spike** — prove the Per-User Browser Flow External Service action in - Slack (real `sub` reaches the server, consent UX, token refresh; §6 verify gate) — **gates everything, do - first**; **build the transport-agnostic core + OpenAPI adapter** (D12), MCP adapter deferred to its trigger; - Data Cloud Data Library + repoint `notion-sync`; retire `po-sync`/`workorder-sync` (D7); Ops agent build + - parity; Finance agent + audit; Lauren Exec + per-user Google grant; **author SA8000 + employee-handbook - corpus** (G8, fund when needed); Testing Center suites; teardown (Bedrock agents/KB/AOSS/guardrail/sync - Lambdas). Every IAM/auth/connection story carries the **mandatory cross-review** label (design.md §8/§10). +- Stories, in dependency order: (0a) **per-user auth spike — gates everything, do first** (Per-User Browser Flow + ES action in Slack: real `sub`, consent, refresh, Testing-Center identity; §6 verify; fallback = custom Bolt); + (0b) foundation — Cognito+Google federation + **per-audience pre-token scoping** (B4) + 5-min group-sync; + **build the transport-agnostic core + OpenAPI adapter + per-tier API GW/WAF** (D12/B3/F4); `sh-agentforce` repo + (full new-repo checklist) + **design & verify `cd-sfdx` + connected-app JWT key in Secrets Manager** (G10/F3); + Data Cloud Data Library + **`notion-sync` dual-feed** (B1); (1) Ops agent + parity; (2) Finance agent + audit; + (3) Lauren Exec + per-user Google grant; **author SA8000 + employee-handbook corpus** (G8, fund when needed); + Testing Center suites; (4) **teardown** after parity sign-off + shared-consumer audit — Bedrock agents/KB/AOSS/ + guardrail and `po-sync`/`workorder-sync`/`notion-sync`-Bedrock-feed (B1). Every IAM/auth/connection story + (incl. the per-audience pre-token Lambda and the `cd-sfdx` connected-app trust) carries the **mandatory + cross-review** label (design.md §8/§10). --- @@ -511,7 +645,7 @@ Severity: **C**ritical / **M**edium / **L**ow. "Verify" = check against current | # | Gap | Sev | Impact | Workaround / status | Verify | |---|-----|-----|--------|---------------------|--------| -| **G1** | **RESOLVED 2026-06-11 (research wf_1bf9e142).** Per-user OAuth Browser Flow on the **native custom remote-MCP connector is unconfirmed**; documented per-user identity exists only for SF-**hosted** MCP. | **C→ mitigated** | If we'd relied on the MCP connector and it bound only a Named Principal: per-user `sub`, ABAC Gmail isolation, and trifecta all break. | **Decided (D4):** do NOT use the native MCP connector for the per-user hop. Use **GA External Service (OpenAPI)/Apex actions + Per-User OAuth Browser Flow External Credential → Cognito** (per-user OAuth is GA, ships in `UserExternalCredential`). | In-org: build one ES action on a Per-User Browser Flow cred → confirm first-call consent, real `sub` reaches the server, and token auto-refresh (§6 verify list) | +| **G1** | **MITIGATION CHOSEN — UNVERIFIED, SPIKE-GATED (B5-relabel).** Per-user OAuth Browser Flow on the native custom remote-MCP connector is unconfirmed (per-user identity is documented only for SF-**hosted** MCP). The chosen GA path (D4) is *plausible* but **not yet proven** to carry our Cognito `sub` end-to-end in Slack. | **C (open until Phase-0a)** | If the mitigation also fails to propagate the real `sub`: per-user identity, ABAC Gmail isolation, and the trifecta all break — and there is no Agentforce-native surface. | **Decided (D4):** ES (OpenAPI)/Apex actions + Per-User OAuth Browser Flow External Credential → Cognito. **Gated by the Phase-0a spike; if it fails → custom-Bolt fallback (§3).** Status is *mitigation chosen, not verified* — do not call this resolved until 0a passes. | Phase-0a: real `sub` reaches the server, first-call consent in Slack, refresh, Testing-Center identity (§6 verify #1/#3/#5) | | **G2** | **RESOLVED 2026-06-11.** Custom remote MCP client is **Beta** (Pilot Jul 2025 → Beta Jan 2026); SF-*hosted* MCP is GA (Apr 2026) but that's not our remote servers. | **C→ avoided** | Schedule/stability risk on the native-MCP path. | **Avoided by D4** — the GA ES/Apex action path is the per-user transport; the Beta MCP connector is off the critical path. Revisit native MCP once remote-MCP per-user binding reaches GA. | SF release notes for remote-MCP GA date | | **G3** | **Trust Layer masking may over-mask.** Agentforce Trust Layer can mask PII; legacy Alex **deliberately left vendor names/contacts unmasked** (lookup is the bot's job). | M | Vendor lookups could be degraded if Trust Layer masks names. | Configure Trust Layer masking to exclude names/contacts; keep MCP-layer redaction authoritative for bank/routing/card/SSN (design.md §2.5). | Trust Layer data-masking config | | **G4** | **Loss of free-text search over WO comments** (old KB feature) — Data Libraries are unstructured-only, WO/PO are structured MCP lookups. | L | Can't fuzzy-search WO comment text. | Lookup-by-id via MCP tools; or a dedicated retriever over a comment text export if demand appears. | — | @@ -519,9 +653,13 @@ Severity: **C**ritical / **M**edium / **L**ow. "Verify" = check against current | **G6** | **RESOLVED 2026-06-11 (research wf_1bf9e142): BYOLLM cannot drive the planner.** Only Salesforce Default / AWS-Hosted options drive the Atlas reasoning engine; BYOLLM is custom-action-only and still routes through SF's Models API/Trust Layer. | M→ resolved | Our-account inference + our guardrail are **unsatisfiable on the planner path**. | **Decided (D5):** use **AWS-Hosted Claude** as the planner; move our guardrail + PII redaction to the **MCP/action layer** (revises D11); rely on the Trust Layer for planner-path moderation. Note AWS-Hosted model is drifting Sonnet 4 → 4.6/Haiku 4.5 (May 2026). | In-org: confirm the reasoning-model selector offers only Default/AWS-Hosted; confirm current AWS-Hosted model name | | **G7** | **Amazon/AMOC corpus is stale** ("migrated from BookStack, may need updating") and access-restricted (AMOC comms = Adam & Robert only). | M | Agent could give outdated Amazon-site guidance. | Audience-restrict the AMOC topic; flag content as stale; re-author before exposing widely. | Notion *AMOC* `33a2ecdd…8162` currency | | **G8** | **SA8000 docs + employee handbook + company policies missing** from Notion (referenced as KB inputs, not found). | M | SA8000/compliance + HR Q&A **blocked**. | **Blocked pending content authoring** — author, stage in S3/Drive, ingest into Ops Knowledge Data Library; until then the agent must decline (without suppressing legitimate misconduct questions). | design.md corpus notes; locate any S3/Drive originals | -| **G9** | **Agentforce + Data Cloud licensing / Einstein-Request consumption.** | M | Cost; BYOLLM cuts ~30% of Einstein Requests but Data Cloud + Agentforce licensing still applies. | Budget; BYOLLM to reduce request spend. | Salesforce contract / Einstein Request limits | -| **G10** | **No SFDX reusable CI/CD workflow** — org reusable workflows are AWS/CDK/OIDC-shaped (design.md §7.1). | M | `sh-agentforce` can't deploy via the existing pattern. | Author a new reusable **`cd-sfdx`** workflow (deploy-then-merge, scratch-org validate). | handbook `cicd.md` | +| **G9** | **Agentforce + Data Cloud licensing / Einstein-Request consumption.** | M | Cost. (The "~30% fewer Einstein Requests via BYOLLM" claim is **dropped — uncorroborated**, §0.1; and BYOLLM is ruled out for the planner anyway, G6.) | Budget Agentforce + Data Cloud licensing + Einstein-Request limits explicitly; see G15 (per-employee licensing). | Salesforce contract / Einstein Request limits | +| **G10** | **No SFDX reusable CI/CD workflow + no OIDC for Salesforce deploy** — org reusable workflows are AWS/CDK/OIDC-shaped (design.md §7.1). | M | `sh-agentforce` can't deploy via the existing pattern; the JWT-bearer connected-app key is a long-lived prod-deploy secret. | Author reusable **`cd-sfdx`** (deploy-then-merge, scratch-org validate); store the connected-app **JWT key in Secrets Manager** (`sh-agentforce/sfdx-jwt-key`), sandbox/prod-scoped, rotated (F3, §1g). Its own design/verify story. | handbook `cicd.md`; SFDX JWT-bearer flow | | **G11** | **Two access-control planes.** Agentforce-in-Slack assigns member access via **Salesforce permissions**; our authz is **Google Groups → Cognito → scopes**. | M | Drift: a user could see an agent but lack the MCP scope, or vice-versa. | Provision Agentforce/Salesforce user access **from the same Google Groups** (SCIM/identity sync) so Google Groups stays the single source of truth (design.md §2.3). | Salesforce SCIM/Google provisioning | +| **G12** | **Parity gate runs on Beta tooling + per-user identity (F1).** Testing Center Test Suites is Beta; batch/AI-generated runs invoke actions under an unclear identity when the action uses a Per-User Browser Flow credential. | M | If batch runs can't invoke per-user actions, the behavioral suite (which authorizes teardown) can't exercise the real tools. | Verify Testing-Center identity under per-user creds in Phase-0a (§6 #5); if unworkable, drive parity via scripted per-user sessions instead of batch. | help.salesforce.com Testing Center; in-org | +| **G13** | **New data processor: Salesforce now processes our tool payloads (F5).** Finance/Gmail tool responses transit the Agentforce planner / Einstein Trust Layer; grounding flows through Data Cloud. design.md's data boundary was AWS + Slack. | M | Vendor/payment/inbox content (PII is masked, but business content is not) is processed by Salesforce — a new processor not in the original threat model. | Explicit acceptance; verify Trust Layer **retention + zero-training** guarantees and Data Cloud data-residency; keep MCP-layer PII redaction authoritative. | Einstein Trust Layer retention/zero-training docs; DPA | +| **G14** | **Per-user tool visibility degraded (F7).** MCP `list_tools` hid tools per-*user* by scope; per-agent action assignment is per-*agent*, so a user lacking a scope still sees the action and fails server-side (403). | L | UX degradation (confusing failures), not a security hole — the per-tool scope check is the boundary. | Graceful-refusal copy in agent instructions; behavioral test for a clean "no access" message (§1h). | — | +| **G15** | **Three unverified-load-bearing platform assumptions (Q1/Q2/Q3).** (a) Sea Haven's Slack plan actually supports **Agentforce Employee Agents**; (b) per-employee Agentforce **user+license** required for the all-staff Ops agent (drives G9 cost + G11 SCIM); (c) Agentforce exposes **`$User.GoogleGroups`** and **`$Session.Channel`** to agent instructions (Lauren Exec's DM-only + scope gating depend on these). | M | If (a) wrong, the surface decision (D2) collapses; if (c) wrong, privacy/DM controls need a different enforcement point (e.g. a context Apex action). | Verify all three in Phase-0 (§6); for (c), design a fallback enforcement now so it's not on the critical path. | slack.com/help Agentforce-in-Slack plan reqs; Salesforce licensing; Agentforce agent variables doc | `FACT` for G1/G2: Agentforce MCP support (Pilot Jul 2025, Beta Jan 2026; SF-hosted MCP GA Apr 2026; OAuth 2.0, JSON-RPC over Streamable HTTP) and Named/External Credential **Per User** identity type — verified via Salesforce @@ -535,8 +673,11 @@ sources (§Sources). The **specific** per-user binding for MCP connectors is the > licensing **accepted**; (2) reasoning model = **AWS-Hosted Claude** (BYOLLM ruled out for the planner); > (3) **D3 accepted** (finance stripped from Lauren Exec); (4) per-user auth = **ES/Apex actions + Per-User > Browser Flow → Cognito** (native MCP connector off the per-user path); (5) corpus authoring (G8) **deferred — -> fund when needed**. The hands-on verify-in-org tests below remain the build-time gate for #2 and #4. Original -> recommendations retained for the record: +> fund when needed**. The hands-on verify-in-org tests below remain the build-time gate. **The historical +> recommendations 1–5 below predate the research (wf_1bf9e142) and retain superseded wording** (e.g. the +> "~30% fewer Einstein Requests" figure in #2, since **dropped as uncorroborated** — §0.1, G9; and "BYOLLM if it +> can be the planner," since **ruled out** — G6). Read §0.1 for the corrected, committed position; the text below +> is kept only as a decision trail. 1. **Surface + licensing.** Confirm **Agentforce Employee Agents in Slack** as the surface and accept Agentforce + Data Cloud licensing/Einstein-Request cost (G9). *Recommend: yes* — Slack is already our hub and only @@ -572,6 +713,15 @@ Groups to reconcile the two control planes (G11, recommend SCIM). 4. **(Model)** Confirm the reasoning-engine selector (Setup → Agentforce Agents; `model_config` in Agent Script) offers only **Salesforce Default** / **AWS-Hosted**, and confirm the current model name behind AWS-Hosted (Sonnet 4 vs 4.6 vs Haiku 4.5). +5. **(Testing Center identity, F1/G12)** Confirm a Testing Center batch/AI-generated run can invoke a Per-User + Browser Flow action and under whose identity. If it can't exercise per-user tools, the behavioral parity gate + must run as scripted per-user sessions instead of batch — decide before Phase 1. +6. **(Slack plan + licensing, Q1/Q2/G15)** Confirm Sea Haven's Slack plan supports **Agentforce Employee Agents**, + and whether the all-staff Ops agent needs a provisioned Agentforce **user + license per employee** (prices G9, + feeds G11 SCIM scope). +7. **(Agent variables, Q3/G15)** Confirm Agentforce exposes **`$User.GoogleGroups`** and **`$Session.Channel`** to + agent instructions. If not, design an alternate enforcement for Lauren Exec's DM-only and per-scope gating + (e.g. a context Apex action that returns channel + group facts) before Phase 1/3. ---