Remediate sh-plan-review findings (B1-B5, F1-F7, NITs, Qs)

Address the Fable plan-review gate (REQUEST CHANGES):
- B1: move ALL teardown to Phase 4 + shared-consumer audit; notion-sync dual-feeds;
  rollback never targets a deleted/starved resource.
- B2: propagate the §0.1 resolutions through D5/D11/§1d/§2.x/§3/G6/G9/§6 (model,
  guardrail, dropped ~30% claim, Phase-0 'verify' not 'decide').
- B3: pin D12 to one facade per trust tier (per-server Cognito audience at the edge);
  reconcile the §1h list_tools test (deferred with MCP adapter).
- B4: specify per-agent External Credential + Cognito app client minting only its tier's
  scopes; add the token-layer trifecta test.
- B5: split Phase 0 (0a auth spike gates 0b); add custom-Bolt fallback; relabel G1 as
  mitigation-chosen/unverified.
- F1-F7: Testing-Center identity (G12); design.md amendments reframed (F2); sh-agentforce
  new-repo checklist + cd-sfdx JWT-key auth (F3); platform-build phase (F4); Salesforce
  data-processor gap (G13/F5); memory+README obligations (F6); per-user visibility
  degradation (G14/F7).
- NITs: S3->Data Cloud ingestion pinned; facade+notion-sync ALARM monitoring; DevName
  naming boundary. Qs Q1-Q3 registered as G15 + verify items.
- §0.3 remediation log + gap count 11->15.
This commit is contained in:
Adam Moussa 2026-06-11 12:44:00 -04:00
parent f4a11c558f
commit 74e834a7c4

View file

@ -46,18 +46,26 @@ off the per-user hop, so the risk is mitigated rather than load-bearing.
| D8 | Proactive jobs (`fetch-classify`, `daily-digest`, `reminder`, `notion-sync`) **stay as our scheduled Lambdas**, not Agentforce; HIGH-priority alerts continue as **Slack DMs** | No conversational equivalent; keeps Haiku classifier in our Bedrock account; cheapest reliable path | design.md §5; §1f below | | D8 | Proactive jobs (`fetch-classify`, `daily-digest`, `reminder`, `notion-sync`) **stay as our scheduled Lambdas**, not Agentforce; HIGH-priority alerts continue as **Slack DMs** | No conversational equivalent; keeps Haiku classifier in our Bedrock account; cheapest reliable path | design.md §5; §1f below |
| D9 | New separate **`sh-agentforce` SFDX repo** for agent metadata (Bot/GenAiPlannerBundle/GenAiPlugin/GenAiFunction) under Agentforce DX | Different toolchain (sf CLI, scratch orgs, deploy-to-org) than the CDK monorepo; one deploy target per repo | `FACT` developer.salesforce.com Agent DX metadata; handbook | | D9 | New separate **`sh-agentforce` SFDX repo** for agent metadata (Bot/GenAiPlannerBundle/GenAiPlugin/GenAiFunction) under Agentforce DX | Different toolchain (sf CLI, scratch orgs, deploy-to-org) than the CDK monorepo; one deploy target per repo | `FACT` developer.salesforce.com Agent DX metadata; handbook |
| D10 | Build **Ops + Finance agents now**; HR/Gusto, SA8000 Q&A, IT/onboarding **later, corpus-gated**; add near-term capabilities as **Topics**, not new agents | Front/ops corpus is rich and current; HR/Amazon/SA8000 corpus is thin/stale/missing | Notion (below); design.md corpus notes | | D10 | Build **Ops + Finance agents now**; HR/Gusto, SA8000 Q&A, IT/onboarding **later, corpus-gated**; add near-term capabilities as **Topics**, not new agents | Front/ops corpus is rich and current; HR/Amazon/SA8000 corpus is thin/stale/missing | Notion (below); design.md corpus notes |
| D11 | Keep the **Bedrock guardrail on the BYOLLM model path** + MCP-layer PII redaction (defense in depth) | Agentforce Trust Layer does not redact OUR tool-response payloads | design.md §2.5, §11.2 | | D11 | **REVISED (2026-06-11, §0.1):** guardrail (prompt-attack + PII redaction) lives **entirely at the MCP/action layer**, not on any model path; the Einstein Trust Layer covers planner-path moderation | The planner runs AWS-Hosted Claude inside Salesforce's boundary (D5) — there is no BYOLLM model path to attach our guardrail to; Trust Layer does not redact OUR tool-response payloads, so MCP-layer redaction is authoritative | design.md §2.5, §11.2; §0.1 |
| D12 | **RESOLVED (2026-06-11, §0.1):** **OpenAPI-first, MCP-ready transport-agnostic core.** Tool handlers + the shared auth/scope/PII/audit guard are built independent of transport; ship the **OpenAPI adapter now** (Agentforce External Service actions; Apex only where an action needs request/response shaping); **defer the MCP adapter** to a named trigger. | Agentforce's per-user hop needs OpenAPI regardless (D4); no committed MCP consumer today, so a second public interface + its conformance/hardening cost isn't justified yet; a transport-agnostic core makes the MCP adapter a cheap later add, not a re-platform; deferring it also drops the IDE-session trifecta nuance until then. | Adam 2026-06-11; research wf_1bf9e142 (D4) | | D12 | **RESOLVED (2026-06-11, §0.1):** **OpenAPI-first, MCP-ready transport-agnostic core.** Tool handlers + the shared auth/scope/PII/audit guard are built independent of transport; ship the **OpenAPI adapter now** (Agentforce External Service actions; Apex only where an action needs request/response shaping); **defer the MCP adapter** to a named trigger. | Agentforce's per-user hop needs OpenAPI regardless (D4); no committed MCP consumer today, so a second public interface + its conformance/hardening cost isn't justified yet; a transport-agnostic core makes the MCP adapter a cheap later add, not a re-platform; deferring it also drops the IDE-session trifecta nuance until then. | Adam 2026-06-11; research wf_1bf9e142 (D4) |
**Agent roster (the answer to "how many"):** **3** — Seahaven Ops, Seahaven Finance, Lauren Exec (§1a, §2). **Agent roster (the answer to "how many"):** **3** — Seahaven Ops, Seahaven Finance, Lauren Exec (§1a, §2).
**Top 5 decisions for Adam (§6):** (1) confirm Agentforce-in-Slack as the surface + accept Agentforce/Data **Top 5 decisions for Adam — ALL RESOLVED 2026-06-11 (§0.1):** (1) Agentforce-in-Slack + licensing **accepted**;
Cloud licensing; (2) BYOLLM-on-our-Bedrock vs AWS-Hosted Claude Sonnet 4; (3) accept D3 (strip finance from (2) reasoning model = **AWS-Hosted Claude** (BYOLLM ruled out for the planner); (3) **D3 accepted** (strip finance
Lauren's exec agent); (4) accept the Apex/External-Service per-user wrapper (D4) rather than waiting for native from Lauren Exec); (4) **D4** per-user auth = ES/Apex actions + Per-User Browser Flow → Cognito; (5) corpus
MCP per-user GA; (5) fund authoring the missing SA8000 + employee-handbook corpus. authoring **deferred** (fund when needed).
**Capability-gap count: 11** (§5). Severity: **2 critical** (G1 per-user MCP auth, G2 remote-MCP Beta), **6 **Capability-gap count: 15** (§5). Severity: **1 critical** (G1 per-user auth — mitigation chosen, spike-gated,
medium**, **3 low**. **unverified**), **9 medium**, **2 low**, **3 resolved/avoided** (G2, G6, and the BYOLLM cost claim).
> **REVISION — plan-review remediation (2026-06-11, §0.3).** This plan was audited by the `sh-plan-review` Fable
> gate; §0.3 logs every BLOCK/FIX addressed. Key structural changes since the first commit: teardown moved wholly
> to Phase 4 with a shared-consumer audit (was contradictory); §0.1 propagated into D5/D11/§1d/§2.x/§3/G6/G9/§6
> (model/guardrail consistency); D12 transport pinned to **one facade per trust tier** (per-server audience
> preserved); per-agent External Credential / Cognito app-client scoping specified (token-layer trifecta);
> Phase 0 split so the auth spike gates the rest, with a **custom-Bolt fallback** if it fails; a platform-build
> phase added; new gaps G12–G15 registered.
--- ---
@ -78,12 +86,13 @@ provisional D4/D5 wording above and the open items in §6.
**Two consequences that ripple into the design (must be honored downstream):** **Two consequences that ripple into the design (must be honored downstream):**
1. **The Agentforce→tool hop is NOT MCP.** Because per-user identity is only achievable on the GA External 1. **The Agentforce→tool hop is NOT MCP.** Because per-user identity is only achievable on the GA External
Service/Apex action path (not the Beta MCP connector), Agentforce reaches our tools through an **OpenAPI/Apex Service/Apex action path (not the Beta MCP connector), Agentforce reaches our tools over **OpenAPI/Apex
facade in front of each MCP server**, authenticated per-user to Cognito. The `sh-mcp-ops`/`-finance` servers, actions** authenticated per-user to Cognito. **Day-one architecture (B3-resolved):** there is NO MCP wire
their JWT validation, scopes, audience-binding, and PII redaction are **unchanged** — only the Agentforce-side protocol at launch — the trust-tier servers (`sh-mcp-ops`, `sh-mcp-finance`) expose their tools as **OpenAPI
transport changes from MCP to OpenAPI/Apex actions. (Other MCP clients can still speak MCP to the same servers.) endpoints** via the transport-agnostic core (D12); the MCP adapter is deferred (D12 trigger). The servers'
Browser Flow requires an interactive first-auth per user (fine for Slack users; the proactive Lambdas in §1f JWT validation, per-tool scope enforcement, audience-binding, PII redaction, and audit are **unchanged** —
are unaffected — they use offline creds). they sit below the adapter and are transport-agnostic. Browser Flow requires an interactive first-auth per
user (fine for Slack users; the proactive Lambdas in §1f are unaffected — they use offline creds).
2. **Our Bedrock guardrail cannot sit on the planner path.** Both AWS-Hosted and BYOLLM keep Salesforce in the 2. **Our Bedrock guardrail cannot sit on the planner path.** Both AWS-Hosted and BYOLLM keep Salesforce in the
inference path, so keeping inference + our own guardrail in account `328440206208` is **unsatisfiable on the inference path, so keeping inference + our own guardrail in account `328440206208` is **unsatisfiable on the
planner path today**. The Einstein **Trust Layer** covers planner-path moderation; **our guardrail + PII planner path today**. The Einstein **Trust Layer** covers planner-path moderation; **our guardrail + PII
@ -99,8 +108,14 @@ Agentforce per-user path (D4) is OpenAPI/Apex, not MCP. Decision:
tool list are generated from one tool registry, so the two can never drift. tool list are generated from one tool registry, so the two can never drift.
- **Ship the OpenAPI adapter now** — Agentforce **External Service (OpenAPI) actions** are the primary transport; - **Ship the OpenAPI adapter now** — Agentforce **External Service (OpenAPI) actions** are the primary transport;
use **Apex `@InvocableMethod`** only for actions needing request/response shaping (e.g. extra redaction, use **Apex `@InvocableMethod`** only for actions needing request/response shaping (e.g. extra redaction,
pagination). This is the one interface to build, secure (one API Gateway + Cognito authorizer + WAF web ACL), pagination).
and test on the critical path. - **One facade per trust tier (B3-resolved — the audience boundary stays in infrastructure, not in shared
handlers).** Deploy a **separate API Gateway + Cognito authorizer per server**: `sh-mcp-ops` facade validates
`aud=sh-mcp-ops`, `sh-mcp-finance` facade validates `aud=sh-mcp-finance`. A token minted for ops is rejected at
the finance gateway's authorizer (and vice-versa) — the audience-binding rejection design.md §2.5 requires,
enforced at the edge. A **shared AWS WAF web ACL** fronts both. This is deliberately NOT one shared gateway with
audience checks pushed into application code: keeping per-tier audiences at separate authorizers preserves the
trust-tier firewall even if a handler bug slips through.
- **Defer the MCP adapter** until a **named trigger**: (a) Claude Code / IDE consumption becomes a real recurring - **Defer the MCP adapter** until a **named trigger**: (a) Claude Code / IDE consumption becomes a real recurring
workflow (not nice-to-have), **or** (b) Salesforce native remote-MCP per-user binding reaches GA (then MCP-native workflow (not nice-to-have), **or** (b) Salesforce native remote-MCP per-user binding reaches GA (then MCP-native
could also collapse the Agentforce-side OpenAPI facade). When triggered, the MCP adapter is a thin add over the could also collapse the Agentforce-side OpenAPI facade). When triggered, the MCP adapter is a thin add over the
@ -113,11 +128,55 @@ Agentforce per-user path (D4) is OpenAPI/Apex, not MCP. Decision:
"don't co-connect finance with Gmail-read in one IDE session." The `sh-mcp` name stays accurate — MCP remains the "don't co-connect finance with Gmail-read in one IDE session." The `sh-mcp` name stays accurate — MCP remains the
strategic protocol, just not the day-one transport. strategic protocol, just not the day-one transport.
**Per-agent credential & token-layer trifecta (B4-resolved — the guarantee is designed, not asserted).** The
claim "separate audiences → no single token spans tiers" is only true if each agent's tokens are minted from a
credential that requests **only its tier's scopes**. Mechanism (specified, testable):
- **One Cognito app client + one External Credential per agent**, each registered with **only its tier's
resource-server scopes**: Seahaven Ops client → `ops:read [ops:tasks]` (`aud=sh-mcp-ops`); Seahaven Finance
client → `finance:read` (`aud=sh-mcp-finance`); Lauren Exec client → `ops:read ops:tasks gmail:self
calendar:self` (`aud=sh-mcp-ops`, **never finance**).
- design.md §2.3 mints the **union** of a user's group-scopes into a token *only for the audience requested*.
Because Lauren Exec's app client requests only the ops audience and ops scopes, **Lauren-the-person's
`finance:read` (from `-finance@`) is never present in any token the Exec agent obtains** — even though she
holds it. She reaches finance only through the Finance agent's separate client/audience. The boundary is the
per-agent credential, not per-agent action assignment (which the plan concedes is "a convenience, never the
boundary," §1h).
- The Cognito pre-token Lambda must therefore scope the minted scopes to the requested `aud` (resource server),
not blanket-union across all audiences. **This is a per-server IAM/authz behavior → mandatory cross-review
(GPT-4.1) before commit** (design.md §8/§10).
- Tested by: "per-agent credential mints only its tier's scopes" (a token obtained via the Exec agent's
credential never contains `finance:read`) — added to the §1h security suite.
**Dropped claim:** the "BYOLLM = ~30% fewer Einstein Requests" figure was **not corroborated** by any verified **Dropped claim:** the "BYOLLM = ~30% fewer Einstein Requests" figure was **not corroborated** by any verified
source — removed from cost modeling. source — removed from cost modeling.
--- ---
## 0.3 Plan-review remediation log — 2026-06-11 (Fable `sh-plan-review` gate)
The first commit of this plan was audited by the Fable adversarial gate (verdict: REQUEST CHANGES). Every BLOCK
and FIX is addressed below; this revision supersedes the pre-audit text wherever they differ.
| Finding | Resolution (where) |
|---------|--------------------|
| **B1** Teardown self-contradiction (Alex deleted Phase 1 *and* the Phase 2 rollback target; shared AOSS deleted before exec-aide retires; KB starved during parallel run) | §3 rewritten: **all** teardown moved to Phase 4 with a shared-consumer audit; legacy bots + KB + AOSS + all sync jobs run untouched through Phase 3; `notion-sync` **dual-feeds** (Bedrock KB + Data Cloud) until Phase 4; rollback never targets a deleted resource. |
| **B2** Plan body still stated pre-resolution decisions (BYOLLM preferred; ~30%; guardrail-on-BYOLLM-path; "decide D4") | §0.1 propagated into **D5, D11, §1d, §2.1–2.3 model lines, §3 guardrail-teardown condition, G6, G9, §6, Phase 0**. |
| **B3** D12 launch architecture ambiguous; one-gateway vs per-server audience | §0.1: **no MCP at launch**; **one facade (API GW + Cognito authorizer) per trust tier**, audience rejection at the edge; §1h list_tools test moved to deferred set. |
| **B4** Token-layer trifecta asserted, not designed | §0.1 per-agent credential block above; §2 Connections name each per-agent app client/credential; new §1h test. |
| **B5** Phase-0 spike "gates everything" but no fallback; spend precedes it; G1 overstated | §3 **Phase 0 split (0a auth spike → 0b platform/Data-Cloud/SFDX)**; explicit **custom-Bolt fallback** if the spike fails (design.md §12); G1 relabelled "mitigation chosen — UNVERIFIED, spike-gated." |
| **F1** Parity gate runs on Beta tooling; Testing-Center identity under per-user creds | §6 verify-#5 added; **G12** registers the Beta/identity dependency. |
| **F2** design.md "annotation, not reopening" inaccurate | §4 reframed as **amendments to locked decisions** (design.md §1 item 4 now false; D3 reverses §12). |
| **F3** `sh-agentforce` new-repo obligations + `cd-sfdx` Salesforce auth missing | §1g + §4 expanded: full new-repo checklist; **JWT-bearer connected-app cert/key in Secrets Manager**, sandbox/prod-scoped, rotation; `cd-sfdx` gets its own design/verify story. |
| **F4** No phase builds the servers | §3 **Phase 0b "Platform build"** added with deploy-role-first, CI/CD, coverage-gate exit criteria. |
| **F5** New trust boundary: Salesforce now processes tool payloads | **G13** registered (Trust Layer / Data Cloud as new data processor; verify retention + zero-training). |
| **F6** Memory/README obligations omitted | §4 **Repo docs** expanded: `sh-agentforce` project memory, `project_sh_mcp.md` update, cross-project refs, both READMEs. |
| **F7** Per-user tool visibility degraded with no replacement | §1h states it; behavioral test + graceful-refusal UX added; **G14**. |
| **N1/N2/N3** | §1f ingestion pinned to **S3 → Data Cloud**; ALARM-only monitoring for the facades + repointed `notion-sync`; Salesforce DevName/kebab boundary noted. |
| **Q1/Q2/Q3** | **G15** (Slack plan supports Employee Agents; per-employee Agentforce licensing; `$User.GoogleGroups`/`$Session.Channel` existence) — all unverified-load-bearing, on the §6 verify gate. |
---
## 1. Platform-level design ## 1. Platform-level design
### 1a. Agent roster — how many, and why ### 1a. Agent roster — how many, and why
@ -182,26 +241,28 @@ Physical tier (`physical:*`) remains **deferred / admin-out-of-band**, no scope
### 1d. AI models — which, and where ### 1d. AI models — which, and where
`FACT` (developer.salesforce.com supported-models; Salesforce "Agentforce 360 for AWS"; Salesforce×Anthropic `FACT` (research wf_1bf9e142, 25/25 verified; developer.salesforce.com supported-models; "Agentforce 360 for
Oct-2025 partnership): Agentforce's **Atlas reasoning engine is model-agnostic**; the default is a AWS"; Salesforce×Anthropic Oct-2025): Agentforce's reasoning engine/planner is driven by the **org/agent model
Salesforce-managed mix (incl. GPT-4o); an **AWS-Hosted option runs Anthropic Claude Sonnet 4 on Amazon Bedrock** selection**, whose **only two documented options are "Salesforce Default"** (managed mix, currently GPT-4o) **and
and can power Atlas; **BYOLLM** (Models API) supports **Amazon Bedrock, Azure OpenAI, OpenAI, Vertex**, runs on "AWS-Hosted"** (Anthropic Claude Sonnet 4 on Bedrock, **inside the Salesforce Trust Boundary**, not our account).
your own credentials/instance, keeps the **Trust Layer**, and consumes **~30% fewer Einstein Requests**. **BYOLLM (Models API) is documented ONLY for custom actions** (prompt templates/Apex/Models-API calls), is **not a
reasoning-engine option**, and even on its action path inference still routes **through Salesforce's Models API /
Trust Layer** — so BYOLLM cannot keep inference in our boundary either.
**Decision (D5):** **Decision (D5) — RESOLVED, §0.1:**
- **Agent reasoning / planner model = Anthropic Claude on Bedrock.** Preferred path: **BYOLLM pointed at our - **Agent reasoning/planner model = AWS-Hosted Claude (Salesforce-managed on Bedrock).** This is the **only**
Bedrock** (account `328440206208`, `us-east-1`) so inference stays in our trust boundary, we keep our own documented way to keep Claude as the planner. **BYOLLM is ruled out for the planner** (G6, resolved). Consequence:
guardrail on the model path (D11), and we cut Einstein-Request spend. **Gap G6:** confirm a BYOLLM endpoint can inference and our own Bedrock guardrail **cannot sit on the planner path** (it runs in Salesforce's boundary) —
be the *agent reasoning model* (Atlas planner), not only a prompt-template/Models-API call. If not, fall back our guardrail + PII redaction move **entirely to the MCP/action layer** (D11) and the **Einstein Trust Layer**
to the **AWS-Hosted Claude Sonnet 4** managed option (confirmed to power Atlas) — same model family, less covers planner-path moderation. We still retain the Claude lineage of the legacy bots (Sonnet 4.5/4.6).
control. Either way we retain the Claude lineage of the legacy bots (Sonnet 4.5 for Alex, Sonnet 4.6 for `ASSUMPTION`: AWS-Hosted's model name is drifting (Sonnet 4 → 4.6 / Haiku 4.5 as of May 2026) — pin the current
Lauren's conversation loop). name in the Phase-0 spike (§6 verify #4); the structural fact (Claude as planner, SF-managed) holds.
- **Classification stays out of Agentforce.** The 15-min `fetch-classify` job keeps using **Bedrock Haiku 4.5** - **Classification stays out of Agentforce.** The 15-min `fetch-classify` job keeps using **Bedrock Haiku 4.5**
in our account (design.md §5, D8). It is proactive/event-driven, has no conversational surface, and shouldn't in our account (design.md §5, D8) — proactive/event-driven, no conversational surface, no Einstein-Request spend.
consume Einstein Requests. - **Per-agent model selection** is set in Setup → Agentforce Agents / `model_config` in Agent Script (`FACT`).
- **Per-agent model selection** is set in Setup → Agentforce Agents (`FACT`). All three agents use the same All three agents use the AWS-Hosted Claude reasoning model.
Claude reasoning model; Finance's lower latency tolerance is fine. - Licensing/capability flag: Agentforce + Data Cloud carry consumption/licensing cost (Gap G9). The earlier
- Licensing/capability flag: BYOLLM and Data Cloud both carry consumption/licensing cost (Gap G9, O5). "~30% fewer Einstein Requests" figure is **dropped — uncorroborated** (§0.1).
### 1e. Prompt Builder / Template Library structure ### 1e. Prompt Builder / Template Library structure
@ -232,7 +293,7 @@ content to **Data Cloud**, which **auto-creates a search index** (chunked + vect
| Corpus | → Data Library | → Retriever | Notes | | Corpus | → Data Library | → Retriever | Notes |
|--------|----------------|-------------|-------| |--------|----------------|-------------|-------|
| Notion How-To/Front SOPs + intake SOPs (unstructured) | **Seahaven Ops Knowledge** | `Seahaven_Ops_Knowledge_Retriever` | Highest-value, current. Source via `notion-sync` repointed to Data Cloud ingestion (S3 → Data Cloud, or Notion connector) | | Notion How-To/Front SOPs + intake SOPs (unstructured) | **Seahaven Ops Knowledge** | `Seahaven_Ops_Knowledge_Retriever` | Highest-value, current. **Ingestion = S3 → Data Cloud (N1, pinned):** `notion-sync` keeps writing `s3://seahaven-kb-docs-328440206208` and Data Cloud ingests from S3 — reuses the existing sync write path, avoids a new Notion-connector dependency, and lets the **same S3 dual-feed** the legacy Bedrock KB during parallel run (B1). |
| Amazon/AMOC subtree (unstructured, **stale**) | same library, separate index segment or tagged | same | Audience-restrict AMOC content; flag staleness (Gap G7) | | Amazon/AMOC subtree (unstructured, **stale**) | same library, separate index segment or tagged | same | Audience-restrict AMOC content; flag staleness (Gap G7) |
| SA8000 docs + employee handbook + company policies | **must be authored**, then ingested | same | **Not in Notion** (Gap G8) — stage in `s3://seahaven-kb-docs-328440206208` or Drive, then ingest | | SA8000 docs + employee handbook + company policies | **must be authored**, then ingested | same | **Not in Notion** (Gap G8) — stage in `s3://seahaven-kb-docs-328440206208` or Drive, then ingest |
| WorkOrders / purchase-orders / SiteAssignments / payments (**structured, live**) | **NOT a Data Library** | n/a | Stay **MCP lookup tools** over DDB (design.md §3); Data Libraries can't hold structured data (D7) | | WorkOrders / purchase-orders / SiteAssignments / payments (**structured, live**) | **NOT a Data Library** | n/a | Stay **MCP lookup tools** over DDB (design.md §3); Data Libraries can't hold structured data (D7) |
@ -267,12 +328,25 @@ handbook is one-deploy-target-per-repo. Coexistence:
- `sh-mcp` (existing): MCP servers, Cognito/auth, jobs — TypeScript/CDK, OIDC-into-AWS, `ci / ci` required check - `sh-mcp` (existing): MCP servers, Cognito/auth, jobs — TypeScript/CDK, OIDC-into-AWS, `ci / ci` required check
(design.md §7). Unchanged. (design.md §7). Unchanged.
- `sh-agentforce` (new): agent metadata. CI runs `sf` validate-deploy against a scratch org; CD does - `sh-agentforce` (new): agent metadata. CI runs `sf` validate-deploy against a scratch org; CD does
**deploy-then-merge** to sandbox → prod org. **Gap G10:** the org's reusable workflows are AWS/CDK-shaped; **deploy-then-merge** to sandbox → prod org.
we need a **new reusable `cd-sfdx` workflow** (Jira story). Naming kebab-case; Dependabot N/A (no npm), but pin - **`sh-agentforce` new-repo checklist (F3 — was missing; per `feedback_new_repo_checklist`):** private repo in
`@salesforce/cli` version. org; **branch protection + a required check** (the SFDX analog of `ci / ci`); **enable vulnerability-alerts +
`dependabot_security_updates` + secret_scanning** — Dependabot is **NOT N/A** (the `github-actions` ecosystem
applies even without npm; also watch `@salesforce/cli`); pin `@salesforce/cli`; kebab-case repo/workflow names.
- **`cd-sfdx` deploy auth (F3 — the prod-deploy key for the entire agent surface; design/verify story of its
own, NOT one Jira bullet).** Salesforce has **no OIDC**; CD authenticates via the **JWT-bearer flow of a
connected app** using a private key + cert. That key is a **long-lived secret** — store it in **AWS Secrets
Manager** (`sh-agentforce/sfdx-jwt-key`, per secrets-and-config.md; NOT a GitHub secret blob, NOT an env var),
pull it in the workflow, **scope separate connected apps for sandbox vs prod**, and define a **rotation cadence**.
`cd-sfdx` is a brand-new org-wide reusable workflow (Gap G10) implementing deploy-then-merge against
sandbox→prod orgs — it gets its **own design + verification step** before any agent metadata depends on it.
- **Cross-review gate extends to Agentforce metadata** that changes tool exposure, audience, or scope binding — - **Cross-review gate extends to Agentforce metadata** that changes tool exposure, audience, or scope binding —
those are security-relevant just like an IAM diff (design.md §8). `ASSUMPTION`: GenAiPlannerBundle/connection security-relevant just like an IAM diff (design.md §8). `ASSUMPTION`: GenAiPlannerBundle/connection changes and
changes go through `cross_reviewer` the same as IAM. the `cd-sfdx` connected-app trust go through `cross_reviewer` (GPT-4.1) the same as IAM.
- **Naming boundary (N3):** Salesforce Developer/API names (`Seahaven_Ops`, Flex template names) **cannot contain
hyphens** and conventionally use PascalCase/snake. The handbook kebab-case rule governs **AWS resources + repo
names** (the repo `sh-agentforce` IS kebab); it does NOT apply to Salesforce-side metadata API names. Recorded so
a future compliance pass doesn't "fix" a non-violation.
### 1h. Test suite — Agentforce Testing Center + the MCP-layer security tests ### 1h. Test suite — Agentforce Testing Center + the MCP-layer security tests
@ -300,21 +374,35 @@ real home:
8. **Parity golden-transcripts** — replay real Alex/Lauren interactions; assert equivalent answers **before** 8. **Parity golden-transcripts** — replay real Alex/Lauren interactions; assert equivalent answers **before**
deprecating each bot (design.md §6, §7.3 parity gate). deprecating each bot (design.md §6, §7.3 parity gate).
**MCP-layer security tests (authoritative, in `sh-mcp`, design.md §7.3 — unchanged):** **Core security tests (authoritative, transport-agnostic, in `sh-mcp`, design.md §7.3):**
audience-binding rejection (an `ops` token rejected by `finance`); per-tool **server-side** scope enforcement; **audience-binding rejection at the per-tier facade** (an `aud=sh-mcp-ops` token rejected by the `sh-mcp-finance`
`list_tools` tool-hiding reflects caller scopes; deny-list **hard revocation**; minimal-scope Google client gateway authorizer, and vice-versa — B3); per-tool **server-side** scope enforcement; **per-agent credential
(a `gmail:self` token can't mint a Calendar token); **per-user refresh-token ABAC isolation**; mints only its tier's scopes** (a token obtained via the Lauren Exec app client never contains `finance:read`,
**finance PII redaction** (bank/routing/card/SSN masked before egress) **while leaving vendor names/contacts even though the person holds it — B4); deny-list **hard revocation**; minimal-scope Google client (a `gmail:self`
UNMASKED** (legacy Alex deliberately left names unmasked — Trust Layer must not re-mask them, Gap G3); per-tool token can't mint a Calendar token); **per-user refresh-token ABAC isolation**; **finance PII redaction**
rate limit + per-session cap; finance audit-record shape. (bank/routing/card/SSN masked before egress) **while leaving vendor names/contacts UNMASKED** (legacy Alex
deliberately left names unmasked — Trust Layer must not re-mask them, Gap G3); per-tool rate limit + per-session
cap; finance audit-record shape.
**Transport-test note (D12):** the active interface is the **OpenAPI adapter**, so the design.md §7.3 **Transport-test note (D12, B3-reconciled):** the active interface is the **OpenAPI adapter**, so the design.md
"Contract / MCP conformance" layer is split — **OpenAPI contract tests** (schema validation, the generated spec §7.3 "Contract / MCP conformance" layer is split — **OpenAPI contract tests** (schema validation, the generated
matches the tool registry) run now; **MCP protocol-conformance tests defer with the MCP adapter**. The security spec matches the tool registry) run now; **MCP `list_tools` tool-hiding and MCP protocol-conformance tests defer
tests above are transport-agnostic and unchanged: server-side per-tool scope enforcement is the authoritative WITH the MCP adapter** (not in the launch suite — they were MCP-only). On the OpenAPI path the per-*user* tool
boundary on either transport, and `list_tools` tool-hiding (MCP-only) is replaced on the OpenAPI path by visibility that `list_tools` gave is replaced by per-*agent* action assignment (each agent granted only its
**per-agent action assignment** (each Agentforce agent is granted only its tier's actions) — still a convenience, tier's actions) — a *convenience, never the boundary*; the per-tool server-side scope check is the boundary
never the boundary (design.md §2.5). (design.md §2.5).
**Per-user visibility degradation (F7 — stated, not silently dropped):** per-agent action assignment is
per-*agent*, not per-*user*. A user **without** `ops:tasks` talking to Seahaven Ops still *sees* the task actions
offered and they fail server-side (403). Boundary holds; UX degrades. Handle with **graceful-refusal copy** in
the agent instructions ("you may not have access to that — server will confirm") and a behavioral test:
a no-`ops:tasks` user gets a clean "you don't have access" rather than a raw error.
**Testing-Center identity caveat (F1):** Testing Center batch/AI-generated runs invoke actions — **under whose
identity** when the action uses a Per-User Browser Flow credential? Headless batch may fail (no interactive
first-auth) or run as a different principal, which would make the behavioral suite unable to exercise the real
per-user tools (the suite authorizes teardown — design.md §6). Resolve in the Phase-0 spike (§6 verify #5);
registered as Gap **G12** (Beta dependency).
Eval criteria: behavioral suite ≥ agreed pass rate before each cutover; MCP/core suite at design.md coverage gate Eval criteria: behavioral suite ≥ agreed pass rate before each cutover; MCP/core suite at design.md coverage gate
(80% lines, 100% on the shared auth/scope guard) — both green are the parity gate for retiring a bot. (80% lines, 100% on the shared auth/scope guard) — both green are the parity gate for retiring a bot.
@ -343,13 +431,17 @@ Eval criteria: behavioral suite ≥ agreed pass rate before each cutover; MCP/co
- **Error Message (≤255):** "Sorry — I hit a problem reaching that information. Please try again in a moment; if - **Error Message (≤255):** "Sorry — I hit a problem reaching that information. Please try again in a moment; if
it keeps failing, post in #it-help and we'll take a look." it keeps failing, post in #it-help and we'll take a look."
- **Languages:** English (US). `ASSUMPTION`: no multilingual requirement today. - **Languages:** English (US). `ASSUMPTION`: no multilingual requirement today.
- **Variables:** `$User.Email`, `$User.GoogleGroups` (for scope context), `$Session.Channel` (public vs DM, for - **Variables:** `$User.Email`; `$User.GoogleGroups` (scope context); `$Session.Channel` (public vs DM, privacy
privacy gating). gating). `ASSUMPTION` (Q3, unverified-load-bearing): that Agentforce exposes Google-group membership and a
- **Connections:** `sh-mcp-ops` MCP server — trust tier **ops**, `aud=sh-mcp-ops`, scopes **`ops:read`** channel-visibility variable to agent instructions — **verify in Phase-0**; if absent, gate privacy/scope via a
(+ `ops:tasks` only when the caller's token carries it). Per-user OAuth 2.1 → Cognito (D4). different mechanism (e.g. a context Apex action). Registered as Gap **G15**.
- **Connections:** `sh-mcp-ops` **tier (OpenAPI facade, `aud=sh-mcp-ops`)** via **External Service actions** over
a **per-agent External Credential `ec-seahaven-ops` (Per-User OAuth Browser Flow → Cognito app client
`sh-agentforce-ops`, scopes `ops:read` [+`ops:tasks` when the caller's token carries it])** — D4/B4. The app
client requests **only ops-tier scopes**; no finance/gmail scope is reachable through this agent.
- **Data:** Data Library **Seahaven Ops Knowledge** via `Seahaven_Ops_Knowledge_Retriever` (Front SOPs, intake - **Data:** Data Library **Seahaven Ops Knowledge** via `Seahaven_Ops_Knowledge_Retriever` (Front SOPs, intake
SOPs, Amazon subtree [restricted], SA8000/handbook once authored). SOPs, Amazon subtree [restricted], SA8000/handbook once authored).
- **Model:** Claude on Bedrock (BYOLLM preferred; AWS-Hosted Claude Sonnet 4 fallback) — D5. - **Model:** AWS-Hosted Claude (Salesforce-managed) — D5.
- **Topics / Subagents:** - **Topics / Subagents:**
- **Work Orders & Sites** — *Description:* WO/PO/site-assignment lookups. *Reasoning:* identify the record id - **Work Orders & Sites** — *Description:* WO/PO/site-assignment lookups. *Reasoning:* identify the record id
or natural-language key; call the lookup tool; summarize with the WorkOrderSummary Flex template. *Actions:* or natural-language key; call the lookup tool; summarize with the WorkOrderSummary Flex template. *Actions:*
@ -383,12 +475,15 @@ Eval criteria: behavioral suite ≥ agreed pass rate before each cutover; MCP/co
- **Error Message (≤255):** "I couldn't complete that finance lookup. Please retry; if it persists, contact Adam - **Error Message (≤255):** "I couldn't complete that finance lookup. Please retry; if it persists, contact Adam
or accounting. (All lookups are audited.)" or accounting. (All lookups are audited.)"
- **Languages:** English (US). - **Languages:** English (US).
- **Variables:** `$User.Email`, `$Session.Channel` (decline sensitive detail in public channels). - **Variables:** `$User.Email`, `$Session.Channel` (decline sensitive detail in public channels — Q3/G15 caveat
- **Connections:** `sh-mcp-finance` MCP server — trust tier **finance**, `aud=sh-mcp-finance`, scope as in §2.1).
**`finance:read`**, **15-min token TTL + deny-list** hard revocation (design.md §2.3/§2.5). Per-user OAuth → - **Connections:** `sh-mcp-finance` **tier (separate OpenAPI facade, `aud=sh-mcp-finance`)** via External Service
Cognito (D4). **No ops/gmail connection on this agent.** actions over a **per-agent External Credential `ec-seahaven-finance` (Per-User OAuth Browser Flow → Cognito app
client `sh-agentforce-finance`, scope `finance:read` only)** — D4/B4. **15-min token TTL + deny-list** hard
revocation (design.md §2.3/§2.5). **No ops/gmail connection, and the app client cannot request ops/gmail
scopes** — the finance audience is reached only here.
- **Data:** none (structured lookups only; no Data Library grounding). - **Data:** none (structured lookups only; no Data Library grounding).
- **Model:** Claude on Bedrock (D5). - **Model:** AWS-Hosted Claude (Salesforce-managed) — D5.
- **Topics / Subagents:** - **Topics / Subagents:**
- **Vendor Search** — *Description:* QBO vendor lookup. *Reasoning:* query QBO; return vetted vendor records. - **Vendor Search** — *Description:* QBO vendor lookup. *Reasoning:* query QBO; return vetted vendor records.
*Actions:* `search_vendors` (MCP, `finance:read`, QBO server-held OAuth). *Actions:* `search_vendors` (MCP, `finance:read`, QBO server-held OAuth).
@ -416,15 +511,18 @@ Eval criteria: behavioral suite ≥ agreed pass rate before each cutover; MCP/co
- **Error Message (≤255):** "I couldn't complete that. Please try again in DM; if it keeps failing, let Adam - **Error Message (≤255):** "I couldn't complete that. Please try again in DM; if it keeps failing, let Adam
know. I only operate in direct messages." know. I only operate in direct messages."
- **Languages:** English (US). - **Languages:** English (US).
- **Variables:** `$User.Email` (the Google identity to act as), `$Session.Channel` (must be DM), per-user Google - **Variables:** `$User.Email` (the Google identity to act as), `$Session.Channel` (must be DM — **the DM-only
grant status. control depends on this variable existing; Q3/G15 — verify in Phase-0, else enforce DM-scoping another way**),
- **Connections:** `sh-mcp-ops` MCP server — trust tier **ops**, `aud=sh-mcp-ops`, scopes per-user Google grant status.
**`ops:read ops:tasks gmail:self calendar:self`**. Gmail/Calendar act as the user via a **separate per-user - **Connections:** `sh-mcp-ops` **tier (OpenAPI facade, `aud=sh-mcp-ops`)** via External Service actions over a
Google OAuth grant** (design.md §2.4), refresh tokens KMS-encrypted, **ABAC-partitioned by `sub`**. Per-user **per-agent External Credential `ec-lauren-exec` (Per-User OAuth Browser Flow → Cognito app client
Cognito OAuth (D4) is what carries the real `sub` so the server selects the right Google token — **directly `sh-agentforce-exec`, scopes `ops:read ops:tasks gmail:self calendar:self`)** — D4/B4. **The app client requests
dependent on Gap G1**. **No finance connection.** NO finance scope**, so even though Lauren-the-person holds `finance:read`, no token this agent obtains carries
it (§0.1 B4). Gmail/Calendar act as the user via a **separate per-user Google OAuth grant** (design.md §2.4),
refresh tokens KMS-encrypted, **ABAC-partitioned by `sub`**; the per-user Cognito token carries the real `sub`
so the server selects the right Google token — **dependent on the Phase-0 auth spike (G1, unverified)**.
- **Data:** none (operates on the user's live Gmail/Calendar, not a Data Library). - **Data:** none (operates on the user's live Gmail/Calendar, not a Data Library).
- **Model:** Claude on Bedrock (D5). - **Model:** AWS-Hosted Claude (Salesforce-managed) — D5.
- **Topics / Subagents:** - **Topics / Subagents:**
- **Inbox Triage** — *Description:* summaries, high-priority, unanswered threads, bypassed work orders, - **Inbox Triage** — *Description:* summaries, high-priority, unanswered threads, bypassed work orders,
search-by-sender. *Reasoning:* call read tools as the user; compose with the InboxDigest Flex template; search-by-sender. *Reasoning:* call read tools as the user; compose with the InboxDigest Flex template;
@ -447,27 +545,42 @@ Follows design.md §6 phasing; each legacy bot is deprecated **only at proven pa
| Phase | Work | Parity / exit gate | Rollback | | Phase | Work | Parity / exit gate | Rollback |
|-------|------|--------------------|----------| |-------|------|--------------------|----------|
| **0 — Foundation** | Stand up Cognito + Google federation + pre-token + 5-min group-sync (design.md §6.1). Create `sh-agentforce` SFDX repo + `cd-sfdx` reusable workflow (Gap G10). Stand up the Data Cloud org + **Seahaven Ops Knowledge** Data Library; repoint `notion-sync`. Decide D4 wrapper (Apex/External Service per-user named credential vs native MCP connector) after verifying G1/G2. | SSO login end-to-end; per-user JWT reaches a smoke-test MCP tool carrying the real `sub`; Data Library retriever returns Front SOP answers. | N/A (legacy still running) | | **0a — Auth spike (GATING, B5)** | **Before any other Phase-0 spend.** Prove the **Per-User OAuth Browser Flow External Service action in Slack** carries the real `sub` to a Cognito-fronted smoke-test endpoint (§6 verify #1, #3, #5). Confirm reasoning-model selector + AWS-Hosted model name (#4). | Real `sub` reaches the server; first-call consent works in Slack; token refresh survives; Testing-Center can invoke per-user actions. **If the spike FAILS → STOP and switch to the fallback (below); do not proceed to 0b.** | N/A (legacy untouched) |
| **1 — Seahaven Ops** | Wire Ops agent → `sh-mcp-ops`; vendors (KB/Maps) + WO/PO/site + knowledge + tasks. | Behavioral suite + golden-transcripts vs **Alex** green; per-user auth + tool-hiding verified at MCP layer. | Keep Alex running in parallel; flip Slack default back to Alex. | | **0b — Platform build (F4)** | Only after 0a passes. **Build the servers:** monorepo + `shared` auth/scope/PII/audit guard + per-service packages + **transport-agnostic core + OpenAPI adapter** + **per-tier API Gateway + Cognito authorizer + shared WAF** (B3); Cognito + Google federation + pre-token (per-audience scoping, B4) + 5-min group-sync (design.md §6.1); **deploy-role-first** (`githubdeploy-sh-mcp` already exists; add for `sh-agentforce`), CI/CD reusable workflows, coverage gate. Create `sh-agentforce` SFDX repo + **design & verify `cd-sfdx`** (G10, F3). Stand up Data Cloud + **Seahaven Ops Knowledge** Data Library; **`notion-sync` DUAL-FEEDS** Bedrock KB **and** Data Cloud (B1 — legacy KB stays fresh for rollback). | SSO end-to-end; per-user JWT reaches a smoke-test tool with real `sub`; **core + OpenAPI adapter green at the design.md §7.3 coverage gate** (80% lines, 100% shared auth guard); Data Library retriever returns Front SOP answers; legacy Bedrock KB still fed. | N/A (legacy untouched) |
| **2 — Seahaven Finance** | Wire Finance agent → `sh-mcp-finance`; full audit logging; 15-min TTL + deny-list. | Audit records emitted; PII-redaction tests green; audience-binding rejection verified. | Finance lookups revert to Alex's QBO action group temporarily. | | **1 — Seahaven Ops** | Wire Ops agent → `sh-mcp-ops` facade; vendors (KB/Maps) + WO/PO/site + knowledge + tasks. | Behavioral suite + golden-transcripts vs **Alex** green; per-user auth + per-tool scope verified at the server. | **Flip Slack default back to Alex** (Alex still running — NOT torn down until Phase 4). |
| **3 — Lauren Exec** | Wire Exec agent → `sh-mcp-ops` `*:self`; per-user Google grant; DM-scoped. Refactor `fetch-classify`/`daily-digest`/`reminder` to import shared packages (design.md §6.4). | Channel-aware-privacy + golden-transcripts vs **Lauren** green; ABAC Gmail isolation proven. | Keep exec-aide running; Lauren's workflow needs explicit sign-off before retiring exec-aide (design.md §6.6). | | **2 — Seahaven Finance** | Wire Finance agent → `sh-mcp-finance` facade; full audit logging; 15-min TTL + deny-list. | Audit records emitted; PII-redaction tests green; **per-tier audience-binding rejection verified** (B3). | **Flip back to Alex** (whose QBO action group is still live — Alex runs through Phase 4). |
| **4 — Teardown** | Only after each replacement is signed off at parity. | — | — | | **3 — Lauren Exec** | Wire Exec agent → `sh-mcp-ops` facade `*:self`; per-user Google grant; DM-scoped. Refactor `fetch-classify`/`daily-digest`/`reminder` to import shared packages (design.md §6.4). | Channel-aware-privacy + golden-transcripts vs **Lauren** green; ABAC Gmail isolation proven. | **Keep exec-aide running**; Lauren's workflow needs explicit sign-off before retiring exec-aide (design.md §6.6). |
| **4 — Teardown** | Only after ALL three replacements are signed off at parity, and after the shared-consumer audit below. | Consumer audit passes (no live reader of the KB/AOSS remains). | — |
**What gets torn down, and when (design.md §1, §5):** **Fallback if the Phase-0a auth spike fails (B5).** The plan does not discard design.md §12's alternatives without
- **Bedrock agents** `seahaven-alex` (`QVL5GEJN9B`) and the exec-aide Sonnet loop — after their respective a Plan B. If Per-User Browser Flow cannot carry the real `sub` to our server (or Slack consent / refresh /
agent's parity sign-off (Alex after Phase 1; Lauren after Phase 3). Testing-Center identity is unworkable), fall back to a **custom Bolt assistant** as the surface (design.md §12) —
- **Bedrock KB `LSDCNHTH6O`** + **OpenSearch Serverless `gv1540frh1crb79gtr4b`** (AOSS, INFRA-92) — after the it keeps the **MCP transport + per-user OAuth** design intact end-to-end and therefore **un-defers the MCP adapter
Data Cloud Data Library is proven in Phase 1 (both Bedrock retrieval Lambdas are non-VPC consumers of this (D12)**. This is a surface change, not an auth/scope/trust-tier redesign — the `sh-mcp` core, Cognito, scopes, and
collection; confirm no other consumer before delete). trust tiers are unchanged. Decision point escalates to Adam; it does NOT mean restarting.
- **Bedrock guardrail `seahaven-alex-guardrail`** — **not** deleted until the MCP-layer PII redaction + (D11)
BYOLLM-path guardrail are live and tested (design.md §5: the guardrail must be explicitly replaced before
deletion — no parity assumed).
- **Sync Lambdas:** `po-sync` + `workorder-sync` decommissioned at Phase 1 (D7); `notion-sync` repointed, not
deleted; `fetch-classify`/`daily-digest`/`reminder` rebuilt, old exec-aide versions decommissioned at Phase 3.
- Archive `seahaven-slack-bot` + `exec-aide` repos; decommission their CDK stacks (design.md §1).
**Rollback principle:** legacy and replacement run **in parallel** through each phase; the Slack default agent is **What gets torn down — ALL IN PHASE 4 (B1-corrected; nothing shared is deleted while a consumer or rollback path
the single flip point; no legacy component is deleted until the corresponding parity gate is signed off. still needs it):**
- **Pre-req: shared-consumer audit.** The AOSS collection `gv1540frh1crb79gtr4b` / KB `LSDCNHTH6O` is a shared
dependency of **BOTH** `seahaven-slack-bot` **and** `exec-aide` (INFRA-92). Before any KB/AOSS delete, confirm
**zero** live consumers remain (both legacy stacks retired; no other reader). Run the INFRA-92 consumer check.
- **Bedrock agents** `seahaven-alex` (`QVL5GEJN9B`) and the exec-aide Sonnet loop — torn down in **Phase 4**, after
their parity sign-off AND after they are no longer any phase's rollback target. (Alex is the Phase 1 *and* Phase
2 rollback target, so it must survive to Phase 4 — this was the B1 contradiction.)
- **Bedrock KB `LSDCNHTH6O`** + **AOSS `gv1540frh1crb79gtr4b`** — Phase 4, after the shared-consumer audit.
- **Bedrock guardrail `seahaven-alex-guardrail`** — **not** deleted until the **MCP/action-layer PII redaction +
prompt-attack replacement are live and tested** (D11, revised — there is no BYOLLM model path; design.md §5
requires the guardrail be explicitly replaced before deletion, no parity assumed).
- **Sync Lambdas:** `notion-sync` **dual-feeds through Phase 3**, then drops the Bedrock-KB feed at Phase 4 (keeps
the Data Cloud feed). `po-sync` + `workorder-sync` **keep feeding the legacy KB until Phase 4** (so a rollback to
Alex is never stale — corrects the original "decommission at Phase 1"); they are retired as KB feeds at teardown
(D7 — the NEW platform never used them; structured data is live MCP tools). `fetch-classify`/`daily-digest`/
`reminder` rebuilt in Phase 3; old exec-aide versions decommissioned at Phase 4.
- Archive `seahaven-slack-bot` + `exec-aide` repos; decommission their CDK stacks (design.md §1) — Phase 4.
**Rollback principle:** legacy and replacement run **in parallel** through Phases 1–3; the Slack default agent is
the single flip point; **no legacy component — especially the shared KB/AOSS and any rollback-target bot — is
deleted, repointed-away, or starved of its sync feed until Phase 4.**
--- ---
@ -484,24 +597,45 @@ the single flip point; no legacy component is deleted until the corresponding pa
per-user auth model + Gap G1, Data Library/retriever design, model choice, Testing Center plan); "Agentforce per-user auth model + Gap G1, Data Library/retriever design, model choice, Testing Center plan); "Agentforce
DX & Release Process" (the `sh-agentforce` repo, `cd-sfdx`, scratch-org flow). DX & Release Process" (the `sh-agentforce` repo, `cd-sfdx`, scratch-org flow).
**Repo docs:** **design.md amendments (F2 — these AMEND locked decisions; record them as amendments, not "annotations"):**
- **`sh-mcp/docs/design.md` §4 (annotation, D12):** design.md §4 frames each server as a remote *MCP* server Two locked design.md decisions are genuinely changed by this plan (Adam-approved) and must be written into
(assumed Slack MCP-client consumer). Annotate it to record the **transport-agnostic core** decision — OpenAPI design.md as dated amendments so the next builder isn't misled by stale locked text:
adapter shipped now for Agentforce, MCP adapter deferred to its named trigger — so the build doesn't couple - **design.md §1 item 4 / §11 / §12** ("Slack's built-in agent is the MCP client; per-user OAuth via Slack's MCP
tool logic to either protocol. (Annotation, not a reopening — the locked auth/scope/trust-tier decisions are client; endpoint reasoning based on Slack egress") is **now false** — the consumer is Agentforce via OpenAPI
unchanged.) per-user actions; the MCP transport is deferred (D12); WAF egress reasoning re-scopes to **Salesforce/Hyperforce**
egress, not Slack. Amend §4 (transport-agnostic core, OpenAPI-first) and §11 (endpoint/egress).
- **design.md §9/§12** ("Lauren gets `finance:read`" *as an agent mapping*) is **reversed** by **D3** — the person
keeps the scope, the Exec agent does not. Amend the §12 agent mapping.
The locked auth/scope/trust-tier *primitives* (Cognito federation, audience-binding, per-tool scope, PII at the
MCP layer) are unchanged — those stay locked.
**Repo READMEs & memory (F6 — global-instruction obligations, were omitted):**
- `sh-mcp` README updated to **OpenAPI-first** (servers expose OpenAPI; MCP adapter deferred).
- `sh-agentforce` README created (agent metadata repo, `cd-sfdx`, scratch-org flow).
- **Project memory:** create `project_sh_agentforce` before the repo-creating conversation ends; update
`project_sh_mcp` for the transport/auth changes; add **cross-project references** `sh-mcp` ↔ `sh-agentforce`
↔ `project_seahaven_slack_bot` ↔ `project_exec_aide` (each side links the other).
**Monitoring (N2 — ALARM-state-only convention, design.md alarm preferences):**
- Each per-tier **API Gateway facade**: 5xx-rate + auth-failure-spike alarms → site-alerts SNS (ALARM action
only, no OK/no-data).
- Repointed **`notion-sync`** carries the design.md ALARM-on-sync-failure through to the dual-feed (two-alarm
pattern: Errors≥1 missing=notBreaching; Invocations<1 over cadence missing=breaching).
**Jira (INFRA project — no migration epic exists yet; related done: INFRA-92 AOSS lockdown, INFRA-37 reminder **Jira (INFRA project — no migration epic exists yet; related done: INFRA-92 AOSS lockdown, INFRA-37 reminder
removal):** removal):**
- **New epic:** "Agentforce migration & MCP platform." - **New epic:** "Agentforce migration & MCP platform."
- Stories (representative): foundation/Cognito+Google federation; `sh-agentforce` repo + `cd-sfdx` reusable - Stories, in dependency order: (0a) **per-user auth spike — gates everything, do first** (Per-User Browser Flow
workflow (G10); **Phase-0 per-user auth spike** — prove the Per-User Browser Flow External Service action in ES action in Slack: real `sub`, consent, refresh, Testing-Center identity; §6 verify; fallback = custom Bolt);
Slack (real `sub` reaches the server, consent UX, token refresh; §6 verify gate) — **gates everything, do (0b) foundation — Cognito+Google federation + **per-audience pre-token scoping** (B4) + 5-min group-sync;
first**; **build the transport-agnostic core + OpenAPI adapter** (D12), MCP adapter deferred to its trigger; **build the transport-agnostic core + OpenAPI adapter + per-tier API GW/WAF** (D12/B3/F4); `sh-agentforce` repo
Data Cloud Data Library + repoint `notion-sync`; retire `po-sync`/`workorder-sync` (D7); Ops agent build + (full new-repo checklist) + **design & verify `cd-sfdx` + connected-app JWT key in Secrets Manager** (G10/F3);
parity; Finance agent + audit; Lauren Exec + per-user Google grant; **author SA8000 + employee-handbook Data Cloud Data Library + **`notion-sync` dual-feed** (B1); (1) Ops agent + parity; (2) Finance agent + audit;
corpus** (G8, fund when needed); Testing Center suites; teardown (Bedrock agents/KB/AOSS/guardrail/sync (3) Lauren Exec + per-user Google grant; **author SA8000 + employee-handbook corpus** (G8, fund when needed);
Lambdas). Every IAM/auth/connection story carries the **mandatory cross-review** label (design.md §8/§10). Testing Center suites; (4) **teardown** after parity sign-off + shared-consumer audit — Bedrock agents/KB/AOSS/
guardrail and `po-sync`/`workorder-sync`/`notion-sync`-Bedrock-feed (B1). Every IAM/auth/connection story
(incl. the per-audience pre-token Lambda and the `cd-sfdx` connected-app trust) carries the **mandatory
cross-review** label (design.md §8/§10).
--- ---
@ -511,7 +645,7 @@ Severity: **C**ritical / **M**edium / **L**ow. "Verify" = check against current
| # | Gap | Sev | Impact | Workaround / status | Verify | | # | Gap | Sev | Impact | Workaround / status | Verify |
|---|-----|-----|--------|---------------------|--------| |---|-----|-----|--------|---------------------|--------|
| **G1** | **RESOLVED 2026-06-11 (research wf_1bf9e142).** Per-user OAuth Browser Flow on the **native custom remote-MCP connector is unconfirmed**; documented per-user identity exists only for SF-**hosted** MCP. | **C→ mitigated** | If we'd relied on the MCP connector and it bound only a Named Principal: per-user `sub`, ABAC Gmail isolation, and trifecta all break. | **Decided (D4):** do NOT use the native MCP connector for the per-user hop. Use **GA External Service (OpenAPI)/Apex actions + Per-User OAuth Browser Flow External Credential → Cognito** (per-user OAuth is GA, ships in `UserExternalCredential`). | In-org: build one ES action on a Per-User Browser Flow cred → confirm first-call consent, real `sub` reaches the server, and token auto-refresh (§6 verify list) | | **G1** | **MITIGATION CHOSEN — UNVERIFIED, SPIKE-GATED (B5-relabel).** Per-user OAuth Browser Flow on the native custom remote-MCP connector is unconfirmed (per-user identity is documented only for SF-**hosted** MCP). The chosen GA path (D4) is *plausible* but **not yet proven** to carry our Cognito `sub` end-to-end in Slack. | **C (open until Phase-0a)** | If the mitigation also fails to propagate the real `sub`: per-user identity, ABAC Gmail isolation, and the trifecta all break — and there is no Agentforce-native surface. | **Decided (D4):** ES (OpenAPI)/Apex actions + Per-User OAuth Browser Flow External Credential → Cognito. **Gated by the Phase-0a spike; if it fails → custom-Bolt fallback (§3).** Status is *mitigation chosen, not verified* — do not call this resolved until 0a passes. | Phase-0a: real `sub` reaches the server, first-call consent in Slack, refresh, Testing-Center identity (§6 verify #1/#3/#5) |
| **G2** | **RESOLVED 2026-06-11.** Custom remote MCP client is **Beta** (Pilot Jul 2025 → Beta Jan 2026); SF-*hosted* MCP is GA (Apr 2026) but that's not our remote servers. | **C→ avoided** | Schedule/stability risk on the native-MCP path. | **Avoided by D4** — the GA ES/Apex action path is the per-user transport; the Beta MCP connector is off the critical path. Revisit native MCP once remote-MCP per-user binding reaches GA. | SF release notes for remote-MCP GA date | | **G2** | **RESOLVED 2026-06-11.** Custom remote MCP client is **Beta** (Pilot Jul 2025 → Beta Jan 2026); SF-*hosted* MCP is GA (Apr 2026) but that's not our remote servers. | **C→ avoided** | Schedule/stability risk on the native-MCP path. | **Avoided by D4** — the GA ES/Apex action path is the per-user transport; the Beta MCP connector is off the critical path. Revisit native MCP once remote-MCP per-user binding reaches GA. | SF release notes for remote-MCP GA date |
| **G3** | **Trust Layer masking may over-mask.** Agentforce Trust Layer can mask PII; legacy Alex **deliberately left vendor names/contacts unmasked** (lookup is the bot's job). | M | Vendor lookups could be degraded if Trust Layer masks names. | Configure Trust Layer masking to exclude names/contacts; keep MCP-layer redaction authoritative for bank/routing/card/SSN (design.md §2.5). | Trust Layer data-masking config | | **G3** | **Trust Layer masking may over-mask.** Agentforce Trust Layer can mask PII; legacy Alex **deliberately left vendor names/contacts unmasked** (lookup is the bot's job). | M | Vendor lookups could be degraded if Trust Layer masks names. | Configure Trust Layer masking to exclude names/contacts; keep MCP-layer redaction authoritative for bank/routing/card/SSN (design.md §2.5). | Trust Layer data-masking config |
| **G4** | **Loss of free-text search over WO comments** (old KB feature) — Data Libraries are unstructured-only, WO/PO are structured MCP lookups. | L | Can't fuzzy-search WO comment text. | Lookup-by-id via MCP tools; or a dedicated retriever over a comment text export if demand appears. | — | | **G4** | **Loss of free-text search over WO comments** (old KB feature) — Data Libraries are unstructured-only, WO/PO are structured MCP lookups. | L | Can't fuzzy-search WO comment text. | Lookup-by-id via MCP tools; or a dedicated retriever over a comment text export if demand appears. | — |
@ -519,9 +653,13 @@ Severity: **C**ritical / **M**edium / **L**ow. "Verify" = check against current
| **G6** | **RESOLVED 2026-06-11 (research wf_1bf9e142): BYOLLM cannot drive the planner.** Only Salesforce Default / AWS-Hosted options drive the Atlas reasoning engine; BYOLLM is custom-action-only and still routes through SF's Models API/Trust Layer. | M→ resolved | Our-account inference + our guardrail are **unsatisfiable on the planner path**. | **Decided (D5):** use **AWS-Hosted Claude** as the planner; move our guardrail + PII redaction to the **MCP/action layer** (revises D11); rely on the Trust Layer for planner-path moderation. Note AWS-Hosted model is drifting Sonnet 4 → 4.6/Haiku 4.5 (May 2026). | In-org: confirm the reasoning-model selector offers only Default/AWS-Hosted; confirm current AWS-Hosted model name | | **G6** | **RESOLVED 2026-06-11 (research wf_1bf9e142): BYOLLM cannot drive the planner.** Only Salesforce Default / AWS-Hosted options drive the Atlas reasoning engine; BYOLLM is custom-action-only and still routes through SF's Models API/Trust Layer. | M→ resolved | Our-account inference + our guardrail are **unsatisfiable on the planner path**. | **Decided (D5):** use **AWS-Hosted Claude** as the planner; move our guardrail + PII redaction to the **MCP/action layer** (revises D11); rely on the Trust Layer for planner-path moderation. Note AWS-Hosted model is drifting Sonnet 4 → 4.6/Haiku 4.5 (May 2026). | In-org: confirm the reasoning-model selector offers only Default/AWS-Hosted; confirm current AWS-Hosted model name |
| **G7** | **Amazon/AMOC corpus is stale** ("migrated from BookStack, may need updating") and access-restricted (AMOC comms = Adam & Robert only). | M | Agent could give outdated Amazon-site guidance. | Audience-restrict the AMOC topic; flag content as stale; re-author before exposing widely. | Notion *AMOC* `33a2ecdd…8162` currency | | **G7** | **Amazon/AMOC corpus is stale** ("migrated from BookStack, may need updating") and access-restricted (AMOC comms = Adam & Robert only). | M | Agent could give outdated Amazon-site guidance. | Audience-restrict the AMOC topic; flag content as stale; re-author before exposing widely. | Notion *AMOC* `33a2ecdd…8162` currency |
| **G8** | **SA8000 docs + employee handbook + company policies missing** from Notion (referenced as KB inputs, not found). | M | SA8000/compliance + HR Q&A **blocked**. | **Blocked pending content authoring** — author, stage in S3/Drive, ingest into Ops Knowledge Data Library; until then the agent must decline (without suppressing legitimate misconduct questions). | design.md corpus notes; locate any S3/Drive originals | | **G8** | **SA8000 docs + employee handbook + company policies missing** from Notion (referenced as KB inputs, not found). | M | SA8000/compliance + HR Q&A **blocked**. | **Blocked pending content authoring** — author, stage in S3/Drive, ingest into Ops Knowledge Data Library; until then the agent must decline (without suppressing legitimate misconduct questions). | design.md corpus notes; locate any S3/Drive originals |
| **G9** | **Agentforce + Data Cloud licensing / Einstein-Request consumption.** | M | Cost; BYOLLM cuts ~30% of Einstein Requests but Data Cloud + Agentforce licensing still applies. | Budget; BYOLLM to reduce request spend. | Salesforce contract / Einstein Request limits | | **G9** | **Agentforce + Data Cloud licensing / Einstein-Request consumption.** | M | Cost. (The "~30% fewer Einstein Requests via BYOLLM" claim is **dropped — uncorroborated**, §0.1; and BYOLLM is ruled out for the planner anyway, G6.) | Budget Agentforce + Data Cloud licensing + Einstein-Request limits explicitly; see G15 (per-employee licensing). | Salesforce contract / Einstein Request limits |
| **G10** | **No SFDX reusable CI/CD workflow** — org reusable workflows are AWS/CDK/OIDC-shaped (design.md §7.1). | M | `sh-agentforce` can't deploy via the existing pattern. | Author a new reusable **`cd-sfdx`** workflow (deploy-then-merge, scratch-org validate). | handbook `cicd.md` | | **G10** | **No SFDX reusable CI/CD workflow + no OIDC for Salesforce deploy** — org reusable workflows are AWS/CDK/OIDC-shaped (design.md §7.1). | M | `sh-agentforce` can't deploy via the existing pattern; the JWT-bearer connected-app key is a long-lived prod-deploy secret. | Author reusable **`cd-sfdx`** (deploy-then-merge, scratch-org validate); store the connected-app **JWT key in Secrets Manager** (`sh-agentforce/sfdx-jwt-key`), sandbox/prod-scoped, rotated (F3, §1g). Its own design/verify story. | handbook `cicd.md`; SFDX JWT-bearer flow |
| **G11** | **Two access-control planes.** Agentforce-in-Slack assigns member access via **Salesforce permissions**; our authz is **Google Groups → Cognito → scopes**. | M | Drift: a user could see an agent but lack the MCP scope, or vice-versa. | Provision Agentforce/Salesforce user access **from the same Google Groups** (SCIM/identity sync) so Google Groups stays the single source of truth (design.md §2.3). | Salesforce SCIM/Google provisioning | | **G11** | **Two access-control planes.** Agentforce-in-Slack assigns member access via **Salesforce permissions**; our authz is **Google Groups → Cognito → scopes**. | M | Drift: a user could see an agent but lack the MCP scope, or vice-versa. | Provision Agentforce/Salesforce user access **from the same Google Groups** (SCIM/identity sync) so Google Groups stays the single source of truth (design.md §2.3). | Salesforce SCIM/Google provisioning |
| **G12** | **Parity gate runs on Beta tooling + per-user identity (F1).** Testing Center Test Suites is Beta; batch/AI-generated runs invoke actions under an unclear identity when the action uses a Per-User Browser Flow credential. | M | If batch runs can't invoke per-user actions, the behavioral suite (which authorizes teardown) can't exercise the real tools. | Verify Testing-Center identity under per-user creds in Phase-0a (§6 #5); if unworkable, drive parity via scripted per-user sessions instead of batch. | help.salesforce.com Testing Center; in-org |
| **G13** | **New data processor: Salesforce now processes our tool payloads (F5).** Finance/Gmail tool responses transit the Agentforce planner / Einstein Trust Layer; grounding flows through Data Cloud. design.md's data boundary was AWS + Slack. | M | Vendor/payment/inbox content (PII is masked, but business content is not) is processed by Salesforce — a new processor not in the original threat model. | Explicit acceptance; verify Trust Layer **retention + zero-training** guarantees and Data Cloud data-residency; keep MCP-layer PII redaction authoritative. | Einstein Trust Layer retention/zero-training docs; DPA |
| **G14** | **Per-user tool visibility degraded (F7).** MCP `list_tools` hid tools per-*user* by scope; per-agent action assignment is per-*agent*, so a user lacking a scope still sees the action and fails server-side (403). | L | UX degradation (confusing failures), not a security hole — the per-tool scope check is the boundary. | Graceful-refusal copy in agent instructions; behavioral test for a clean "no access" message (§1h). | — |
| **G15** | **Three unverified-load-bearing platform assumptions (Q1/Q2/Q3).** (a) Sea Haven's Slack plan actually supports **Agentforce Employee Agents**; (b) per-employee Agentforce **user+license** required for the all-staff Ops agent (drives G9 cost + G11 SCIM); (c) Agentforce exposes **`$User.GoogleGroups`** and **`$Session.Channel`** to agent instructions (Lauren Exec's DM-only + scope gating depend on these). | M | If (a) wrong, the surface decision (D2) collapses; if (c) wrong, privacy/DM controls need a different enforcement point (e.g. a context Apex action). | Verify all three in Phase-0 (§6); for (c), design a fallback enforcement now so it's not on the critical path. | slack.com/help Agentforce-in-Slack plan reqs; Salesforce licensing; Agentforce agent variables doc |
`FACT` for G1/G2: Agentforce MCP support (Pilot Jul 2025, Beta Jan 2026; SF-hosted MCP GA Apr 2026; OAuth 2.0, `FACT` for G1/G2: Agentforce MCP support (Pilot Jul 2025, Beta Jan 2026; SF-hosted MCP GA Apr 2026; OAuth 2.0,
JSON-RPC over Streamable HTTP) and Named/External Credential **Per User** identity type — verified via Salesforce JSON-RPC over Streamable HTTP) and Named/External Credential **Per User** identity type — verified via Salesforce
@ -535,8 +673,11 @@ sources (§Sources). The **specific** per-user binding for MCP connectors is the
> licensing **accepted**; (2) reasoning model = **AWS-Hosted Claude** (BYOLLM ruled out for the planner); > licensing **accepted**; (2) reasoning model = **AWS-Hosted Claude** (BYOLLM ruled out for the planner);
> (3) **D3 accepted** (finance stripped from Lauren Exec); (4) per-user auth = **ES/Apex actions + Per-User > (3) **D3 accepted** (finance stripped from Lauren Exec); (4) per-user auth = **ES/Apex actions + Per-User
> Browser Flow → Cognito** (native MCP connector off the per-user path); (5) corpus authoring (G8) **deferred — > Browser Flow → Cognito** (native MCP connector off the per-user path); (5) corpus authoring (G8) **deferred —
> fund when needed**. The hands-on verify-in-org tests below remain the build-time gate for #2 and #4. Original > fund when needed**. The hands-on verify-in-org tests below remain the build-time gate. **The historical
> recommendations retained for the record: > recommendations 1–5 below predate the research (wf_1bf9e142) and retain superseded wording** (e.g. the
> "~30% fewer Einstein Requests" figure in #2, since **dropped as uncorroborated** — §0.1, G9; and "BYOLLM if it
> can be the planner," since **ruled out** — G6). Read §0.1 for the corrected, committed position; the text below
> is kept only as a decision trail.
1. **Surface + licensing.** Confirm **Agentforce Employee Agents in Slack** as the surface and accept Agentforce 1. **Surface + licensing.** Confirm **Agentforce Employee Agents in Slack** as the surface and accept Agentforce
+ Data Cloud licensing/Einstein-Request cost (G9). *Recommend: yes* — Slack is already our hub and only + Data Cloud licensing/Einstein-Request cost (G9). *Recommend: yes* — Slack is already our hub and only
@ -572,6 +713,15 @@ Groups to reconcile the two control planes (G11, recommend SCIM).
4. **(Model)** Confirm the reasoning-engine selector (Setup → Agentforce Agents; `model_config` in Agent Script) 4. **(Model)** Confirm the reasoning-engine selector (Setup → Agentforce Agents; `model_config` in Agent Script)
offers only **Salesforce Default** / **AWS-Hosted**, and confirm the current model name behind AWS-Hosted offers only **Salesforce Default** / **AWS-Hosted**, and confirm the current model name behind AWS-Hosted
(Sonnet 4 vs 4.6 vs Haiku 4.5). (Sonnet 4 vs 4.6 vs Haiku 4.5).
5. **(Testing Center identity, F1/G12)** Confirm a Testing Center batch/AI-generated run can invoke a Per-User
Browser Flow action and under whose identity. If it can't exercise per-user tools, the behavioral parity gate
must run as scripted per-user sessions instead of batch — decide before Phase 1.
6. **(Slack plan + licensing, Q1/Q2/G15)** Confirm Sea Haven's Slack plan supports **Agentforce Employee Agents**,
and whether the all-staff Ops agent needs a provisioned Agentforce **user + license per employee** (prices G9,
feeds G11 SCIM scope).
7. **(Agent variables, Q3/G15)** Confirm Agentforce exposes **`$User.GoogleGroups`** and **`$Session.Channel`** to
agent instructions. If not, design an alternate enforcement for Lauren Exec's DM-only and per-scope gating
(e.g. a context Apex action that returns channel + group facts) before Phase 1/3.
--- ---