diff --git a/docs/agentforce-plan.md b/docs/agentforce-plan.md index 4c54285..3fa0b38 100644 --- a/docs/agentforce-plan.md +++ b/docs/agentforce-plan.md @@ -56,9 +56,10 @@ off the per-user hop, so the risk is mitigated rather than load-bearing. from Lauren Exec); (4) **D4** per-user auth = ES/Apex actions + Per-User Browser Flow → Cognito; (5) corpus authoring **deferred** (fund when needed). -**Capability-gap count: 15 registered** (§5). Severity: **1 critical** (G1 per-user auth — mitigation chosen, -spike-gated, **unverified**), **9 medium** (G3, G7–G13, G15), **3 low** (G4, G5, G14), **2 resolved/avoided** -(G2, G6). (The dropped BYOLLM ~30% cost claim is not a registered gap.) +**Capability-gap count: 16 registered** (§5). Severity: **2 critical** — **G16 Employee-Agent callout identity +(the new top risk, research Finding 4)** + G1 per-user auth (mitigation chosen, spike-gated); **8 medium** +(G3, G7–G11, G13, G15), **3 low** (G4, G5, G14), **3 resolved** (G2, G6, G12). Phase-0a research (§0.4) resolved +the Cognito mechanics and shifted the controlling risk to G16. (The dropped BYOLLM ~30% cost claim is not a registered gap.) > **REVISION — plan-review remediation (2026-06-11, §0.3).** This plan was audited by the `sh-plan-review` Fable > gate; §0.3 logs every BLOCK/FIX addressed. Key structural changes since the first commit: teardown moved wholly @@ -138,34 +139,39 @@ Agentforce per-user path (D4) is OpenAPI/Apex, not MCP. Decision: claim "separate audiences → no single token spans tiers" is only true if each agent's tokens are minted from a credential that requests **only its tier's scopes**. Mechanism (specified, testable): -- **PRIMARY mechanism — per-app-client scope restriction (vendor-supported, airtight regardless of group union, - BLOCK-1-resolved).** Each agent gets **one Cognito app client + one External Credential**, and the app client's - **`AllowedOAuthScopes` is set to only its tier's resource-server scopes**: Ops client → `ops:read [ops:tasks]`; - Finance client → `finance:read`; Lauren Exec client → `ops:read ops:tasks gmail:self calendar:self` (**never - finance**). Cognito's token endpoint **cannot issue a scope outside a client's `AllowedOAuthScopes`** — so even - though Lauren-the-person holds `finance:read` (via `-finance@`), the Exec client can never obtain a token - carrying it. This is the intended boundary — independent of design.md §2.3's group→scope union. - **CAVEAT (Gemini round-3 BLOCK — UNVERIFIED, fatal Phase-0a check):** this rests on Cognito **strictly - filtering the issued token's scopes to the client's `AllowedOAuthScopes` even when a pre-token V2 Lambda - runs**. In some Cognito configs a V2 pre-token Lambda can *return scopes exceeding* the client's allow-list — if - so, the app-client control is soft and the pre-token Lambda becomes load-bearing after all. Verify "issued - scopes ⊆ `AllowedOAuthScopes` regardless of Lambda output" as a **fatal** 0a gate item (§6 #8); if it does NOT - hold, the pre-token fail-closed suppression below is promoted from backstop to primary and gets the heavier test. +- **PRIMARY mechanism — per-app-client scopes + a SUPPRESS-ONLY pre-token Lambda (RESOLVED by research + wf_3e88d8a8, Finding 1).** Each agent gets one Cognito app client + one External Credential whose + **`AllowedOAuthScopes` = only its tier's scopes** (Ops → `ops:read [ops:tasks]`; Finance → `finance:read`; + Lauren Exec → `ops:read ops:tasks gmail:self calendar:self`, **never finance**). The research **CONFIRMED the + Gemini caveat**: a pre-token V2 Lambda's `scopesToAdd` is **NOT** bounded by `AllowedOAuthScopes` (only a + no-blank-space rule applies — AWS docs). Resolution: **our pre-token Lambda uses `scopesToSuppress` ONLY — + never `scopesToAdd` for a tier scope.** Because the OAuth-flow base token is already bounded by + `AllowedOAuthScopes` (a client cannot *request* a scope it lacks — auth fails) and a suppress-only Lambda can + only *remove*, **issued scopes ⊆ `AllowedOAuthScopes` always holds.** `AllowedOAuthScopes` is the genuine + per-tier ceiling; the Lambda enforces *group entitlement* by suppressing scopes the user's groups don't grant + (e.g. strip `ops:tasks` for a non-`-assistant@` user). Lauren Exec's client never lists finance, so no Exec + token can carry it regardless of the Lambda. +- **HARD RULE (cross-review-gated, Finding 1): the pre-token Lambda MUST be suppress-only.** If anyone ever adds + `scopesToAdd` for a tier scope, `AllowedOAuthScopes` stops being a ceiling and the trifecta boundary goes soft. + Enforced by (a) a unit test asserting the Lambda emits **no `scopesToAdd`**; (b) the Lambda also reads + `event.callerContext.clientId` and refuses to emit any scope outside that client's tier (belt-and-suspenders); + (c) the pool runs on the **Essentials** feature plan (default; **Lite silently ignores V2** — Finding 2). Its + IAM/integrity is load-bearing → mandatory cross-review. **0a fatal check:** empirically confirm issued scopes ⊆ + `AllowedOAuthScopes` with the suppress-only Lambda live. - **AMENDS design.md §2.3 (BLOCK-1 / F2).** §2.3 describes a plain **union** of a user's group-scopes (and even shows Lauren's token *with* `finance:read`); the substrate is ambiguous on per-audience subsetting. This plan **amends** §2.3: group→scope mapping still defines what a user *may* hold, but the **issued token is bounded by the requesting app client's allowed scopes** (per-tier). Added to the §4 amendment list; not attributed to locked text. -- **BACKSTOP — pre-token suppression, fail-closed.** The pre-token Lambda additionally suppresses any scope not - belonging to the requested resource server and **fails closed + alarms** if a cross-audience scope ever appears - (CR-3) — defense in depth behind the app-client control, not the primary boundary. **Per-server IAM/authz - behavior → mandatory cross-review (GPT-4.1) before commit** (design.md §8/§10). -- **UNVERIFIED-LOAD-BEARING — the `aud`/authorizer mechanism (BLOCK-1, spike-gated like G1).** Cognito access - tokens natively carry `client_id` + resource-server-prefixed scopes, **not a per-resource-server `aud` claim by - default**, and customizing the *access* token (pre-token **V2** trigger) may require a paid Cognito feature - plan. So "the per-tier authorizer validates `aud=sh-mcp-ops`" needs a **concrete, chosen mechanism** — scope- - prefix check at the authorizer, `client_id` allow-list per gateway, or a custom `aud` injected via pre-token V2 - — and confirmation the trigger version is available. **Verify in Phase-0 (new §6 #8); on the 0b exit gate.** +- **`aud`/authorizer mechanism (RESOLVED by research wf_3e88d8a8, Finding 3).** Cognito access tokens carry + `client_id` + resource-server-prefixed scopes and **no `aud` by default** (an `aud` appears only with + managed-login "resource binding," one resource per request). Chosen mechanism: the per-tier API Gateway **JWT + authorizer validates `client_id` against a per-gateway allow-list** (the documented fallback — API GW checks + `client_id` when `aud` is absent) **AND** the server checks the **resource-server-prefixed scope** + (`{resourceServerId}/{scope}`) — together these are the audience proxy. (Optional, since our flow IS + managed-login: request a `resource` binding to also get a real `aud` — test in 0a.) The `client_id` allow-list + is **injected as config (SSM/CDK env), not hardcoded** (decoupling, §6 #8d). Server still re-validates issuer + + client_id + scope on every call (CR-1, design.md §2.5). - **Backend `sub` validation (CR-2).** Do not assume Salesforce's External Credential forwards the Google `sub` unchanged — verify in Phase-0a that the JWT our endpoint receives carries the **Google Workspace `sub`** (not a Salesforce/Cognito-internal id), and have each server **reject any token whose `sub` is not a valid Workspace @@ -272,6 +278,24 @@ Re-run `cross_reviewer` on the concrete pre-token Lambda + IAM/Cognito resource- --- +## 0.4 Phase-0a research findings — 2026-06-11 (workflow wf_3e88d8a8, Sonnet ×14, adversarially verified) + +The doc-answerable 0a unknowns were resolved by web research *before* the live spike. Impacts: + +| # | Finding | Confidence | Plan impact | +|---|---------|-----------|-------------| +| 1 | **Cognito `AllowedOAuthScopes` does NOT cap a pre-token V2 Lambda's `scopesToAdd`** (only a no-blank-space rule). | med-high | **Design changed:** the pre-token Lambda is now **suppress-only** (never `scopesToAdd` for a tier scope), which keeps `AllowedOAuthScopes` as the genuine per-tier ceiling. Hard rule + test added (§0.1). | +| 2 | Pre-token **V2 requires the Essentials/Plus feature plan**; new pools default to Essentials; **Lite silently ignores V2**. ~2.7× MAU cost vs Lite (negligible at our scale). | high | Provision the pool on Essentials; confirm not Lite. Cost negligible. (§0.1 hard rule.) | +| 3 | Cognito access tokens carry `client_id` + scopes, **no `aud`** by default (only via managed-login resource binding). API GW authorizer checks `client_id` when `aud` absent. | high | **Mechanism chosen:** per-gateway `client_id` allow-list + resource-server scope-prefix as the audience proxy; optional resource-binding for a real `aud` (test in 0a). (§0.1.) | +| 4 | **Employee-Agent callouts may carry SERVICE identity, not the user's** — "user identity is lost by default" for agent→external client-credentials callouts (primary SF Architects doc). No primary example of per-user Browser Flow + Cognito + `sub` end-to-end. `offline_access` needed or tokens die at 1h. | med (doc) | **NEW CRITICAL gap G16 — now THE top risk** (above the Cognito mechanism). 0a fatal #1 = prove the Slack Employee-Agent callout carries the user's token. Fallbacks: MuleSoft Trusted Agent Identity (RFC 8693), signed user-header, or custom-Bolt. | +| 5 | **Per-user actions are UNTESTABLE via batch Testing Center** — batch/Testing-API run as the client-credentials "Run As" service account (no `UserExternalCredential`). | high | G12 resolved: per-user parity runs as **scripted interactive sessions** (Agent Builder preview / Agent API `bypassUser=false`), not batch. | +| 6 | All-staff agent is **not free** — Grid not required, Identity license waives the *seat* only; **Employee agents consume Flex Credits (~$0.10/action) or the $125/user/mo add-on**. | med | G9 updated with the real cost model; verify the exact unmetered PSL name in-org. | +| 7 | Org-wide planner = 2 options (Default GPT-4o / AWS-Hosted Claude Sonnet 4 → routed to 4.6). **BUT Agent Script `model_config` (Summer '26) can override per-agent to ANY supported model** (50+, incl. Gemini 3.1 Pro, Claude Haiku 4.5); whether a **BYOLLM endpoint** can be a per-agent planner is **unresolved/possibly favorable**. | med | D5 (org-level AWS-Hosted Claude) **stands**, but **reopen as a verify**: can `model_config` place our BYO Bedrock Claude as a per-agent planner? Added to §6. Does NOT block; potential upside. | + +**Net:** the auth design is now research-grounded (suppress-only Lambda; client_id/scope-prefix audience; Essentials plan), and the **single biggest risk shifted** from "does Cognito bound scopes" to **"does an Employee-Agent-in-Slack callout even carry the user's identity" (G16)** — which only the live 0a spike can settle, with named fallbacks if it doesn't. + +--- + ## 1. Platform-level design ### 1a. Agent roster — how many, and why @@ -788,10 +812,11 @@ Severity: **C**ritical / **M**edium / **L**ow. "Verify" = check against current | **G6** | **RESOLVED 2026-06-11 (research wf_1bf9e142): BYOLLM cannot drive the planner.** Only Salesforce Default / AWS-Hosted options drive the Atlas reasoning engine; BYOLLM is custom-action-only and still routes through SF's Models API/Trust Layer. | M→ resolved | Our-account inference + our guardrail are **unsatisfiable on the planner path**. | **Decided (D5):** use **AWS-Hosted Claude** as the planner; move our guardrail + PII redaction to the **MCP/action layer** (revises D11); rely on the Trust Layer for planner-path moderation. Note AWS-Hosted model is drifting Sonnet 4 → 4.6/Haiku 4.5 (May 2026). | In-org: confirm the reasoning-model selector offers only Default/AWS-Hosted; confirm current AWS-Hosted model name | | **G7** | **Amazon/AMOC corpus is stale** ("migrated from BookStack, may need updating") and access-restricted (AMOC comms = Adam & Robert only). | M | Agent could give outdated Amazon-site guidance. | Audience-restrict the AMOC topic; flag content as stale; re-author before exposing widely. | Notion *AMOC* `33a2ecdd…8162` currency | | **G8** | **SA8000 docs + employee handbook + company policies missing** from Notion (referenced as KB inputs, not found). | M | SA8000/compliance + HR Q&A **blocked**. | **Blocked pending content authoring** — author, stage in S3/Drive, ingest into Ops Knowledge Data Library; until then the agent must decline (without suppressing legitimate misconduct questions). | design.md corpus notes; locate any S3/Drive originals | -| **G9** | **Agentforce + Data Cloud licensing / Einstein-Request consumption + parallel-run double-spend (FIX-4).** | M | Cost. The ~30% BYOLLM claim is **dropped** (§0.1, G6). Plus: Phases 1–3 run legacy (2 Bedrock agents + KB + AOSS standing cost + 3 sync feeds) AND Agentforce/Data Cloud **simultaneously**. | Budget Agentforce + Data Cloud + Einstein-Request limits (G15 per-employee licensing) **and a dual-run AWS line**; **time-box Phases 1–3 (~6–8 wks)** so parallel-run doesn't drift open-ended (§3). | Salesforce contract / Einstein limits; AWS cost explorer dual-run line | +| **G9** | **Cost — all-staff Ops agent is NOT free (research Finding 6) + parallel-run double-spend.** Agentforce in Slack works on **all paid Slack plans (Grid not required)**, and a **no-cost Salesforce Identity license** covers non-CRM users — BUT that only waives the *CRM seat*; **Employee agents consume Flex Credits (~$0.10/action, 20 credits/action)** OR need the **$125/user/mo Agentforce add-on**. "Free for all staff" is marketing framing about seats, not consumption. Plus dual-run: Phases 1–3 run legacy (2 Bedrock agents + KB + AOSS) AND Agentforce/Data Cloud at once. | M | Real per-action or per-user cost at all-staff scale; the unmetered PSL name is uncertain (docs vs blog differ) — wrong name silently meters everything. | Size a Flex-Credit pool or budget the add-on; **verify the exact unmetered PSL name in-org**; budget a dual-run AWS line; **time-box Phases 1–3 (~6–8 wks)**. | slack.com pricing; Trailhead Agentforce-for-Employees (Flex Credits); Salesforce account team | | **G10** | **No SFDX reusable CI/CD workflow + no OIDC for Salesforce deploy** — org reusable workflows are AWS/CDK/OIDC-shaped (design.md §7.1). | M | `sh-agentforce` can't deploy via the existing pattern; the JWT-bearer connected-app key is a long-lived prod-deploy secret. | Author reusable **`cd-sfdx`** (deploy-then-merge, scratch-org validate); store the connected-app **JWT key in Secrets Manager** (`sh-agentforce/sfdx-jwt-key`), sandbox/prod-scoped, rotated (F3, §1g). Its own design/verify story. | handbook `cicd.md`; SFDX JWT-bearer flow | | **G11** | **Two access-control planes.** Agentforce-in-Slack assigns member access via **Salesforce permissions**; our authz is **Google Groups → Cognito → scopes**. | M | Drift: a user could see an agent but lack the MCP scope, or vice-versa. | Provision Agentforce/Salesforce user access **from the same Google Groups** (SCIM/identity sync) so Google Groups stays the single source of truth (design.md §2.3). | Salesforce SCIM/Google provisioning | -| **G12** | **Parity gate runs on Beta tooling + per-user identity (F1).** Testing Center Test Suites is Beta; batch/AI-generated runs invoke actions under an unclear identity when the action uses a Per-User Browser Flow credential. | M | If batch runs can't invoke per-user actions, the behavioral suite (which authorizes teardown) can't exercise the real tools. | Verify Testing-Center identity under per-user creds in Phase-0a (§6 #5); if unworkable, drive parity via scripted per-user sessions instead of batch. | help.salesforce.com Testing Center; in-org | +| **G12** | **RESOLVED by research (Finding 5): per-user actions are UNTESTABLE via batch Testing Center.** Testing Center batch/AI-generated runs and the Testing API execute under the External Client App's **client-credentials "Run As" SERVICE account** — which has no `UserExternalCredential`, so any Per-User Browser Flow action fails at callout. | M | The behavioral/parity suite (which authorizes teardown) cannot exercise per-user (Lauren Exec / Gmail) tools via batch. | **Decided:** drive per-user parity via **scripted interactive sessions** (Agent Builder preview, or Agent API with `bypassUser=false` + a pre-authorized real user token) — NOT batch. Non-per-user (Ops/Finance read) tools can still use batch. | developer.salesforce.com testing-api-connect; agent-api-get-started (`bypassUser`) | +| **G16** | **CRITICAL — Employee-Agent callout may carry SERVICE identity, not the user's (research Finding 4, the new top risk).** Salesforce Architects docs state verbatim: "User identity is lost by default. When an agent calls an MCP server or downstream API using client credentials, the request carries the agent's service identity and not the end user's." Per-user External Credentials forward the user's token ONLY in a **user-session** context; an autonomously-invoked agent uses the service identity. Whether an **Employee Agent in Slack** runs each tool callout in the end-user's session (so the per-user Cognito token + real `sub` flow) or autonomously (service identity) is **undocumented and unproven** — there is NO primary Salesforce example of per-user Browser Flow + Cognito + `sub` propagation end-to-end. | **C** | If callouts use service identity, **per-user `sub`, ABAC Gmail isolation, and the lethal-trifecta all collapse** — every user looks like one service account to our MCP servers. This is now THE controlling risk, above the Cognito mechanism. | **0a fatal #1:** prove an Employee-Agent-in-Slack tool callout carries the signed-in user's Cognito token (real Google `sub`) to our server. If it does NOT: fallbacks are (a) MuleSoft Flex Gateway "Trusted Agent Identity" (RFC 8693 token exchange), (b) a signed user-identity header our server trusts, or (c) the custom-Bolt surface (design.md §12, keeps per-user OAuth intact). Also: grant **`offline_access`** on the Cognito client or per-user tokens die at 1h with no documented auto-refresh (Finding 4). | architect.salesforce.com end-user-identity-propagation; nc-use-oauth-cred-in-callout; in-org spike | | **G13** | **New data processor: Salesforce now processes our tool payloads (F5).** Finance/Gmail tool responses transit the Agentforce planner / Einstein Trust Layer; grounding flows through Data Cloud. design.md's data boundary was AWS + Slack. | M | Vendor/payment/inbox content (PII is masked, but business content is not) is processed by Salesforce — a new processor not in the original threat model. | Explicit acceptance; verify Trust Layer **retention + zero-training** guarantees and Data Cloud data-residency; keep MCP-layer PII redaction authoritative. | Einstein Trust Layer retention/zero-training docs; DPA | | **G14** | **Per-user tool visibility degraded (F7).** MCP `list_tools` hid tools per-*user* by scope; per-agent action assignment is per-*agent*, so a user lacking a scope still sees the action and fails server-side (403). | L | UX degradation (confusing failures), not a security hole — the per-tool scope check is the boundary. | Graceful-refusal copy in agent instructions; behavioral test for a clean "no access" message (§1h). | — | | **G15** | **Three unverified-load-bearing platform assumptions (Q1/Q2/Q3).** (a) Sea Haven's Slack plan actually supports **Agentforce Employee Agents**; (b) per-employee Agentforce **user+license** required for the all-staff Ops agent (drives G9 cost + G11 SCIM); (c) Agentforce exposes **`$User.GoogleGroups`** and **`$Session.Channel`** to agent instructions (Lauren Exec's DM-only + scope gating depend on these). | M | If (a) wrong, the surface decision (D2) collapses; if (c) wrong, privacy/DM controls need a different enforcement point (e.g. a context Apex action). | Verify all three in Phase-0 (§6); for (c), design a fallback enforcement now so it's not on the critical path. | slack.com/help Agentforce-in-Slack plan reqs; Salesforce licensing; Agentforce agent variables doc | @@ -858,14 +883,22 @@ Groups to reconcile the two control planes (G11, recommend SCIM). 7. **(Agent variables, Q3/G15)** Confirm Agentforce exposes **`$User.GoogleGroups`** and **`$Session.Channel`** to agent instructions. If not, design an alternate enforcement for Lauren Exec's DM-only and per-scope gating (e.g. a context Apex action that returns channel + group facts) before Phase 1/3. -8. **(Cognito token-mint mechanism, BLOCK-1 + Gemini round-3 — FATAL on the 0a gate)** Three things to prove in - the 0a throwaway kit before 0b builds real facades: - (a) **Confirm `AllowedOAuthScopes` strictly filters the issued token** — i.e. issued scopes ⊆ the client's - allow-list **even when a pre-token V2 Lambda runs** (Gemini BLOCK). If a V2 Lambda can widen beyond the - allow-list, the app-client control is soft → promote pre-token fail-closed suppression to primary. +0. **(Employee-Agent callout identity — FATAL #1, G16, research Finding 4)** Prove that an **Employee Agent in + Slack**, when a user invokes a tool, makes the downstream callout carrying **that user's** Cognito token (real + Google `sub`) — not the agent's service identity. This is the single highest risk; if it fails, escalate to the + G16 fallbacks (MuleSoft RFC 8693 / signed user-header / custom-Bolt). Also grant `offline_access` on the Cognito + client (else per-user tokens die at 1h). +8. **(Cognito token-mint mechanism — FATAL, research Finding 1/2/3 resolved; confirm empirically)** In the 0a kit: + (a) **Confirm issued scopes ⊆ `AllowedOAuthScopes` with a SUPPRESS-ONLY pre-token Lambda** (research says + `scopesToAdd` is NOT capped — so we forbid it; prove suppress-only keeps the ceiling). (b) **Confirm the `pre-token V2` trigger is available** on our Cognito plan (it may be a paid feature). - (c) **Pick + prove the per-tier rejection mechanism** — `client_id` allow-list per gateway, resource-server - scope-prefix, or custom `aud` via V2 — since Cognito access tokens carry no native per-resource-server `aud`. + (b) Confirm the pool is on **Essentials** (not Lite — Lite ignores V2). (c) **Use the resolved per-tier + rejection mechanism** — `client_id` allow-list per gateway + resource-server scope-prefix (research Finding 3); + optionally test managed-login resource-binding for a real `aud`. +9. **(Model — reopen, research Finding 7)** Confirm whether Agent Script **`model_config` can place a BYOLLM / + our-Bedrock Claude endpoint as a per-agent planner** (Summer '26 allows per-agent overrides to any supported + model). If yes, it reopens keeping inference + our guardrail on the planner path for specific agents — upside + over the org-level AWS-Hosted default (D5). Non-blocking. (d) **The client/audience matrix is injected as CONFIG, not hardcoded (decoupling FIX/QUESTION).** The transport-agnostic core must NOT embed gateway/topology identifiers in application logic (that would violate the D12/B3 decoupling). It validates against an allow-list **supplied via SSM Parameter Store or a CDK-injected env