First-time deploy procedure for the Phase 2a Cognito auth substrate: the five one-time prerequisites (CDK bootstrap, the two Google secrets with exact JSON shapes, the githubdeploy-sh-mcp OIDC role, the managed Google Groups), the synth → diff → watched manual first deploy → CD hand-off flow, post-deploy validation (ESSENTIALS + V2 trigger, finance client ceiling, group-sync run, deny-list hard-revocation smoke test), and rollback/teardown (RETAIN tables; 0a spike teardown deferred until sh-mcp-auth is validated). Notes that CD (deploy.yaml) is already wired and red on every merge until the OIDC deploy role exists, and that no Cognito hosted-UI domain ships in 2a (OAuth code flow deferred to 2b surface wiring).
10 KiB
Runbook — deploy sh-mcp-auth
First-time deploy of the Cognito auth substrate (Phase 2a). The stack is on main
(infra/, CDK app cdk.json → tsx infra/bin/app.ts) but synth-only / not yet
deployed. This runbook takes it from "synthesizes in CI" to "live in AWS and
validated."
- Account / region:
328440206208/us-east-1(mgmt == prod). - Stack name:
sh-mcp-auth. - Owner of the manual prerequisites: Adam (Google + IAM steps cannot be done from CDK).
- Toolchain: aws-cdk
2.1128.1(CLI, root devDep), aws-cdk-lib2.260.0, constructs10.6.0.
CD is already wired and currently red.
.github/workflows/deploy.yaml(push-to-main→ org reusablecd-cdk.yaml@main, OIDC) has failed on every merge so far (~4s) because thegithubdeploy-sh-mcpdeploy role does not exist yet (Prereq 4). Once the prerequisites below are in place, CD deploys on the next push tomain; the first deploy is done manually (Step 2) so a human watches the initial resource creation.
What this deploys
- Cognito user pool
sh-mcp(FeaturePlan ESSENTIALS — required for the V2 pre-token trigger; Lite silently ignores it). - Google external OIDC IdP (
client_id/secretresolved from Secrets Manager at deploy via a CFN dynamic reference — never inlined). - Resource servers
sh-mcp-ops/sh-mcp-finance+ three app clients whoseAllowedOAuthScopesare the trust-tier ceiling:sh-agentforce-ops,-exec(ops + gmail/calendar, no finance),-finance(finance:readonly, 15-min access token, 8h refresh). - Cognito groups
sh-mcp-ops/-assistant/-finance/-admin. - DynamoDB
sh-mcp-sync-state+sh-mcp-deny-list(bothRETAIN, logical IDs pinned). - Lambdas
sh-mcp-pre-token-gen(suppress-only, V2 trigger) andsh-mcp-group-sync(5-min EventBridge schedule), arm64, 60-day log groups. - SNS alarm topic
sh-mcp-alarms(CMKalias/seahaven-alarm-topics) + 3 alarms.
Out of scope (deferred to Phase 2b / surface wiring)
- No Cognito hosted-UI domain and the app-client
callbackUrlsare a placeholder. The OAuth authorization-code flow (and therefore the Google redirect URIhttps://<domain>/oauth2/idpresponse) is not exercisable until the surface is chosen and aUserPoolDomain+ real callback URLs are added. The substrate deploys and group-sync runs without them. - API Gateway / WAF / server hosting / real integration clients (
jobs/, QBO, Maps, DynamoDB data clients) — Phase 2b.
Prerequisites (one-time, Adam)
1. Confirm CDK bootstrap
The account is used by other stacks, so it is almost certainly bootstrapped with
the modern (newStyleStackSynthesis) bootstrap. Verify:
aws cloudformation describe-stacks --stack-name CDKToolkit --region us-east-1 \
--query 'Stacks[0].Outputs[?OutputKey==`BootstrapVersion`].OutputValue' --output text
# expect a version >= 6. If the stack is missing: npx cdk bootstrap aws://328440206208/us-east-1
2. Google OAuth 2.0 web client → secret sh-mcp/google-oidc
In Google Cloud Console (the Workspace's project) → APIs & Services →
Credentials → Create OAuth client ID → Web application. The authorized redirect
URI is the Cognito domain callback — deferred (added in 2b when the hosted-UI
domain exists: https://<cognito-domain>/oauth2/idpresponse). Create the client
now to capture client_id / client_secret; the redirect URI can be edited later.
Store it (JSON, exact keys — read by the Google IdP providerDetails):
aws secretsmanager create-secret \
--name sh-mcp/google-oidc --region us-east-1 \
--description "Google OAuth web client for Cognito federation (sh-mcp-auth)" \
--secret-string '{"client_id":"<CLIENT_ID>.apps.googleusercontent.com","client_secret":"<CLIENT_SECRET>"}'
3. Google Workspace service account + domain-wide delegation → secret sh-mcp/google-directory-sa
The group-sync Lambda reads Google Group membership via a service account with domain-wide delegation impersonating an admin.
- Create a service account (GCP) and a JSON key.
- Admin console → Security → API controls → Domain-wide delegation → add the
SA's client ID with scope:
https://www.googleapis.com/auth/admin.directory.group.readonly - Pick an admin user for the SA to impersonate (the
subject).
Store it (JSON — fields consumed by GoogleServiceAccount in auth/group-sync):
aws secretsmanager create-secret \
--name sh-mcp/google-directory-sa --region us-east-1 \
--description "Google Workspace SA (domain-wide delegation, Directory groups.readonly) for sh-mcp group-sync" \
--secret-string '{
"client_email":"sh-mcp-group-sync@<project>.iam.gserviceaccount.com",
"private_key":"-----BEGIN PRIVATE KEY-----\n...\n-----END PRIVATE KEY-----\n",
"subject":"admin@seahavenind.com"
}'
The stack's IAM grant is scoped to
secret:sh-mcp/google-directory-sa-*, which matches the 6-char suffix Secrets Manager appends. Keep the name exactlysh-mcp/google-directory-sa.
4. OIDC deploy role githubdeploy-sh-mcp
CD assumes this role via GitHub OIDC. Add the repo to
Sea-Haven-Industries/.github/oidc-deploy-roles.yaml (pattern githubdeploy-<repo>),
trust limited to repo:Sea-Haven-Industries/sh-mcp:ref:refs/heads/main, with
permissions to deploy the stack (CloudFormation + the resource set: Cognito,
Lambda, DynamoDB, IAM PassRole for the Lambda roles, Logs, Events, SNS, and
secretsmanager:GetSecretValue on the two sh-mcp/google-* secrets for the
dynamic reference). This is an IAM change → run the mandatory GPT-4.1 IAM
cross-review on the role JSON before applying.
5. Google Groups exist
The four managed groups must exist as Google Groups, or group-sync's first run errors (and correctly keeps the freshness marker stale → pre-token fails closed):
sh-mcp-ops@seahavenind.com, sh-mcp-assistant@seahavenind.com,
sh-mcp-finance@seahavenind.com, sh-mcp-admin@seahavenind.com
Step 1 — synth + diff (no AWS writes)
cd ~/Documents/repositories/sh-mcp
npm ci
npx cdk synth sh-mcp-auth
npx cdk diff sh-mcp-auth # against an empty account this is the full create set
Optional context flags (both safe to omit on the first deploy):
-c logsKmsArn=arn:aws:kms:us-east-1:328440206208:key/<seahaven-logs-key-id>— CMK-encrypt the two Lambda log groups (else AWS-managed encryption). The group-sync log group records reconciliation counts only, not member emails.-c callbackUrls='https://app.example.com/oauth/callback'— overrides the placeholder; only meaningful once the surface domain exists (2b).
Step 2 — first deploy (manual, watched)
npx cdk deploy sh-mcp-auth \
-c logsKmsArn=arn:aws:kms:us-east-1:328440206208:key/<seahaven-logs-key-id> \
--require-approval any-change # review the IAM diff prompt before confirming
Expect a CMK-key + IAM-policy approval prompt (the group-sync Cognito/Secrets
grants). Note the stack outputs (exported for 2b cross-stack import):
sh-mcp-user-pool-id, sh-mcp-ops-client-id, sh-mcp-exec-client-id,
sh-mcp-finance-client-id.
CIS Section-4 alarms in this account fire on manual IAM changes — expect an alarm notification from the role/policy creation; it is benign here.
Step 3 — hand off to CD
After the manual first deploy succeeds and Prereq 4 is in place, every push to
main deploys via .github/workflows/deploy.yaml. Re-run the last failed deploy
to confirm it now goes green:
gh workflow run deploy.yaml --ref main # or push any commit to main
gh run watch $(gh run list --workflow=deploy.yaml --limit 1 --json databaseId --jq '.[0].databaseId')
Post-deploy validation
POOL=$(aws cloudformation list-exports --region us-east-1 \
--query "Exports[?Name=='sh-mcp-user-pool-id'].Value" --output text)
# Pool is ESSENTIALS and the V2 pre-token trigger is attached
aws cognito-idp describe-user-pool --user-pool-id "$POOL" --region us-east-1 \
--query 'UserPool.{Tier:UserPoolTier,PreToken:LambdaConfig.PreTokenGenerationConfig}'
# expect Tier=ESSENTIALS, PreToken.LambdaVersion=V2_0
# The four managed groups exist
aws cognito-idp list-groups --user-pool-id "$POOL" --region us-east-1 \
--query 'Groups[].GroupName'
# Finance client ceiling: finance:read ONLY, 15-min access token
aws cognito-idp describe-user-pool-client --user-pool-id "$POOL" --region us-east-1 \
--client-id <finance-client-id> \
--query 'UserPoolClient.{Scopes:AllowedOAuthScopes,Access:AccessTokenValidity,Refresh:RefreshTokenValidity,Units:TokenValidityUnits}'
# Group-sync ran and wrote the freshness marker (within the last 5 min)
aws lambda invoke --function-name sh-mcp-group-sync --region us-east-1 /dev/stdout | tail -1
aws dynamodb get-item --table-name sh-mcp-sync-state --region us-east-1 \
--key '{"pk":{"S":"group-sync"}}' --query 'Item.lastSuccessfulSyncMs'
# Tail the structured logs:
aws logs tail /aws/lambda/sh-mcp-group-sync --region us-east-1 --since 10m
Hard-revocation smoke test (deny-list overlay) — add a sub, confirm the
pre-token Lambda strips all tier scopes on the next mint, then remove it:
aws dynamodb put-item --table-name sh-mcp-deny-list --region us-east-1 \
--item '{"sub":{"S":"<test-sub>"},"expiresAt":{"N":"'$(($(date +%s)+600))'"}}'
# (mint a token for that user via the hosted UI once 2b adds the domain; expect base scopes only)
aws dynamodb delete-item --table-name sh-mcp-deny-list --region us-east-1 \
--key '{"sub":{"S":"<test-sub>"}}'
Rollback / teardown
- Tables
RETAINon stack delete —cdk destroy sh-mcp-authleavessh-mcp-sync-stateandsh-mcp-deny-list(and their logical IDs are pinned, so a later redeploy re-adopts them). Delete the tables manually only if you intend to lose revocation/freshness state. - A bad deploy rolls back automatically (CloudFormation). To revert code, revert
the commit on
main; CD redeploys the prior template. - 0a spike teardown is the LAST step, not part of this deploy. The live 0a
spike kit (Cognito pool
us-east-1_GsDbGe0pa, probe Lambda/API, the SFsh_mcp_0aobjects) stays up untilsh-mcp-authis deployed and validated — only then run the spike teardown so the proven chain isn't lost prematurely.
Owed once deployed (per CLAUDE.md)
- Confluence "AWS Architecture Map" (id 1540098) — add the
sh-mcp-authMermaid subgraph. - Update
project_sh_mcpmemory's deployed-resources list (pool id, client ids, table/Lambda names) from synth-only to live.