open-swe/infra/lib/constructs/github-deploy-roles.ts
Adam Moussa a33aaec495
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54)
* fix: enforce a replay window on Linear webhooks (AUTHZ-001)

verify_linear_signature accepted any correctly-signed body with no freshness
check, so a captured request could be replayed indefinitely. Parse the
signed webhookTimestamp (Unix ms) and reject requests outside a 60s window,
failing closed when the field is missing or malformed — mirroring the Slack
verifier.

* fix: stop leaking upstream auth-error bodies into user comments

get_github_token_for_user folded the raw upstream response text into the
error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the
full body server-side only and return a generic "GitHub auth failed (status
<code>)". Also document the accepted shared-installation-token blast radius on
the bot-token-only path (AUTHZ-003).

* fix: bind sandbox and token caches to repo to prevent thread-id collision

A PR head-branch name is attacker-controllable and get_thread_id_from_branch
derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01).
The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on
thread_id alone, and a cached sandbox was reused after only an echo-ping, so a
different repo's webhook could bind to another thread's sandbox or token.

Without changing the persistent thread-id scheme:
- Persist the bound repo (owner/name) in thread metadata on sandbox creation and
  refuse to reuse a sandbox whose bound repo does not match the current event
  (SandboxRepoMismatchError); the in-memory proxy also carries the binding.
- Bind the GitHub-token cache entries to their repo and evict on a cross-repo
  read so a colliding thread_id cannot be served another repo's token.
- Thread repo through the reviewer and the webhook token resolvers.

* fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04)

The instance role and the GitHub deploy app role granted s3:ListBucket on the
whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts)
only ever lists under releases/, so add a StringLike s3:prefix=releases/*
condition. GetBucketLocation has no s3:prefix in its request context, so it
moves to its own unconditioned statement. Also document the accepted F-2
cross-env existence-oracle residual on BatchGetSecretValue.

* chore: suppress test-fixture credential false positive; document AUTHZ-002

Add a machine-level suppression for the fake Datadog key in the
test_team_credentials encryption-roundtrip fixture (CWE-798, not a real
credential). Clarify that the within-org thread-write path is intentional by
design (AUTHZ-002) — comment only, no behavior change.

* fix: casefold repo-binding keys to avoid spurious cross-repo mismatch

GitHub owner/name are case-insensitive. Casefold the owner/name key on both the
write (binding) and read (compare) sides — repo_cache_key and the metadata
bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate
same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2).

* fix: stop leaking upstream auth body in unexpected-result branch

The 2xx-but-missing-token/url branch echoed the parsed upstream response body
into the user-facing error. Return a generic message and log response_data
server-side only, mirroring the existing HTTPStatusError fix (Gap 4).

* fix: fail closed for unbound-legacy sandboxes and catch repo mismatch

Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no
recorded bound_repo (a pre-binding legacy thread, post-deploy) previously
reconnected-and-served the sandbox to the current repo, then rebound it. Now
fail closed: drop the stale id and recreate a fresh sandbox bound to this repo,
logging a reconnect-with-missing-binding event. A sandbox is never served to a
repo unless its binding is known and matches; new threads bind on first run
unchanged.

Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints,
log it for alarming, and surface a clean sanitized error instead of letting an
opaque deep-stack exception crash-loop the worker.

* chore: suppress test-fixture credential false positive in token-TTL tests

Add a machine-level suppression for the fake "ghp_secret" GitHub token used by
the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and
not a valid PAT; scoped to the unit test only.
2026-06-29 12:21:19 -04:00

173 lines
7.6 KiB
TypeScript

import * as iam from "aws-cdk-lib/aws-iam";
import { Construct } from "constructs";
import {
ACCOUNT,
EnvName,
GITHUB_OIDC_PROVIDER_ARN,
REGION,
oidcSubject,
} from "../config";
/**
* Per-ENV GitHub Actions OIDC deploy roles. Created ONCE per env in the
* dedicated `open-swe-iam` stack. Two roles per env, per the locked architecture's
* "dual OIDC roles":
*
* - githubdeploy-open-swe-infra-<env> → CFN/IAM (CDK) deploys of that env's stack
* - githubdeploy-open-swe-app-<env> → app deploys (env-tag-scoped SSM + S3 read)
*
* T5 OSWE-IAC-01/02 fix: roles are split per env and the trust subject is
* env-scoped (dev = dev branch ref; prod = the GitHub `prod` Environment subject,
* so the manual-approval gate is IAM-enforced). A dev-branch token therefore
* cannot SendCommand to the prod box nor assume a prod deploy role.
*
* Reviewed at T4 (GPT-4.1 IAM cross-review) + T5 (/sh-security-review) and
* deployed FIRST (BLOCK#3 "OIDC-role-first" ordering) before any other infra or
* secrets CI step.
*/
export class GithubDeployRoles extends Construct {
public readonly infraRole: iam.Role;
public readonly appRole: iam.Role;
constructor(scope: Construct, id: string, envName: EnvName) {
super(scope, id);
// The provider already exists account-wide — reference, never re-create.
const provider = iam.OpenIdConnectProvider.fromOpenIdConnectProviderArn(
this,
"GithubOidcProvider",
GITHUB_OIDC_PROVIDER_ARN,
);
// T4 BLOCK#2 + T5 IAC-01/02: exact env-scoped subject via StringEquals (no
// StringLike, no `*`). prod = environment:prod (manual-approval gate),
// dev = the dev branch ref.
const trust = new iam.WebIdentityPrincipal(provider.openIdConnectProviderArn, {
StringEquals: {
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
"token.actions.githubusercontent.com:sub": oidcSubject(envName),
},
});
// ---- githubdeploy-open-swe-infra-<env> --------------------------------
this.infraRole = new iam.Role(this, "InfraDeployRole", {
roleName: `githubdeploy-open-swe-infra-${envName}`,
assumedBy: trust,
description: `GitHub OIDC role for CDK deploys of the open-swe-${envName} infra stack (assumes CDK bootstrap roles).`,
});
// Org-standard CDK deploy pattern (mirrors githubdeploy-seahaven-account-
// baseline / -forgejo / -apm-wo-analysis): the deploy role only needs to
// assume the CDK bootstrap roles. The actual CloudFormation + IAM + resource
// permissions are exercised by the bootstrap `cfn-exec-role`, whose scope is
// owned by the CDKToolkit stack — NOT granted directly here.
//
// T4 BLOCK#1: GPT-4.1 flagged the `cdk-hnb659fds-*` wildcard and recommended
// enumerating the four exact ARNs. ACCEPTED EXCEPTION (Adam, 2026-06-26): kept
// as the verified org-wide convention (githubdeploy-seahaven-account-baseline
// uses the identical wildcard). Only `cdk bootstrap` creates roles with this
// prefix, so practical escalation risk is low.
// T5 residual (OSWE-IAC-02): the single account-wide cfn-exec-role means the
// dev infra role can technically deploy any stack; per-env trust gates WHO can
// assume, and the prod role requires the environment:prod approval. Per-env
// bootstrap qualifiers would close the residual fully (future hardening).
this.infraRole.addToPolicy(
new iam.PolicyStatement({
sid: "AssumeCdkBootstrapRoles",
actions: ["sts:AssumeRole"],
resources: [`arn:aws:iam::${ACCOUNT}:role/cdk-hnb659fds-*`],
}),
);
// ---- githubdeploy-open-swe-app-<env> ----------------------------------
this.appRole = new iam.Role(this, "AppDeployRole", {
roleName: `githubdeploy-open-swe-app-${envName}`,
assumedBy: trust,
description: `GitHub OIDC role for open-swe-${envName} app deploys: env-tag-scoped ssm:SendCommand + read of the ${envName} S3 artifact bucket.`,
});
// T5 OSWE-IAC-01 fix: SendCommand only to instances tagged project=open-swe
// AND env=<this env> (a SINGLE value, not {dev,prod}). The dev app role can
// never command the prod box and vice versa — env isolation in IAM.
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "SsmSendCommandTagScoped",
actions: ["ssm:SendCommand"],
resources: [`arn:aws:ec2:${REGION}:${ACCOUNT}:instance/*`],
conditions: {
StringEquals: {
"ssm:resourceTag/project": "open-swe",
"ssm:resourceTag/env": envName,
},
},
}),
);
// SendCommand also has to reference the command document. Scope to this env's
// open-swe deploy document ONLY.
// T4 BLOCK#3 (CLOSED at T19): GPT-4.1 flagged AWS-RunShellScript as an
// arbitrary-shell escalation path. The dedicated `open-swe-${envName}-deploy`
// SSM document (app-service.ts) now runs the fixed, parameter-less command
// `bash /opt/open-swe/bin/deploy.sh`, so AWS-RunShellScript is dropped here:
// this role can run ONLY that one document, and only on its own env's box
// (tag-scoped by the SsmSendCommandTagScoped statement above).
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "SsmSendCommandDocuments",
actions: ["ssm:SendCommand"],
resources: [`arn:aws:ssm:${REGION}:${ACCOUNT}:document/open-swe-${envName}-deploy`],
}),
);
// Poll command results. These read actions do not support resource-level
// scoping, so `*` is required by the API (T4 FIX: API limitation, documented).
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "SsmReadCommandStatus",
actions: [
"ssm:GetCommandInvocation",
"ssm:ListCommands",
"ssm:ListCommandInvocations",
],
resources: ["*"],
}),
);
// Read+WRITE access to THIS env's artifact bucket only (T19): the
// build-artifacts workflow uploads app.tar.gz / spa.tar.gz under releases/*,
// then fires the deploy document so the box pulls them via its instance role.
// Object actions are scoped to releases/* (the only prefix CI writes), and to
// THIS env's bucket — a dev token can never write the prod bucket. No
// bucket-level mutation (no PutBucket*/Delete bucket) — that stays with CDK.
// GetObject + PutObject (S3-to-S3 copy = Get source + Put dest) is all the
// publish/rollback path uses; s3:DeleteObject is deliberately NOT granted so a
// CI token cannot erase an immutable release or the releases/last-good rollback
// fallback (lifecycle expiry handles old-version cleanup, not CI).
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "ReadWriteArtifactObjects",
actions: ["s3:GetObject", "s3:PutObject"],
resources: [`arn:aws:s3:::open-swe-${envName}-assets/releases/*`],
}),
);
// ListBucket is constrained to the releases/ prefix (F-1/IAC-04): the
// publish/rollback scripts only ever list under releases/, so a leaked CI
// token cannot enumerate anything else in the bucket. GetBucketLocation
// carries no s3:prefix, so it stays a separate, unconditioned statement.
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "ListArtifactBucket",
actions: ["s3:ListBucket"],
resources: [`arn:aws:s3:::open-swe-${envName}-assets`],
conditions: { StringLike: { "s3:prefix": ["releases/*"] } },
}),
);
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "GetArtifactBucketLocation",
actions: ["s3:GetBucketLocation"],
resources: [`arn:aws:s3:::open-swe-${envName}-assets`],
}),
);
}
}