mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 09:13:14 +00:00
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)
Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.
- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module
* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8
The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.
Map profile effort to additional_model_request_fields:
{thinking: {type: adaptive, display: summarized},
output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.
* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids
Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
(set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.
Surfaced by the cross-family review + verified against deploy/.
* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip
From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
error code only, so the role ARN + account id in the raw botocore message never
reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
(Converse emits reasoning_content, not thinking) so the middleware is not a no-op
on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)
* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids
Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
each routed region (us-east-1/2, us-west-2). The model runs in the server process
on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
seed_store.sh's default via pick precedence, so the seed-script fix alone was
insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
ids to the Bedrock id (config.toml's model_id was an active, now-broken value).
AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.
* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)
Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
(eval judge only — Bedrock builder/reviewer auth via the host IAM role).
REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
160 lines
7.6 KiB
TypeScript
160 lines
7.6 KiB
TypeScript
import * as iam from "aws-cdk-lib/aws-iam";
|
|
import { Construct } from "constructs";
|
|
import { ACCOUNT, EnvName, REGION, prefix } from "../config";
|
|
|
|
/**
|
|
* Least-privilege EC2 instance role for the open-swe box (one per env).
|
|
*
|
|
* Grants exactly what the boot/runtime flow needs and NOTHING ELSE — no admin,
|
|
* no `*` resources except where the AWS action genuinely has no resource-level
|
|
* scoping. Per-env so the dev box can never read prod secrets/config and vice
|
|
* versa. Reviewed at T4 (GPT-4.1 IAM cross-review) / T5 (/sh-security-review)
|
|
* before it is ever deployed (T6).
|
|
*/
|
|
export class InstanceRole extends Construct {
|
|
public readonly role: iam.Role;
|
|
|
|
constructor(scope: Construct, id: string, env: EnvName) {
|
|
super(scope, id);
|
|
const p = prefix(env);
|
|
|
|
this.role = new iam.Role(this, "Role", {
|
|
roleName: `${p}-instance-role`,
|
|
assumedBy: new iam.ServicePrincipal("ec2.amazonaws.com"),
|
|
description: `EC2 instance role for the ${p} open-swe box (least-privilege).`,
|
|
});
|
|
|
|
// AWS-managed: lets the SSM agent register the instance and RECEIVE the
|
|
// app-deploy `ssm:SendCommand` from githubdeploy-open-swe-app. This is the
|
|
// standard Session-Manager / RunCommand grant and is the only managed
|
|
// policy on the role. DELIBERATE — flag for T4 confirmation.
|
|
this.role.addManagedPolicy(
|
|
iam.ManagedPolicy.fromAwsManagedPolicyName("AmazonSSMManagedInstanceCore"),
|
|
);
|
|
|
|
// Read the build artifact from the env's S3 asset bucket (deploy = pull).
|
|
// Scoped to releases/* — the only prefix CI writes and the box pulls — so a
|
|
// compromised box (or stolen IMDS creds) cannot read anything else that might
|
|
// ever land in the bucket (least-privilege; mirrors the app role's write scope).
|
|
this.role.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "ReadArtifactObjects",
|
|
actions: ["s3:GetObject"],
|
|
resources: [`arn:aws:s3:::${p}-assets/releases/*`],
|
|
}),
|
|
);
|
|
// ListBucket is constrained to the releases/ prefix (F-1/IAC-04) — the box
|
|
// only ever lists release artifacts, so a compromised box cannot enumerate
|
|
// any other object that might land in the bucket. GetBucketLocation has no
|
|
// s3:prefix in its request context, so it stays a separate, unconditioned
|
|
// statement (the condition would otherwise AccessDeny it).
|
|
this.role.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "ListArtifactBucket",
|
|
actions: ["s3:ListBucket"],
|
|
resources: [`arn:aws:s3:::${p}-assets`],
|
|
conditions: { StringLike: { "s3:prefix": ["releases/*"] } },
|
|
}),
|
|
);
|
|
this.role.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "GetArtifactBucketLocation",
|
|
actions: ["s3:GetBucketLocation"],
|
|
resources: [`arn:aws:s3:::${p}-assets`],
|
|
}),
|
|
);
|
|
|
|
// Read non-sensitive config from SSM Parameter Store under /open-swe-<env>/*.
|
|
this.role.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "ReadSsmConfig",
|
|
actions: ["ssm:GetParameter", "ssm:GetParameters", "ssm:GetParametersByPath"],
|
|
resources: [`arn:aws:ssm:${REGION}:${ACCOUNT}:parameter/${p}/*`],
|
|
}),
|
|
);
|
|
|
|
// VALUE access — Secrets Manager under open-swe-<env>/*. Secret ARNs carry a
|
|
// random 6-char suffix, hence the trailing `*`. This is the statement that
|
|
// actually gates which secret VALUES the box can read: prefix-scoped, so the
|
|
// dev box can never read prod secret values (and vice versa). GetSecretValue is
|
|
// checked per-secret even when the value is returned via the batch call below.
|
|
this.role.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "ReadSecretValues",
|
|
actions: ["secretsmanager:GetSecretValue", "secretsmanager:DescribeSecret"],
|
|
resources: [`arn:aws:secretsmanager:${REGION}:${ACCOUNT}:secret:${p}/*`],
|
|
}),
|
|
);
|
|
|
|
// BatchGetSecretValue MUST be granted on `*` — it is a collection action that
|
|
// AWS authorizes against the account, NOT the per-secret ARN, REGARDLESS of
|
|
// whether the caller uses `--filters` or `--secret-id-list`. A prefix-scoped
|
|
// BatchGetSecretValue AccessDenies the whole call ("no identity-based policy
|
|
// allows the secretsmanager:BatchGetSecretValue action") — VERIFIED on the live
|
|
// dev box 2026-06-29 (the OSWE-IAC-SECRETS-LIST-01 attempt to prefix-scope it
|
|
// crash-looped the box once the prior broad grant's eventual-consistency lapsed).
|
|
// This `*` does NOT widen VALUE access: a secret value is only returned when the
|
|
// prefix-scoped GetSecretValue above also allows it, so cross-env value isolation
|
|
// holds. The win that DID survive: fetch-config uses `--secret-id-list` (explicit
|
|
// names, no name filter), so `secretsmanager:ListSecrets` is NOT needed and is
|
|
// intentionally omitted — the box cannot enumerate secret names account-wide.
|
|
// F-2 (accepted residual): because the grant is `*`, a caller naming a secret
|
|
// in ANOTHER env's prefix learns whether that name EXISTS (an existence oracle
|
|
// via the per-secret AccessDenied-vs-not signal) even though the VALUE stays
|
|
// gated by the prefix-scoped GetSecretValue above. Accepted within Sea Haven's
|
|
// single-tenant account 328440206208 — cross-env VALUE isolation is preserved.
|
|
this.role.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "BatchGetSecretValues",
|
|
actions: ["secretsmanager:BatchGetSecretValue"],
|
|
resources: ["*"],
|
|
}),
|
|
);
|
|
|
|
// NOTE (T11): SSM SecureString + Secrets Manager here are assumed to use the
|
|
// AWS-managed keys (alias/aws/ssm, alias/aws/secretsmanager) for which the
|
|
// service grants Decrypt implicitly — so NO kms:Decrypt is granted. If T11
|
|
// moves these to a customer CMK, add a scoped `kms:Decrypt` on that key ARN
|
|
// ONLY (not `*`).
|
|
|
|
// Invoke the Bedrock Claude model. DEFAULT_MODEL_ID is
|
|
// `bedrock_converse:us.anthropic.claude-opus-4-8`, and the model runs in the
|
|
// LangGraph server PROCESS on this box (not in the sandbox), so the EC2
|
|
// instance role is the calling principal. The `us.` cross-region inference
|
|
// profile fans out to us-east-1 / us-east-2 / us-west-2, and Bedrock authorizes
|
|
// InvokeModel against BOTH the inference-profile ARN AND the underlying
|
|
// foundation-model ARN in each routed region — all four resources are required
|
|
// or the call AccessDenies. Scoped to opus-4-8 ONLY (least-privilege): adding a
|
|
// new Bedrock model to SUPPORTED_MODELS means extending this resource list.
|
|
// IAM change — flag for T4 (GPT-4.1 IAM cross-review) / T5 (/sh-security-review).
|
|
this.role.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "InvokeBedrockClaude",
|
|
actions: ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"],
|
|
resources: [
|
|
`arn:aws:bedrock:${REGION}:${ACCOUNT}:inference-profile/us.anthropic.claude-opus-4-8`,
|
|
"arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-opus-4-8",
|
|
"arn:aws:bedrock:us-east-2::foundation-model/anthropic.claude-opus-4-8",
|
|
"arn:aws:bedrock:us-west-2::foundation-model/anthropic.claude-opus-4-8",
|
|
],
|
|
}),
|
|
);
|
|
|
|
// Ship application logs to CloudWatch Logs under /open-swe/<env>/*.
|
|
this.role.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "PutAppLogs",
|
|
actions: [
|
|
"logs:CreateLogGroup",
|
|
"logs:CreateLogStream",
|
|
"logs:PutLogEvents",
|
|
"logs:DescribeLogStreams",
|
|
],
|
|
resources: [
|
|
`arn:aws:logs:${REGION}:${ACCOUNT}:log-group:/open-swe/${env}/*`,
|
|
`arn:aws:logs:${REGION}:${ACCOUNT}:log-group:/open-swe/${env}/*:*`,
|
|
],
|
|
}),
|
|
);
|
|
}
|
|
}
|