The built-in GITHUB_TOKEN cannot open the sync PR: the enterprise policy
blocks GitHub Actions from creating/approving pull requests, which
overrides the org and repo settings. That restriction applies only to
github-actions[bot], so mint a PROMOTE_APP installation token and pass
it to create-pull-request, mirroring promote-to-main.yml.
Add upstream-ledger-sync action to run daily to pull upstream commits, update triage.jsonl, and open a PR against dev so the list can be automatically updated to stay in line with upstream.
* chore(deps): bump vite-tsconfig-paths to ^6.1.1 and jsdom to ^29.1.1
Reconcile two Dependabot PRs (#107, #108) into one branch with a
single bun install so package.json and bun.lock stay consistent.
vite-tsconfig-paths: ^5.1.4 -> ^6.1.1 (dependencies)
jsdom: ^27.4.0 -> ^29.1.1 (devDependencies)
* Add --frozen-lockfile to CI Checks
Normal bun install treats bun.lock as updatable. If package.json requests a requirement bun.lock doesn't satisfy, bun quietly rewrites the lock and moves on. The committed lock is never updated.
* chore: Add trailing newline on new last run
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
Adds a triage-ledger job running `make triage-check` so a ledger edit that
forgets to regenerate the markdown (make triage-render) fails CI instead of
drifting silently. Stdlib-only, no deps.
* ci: allow deps-dev, ui, deploy, ci scopes in PR title lint
Dependabot dev-dependency PRs use the deps-dev scope and humans commonly
use ui/deploy/ci, none of which were in the allowed list, so lint-pr-title
failed on otherwise-valid PRs.
* ci: allow infra scope in PR title lint
* chore: decommission self-hosted AWS LangGraph stack
Removes the now-dead self-host IaC and AWS-only CI/CD after destroying the
dev + prod CloudFormation stacks (open-swe-dev, open-swe-prod, open-swe-iam,
and the dev-exclusive CDKToolkit-oswedev bootstrap) in account 328440206208,
us-east-1. The deployment is now managed (LangGraph Cloud + Vercel).
- remove infra/ (CDK app: app + IAM stacks, constructs, aspects, tests)
- remove deploy/ami (Packer AMI build) and deploy/seahaven (boot/config
scripts, DEPLOYMENT/ROTATION runbooks)
- remove AWS-only workflows: cd-infra, ci-infra, build-artifacts, rollback
- README: rewrite the Deployment section to the managed LangGraph Cloud +
Vercel view; drop dead links to infra/ and deploy/seahaven
Preserved: the shared default CDKToolkit bootstrap and promote-dev-to-prod.yml.
The RETAIN'd Secrets Manager shells and open-swe-<env>-assets S3 buckets
survive cdk destroy by design (orphaned) and need a separate deliberate cleanup.
* chore: clean up dangling references left by the AWS decommission
Folds in the FIX-level items from the #64 review gates (GPT-4.1 cross-review +
/sh-security-review), none of which were blockers:
- delete orphaned .github/scripts/{package-artifacts,publish-and-deploy,roll-box,
rollback}.sh — their only callers were the removed AWS deploy workflows
- drop the deleted /infra dir from dependabot.yml npm directories (was producing
a recurring Dependabot config error)
- remove the stale OSWE-IAC-SECRETS-LIST-01 suppression (referenced the deleted
infra/lib/constructs/instance-role.ts)
- repoint the README promotion link to promote-to-main.yml (renamed in #63)
The promote-dev-to-prod.yml comment in check-dev-green.sh is intentionally left
to #63, which rewrites that same line.
* ci: re-home prod promotion into a gated promote-to-main workflow
Migrate the prod-deploy gate to managed LangGraph Cloud (git-connected to
`main`) + Vercel. Under managed, a push to `main` auto-deploys prod, so the
dev -> main fast-forward IS the prod deploy trigger -- the bespoke AWS CD step
is obsolete and already gone from this workflow.
Re-home `promote-dev-to-prod.yml` -> `promote-to-main.yml`:
- gate the promote job on the `prod` GitHub Environment (required reviewer
amoussa1229), restoring the manual prod-approval that the retired AWS CD job
used to carry;
- drop the nightly auto-promote cron -- a scheduled auto-promotion conflicts
with a manual approval gate now that the push deploys prod; promotion is
workflow_dispatch only;
- keep the dev-HEAD-fully-green precondition and the seahaven-promotion App
fast-forward push (sole non-admin bypass actor on `main` ruleset 18238334).
Update the companion check-dev-green.sh filename reference.
* ci: refuse promote-to-main dispatch from any ref other than dev
Defense-in-depth atop the already-pinned `ref: dev` checkout: workflow_dispatch
runs the workflow definition from the launched ref, so reject a non-dev dispatch
before the App token is minted. Surfaced by the GPT-4.1 cross-review of #63.
The nightly promote fast-forwards main to a fully-green dev HEAD, but the
push (as github-actions[bot]) is rejected by the main ruleset: it requires
PRs + a status check and the default token is not a bypass actor, so a
direct ref push can never land regardless of fast-forwardability. The
prior comment claiming protection 'only rejects non-FF' was wrong.
Mint a GitHub App installation token (actions/create-github-app-token,
SHA-pinned) and push with it; the App must be added to the main ruleset's
bypass actors out-of-band. The promoted commit already passed every check
on dev (gated by check-dev-green.sh), so re-gating it via a PR on main is
redundant.
Also fix a gate self-poison: a stale failed 'promote' check-run from a
prior run on the same dev HEAD blocked every subsequent gate run (it was
excluded only by the current run_id). Exclude prior promote check-runs
too, scoped to name=='promote' AND a /actions/runs/ details_url so an
external app cannot hide a real failing check by naming it 'promote'; the
positive REQUIRED_CHECKS allow-list stays authoritative.
Gates: GPT-4.1 cross-review APPROVE (no security regression). Unit-tested:
stale promote ignored -> PASS; real failure / external promote / missing
required check -> BLOCK. shellcheck clean (also fixed a pre-existing
SC2295 on the run_id match).
* Align workflows with Sea Haven CI/CD handbook
Bring the workflow suite in line with the handbook: bump
actions/checkout to v7 (Node 24 runtime, already standardized),
kebab-case the two snake_case workflow filenames, and add the
org-standard Labeler caller and Dependency Review gate so vulnerable
or disallowed-license deps and unlabeled PRs are caught automatically.
File renames only — job/check display names are unchanged, so the
promotion gate's REQUIRED_CHECKS and branch-protection required
checks are unaffected.
Refs: INFRA-115
* Drop Agent prefix from CI workflow + job names
The handbook names workflows for what they do (CI, Deploy, Labeler),
not the component they run, matching .github and afterhours-shift-manager.
Rename the suite to CI and its jobs to Lint / Format check / Unit tests,
and keep the promotion gate's REQUIRED_CHECKS in sync.
Refs: INFRA-115
---------
Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
T20 CD safety nets. Two gaps closed before the first real prod deploy:
1. Promotion gate. promote_dev_to_prod.yml previously fast-forwarded
dev→main unconditionally. It now hard-gates on check-dev-green.sh:
every check-run on the dev HEAD must be completed+passing AND the
Agent CI suite (lint/format/unit/E2E) must be present+success, or the
promotion blocks (fails safe on a missing/renamed check). The promote
run excludes its OWN check-run by run-id (unforgeable), never by the
mutable name "promote", so a colliding red check cannot hide. Fields
are read with a 0x1F separator so an empty conclusion (every
in_progress check) cannot shift columns. ci.yml now also runs on
push:dev so dev HEAD actually carries that signal (a PR check alone
can be admin-merged past).
2. Rollback + last-good. publish-and-deploy.sh advances
releases/last-good/ only after a successful roll (deploy.sh gates on
`systemctl is-active`), and makes releases/latest/ transactional —
reverting to the prior release if the roll fails so a replaced box
never self-deploys a broken release. New rollback.yml + rollback.sh
re-point latest at last-good (or an explicit sha) and re-fire the
deploy; prod is gated by the `prod` Environment approval, same as a
deploy. The shared fire/wait/aggregate-gate logic is factored into
roll-box.sh (used by both forward and backward rolls).
Least-privilege: drop the unused s3:DeleteObject from the app deploy
role — publish/rollback/deploy only Get+Put (S3-to-S3 copy), and the
rollback fallback now depends on immutable release history staying
intact. Lifecycle expiry (not CI) handles old-version cleanup.
Gate logic unit-tested (7 cases + jq round-trip). IAM change +
release-safety control cross-reviewed by GPT-4.1: APPROVE, no blocks.
Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi
cd-infra reported failure on every successful deploy: the 'Stack outputs' step ran
'aws cloudformation describe-stacks' with the githubdeploy-open-swe-infra-<env> role,
which intentionally lacks cloudformation:DescribeStacks. The cdk deploy itself succeeds
(it reads outputs via the bootstrap cfn-exec role it assumes). Switch to
'cdk deploy --outputs-file cdk-outputs.json' + cat — no extra IAM grant, and the job
goes green on actual deploy success instead of masking real failures behind a red run.
* fix(infra): grant instance role BatchGetSecretValue + ListSecrets for .env materialization
fetch-config.sh materializes the box's .env via
`secretsmanager batch-get-secret-value --filters Key=name,Values=open-swe-<env>/`,
but the instance role only granted GetSecretValue/DescribeSecret. BatchGetSecretValue
is a distinct IAM action, so the call was AccessDenied and open-swe.service
crash-looped (no .env written -> ExecStartPre exit 1).
- Add secretsmanager:BatchGetSecretValue to the prefix-scoped ReadSecrets statement.
- Add secretsmanager:ListSecrets on * (required by the name-prefix filtered batch
call; the API has no resource-level scoping for the list action — fits the role's
stated exception). Secret VALUES stay prefix-scoped; only names are enumerable.
Reviews: GPT-4.1 IAM cross-review BLOCK=none; /sh-security-review iac-iam one LOW
metadata residual (no critical/high), recorded as OSWE-IAC-SECRETS-LIST-01.
Refs T7/T19 dev bring-up.
* ci: lift Node heap cap for Playwright E2E build (vite OOM)
The E2E job's Playwright globalSetup runs the real `bun run build`, whose vite
bundle exceeds Node's default ~2 GB heap and OOMs (JavaScript heap out of memory) —
the same failure fixed for build-artifacts.yml in #19. Set
NODE_OPTIONS=--max-old-space-size=8192 on the Run E2E step.
The systemd unit booted with a literal `@@OPENSWE_ENV@@` (fetch-config.sh got the
token, not "dev" -> exit 2 -> crash-loop) because user-data.sh is double-templated:
CDK substitutes @@tokens@@ AND user-data seds @@tokens@@ into the baked
systemd/nginx files. CDK's `.replace(/@@OPENSWE_ENV@@/g, "dev")` clobbered the sed
PATTERN (`s|@@OPENSWE_ENV@@|...|` -> `s|dev|...|`, a no-op), so the unit's token
never got replaced. Same collision hit @@SERVER_NAME@@ (masked by nginx
default_server).
Fix: CDK tokens move to a DISTINCT delimiter %%...%% (rendered in app-service.ts);
the @@...@@ tokens stay for the baked-template seds. No AMI rebuild (templates
unchanged). Add a guard test asserting no unresolved %%CDK%% token survives in the
synthesized user-data. Also add .github/scripts/** to build-artifacts paths so
script-only changes trigger a publish.
jest 20/20; tsc + shellcheck clean; rendered user-data: OPENSWE_ENV="dev",
SERVER_NAME="openswe-dev.seahaven.com", @@ sed patterns preserved, 16872 B.
The build-artifacts SPA build hit Node's default ~2 GB heap cap and aborted
(JS heap OOM, exit 134) — the same memory-hungry vite build that needed an 8 GB
swapfile on-box. The runner has ~16 GB, so set NODE_OPTIONS=--max-old-space-size=8192
on both build steps.
* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)
Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.
deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
- the file provisioner failed uploading the templates dir ('scp: …: Is a
directory') — a trailing-slash contents-upload needs the dest dir to exist;
added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
dest trailing slash.
- the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
environment_vars never reached provision.sh (which runs under set -u and
aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.
infra:
- ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
discipline docs.
- app-service.ts: machineImage → bakedOpenSweArm64().
- open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
- cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
- README: Baked AMI + EBS-replacement-discipline section.
tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.
NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).
* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021
Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.
* fix(deploy): GitHub App + Slack required for prod only, not dev
Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.
Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).
* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)
Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.
Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
/opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
baked open-swe-base-arm64 AMI (folds in the held #16).
IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.
Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
target is healthy even before the first release); deploy.sh is base64-rendered
by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).
CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
ui/.output/public -> spa.tar.gz), package the Python source via git archive
(app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
push dev -> dev (auto); push main -> prod (env "prod" approval gate).
Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.
* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard
Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:
- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
not CommandInvocations[0], so a partial failure across the brief 2-instance
replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
(defense in depth over .gitignore; scoped to data extensions so *_credentials.py
source is not a false positive — verified against the real tree).
Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).
shellcheck/tsc/jest(16) clean; both stacks synth offline.
* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard
The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)
- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
GroupDescription + rule descriptions are pure ASCII, so this fails the build
instead of a deploy next time.
jest 18/18; tsc clean.
* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)
The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.
jest 18/18; tsc clean.
* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit
The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.
jest 19/19; minified deploy.sh passes bash -n + shellcheck.
* ci: path-filtered infra CI/CD with dual OIDC roles + prod approval gate (T18)
Add the /infra half of the combined-repo pipeline (the Python agent keeps ci.yml):
- ci-infra.yml — PR check on infra/** : tsc + jest + cdk synth via the org
reusable ci-typescript-cdk.yaml (working-directory: infra).
- cd-infra.yml — push to dev/main on infra/** (or dispatch):
* job 'ci' (reusable) is the CI-green precondition (deploy needs: ci).
* deploy-dev (ref=dev, NO environment) → cdk deploy OpenSweDevStack,
assuming githubdeploy-open-swe-infra-dev (OIDC sub ref:refs/heads/dev). AUTO.
* deploy-prod (ref=main, environment: prod) → cdk deploy OpenSweProdStack,
assuming githubdeploy-open-swe-infra-prod (OIDC sub environment:prod). The
'prod' Environment's required reviewer is the manual-approval gate.
Deliberately self-contained (NOT the reusable cd-cdk.yaml) because that runs
'cdk deploy --all' — from a single-env push it would deploy the other env + the
shared IAM stack, breaking the per-env boundary. CD targets one stack per env;
the shared open-swe-iam stack is human-gated (T6), never deployed by CD.
Infra CI is enforced at the DEPLOY boundary (deploy jobs need ci), not as a
branch-protection required check — path-filtering a required check would deadlock
app-only PRs. Documented in infra/README.md along with the post-T6 prerequisites
(repo vars AWS_DEPLOY_ROLE_INFRA_{DEV,PROD}; a 'prod' Environment w/ reviewer).
Not active until the IAM roles are applied (T6) — assuming a nonexistent role
just fails closed. App-side CD (S3 artifact + SSM) is T19.
* fix(infra): commit jest.config.js (was ignored by *.js → infra CI used Babel)
The infra/.gitignore *.js rule (for compiled CDK output) silently swept up the
hand-authored jest.config.js, so it was never committed. Local jest passed (file
present in the working tree) but CI's fresh checkout lacked it → jest fell back to
the default Babel transform → 'Cannot use import statement outside a module' on the
TypeScript test. Surfaced now because T18 is the first workflow to run infra jest
in CI. Negate the ignore for this one file and commit it.
Aligns the promotion target with the locked migration architecture
(main = PROD, dev = integration). Plain ref push is fast-forward-only;
branch protection on main rejects non-FF pushes, so a diverged main
fails loudly instead of force-moving prod.
* ci: promote dev to prod via fast-forward instead of force-pushing main
Repoints the daily promotion workflow to feed prod from the dev integration
trunk (prod <- dev) rather than force-pushing the upstream mirror (main).
Uses a fast-forward push so a release can never rewrite prod history.
* fix: sort imports in aegra_entry.py to unblock dev CI
The Sea Haven deployment commit left aegra_entry.py with an unsorted import
block (ruff I001), which fails the required 'Agent lint' check and blocks all
merges into dev. Apply the import-sort autofix.
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff
Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.
- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).
Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.
* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright
The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.
Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.
* test(open-swe): record Playwright trace + video on every E2E run
Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
* Run reviewer eval in a GitHub Action; make dashboard a read-only progress view
The dashboard launched the eval as a subprocess inside the serving deployment
worker, so a container recycle killed long runs and discarded results that had
already completed server-side. Move the harness to a workflow_dispatch Action
(run on prod). run_eval now publishes status/progress/log-tail to the LangGraph
store record the dashboard reads, so /admin/evals stays a live view; a killed
Action surfaces as failed via the stale-heartbeat reconcile.
* reviewer_eval workflow: pass inputs via env, no shell interpolation
Addresses the reviewer finding: workflow_dispatch string inputs were
interpolated into the run: block (limit unquoted), allowing shell injection in
a job holding LANGSMITH/ANTHROPIC keys. Pass inputs through env and reference
quoted "$VARS"; validate limit is numeric and build its flag in bash.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Adds a scheduled workflow that force-pushes main onto a prod branch at
08:00 UTC (midnight PST / 1am PDT). The LangGraph deployment will track
prod instead of main, so routine merges no longer trigger redeploys
that cancel in-flight 30-minute coding runs. workflow_dispatch is
retained for hotfixes that need to ship immediately.