Part of the domain-reorg adoption (build plan step C3): fork content,
upstream layout. Adds agent/graphs/{agent,analyzer,chat,reviewer,
scheduler}.py as thin re-export shims delegating to the existing fork
graph factories (agent.server/analyzer/chat/reviewer/scheduler), plus
agent/providers/__init__.py re-exporting agent.utils.model's
make_model/provider_model_kwargs/fallback_model_id_for surface —
verbatim upstream content, verified each import resolves against fork
modules with no name changes needed.
agent/runtime/{constants,execution}.py deviate from upstream's
verbatim shim bodies: rather than duplicating DEFAULT_LLM_MODEL_ID/
DEFAULT_LLM_MAX_TOKENS/DEFAULT_RECURSION_LIMIT/MODEL_CALL_RECURSION_LIMIT
and graph_loaded_for_execution's logic (upstream's shims assume
agent/server.py already had these extracted into runtime/ modules,
which is out of this commit's scope — server.py is untouched), they
import the fork's existing agent.server attributes directly. This
keeps the values/logic single-sourced instead of forking a second
copy that could drift.
agent/runtime/sandbox.py's delegation targets also differ from
upstream: fork's sandbox lifecycle helpers are private
(_get_cached_sandbox_backend, _configure_git_identity,
_recreate_sandbox in agent/server.py) since the fork's sync
4-case `__creating__` sentinel design (AGENTS.md) never made them
public. get_cached_sandbox_backend() also drops upstream's
caller-supplied `reconnect` callback parameter — fork's
_get_cached_sandbox_backend is a plain cache lookup; reconnection is
handled internally by ensure_sandbox_for_thread/
check_or_recreate_sandbox, not via a passed-in callback. No other
signature changes.
Added fork-only agent/graphs/ci_monitor.py (delegates to
agent.ci_monitor:get_ci_monitor) for symmetry, since upstream deleted
its ci-autofix cluster and has no equivalent shim. langgraph.json's
five stock graph entrypoints plus the fork-only ci_monitor now all
point at agent.graphs.<name>; http.app stays agent.webapp:app
(unchanged, per plan).
Deliberately NOT included (owned by build plan step C4, the FastAPI
split, gated on /sh-security-review): agent/api/{__init__,app,
health}.py, agent/webhooks/common.py, and the three
agent/webhooks/{github,linear,slack}_routes.py files. Those aren't
thin structural shims like the 21-file list implies in isolation —
they carry the fork's actual webhook dispatch/verify logic split out
of the still-monolithic webapp.py, which hasn't happened yet.
Building them now against upstream's placeholder content would ship
incomplete auth surface that C4 would just discard and redo.
Pinned oven-sh/setup-bun's bun-version to 1.3.14 (the version
installed locally; ui/ has no .bun-version file or package.json
engines/packageManager field pinning one) across all three CI jobs
that install bun, removing the latest-resolution flake.
Gates: ruff check + ruff format --check (clean), pytest --co -q
(1637 collected, no import errors), a direct import smoke-test of
every new module's public symbols, and a make dev boot check —
langgraph dev registered all six graphs (agent, reviewer, analyzer,
chat, scheduler, ci_monitor) each importing from agent.graphs.*, and
loaded the custom app from agent.webapp:app, before the process was
killed. (The subsequent lifespan failure, "DEFAULT_SANDBOX_SNAPSHOT_ID
must be set when SANDBOX_TYPE=langsmith", is expected with no .env
secrets configured in this environment and unrelated to this commit.)
* feat(models): re-add Fable 5 with admin disable toggle (port of upstream #1677)
* refactor(models): convert re-added Fable 5 to Bedrock model IDs
* fix(open-swe): correct Fable copy to describe provider data sharing, not ZDR
The ported admin toggle description and code comments described Fable 5 as
incompatible with Zero Data Retention. That is backwards: Fable 5 requires
the account to opt into Bedrock provider_data_share — prompts/completions are
retained and shared with Anthropic (up to 30 days, incl. human review). The
old UI copy would lead an admin to believe the opposite of what enabling the
toggle does. Reword the toggle description and the gate_fable_model /
team_settings comments accordingly. Still off by default. Refs #171.
* chore(deps): bump vite-tsconfig-paths to ^6.1.1 and jsdom to ^29.1.1
Reconcile two Dependabot PRs (#107, #108) into one branch with a
single bun install so package.json and bun.lock stay consistent.
vite-tsconfig-paths: ^5.1.4 -> ^6.1.1 (dependencies)
jsdom: ^27.4.0 -> ^29.1.1 (devDependencies)
* Add --frozen-lockfile to CI Checks
Normal bun install treats bun.lock as updatable. If package.json requests a requirement bun.lock doesn't satisfy, bun quietly rewrites the lock and moves on. The committed lock is never updated.
* chore: Add trailing newline on new last run
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
Adds a triage-ledger job running `make triage-check` so a ledger edit that
forgets to regenerate the markdown (make triage-render) fails CI instead of
drifting silently. Stdlib-only, no deps.
* chore: decommission self-hosted AWS LangGraph stack
Removes the now-dead self-host IaC and AWS-only CI/CD after destroying the
dev + prod CloudFormation stacks (open-swe-dev, open-swe-prod, open-swe-iam,
and the dev-exclusive CDKToolkit-oswedev bootstrap) in account 328440206208,
us-east-1. The deployment is now managed (LangGraph Cloud + Vercel).
- remove infra/ (CDK app: app + IAM stacks, constructs, aspects, tests)
- remove deploy/ami (Packer AMI build) and deploy/seahaven (boot/config
scripts, DEPLOYMENT/ROTATION runbooks)
- remove AWS-only workflows: cd-infra, ci-infra, build-artifacts, rollback
- README: rewrite the Deployment section to the managed LangGraph Cloud +
Vercel view; drop dead links to infra/ and deploy/seahaven
Preserved: the shared default CDKToolkit bootstrap and promote-dev-to-prod.yml.
The RETAIN'd Secrets Manager shells and open-swe-<env>-assets S3 buckets
survive cdk destroy by design (orphaned) and need a separate deliberate cleanup.
* chore: clean up dangling references left by the AWS decommission
Folds in the FIX-level items from the #64 review gates (GPT-4.1 cross-review +
/sh-security-review), none of which were blockers:
- delete orphaned .github/scripts/{package-artifacts,publish-and-deploy,roll-box,
rollback}.sh — their only callers were the removed AWS deploy workflows
- drop the deleted /infra dir from dependabot.yml npm directories (was producing
a recurring Dependabot config error)
- remove the stale OSWE-IAC-SECRETS-LIST-01 suppression (referenced the deleted
infra/lib/constructs/instance-role.ts)
- repoint the README promotion link to promote-to-main.yml (renamed in #63)
The promote-dev-to-prod.yml comment in check-dev-green.sh is intentionally left
to #63, which rewrites that same line.
* Align workflows with Sea Haven CI/CD handbook
Bring the workflow suite in line with the handbook: bump
actions/checkout to v7 (Node 24 runtime, already standardized),
kebab-case the two snake_case workflow filenames, and add the
org-standard Labeler caller and Dependency Review gate so vulnerable
or disallowed-license deps and unlabeled PRs are caught automatically.
File renames only — job/check display names are unchanged, so the
promotion gate's REQUIRED_CHECKS and branch-protection required
checks are unaffected.
Refs: INFRA-115
* Drop Agent prefix from CI workflow + job names
The handbook names workflows for what they do (CI, Deploy, Labeler),
not the component they run, matching .github and afterhours-shift-manager.
Rename the suite to CI and its jobs to Lint / Format check / Unit tests,
and keep the promotion gate's REQUIRED_CHECKS in sync.
Refs: INFRA-115
---------
Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
T20 CD safety nets. Two gaps closed before the first real prod deploy:
1. Promotion gate. promote_dev_to_prod.yml previously fast-forwarded
dev→main unconditionally. It now hard-gates on check-dev-green.sh:
every check-run on the dev HEAD must be completed+passing AND the
Agent CI suite (lint/format/unit/E2E) must be present+success, or the
promotion blocks (fails safe on a missing/renamed check). The promote
run excludes its OWN check-run by run-id (unforgeable), never by the
mutable name "promote", so a colliding red check cannot hide. Fields
are read with a 0x1F separator so an empty conclusion (every
in_progress check) cannot shift columns. ci.yml now also runs on
push:dev so dev HEAD actually carries that signal (a PR check alone
can be admin-merged past).
2. Rollback + last-good. publish-and-deploy.sh advances
releases/last-good/ only after a successful roll (deploy.sh gates on
`systemctl is-active`), and makes releases/latest/ transactional —
reverting to the prior release if the roll fails so a replaced box
never self-deploys a broken release. New rollback.yml + rollback.sh
re-point latest at last-good (or an explicit sha) and re-fire the
deploy; prod is gated by the `prod` Environment approval, same as a
deploy. The shared fire/wait/aggregate-gate logic is factored into
roll-box.sh (used by both forward and backward rolls).
Least-privilege: drop the unused s3:DeleteObject from the app deploy
role — publish/rollback/deploy only Get+Put (S3-to-S3 copy), and the
rollback fallback now depends on immutable release history staying
intact. Lifecycle expiry (not CI) handles old-version cleanup.
Gate logic unit-tested (7 cases + jq round-trip). IAM change +
release-safety control cross-reviewed by GPT-4.1: APPROVE, no blocks.
Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi
* fix(infra): grant instance role BatchGetSecretValue + ListSecrets for .env materialization
fetch-config.sh materializes the box's .env via
`secretsmanager batch-get-secret-value --filters Key=name,Values=open-swe-<env>/`,
but the instance role only granted GetSecretValue/DescribeSecret. BatchGetSecretValue
is a distinct IAM action, so the call was AccessDenied and open-swe.service
crash-looped (no .env written -> ExecStartPre exit 1).
- Add secretsmanager:BatchGetSecretValue to the prefix-scoped ReadSecrets statement.
- Add secretsmanager:ListSecrets on * (required by the name-prefix filtered batch
call; the API has no resource-level scoping for the list action — fits the role's
stated exception). Secret VALUES stay prefix-scoped; only names are enumerable.
Reviews: GPT-4.1 IAM cross-review BLOCK=none; /sh-security-review iac-iam one LOW
metadata residual (no critical/high), recorded as OSWE-IAC-SECRETS-LIST-01.
Refs T7/T19 dev bring-up.
* ci: lift Node heap cap for Playwright E2E build (vite OOM)
The E2E job's Playwright globalSetup runs the real `bun run build`, whose vite
bundle exceeds Node's default ~2 GB heap and OOMs (JavaScript heap out of memory) —
the same failure fixed for build-artifacts.yml in #19. Set
NODE_OPTIONS=--max-old-space-size=8192 on the Run E2E step.
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff
Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.
- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).
Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.
* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright
The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.
Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.
* test(open-swe): record Playwright trace + video on every E2E run
Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.