mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 10:23:14 +00:00
1047 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c9e6df76d1
|
chore(deps): bump actions/dependency-review-action from 4 to 5 (#67)
Bumps [actions/dependency-review-action](https://github.com/actions/dependency-review-action) from 4 to 5. - [Release notes](https://github.com/actions/dependency-review-action/releases) - [Commits](https://github.com/actions/dependency-review-action/compare/v4...v5) --- updated-dependencies: - dependency-name: actions/dependency-review-action dependency-version: '5' dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
f9b02a64d6
|
chore(deps): bump actions/setup-node from 4 to 6 (#66)
Bumps [actions/setup-node](https://github.com/actions/setup-node) from 4 to 6. - [Release notes](https://github.com/actions/setup-node/releases) - [Commits](https://github.com/actions/setup-node/compare/v4...v6) --- updated-dependencies: - dependency-name: actions/setup-node dependency-version: '6' dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
2b01652754
|
refactor: adopt modular webhook architecture (#1621) + port fork customizations (#85)
* Adopt upstream modular webhook skeleton (#1621) Apply the durable-interrupt-dispatch refactor: split the monolithic webapp.py into a thin routing layer plus per-source handlers in webhooks/{github,slack,linear}.py, and add completion.py, dispatch.py, and reconcile.py. Reconcile fork divergence by keeping the Bedrock/ Fireworks cross-provider fallback, the no-agent-attribution prompt policy, the dashboard-handoff re-export, and the Slack channel-info cache. ci_autofix is restored on the new dispatch model in a later commit. Refs: #80 * Port fork webhook security delta onto modular handlers Re-apply the fork's security customizations that #1621 did not carry: Linear webhook replay protection (freshness window on the signed webhookTimestamp), per-repo token-cache binding threaded through the thread token resolvers, the INTERNAL_BOT_LOGINS self-check in the review-finding-reply path, and a user-mapping cache refresh before email resolution on the issue and PR-comment paths (multi-replica staleness). Existing fork security tests pass unchanged. Refs: #80 * Restore CI auto-fix on the modular dispatch model Bring back ci_autofix.py and the ci_monitor graph that #1621 deleted, re-wiring the fork's security-reviewed PR-babysitting onto the new structure: the CI-event, autofix-toggle, and review-feedback handlers move into webhooks/github.py and the github_webhook router re-gains the check_run/check_suite/workflow_run/status routing plus the autofix command and actionable-review branches. Auto-fix runs now dispatch through dispatch_agent_run (durability + completion webhook) while keeping the deliberate batch-while-busy skip-rule via get_thread_active_status. Restore langgraph.json's ci_monitor entry and the fork autofix tests (dispatch mock + import paths re-pointed). Refs: #80 * Reformat and update docs for the modular webhook split Point CLAUDE.md and deploy/MIGRATION.md at the new webhooks/ modules and the dispatch/completion/reconcile contract, and mark the user-mapping cache-refresh fix as applied on the GitHub handlers. Refs: #80 * Restore reject backstop for autofix dispatch A burst of near-simultaneous CI events for one head SHA can slip past the busy-check before the dedupe SHA is recorded, so dispatch the autofix path with multitask_strategy=reject (dev's prior platform default) to drop duplicate concurrent creates instead of letting them interrupt each other. Also make the completion failure-reply dedup claim-then-post and drop the unreachable interrupted branch. --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
9f7a1cc481
|
feat: post scheduled-run report to a configured Slack channel (#83)
* Post scheduled-run reports to a Slack channel Scheduled runs previously had no source channel and finished silently in the dashboard. Allow an automation to post its final report to a configured Slack channel as the bot, reusing existing Slack plumbing instead of the deferred run-completion webhook. Refs: #82 * Reconcile Slack report feature with dev merge Dev refactored slack_thread_reply to async and already added post_slack_top_level_message_with_ts; drop the duplicate definition and await the tool in the feature's tests. Refs: #82 * Harden scheduled Slack report channel posting The no-thread_ts top-level path fired for every slack_thread_reply call during a scheduled run, spraying disconnected messages and dead interactive buttons into the report channel. Cap top-level posts at one per run and drop options/plan_approval blocks in that mode, so the mechanism (not just the prompt) enforces a single clean report. Also tighten the channel-ID regex to require a leading letter and document why top-level posts store no run mapping. Refs: #82 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
1f060f2a1d
|
chore: sync upstream/main, defer #1621 modular webhooks (#81)
* chore: bake sfw binary into sandbox image (#1611) sfw only ships a launcher that fetches its real binary at first run and does a daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in the sandbox (restricted egress; the proxy injects the GitHub App installation token, which lacks access to that repo), so `sfw yarn install` errors with "could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline. * feat: editable plan mode + fix review-plan banner overlap (#1610) * feat: editable plan mode + fix review-plan banner overlap Lets the thread owner edit the plan markdown by hand from the plan-review page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id} endpoint that re-publishes the plan and mirrors it into the sandbox plan.md, so approve hands the edited plan to the agent as the source of truth. Also fixes the collapsed git-panel's floating expand button covering the "Review plan ->" banner by reserving space for it. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: abort plan approval when the published plan read fails get_plan_content() swallowed store errors and returned None, so a transient failure during approve would still mark the plan approved and dispatch the generic fallback text — silently dropping an owner's edited plan. Read the plan strictly (raise_on_error=True) so approval aborts instead, matching the comment read. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show message timestamps (#1609) * feat: show message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: suppress fallback message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: stable message + tool-call hover timestamps Stamp a stable client-side arrival time per message and tool call (keyed by id, persisted to localStorage). Messages render the timestamp inline; tool rows reveal a dim timestamp chip on hover. Real backend created_at still takes precedence when present. * fix: hide client-stamped message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add PR trace resolution (#1612) * feat: add PR trace resolution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: inject reviewer trace context as JSON Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review on PR trace resolution Use the documented LangSmith metadata filter syntax (and(eq(metadata_key,...), eq(metadata_value,...))) instead of has(metadata, '{...}'), which does not match runs — _list_thread_runs was silently returning nothing. Bound full-text searches to a 90-day window so they don't hit LangSmith's large-window rate limit. Also folds in the best-effort branch->head-sha resolver (dropping the weighted scoring/threshold + repo/file evidence + GitHub hydration), sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint. The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session were removed; resolution now runs deterministically from the trusted run config with no model-controlled pr_url or thread_id. * fix: scope branch trace search to the repo Branch names like fix-tests aren't unique across repos (or older PRs) in a shared tracing project, so an unscoped branch hit could resolve to an unrelated thread and write its runs into the reviewer sandbox. Require the repo slug to co-occur with the branch in matched runs; the full head SHA stays unscoped since it is globally unique. Addresses open-swe review on PR #1612. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: include plan links in PR descriptions (#1613) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: gate workflow pushes with approval (#1614) * feat: gate workflow pushes with approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve proxy refresh test compatibility Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: bind workflow approvals to pushed ref Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: recover thread work as patch (#1615) * feat: recover thread work as patch Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: search sandbox cwd for recovery patches Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: omit plan link in PR description when no plan exists (#1618) Plan links in PR descriptions were always built from the thread id, so runs that never produced a plan linked to an empty plan-review page. Now the plan content store is consulted first; the link is only added when a plan with non-empty markdown actually exists. A transient store failure degrades gracefully (no link) rather than blocking PR creation. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add filter & grouping menu to agents threads sidebar (#1617) Add a Cursor-style control to the agents sidebar that groups (None/Date/ Status/Project), filters (ownership, status, source, pull request, model, repo, include-resolved), and compacts the threads list. All client-side over already-fetched sidebar threads; preferences persist in localStorage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: update langsmith sdk to 0.9.3 (#1616) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: clickable shared PR header in git panel and reviews (#1620) * feat: clickable shared PR header in git panel and reviews Replace the standalone "View PR" button in the agent git panel with a clickable PR title, matching the reviews view. Extract a shared PrHeader component reused by both the git panel and the review main body. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: drop PrHeader wrapper, use shared component directly The review-side PrHeader was just a thin adapter mapping detail -> the shared component's props. Inline it at the call site and use the shared PrHeader directly so there's a single component. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: durable interrupt dispatch + completion webhook (#1621) * wip(rebuild): core reliability spine - remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring) - dispatch core: agent/dispatch.py with multitask_strategy=interrupt + durability=sync + completion webhook; reroute all webhook + plan triggers; drop the racy in-process lock + is_thread_active busy-check - completion webhook: agent/completion.py + /webhooks/run-complete loopback route for failure/timeout replies (idempotent) Co-authored-by: open-swe[bot] * feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning Parallel batch on top of the reliability spine: - async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden the IP check to 'not is_global' (+ IPv4-mapped unwrap) - reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list -> cancel_many), wired into the scheduler graph via task='reconcile' - shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare httpx.AsyncClient() across utils/dashboard/webapp/middleware - run budget: MODEL_CALL_RECURSION_LIMIT 5000->250 - fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8) - drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls) - confirm tool-result eviction + summarization auto-wired via backend - slim system prompt ~8% (full harness-profile rewrite deferred) Co-authored-by: open-swe[bot] * feat(rebuild): harness-profile prompt + split webhooks out of webapp - prompt.py: own the system prompt via a registered harness profile (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that share it stay safe), registered across all 4 providers; per-thread values stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped ALL-CAPS markers. - webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the routes + tests; moved handlers reach shared helpers via the webapp namespace to preserve the test suite's monkeypatch targets. Full suite: 1168 passing, lint clean. Co-authored-by: open-swe[bot] * Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks Reverts the 250 cap from the run-budget change — long-running tasks legitimately need many model calls. The notify_step_limit_reached safety net still fires if a run does hit the cap, so runs end with a signal either way. Co-authored-by: open-swe[bot] * fix: address PR review (auth, SSRF, interrupted status, redirect headers) - completion.py: drop `interrupted` from failure statuses — with multitask_strategy=interrupt a follow-up ends the prior run as interrupted, which is healthy, not a failure to report. [open-swe] - /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest. [corridor-security] - SSRF: extract the URL validator to agent/utils/url_safety.py and apply it before server-side image fetches in multimodal.fetch_image_block. [corridor-security] - http_request: preserve caller headers/extensions across redirect hops instead of dropping them on the first hop. [open-swe] Co-authored-by: open-swe[bot] * chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo) Co-authored-by: open-swe[bot] * fix: fail closed on run-complete webhook auth when secret unset Corridor follow-up: verify_run_complete_token returns False (not True) when RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never unauthenticated. Logs a startup warning when the secret is absent, and dispatch skips registering the webhook when there's no secret (no rejected callbacks). Co-authored-by: open-swe[bot] --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: restore forced tool call to prevent premature run stops (#1622) Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task. Shipping to test whether it fixes runs that stop halfway through. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619) Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1. - [Release notes](https://github.com/langchain-ai/langgraph/releases) - [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1) --- updated-dependencies: - dependency-name: langgraph-checkpoint dependency-version: 4.1.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * fix: post reviewer resolution notes verbatim (#1624) * fix: post reviewer resolution notes verbatim Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: stabilize dashboard follow-up e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make e2e attribution marker durable Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: only echo found e2e attribution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: check live dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625) Installs hung when prefixed with sfw inside the sandbox (trace 019f0608 stalled on a pending `sfw npm install` execute, never returned). Strip the Socket Firewall guidance from the agent and reviewer prompts so installs run through the project's package manager directly. sfw stays in the Docker image; nothing invokes it now. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make plan view mobile friendly (#1636) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: fall back to vision model for image threads (#1626) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: surface Slack thread errors (#1627) * fix: surface Slack thread errors Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: don't set failure_reply_posted on Slack preprocessing errors The preprocessing error handler was setting failure_reply_posted=True, the same idempotency flag handle_run_completion checks to suppress duplicate run-failure replies. Since preprocessing failures happen before any run exists but the flag persists on the thread, a subsequent run failure on the same thread would be silently ignored. The preprocessing handler already posts its own Slack reply, so the run-completion idempotency flag should not be set here. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: avoid recapping Slack replies (#1629) * chore: avoid recapping Slack replies Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: simplify Slack reply prompt wording Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: update Slack trace reply on web handoff (#1630) * fix: update Slack trace reply on web handoff Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: trigger web handoff on dashboard starts Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: format web handoff as contextual fragment Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve trace_message_ts when overwriting Slack run mapping When store_slack_run_mapping is called without trace_message_ts (e.g. on follow-up Slack mentions), it was unconditionally overwriting the thread-level mapping and clobbering the timestamp captured from the initial trace reply. After that, _notify_slack_web_handoff could not find the original message, so a subsequent move to Web silently skipped the Slack trace update. Now, when trace_message_ts is not passed, the existing thread mapping is read first and its trace_message_ts is preserved. * style: ruff format --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643) * fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures shiki lazy-imports a grammar per language and these libs only live inside lazy route components, so Vite's startup scanner never sees them. They get discovered on first thread navigation, triggering a dep re-optimize + force-reload that aborts the in-flight route-chunk import, surfacing as "Failed to fetch dynamically imported module: .../$threadId.tsx". Pre-bundle them (and the github themes + common code-block languages) via optimizeDeps.include so the optimize happens once at startup. Dev-only; production bundles are unaffected. * fix: pre-bundle canonical shiki docker/make langs instead of aliases --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show queued dashboard follow-ups (#1631) * feat: show queued dashboard follow-ups Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: de-dupe queued follow-ups while streaming --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: notify Slack on plan approval (#1632) * feat: notify Slack on plan approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: post Slack approval notice after successful dispatch Move the _maybe_post_plan_approved_to_slack call until after _dispatch_followup succeeds so the Slack thread is not told implementation is beginning before the LangGraph run is created. Addresses PR review comment. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: include Slack channel context in prompts (#1633) Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> * chore: keep plan guidance high-level (#1634) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: publish plans from sandbox files (#1635) * feat: publish plans from sandbox files Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: avoid fixed plan filenames Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: virtualize local sandbox file paths Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve plan_file_path across set_plan_status set_plan_status was rewriting the content record with only markdown and status, dropping plan_file_path. After a reject, the owner's dashboard edit would mirror to a different file than the agent's original, and the next save_plan could republish the stale file. Preserve plan_file_path when updating status. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: return to thread after plan approval (#1637) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Slack breakout thread tool (#1638) * feat: add Slack breakout thread tool Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: make fake LLM scripts declarative Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: exclude slack_start_new_thread from plan mode The breakout tool can dispatch a fresh agent run that starts outside the current plan-mode state, bypassing the approval flow. Add it to PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating tools while planning. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: require bun for ui agent work (#1639) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: request actions read for sandbox logs (#1642) * fix: request actions read for sandbox logs Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: restore actions:read scope after workflow push After an approved workflow push, the guard was restoring the proxy with BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which includes actions: read) and fall back to BASE if the install hasn't granted Actions read — mirroring the pattern in _create_sandbox_with_proxy. Addresses review comment on PR #1642. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: widen split review diffs (#1647) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: install missing deps before verification (#1646) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm (#1645) * chore: require pnpm for ui agent work Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * ci: use corepack for ui pnpm e2e build Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Sonnet 5 to model picker (#1651) * chore: update Sonnet examples to Sonnet 5 Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: add Sonnet 5 to model picker Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * Remove dead breakout-thread e2e scenario after dropping the tool The merge resolution deferred upstream's Slack breakout-thread tool (slack_start_new_thread, #1638) since it depends on the #1621 dispatch module, but the e2e harness still scripted it. Removing the tool name from fake_llm.py's _tool_step call left a malformed scenario, crashing the langgraph-dev web server at import (TypeError: _tool_step() missing 'call_id') and failing Playwright E2E. Drop the "breakout" script scenario, its _is_breakout_request helper + ScriptRule, and the corresponding full_flow.spec.ts test. * Revert upstream pnpm switch; keep bun for the UI build The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/ global-setup.ts and ui/package.json, but our fork builds the UI with bun (vercel.json + the E2E workflow's setup-bun). That left the Playwright globalSetup running `corepack pnpm install --frozen-lockfile` with no pnpm-lock.yaml, failing E2E at UI build time. Revert global-setup.ts and ui/package.json to the dev (bun) baseline, drop the merge-added ui/pnpm-lock.yaml, and remove the re-added ui/AGENTS.md (our fork had deleted it). * Align plan-review e2e + UI with the HEAD (pre-#1635) backend The merge left a split plan vertical: the backend save_plan/plan_api are HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/ #1637 per #80), but the plan UI and e2e harness were upstream's. The fake_llm scenario called save_plan(plan_file_path=...) — upstream's file-based #1635 contract — while HEAD save_plan takes plan_markdown, so the plan never saved and PlanReview never rendered (E2E failure on the plan-review locator). Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts / $threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the whole plan flow (save -> render -> approve -> implement) is consistent with the HEAD backend. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com> Co-authored-by: Mukil Loganathan <mukil@langchain.dev> |
||
|
|
eb98ff4c30
|
feat: add 3 verified Fireworks models to selectable set (#79)
* Add 10 Fireworks models to selectable set Surface additional Fireworks-served models in the profile editor so they can be chosen per-thread, per-profile, and as team defaults. Each entry carries its recommended efforts and image support; only MiniMax M3 is multimodal. Refs: #78 * Suppress reasoning_effort on non-reasoning models Instruct-only Fireworks ids (kimi-k2-instruct-0905, mistral-large-3-fp8, qwen3-30b-a3b-instruct-2507) don't reason, so sending reasoning_effort either 400s (unusable at default effort) or is a silent no-op. Add a per-model reasoning flag (default True) and omit the param entirely for ids marked non-reasoning. Refs: #78 * Gate out 7 undeployed Fireworks models Account serverless probe returned 404 for 7 of the 10 proposed ids, so only minimax-m3, gpt-oss-120b, and deepseek-v4-flash are callable. Keep those 3 and drop the rest. All 3 survivors are reasoning-capable, so the per-model reasoning-effort suppression added earlier is no longer needed and is reverted. Refs: #78 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
860ce93ad7
|
docs: record final managed-deployment topology in MIGRATION.md (#77)
Add an authoritative "Final, verified topology" section (1a) and reconcile the phased plan with the executed end-state, superseding the spike-era values throughout. Captures: two LangGraph Cloud deployments (dev/prod, one LangSmith workspace) with their final URL hashes; one Vercel project open-swe-prod with production + custom dev environments and per-env LANGGRAPH_BACKEND_URL; the Nitro routeRules proxy mechanism (PR #76, superseding the vercel.json rewrite and PR #75's build-vercel-output.mjs); two dev/prod-isolated GitHub Apps; per-deployment user stores; the Bedrock IAM users; and the seven hard-won operational gotchas. Refs: #65 #74 #76 |
||
|
|
8441bbb2d8
|
docs(migration): record the two Bedrock IAM users + scoped invoke policy (#74)
Document the live execution of MIGRATION.md §10.1 (Bedrock auth via static keys): the customer-managed least-privilege policy open-swe-bedrock-invoke and the open-swe-dev-bedrock / open-swe-prod-bedrock IAM users. Captures the static-key deviation rationale (managed LangGraph Cloud cannot assume a role) and that both mandatory gates (GPT-4.1 IAM cross-review, /sh-security-review) passed with no critical/high. |
||
|
|
49fd48d32f
|
chore(deps): bump the minor-and-patch group across 2 directories with 4 updates (#69)
Bumps the minor-and-patch group with 1 update in the /tests/e2e directory: [@playwright/test](https://github.com/microsoft/playwright). Bumps the minor-and-patch group with 3 updates in the /ui directory: [monaco-editor](https://github.com/microsoft/monaco-editor), [@tanstack/devtools-vite](https://github.com/TanStack/devtools/tree/HEAD/packages/devtools-vite) and [prettier-plugin-tailwindcss](https://github.com/tailwindlabs/prettier-plugin-tailwindcss). Updates `@playwright/test` from 1.61.0 to 1.61.1 - [Release notes](https://github.com/microsoft/playwright/releases) - [Commits](https://github.com/microsoft/playwright/compare/v1.61.0...v1.61.1) Updates `monaco-editor` from 0.52.2 to 0.55.1 - [Release notes](https://github.com/microsoft/monaco-editor/releases) - [Changelog](https://github.com/microsoft/monaco-editor/blob/main/CHANGELOG.md) - [Commits](https://github.com/microsoft/monaco-editor/compare/v0.52.2...v0.55.1) Updates `@tanstack/devtools-vite` from 0.6.1 to 0.8.1 - [Release notes](https://github.com/TanStack/devtools/releases) - [Changelog](https://github.com/TanStack/devtools/blob/main/packages/devtools-vite/CHANGELOG.md) - [Commits](https://github.com/TanStack/devtools/commits/@tanstack/devtools-vite@0.8.1/packages/devtools-vite) Updates `prettier-plugin-tailwindcss` from 0.7.4 to 0.8.0 - [Release notes](https://github.com/tailwindlabs/prettier-plugin-tailwindcss/releases) - [Changelog](https://github.com/tailwindlabs/prettier-plugin-tailwindcss/blob/main/CHANGELOG.md) - [Commits](https://github.com/tailwindlabs/prettier-plugin-tailwindcss/compare/v0.7.4...v0.8.0) --- updated-dependencies: - dependency-name: "@playwright/test" dependency-version: 1.61.1 dependency-type: direct:development update-type: version-update:semver-patch dependency-group: minor-and-patch - dependency-name: "@tanstack/devtools-vite" dependency-version: 0.8.1 dependency-type: direct:development update-type: version-update:semver-minor dependency-group: minor-and-patch - dependency-name: monaco-editor dependency-version: 0.55.1 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: minor-and-patch - dependency-name: prettier-plugin-tailwindcss dependency-version: 0.8.0 dependency-type: direct:development update-type: version-update:semver-minor dependency-group: minor-and-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
5caaf885a4
|
chore(deps): bump fireworks-ai from 1.2.0a75 to 1.2.0a85 (#37)
Bumps [fireworks-ai](https://fireworks.ai) from 1.2.0a75 to 1.2.0a85. --- updated-dependencies: - dependency-name: fireworks-ai dependency-version: 1.2.0a85 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
c75e76c05e
|
fix(ui): drive Vercel dashboard-API proxy through Nitro routeRules (#76)
The Vercel deploy 404'd at `/` because PR #75's `vercel-build` script (`scripts/build-vercel-output.mjs`) ran `rm -rf .vercel/output` and rebuilt it from `.output/public`. On Vercel CI, Nitro's Vercel preset auto-activates (from the VERCEL env var) and emits the Build Output API layout to `.vercel/output` itself during `vite build`; the script then clobbered that correct output with a static-only config that could not resolve the SPA `_shell.html` fallback or the server functions, so production returned `404: NOT_FOUND` even though the build was READY. Stop fighting Nitro and drive the proxy through it instead: - Remove `scripts/build-vercel-output.mjs` and the `vercel-build` script; restore the plain `vite build` for both local and Vercel builds. - Add an env-driven Nitro `routeRules` proxy in `vite.config.ts`. Nitro's Vercel preset compiles a plain external-URL `proxy` rule into a CDN-level rewrite in `.vercel/output/config.json` at build time, reading LANGGRAPH_BACKEND_URL (the per-project Vercel env var). Proxy, not redirect, so the osw_session cookie stays first-party (same-origin). The build fails loudly if LANGGRAPH_BACKEND_URL is missing on Vercel; the rule is omitted for plain local/off-Vercel builds (which use the E2E_HARNESS mock proxy). - `vercel.json`: `buildCommand` back to `bun run build`; drop the `.output/public` outputDirectory so Vercel serves Nitro's `.vercel/output`. Validated with `LANGGRAPH_BACKEND_URL=... NITRO_PRESET=vercel bun run build`: the generated `.vercel/output/config.json` contains the `/dashboard/api/(.*)` -> backend rewrite ahead of `handle: filesystem` and the SPA catch-all, `_shell.html` and the 699 hashed assets are emitted, and typecheck passes. Supersedes the broken approach in #75. |
||
|
|
076e20d7e6
|
feat(ui): derive Vercel proxy backend from per-project env var (#75)
Generate the /dashboard/api/* proxy destination at build time from LANGGRAPH_BACKEND_URL via the Vercel Build Output API, instead of hardcoding it in ui/vercel.json. This lets the dev and main branches stay byte-identical (required by the fast-forward-only prod promotion) while each Vercel project proxies to its own LangGraph backend purely via its own env var — so both the dev and prod dashboards can be git-linked. - ui/vercel.json: drop the hardcoded rewrites; buildCommand -> bun run vercel-build - ui/package.json: add vercel-build = vite build && node scripts/build-vercel-output.mjs - ui/scripts/build-vercel-output.mjs: emit .vercel/output/config.json with [ proxy /dashboard/api/* -> $LANGGRAPH_BACKEND_URL, filesystem, SPA fallback ]; throws if LANGGRAPH_BACKEND_URL is unset so a misconfig fails the build. Keeps the app same-origin (no CORS/auth change); preserves the _shell.html SPA fallback. Set LANGGRAPH_BACKEND_URL per Vercel project (All Environments). |
||
|
|
a30ce2ab40
|
feat: managed LangGraph Cloud + Vercel migration (PR2 — code fixes + docs) (#65)
* fix(dashboard): managed-cloud OAuth hardening + admin user-mapping endpoint Prepare the dashboard backend for the managed LangGraph Cloud + Vercel runtime, where the API is HTTPS and cross-site from the UI. - OAuth redirect_uri (#2): coerce a schemeless DASHBOARD_API_BASE_URL to https:// in _api_base_url() so GitHub stops rejecting login with "redirect_uri not associated with this application". _cookie_security() now treats a schemeless (managed) value as Secure; SameSite=None too, consistent with the coerced scheme. - OAuth state cookie (#3): document that osw_oauth_state is host-only by design (a Domain cookie is unsafe across *.vercel.app, a public suffix), so login must always start on the stable alias to avoid "oauth state mismatch". Operational contract; no behavioral change. - Admin user mappings (#4): add POST /admin/user-mappings so an admin can set the github_login -> work_email link from the dashboard instead of a raw Store write. New "admin" MappingSource provenance value. * fix(webapp): refresh user-mapping cache on GitHub webhook paths On managed LangGraph Cloud the backend runs multiple replicas, so the per-process GitHub<->work-email mapping cache can be stale on the replica handling a webhook (a mapping created on another replica is invisible until refresh). process_github_pr_comment and process_github_issue now refresh the cache from the durable Store before resolving the author's email, matching the existing Slack mention path (process_slack_mention). * perf(webapp): defer deepagents import to speed custom-app cold start The custom FastAPI app (agent.webapp:app, the langgraph.json http.app) pulled deepagents -> langchain_anthropic -> anthropic into its import graph via dashboard.routes, only to build skill/chat seed files. Defer those create_file_data imports into the functions that use them. Removes deepagents/langchain_anthropic/anthropic from app import entirely and roughly halves module-import wall time (~0.6-0.8s -> ~0.35s warm; larger cold-start saving since native anthropic init is skipped). Behavior identical. (reviewer_diff already imports deepagents under TYPE_CHECKING.) * feat(ui): set work_email user mappings from the admin dashboard Add an "Add / update" form to the admin User mappings section and the adminUpsertUserMapping API client method, wiring the new POST /admin/user-mappings endpoint. Admins can now create or update a github_login -> work_email mapping directly instead of waiting for the user to self-connect Slack. * docs: document managed LangGraph Cloud + Vercel deployment - INSTALLATION §10: add the managed production env triad (LANGGRAPH_URL, DASHBOARD_BASE_URL + DASHBOARD_API_BASE_URL with https://, empty VITE_DASHBOARD_API_BASE_URL for same-origin), the stable-alias login and vercel.json stable-deployment-URL requirements, multi-replica cache note, plus redirect_uri-scheme and oauth-state-mismatch troubleshooting. Refresh the langgraph.json snippet to all six graphs. - README: reframe deployment around the managed migration; link the plan. - deploy/MIGRATION.md: import the self-hosted -> managed migration plan. |
||
|
|
7f60324f0c
|
chore: decommission self-hosted AWS LangGraph stack (#64)
* chore: decommission self-hosted AWS LangGraph stack Removes the now-dead self-host IaC and AWS-only CI/CD after destroying the dev + prod CloudFormation stacks (open-swe-dev, open-swe-prod, open-swe-iam, and the dev-exclusive CDKToolkit-oswedev bootstrap) in account 328440206208, us-east-1. The deployment is now managed (LangGraph Cloud + Vercel). - remove infra/ (CDK app: app + IAM stacks, constructs, aspects, tests) - remove deploy/ami (Packer AMI build) and deploy/seahaven (boot/config scripts, DEPLOYMENT/ROTATION runbooks) - remove AWS-only workflows: cd-infra, ci-infra, build-artifacts, rollback - README: rewrite the Deployment section to the managed LangGraph Cloud + Vercel view; drop dead links to infra/ and deploy/seahaven Preserved: the shared default CDKToolkit bootstrap and promote-dev-to-prod.yml. The RETAIN'd Secrets Manager shells and open-swe-<env>-assets S3 buckets survive cdk destroy by design (orphaned) and need a separate deliberate cleanup. * chore: clean up dangling references left by the AWS decommission Folds in the FIX-level items from the #64 review gates (GPT-4.1 cross-review + /sh-security-review), none of which were blockers: - delete orphaned .github/scripts/{package-artifacts,publish-and-deploy,roll-box, rollback}.sh — their only callers were the removed AWS deploy workflows - drop the deleted /infra dir from dependabot.yml npm directories (was producing a recurring Dependabot config error) - remove the stale OSWE-IAC-SECRETS-LIST-01 suppression (referenced the deleted infra/lib/constructs/instance-role.ts) - repoint the README promotion link to promote-to-main.yml (renamed in #63) The promote-dev-to-prod.yml comment in check-dev-green.sh is intentionally left to #63, which rewrites that same line. |
||
|
|
430a1cdff9
|
ci: re-home prod promotion into a gated promote-to-main workflow (PR1: managed-LGC migration) (#63)
* ci: re-home prod promotion into a gated promote-to-main workflow Migrate the prod-deploy gate to managed LangGraph Cloud (git-connected to `main`) + Vercel. Under managed, a push to `main` auto-deploys prod, so the dev -> main fast-forward IS the prod deploy trigger -- the bespoke AWS CD step is obsolete and already gone from this workflow. Re-home `promote-dev-to-prod.yml` -> `promote-to-main.yml`: - gate the promote job on the `prod` GitHub Environment (required reviewer amoussa1229), restoring the manual prod-approval that the retired AWS CD job used to carry; - drop the nightly auto-promote cron -- a scheduled auto-promotion conflicts with a manual approval gate now that the push deploys prod; promotion is workflow_dispatch only; - keep the dev-HEAD-fully-green precondition and the seahaven-promotion App fast-forward push (sole non-admin bypass actor on `main` ruleset 18238334). Update the companion check-dev-green.sh filename reference. * ci: refuse promote-to-main dispatch from any ref other than dev Defense-in-depth atop the already-pinned `ref: dev` checkout: workflow_dispatch runs the workflow definition from the launched ref, so reject a non-dev dispatch before the App token is minted. Surfaced by the GPT-4.1 cross-review of #63. |
||
|
|
a4ed19ba61
|
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)
Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.
- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module
* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8
The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.
Map profile effort to additional_model_request_fields:
{thinking: {type: adaptive, display: summarized},
output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.
* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids
Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
(set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.
Surfaced by the cross-family review + verified against deploy/.
* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip
From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
error code only, so the role ARN + account id in the raw botocore message never
reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
(Converse emits reasoning_content, not thinking) so the middleware is not a no-op
on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)
* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids
Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
each routed region (us-east-1/2, us-west-2). The model runs in the server process
on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
seed_store.sh's default via pick precedence, so the seed-script fix alone was
insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
ids to the Bedrock id (config.toml's model_id was an active, now-broken value).
AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.
* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)
Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
(eval judge only — Bedrock builder/reviewer auth via the host IAM role).
REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
|
||
|
|
8a9974c3c4
|
feat: author Slack/dashboard/schedule commits + PRs as the app by default (#57) (#60)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: default Slack/dashboard/schedule PRs + commits to the app identity (#57) Slack/dashboard/schedule runs now author PRs and run git/gh operations as the GitHub App seahaven-openswe[bot] by default (matching GitHub-issue runs), so the self-review 422 is impossible by construction rather than guarded in the prompt. A profile flag author_prs_as_user restores per-user attribution. - open_pull_request._resolve_pr_author_token + auth.resolve_github_token: default to the installation token for these sources; per-user only when opted in. - authorship: commit identity -> seahaven-openswe[bot] (numeric noreply; accepted Vercel-resolution risk, documented inline). - self-trigger safety: INTERNAL_BOT_LOGINS + webapp/reviewer_reconcile/reply markers recognize seahaven-openswe[bot] (bot-authored events are now ours). Supersedes the prompt-only guard in #58. * fix: author commits as the app bot in the default path (SH-IDSPLIT-01) Security review found the commit identity was NOT actually unified to the bot: resolve_triggering_user_identity got a 403 from the installation token and fell back to configurable['github_login'], so commits were still authored as the triggering user (commit=user, push+PR=bot — a three-way split that missed the stated goal). Now gate the triggering-user identity resolution on the same default-bot decision as the token: slack/dashboard/schedule default to the app bot identity unless author_prs_as_user is set. * docs(security): record AUTHZ-SLACK-BOT-DEFAULT-001 as an accepted residual (#59) Single-user deployment; bounded by App-on-pilot + ALLOWED_GITHUB_REPOS lock. Revisit (add a per-user gate) before expanding users or the App installation. |
||
|
|
134963647b
|
chore(security): suppress pre-existing history scanner false-positives (#56)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Adds repo-local suppressions for 5 verified-FP gitleaks findings that block pushes (forcing --no-verify), all in committed history / docs / CI fixtures: - .env.ci (dev-only e2e values, intentionally committed) - .env.example (placeholders) - .github/ci/fake_github_app_key.pem (throwaway CI test key) - INSTALLATION.md (example GITHUB_APP_PRIVATE_KEY .env block) - README.md (prose mis-matched by the generic-api-key heuristic) Repo-local (not machine-level) so they load in git worktrees too. Also drops the now-obsolete OSWE-IAC-AUDIT-01 suppression (B-1, fixed in #55). |
||
|
|
e9499e49b8
|
fix: isolate dev CDK deploys on their own bootstrap qualifier (B-1/OSWE-IAC-01) (#55)
* fix: isolate dev CDK deploys on their own bootstrap qualifier (B-1) Dev synthesizes against the oswedev qualifier and the dev infra deploy role is scoped to cdk-oswedev-* — it can no longer assume the default hnb659fds bootstrap roles whose admin cfn-exec-role deploys prod, closing the cross-env escalation (OSWE-IAC-01). Prod stays on the default qualifier. * test: assert per-env bootstrap qualifier isolation + document (B-1) |
||
|
|
a33aaec495
|
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54)
* fix: enforce a replay window on Linear webhooks (AUTHZ-001) verify_linear_signature accepted any correctly-signed body with no freshness check, so a captured request could be replayed indefinitely. Parse the signed webhookTimestamp (Unix ms) and reject requests outside a 60s window, failing closed when the field is missing or malformed — mirroring the Slack verifier. * fix: stop leaking upstream auth-error bodies into user comments get_github_token_for_user folded the raw upstream response text into the error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the full body server-side only and return a generic "GitHub auth failed (status <code>)". Also document the accepted shared-installation-token blast radius on the bot-token-only path (AUTHZ-003). * fix: bind sandbox and token caches to repo to prevent thread-id collision A PR head-branch name is attacker-controllable and get_thread_id_from_branch derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01). The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on thread_id alone, and a cached sandbox was reused after only an echo-ping, so a different repo's webhook could bind to another thread's sandbox or token. Without changing the persistent thread-id scheme: - Persist the bound repo (owner/name) in thread metadata on sandbox creation and refuse to reuse a sandbox whose bound repo does not match the current event (SandboxRepoMismatchError); the in-memory proxy also carries the binding. - Bind the GitHub-token cache entries to their repo and evict on a cross-repo read so a colliding thread_id cannot be served another repo's token. - Thread repo through the reviewer and the webhook token resolvers. * fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04) The instance role and the GitHub deploy app role granted s3:ListBucket on the whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts) only ever lists under releases/, so add a StringLike s3:prefix=releases/* condition. GetBucketLocation has no s3:prefix in its request context, so it moves to its own unconditioned statement. Also document the accepted F-2 cross-env existence-oracle residual on BatchGetSecretValue. * chore: suppress test-fixture credential false positive; document AUTHZ-002 Add a machine-level suppression for the fake Datadog key in the test_team_credentials encryption-roundtrip fixture (CWE-798, not a real credential). Clarify that the within-org thread-write path is intentional by design (AUTHZ-002) — comment only, no behavior change. * fix: casefold repo-binding keys to avoid spurious cross-repo mismatch GitHub owner/name are case-insensitive. Casefold the owner/name key on both the write (binding) and read (compare) sides — repo_cache_key and the metadata bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2). * fix: stop leaking upstream auth body in unexpected-result branch The 2xx-but-missing-token/url branch echoed the parsed upstream response body into the user-facing error. Return a generic message and log response_data server-side only, mirroring the existing HTTPStatusError fix (Gap 4). * fix: fail closed for unbound-legacy sandboxes and catch repo mismatch Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no recorded bound_repo (a pre-binding legacy thread, post-deploy) previously reconnected-and-served the sandbox to the current repo, then rebound it. Now fail closed: drop the stale id and recreate a fresh sandbox bound to this repo, logging a reconnect-with-missing-binding event. A sandbox is never served to a repo unless its binding is known and matches; new threads bind on first run unchanged. Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints, log it for alarming, and surface a clean sanitized error instead of letting an opaque deep-stack exception crash-loop the worker. * chore: suppress test-fixture credential false positive in token-TTL tests Add a machine-level suppression for the fake "ghp_secret" GitHub token used by the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and not a valid PAT; scoped to the unit test only. |
||
|
|
9444fd7677
|
docs: document live AWS prod deploy for Open SWE (#53)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Prod went live 2026-06-29 on self-hosted AWS EC2 behind the shared seahaven-com ALB, superseding the on-prem VM model the runbook described. - Rewrite deploy/seahaven/DEPLOYMENT.md as the canonical end-to-end runbook: infra CD (CDK stacks + OIDC roles + prod approval gate), config seeding (put-config.sh, the 13 boot-required prod vars, fetch-config fail-fast), app artifact deploy (S3 + SSM roll + is-active gate), promotion/rollback, live prod facts, and a RETAIN secret-shell troubleshooting entry that cross-references infra/README.md. - Correct retired *.seahavenind.com hosts to *.seahaven.com throughout and document the live GitHub/Slack/Linear webhook + OAuth endpoints. - Add a concise Deployment section to README pointing at the runbook. - Fix the stale host in the retired on-prem nginx/openswe.conf and mark it superseded by the AMI template. |
||
|
|
f379fbdaa9
|
docs: document RETAIN secret-shell orphan gotcha (#52)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
RETAIN + a fixed secret name means a failed FIRST create leaves empty secret shells behind when the stack rolls back. The shells keep the global `open-swe-<env>/<VAR>` names, so every later create fails with `AlreadyExists`, and a plain delete-secret keeps the name reserved for the recovery window rather than freeing it. Record the trap and the force-delete recovery (only for empty shells) in the config-store construct and the infra README so the next teardown/rebuild, secret logical-id change, or new-env stand-up does not rediscover it the hard way. Prod's first deploy hit this on 2026-06-29: 28 orphaned shells from an earlier failed create reserved the names and had to be force-deleted before the stack would create. |
||
|
|
86b4859589
|
fix: env-scope the EC2 launch template name (unblocks prod deploy) (#51)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
requireImdsv2:true makes CDK auto-create a launch template named from the
construct id ('Instance' -> 'InstanceLaunchTemplate') with no env qualifier,
so OpenSweDevStack and OpenSweProdStack both render
LaunchTemplateName: InstanceLaunchTemplate. dev created it first (the live
dev box runs on it); the prod first-deploy then failed with
InvalidLaunchTemplateName.AlreadyExistsException and the whole stack rolled
back.
Force a per-env LT name (open-swe-<env>-lt) via an aspect (the LT is created
at synth time by the requireImdsv2 handling, not in the constructor), and
rename the instance's launch-template REFERENCE in lockstep so CFN still
resolves it. synth-verified: dev=open-swe-dev-lt, prod=open-swe-prod-lt on
both the LT resource and the instance reference; version GetAtt preserved.
NOTE: deploying this renames dev's LT -> one-time dev box replacement
(stateless; boots from the baked AMI + pulls releases/latest). prod then
creates open-swe-prod-lt cleanly.
|
||
|
|
3c69dd9de6
|
fix: BatchGetSecretValue must be granted on * (corrects #48, fixes dev crash-loop) (#50)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
#48 (OSWE-IAC-SECRETS-LIST-01) scoped secretsmanager:BatchGetSecretValue to the open-swe-<env>/* ARN on the theory that an explicit --secret-id-list batch authorizes per-secret. That is FALSE: BatchGetSecretValue is a collection action AWS authorizes against the account (*), regardless of --filters vs --secret-id-list. A prefix-scoped grant AccessDenies the whole call. The dev box passed right after #48 only because the prior broad grant had not finished propagating; once it lapsed, fetch-config got AccessDenied -> loaded 0 secrets -> FAIL-FAST -> open-swe.service crash-loop. Verified on the live dev box (i-0af4e03e8bf70e6c3): the exact call returned 'not authorized to perform: secretsmanager:BatchGetSecretValue'; restoring the * grant recovered it. Move BatchGetSecretValue back to Resource:* (its own statement); keep GetSecretValue + DescribeSecret prefix-scoped (those gate VALUE access, so cross-env isolation holds). The surviving win from #48: --secret-id-list needs no name filter, so ListSecrets stays dropped -> no account-wide name enumeration. fetch-config.sh is unchanged (--secret-id-list is correct). The /sh-security-review finding OSWE-IAC-IAM-01 called this out and was wrongly refuted; the reference_secretsmanager_batch_get memory was wrong. |
||
|
|
9acf071ae4
|
fix: make dev->main promotion push succeed via bypass-actor App token (#49)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
The nightly promote fast-forwards main to a fully-green dev HEAD, but the push (as github-actions[bot]) is rejected by the main ruleset: it requires PRs + a status check and the default token is not a bypass actor, so a direct ref push can never land regardless of fast-forwardability. The prior comment claiming protection 'only rejects non-FF' was wrong. Mint a GitHub App installation token (actions/create-github-app-token, SHA-pinned) and push with it; the App must be added to the main ruleset's bypass actors out-of-band. The promoted commit already passed every check on dev (gated by check-dev-green.sh), so re-gating it via a PR on main is redundant. Also fix a gate self-poison: a stale failed 'promote' check-run from a prior run on the same dev HEAD blocked every subsequent gate run (it was excluded only by the current run_id). Exclude prior promote check-runs too, scoped to name=='promote' AND a /actions/runs/ details_url so an external app cannot hide a real failing check by naming it 'promote'; the positive REQUIRED_CHECKS allow-list stays authoritative. Gates: GPT-4.1 cross-review APPROVE (no security regression). Unit-tested: stale promote ignored -> PASS; real failure / external promote / missing required check -> BLOCK. shellcheck clean (also fixed a pre-existing SC2295 on the run_id match). |
||
|
|
faae9a685b
|
Scope secrets fetch to --secret-id-list; drop ListSecrets grant (#48)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Switch fetch-config.sh from a name-prefix batch-get-secret-value --filters scan to an explicit --secret-id-list (the 28 SECRET_VARS, chunked at the 20/call cap). An id-list batch authorizes per-secret ARN, so the instance role's BatchGetSecretValue moves from Resource:* to the open-swe-<env>/* prefix and the account-wide ListSecrets grant is dropped entirely. The box can no longer enumerate secret names account-wide; cross-env value isolation is unchanged (GetSecretValue was already prefix-scoped). Resolves OSWE-IAC-SECRETS-LIST-01. Also capture each chunk response into a variable and consume the producer via command substitution so a failed AWS call aborts under set -e instead of being swallowed by process substitution and misreported as a missing required var. Refs: OSWE-IAC-SECRETS-LIST-01 |
||
|
|
52cc696431
|
fix(deps): bump PyJWT to >=2.13.0 to close HS256 forgery (#47)
PyJWT <2.13.0 accepts a public-key JWK as an HMAC secret, letting an attacker forge HS256 tokens when mixed key families are allowed (GHSA high-sev Dependabot alert). 2.13.0 rejects the mismatch. Direct dependency; uv.lock re-resolves 2.12.1 -> 2.13.0. |
||
|
|
f79f49c755
|
Make PR title repo-aware and auto-link issues (#46)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Hardcoding the Sea Haven no-type-prefix PR title made every PR fail semantic-PR-title gates (this repo's PR Title Lint, upstream open-swe), forcing manual retitling. Make the title rule detect a conventional-commit gate and conform, falling back to the imperative style otherwise. Also add Closes/Refs issue-linking guidance and the default-branch auto-close caveat. Refs: #41 Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com> |
||
|
|
2533402f6b
|
chore(deps): update langgraph-cli[inmem] requirement (#38)
Updates the requirements on [langgraph-cli[inmem]](https://github.com/langchain-ai/langgraph) to permit the latest version. - [Release notes](https://github.com/langchain-ai/langgraph/releases) - [Commits](https://github.com/langchain-ai/langgraph/compare/cli==0.4.27...cli==0.4.30) --- updated-dependencies: - dependency-name: langgraph-cli[inmem] dependency-version: 0.4.30 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
df96fe49b5
|
chore(deps): bump astral-sh/setup-uv (#32)
Bumps the minor-and-patch group with 1 update in the / directory: [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv).
Updates `astral-sh/setup-uv` from 8.1.0 to 8.2.0
- [Release notes](https://github.com/astral-sh/setup-uv/releases)
- [Commits](
|
||
|
|
485674d941
|
chore(deps): bump python in the minor-and-patch group across 1 directory (#31)
Bumps the minor-and-patch group with 1 update in the / directory: python. Updates `python` from 3.14.5-slim-trixie to 3.14.6-slim-trixie --- updated-dependencies: - dependency-name: python dependency-version: 3.14.6-slim-trixie dependency-type: direct:production update-type: version-update:semver-patch dependency-group: minor-and-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
719741f279
|
Align dependabot with Sea Haven conventions (#39)
Switch all ecosystems to a weekly schedule and drop the redundant patterns: ["*"] selectors. Remove the major group so patch+minor bumps stay grouped into one PR per ecosystem while each major lands as its own PR, matching the org standard. Add npm coverage for the JS/TS surface (root, /infra CDK, /tests/e2e Playwright, /ui dashboard) and assign updates to amoussa1229. Delete the stale ui/yarn.lock so Dependabot tracks ui/bun.lock cleanly. Refs: #34 Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com> |
||
|
|
f3db9f02e3
|
Adopt Sea Haven agent conventions, no attribution (#30)
Codify the box-only #4 customizations into Git so the AWS deployment (which deploys from this repo) actually applies them — previously only the retired sh-openswe box had them. - prompt.py: branch names feature|bug|hotfix/<kebab> (optional <KEY->); imperative PR titles with no conventional-commit type: prefix; PR body Summary/Validation/Tests/Notes; handbook commit format. Rewrite the collaboration template from an attribution MANDATE to a PROHIBITION — no Co-authored-by bot trailer, no "Made by [Open SWE]" footer, no agent/AI notes on any artifact. - github_comments.py: add @seahaven-openswe (the deployed App slug) to the mention triggers. - authorship.py: remove the now-unused attribution helpers (build_pr_attribution_footer, add_bot_coauthor_trailer, add_pr_collaboration_note, PR_ATTRIBUTION_*). Keep OPEN_SWE_BOT_* — server.py still uses them for the sandbox git identity. - Flip the attribution unit tests to assert the no-attribution behavior; drop tests for the removed helpers. Commits stay authored as the triggering user for now — flipping authorship to the bot account depends on the Vercel preview-deploy constraint and is deferred to #11. Refs: #4 #11 Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi |
||
|
|
3af0bd5e16
|
ci: align workflows with Sea Haven CI/CD handbook (#29)
* Align workflows with Sea Haven CI/CD handbook Bring the workflow suite in line with the handbook: bump actions/checkout to v7 (Node 24 runtime, already standardized), kebab-case the two snake_case workflow filenames, and add the org-standard Labeler caller and Dependency Review gate so vulnerable or disallowed-license deps and unlabeled PRs are caught automatically. File renames only — job/check display names are unchanged, so the promotion gate's REQUIRED_CHECKS and branch-protection required checks are unaffected. Refs: INFRA-115 * Drop Agent prefix from CI workflow + job names The handbook names workflows for what they do (CI, Deploy, Labeler), not the component they run, matching .github and afterhours-shift-manager. Rename the suite to CI and its jobs to Lint / Format check / Unit tests, and keep the promotion gate's REQUIRED_CHECKS in sync. Refs: INFRA-115 --------- Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com> |
||
|
|
3224d7abf4
|
ci: gate dev→prod promotion on green checks + add rollback safety net (#28)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
Agent CI / Agent lint (push) Waiting to run
Agent CI / Agent format check (push) Waiting to run
Agent CI / Agent unit tests (push) Waiting to run
Agent CI / Playwright E2E (push) Waiting to run
T20 CD safety nets. Two gaps closed before the first real prod deploy: 1. Promotion gate. promote_dev_to_prod.yml previously fast-forwarded dev→main unconditionally. It now hard-gates on check-dev-green.sh: every check-run on the dev HEAD must be completed+passing AND the Agent CI suite (lint/format/unit/E2E) must be present+success, or the promotion blocks (fails safe on a missing/renamed check). The promote run excludes its OWN check-run by run-id (unforgeable), never by the mutable name "promote", so a colliding red check cannot hide. Fields are read with a 0x1F separator so an empty conclusion (every in_progress check) cannot shift columns. ci.yml now also runs on push:dev so dev HEAD actually carries that signal (a PR check alone can be admin-merged past). 2. Rollback + last-good. publish-and-deploy.sh advances releases/last-good/ only after a successful roll (deploy.sh gates on `systemctl is-active`), and makes releases/latest/ transactional — reverting to the prior release if the roll fails so a replaced box never self-deploys a broken release. New rollback.yml + rollback.sh re-point latest at last-good (or an explicit sha) and re-fire the deploy; prod is gated by the `prod` Environment approval, same as a deploy. The shared fire/wait/aggregate-gate logic is factored into roll-box.sh (used by both forward and backward rolls). Least-privilege: drop the unused s3:DeleteObject from the app deploy role — publish/rollback/deploy only Get+Put (S3-to-S3 copy), and the rollback fallback now depends on immutable release history staying intact. Lifecycle expiry (not CI) handles old-version cleanup. Gate logic unit-tested (7 cases + jq round-trip). IAM change + release-safety control cross-reviewed by GPT-4.1: APPROVE, no blocks. Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi |
||
|
|
5e30bc6be2
|
Point dev agent at seahaven-open-swe-dev org + un-pin the owner guard (#27)
Dev now runs against the dedicated seahaven-open-swe-dev org (repo openswe-dev-sandbox), isolated from the real Sea Haven org. Two changes: - config-store.ts: iacManagedSsm repo targeting is now per-env — dev = seahaven-open-swe-dev/openswe-dev-sandbox, prod stays Sea-Haven-Industries/open-swe-pilot. ALLOWED_GITHUB_ORGS tracks the env owner. Adds a dev-only SEED_USER_MAPPINGS param so the triggering GitHub login (amoussa1229) resolves and @openswe comments aren't skipped. - fetch-config.sh: replace the unconditional hard-pin to Sea-Haven-Industries with an owner GUARD that HONORS the configured owner (OPENSWE_REPO_OWNER override, else the SSM value) but forces a per-env safe org when the normalized owner is blank or the upstream langchain-ai. Normalization (lowercase, strip whitespace, first path segment, drop dots) catches langchain-ai/<repo>, langchain-ai., and case variants without over-blocking legit orgs (e.g. langchain-ai-fork). Fallback org is per-env so dev can't fall back into the real org. Reviews: GPT-4.1 cross-review APPROVE (round 1 found a path/dot bypass -> hardened, round 2 clean); /sh-security-review authz one LOW (env-invariant fallback) -> fixed. Positive org allowlist still enforced by the app via ALLOWED_GITHUB_ORGS. |
||
|
|
082768ff51
|
ci: print stack outputs from CDK --outputs-file (drop describe-stacks) (#26)
cd-infra reported failure on every successful deploy: the 'Stack outputs' step ran 'aws cloudformation describe-stacks' with the githubdeploy-open-swe-infra-<env> role, which intentionally lacks cloudformation:DescribeStacks. The cdk deploy itself succeeds (it reads outputs via the bootstrap cfn-exec role it assumes). Switch to 'cdk deploy --outputs-file cdk-outputs.json' + cat — no extra IAM grant, and the job goes green on actual deploy success instead of masking real failures behind a red run. |
||
|
|
f2633f9bd0
|
fix: don't crash-loop the box when no user mapping is configured (#25)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
seed_store.sh (ExecStartPost) exited 1 when it couldn't resolve a user_mappings entry, which — with Type=simple — fails the whole unit and crash-loops the box. This contradicts the script's own OSWE-SEED-03 precedent (the server-not-ready path exits 0 specifically to avoid a restart loop). A missing user mapping is a seeding gap, not an unhealthy server: the @openswe trigger just won't resolve a commenter, which a deployment-validation env (dev) does not need. Make it non-fatal — warn and skip the user_mappings seed; team_settings is still seeded and the service starts. The empty MAPPINGS array makes the seed loop a no-op. Set SEED_USER_MAPPINGS / CONFIGURED_ADMINS / OPENSWE_OWNER_LOGIN+EMAIL to seed it. Verified on the dev box: service active, langgraph bound to 127.0.0.1:2024 (loopback), /ok 200, /healthz 200. |
||
|
|
a2877d46ff
|
fix: paginate batch-get-secret-value in fetch-config (secrets dropped past page 1) (#24)
fetch-config materialized only the FIRST page of Secrets Manager results: the AWS CLI does NOT auto-paginate batch-get-secret-value (the script's comment claiming it does was wrong; --no-cli-pager only disables the output pager, not API pagination). With 28 secret shells under open-swe-<env>/ the first page returned ~10 items, stranding the rest on later pages. Required secrets that landed past page 1 (DASHBOARD_JWT_SECRET, TOKEN_ENCRYPTION_KEY, LANGSMITH_API_KEY_PROD) were silently dropped, tripping the FAIL-FAST 'missing required var' guard and crash-looping open-swe.service. Follow NextToken across pages (new batch_get_secrets_tsv helper). Verified against the live open-swe-dev secrets: now loads all 5 populated secrets (was 2). Same per-record base64 / exact-prefix / accept_var / emit_var hardening — only the page loop is new. |
||
|
|
cdcb6625b9
|
fix: BatchGetSecretValue must be on * for the filtered batch call (#23)
The prior fix scoped secretsmanager:BatchGetSecretValue to the env-prefixed secret ARN, but the live box still got AccessDenied: batch-get-secret-value invoked WITH a name --filters is a COLLECTION call that AWS authorizes against * (a per-secret ARN does not satisfy it). Split the statement: - GetSecretValue + DescribeSecret stay PREFIX-scoped (secret:open-swe-<env>/*) — this is what gates which secret VALUES the box can read (checked per-secret in the batch). - BatchGetSecretValue + ListSecrets move to a * operation-level statement (the filtered collection call + the list action; neither is resource-scopable for this usage). VALUE isolation preserved (dev box still cannot read prod secret values); only secret NAME/metadata enumeration is widened. GPT-4.1 IAM cross-review: BLOCK none, FIX none. Suppression OSWE-IAC-SECRETS-LIST-01 updated; future hardening (explicit --secret-id-list to drop both * grants) tracked there. |
||
|
|
cfbdcda9b7
|
fix: instance role BatchGetSecretValue + ListSecrets for .env materialization (#22)
* fix(infra): grant instance role BatchGetSecretValue + ListSecrets for .env materialization fetch-config.sh materializes the box's .env via `secretsmanager batch-get-secret-value --filters Key=name,Values=open-swe-<env>/`, but the instance role only granted GetSecretValue/DescribeSecret. BatchGetSecretValue is a distinct IAM action, so the call was AccessDenied and open-swe.service crash-looped (no .env written -> ExecStartPre exit 1). - Add secretsmanager:BatchGetSecretValue to the prefix-scoped ReadSecrets statement. - Add secretsmanager:ListSecrets on * (required by the name-prefix filtered batch call; the API has no resource-level scoping for the list action — fits the role's stated exception). Secret VALUES stay prefix-scoped; only names are enumerable. Reviews: GPT-4.1 IAM cross-review BLOCK=none; /sh-security-review iac-iam one LOW metadata residual (no critical/high), recorded as OSWE-IAC-SECRETS-LIST-01. Refs T7/T19 dev bring-up. * ci: lift Node heap cap for Playwright E2E build (vite OOM) The E2E job's Playwright globalSetup runs the real `bun run build`, whose vite bundle exceeds Node's default ~2 GB heap and OOMs (JavaScript heap out of memory) — the same failure fixed for build-artifacts.yml in #19. Set NODE_OPTIONS=--max-old-space-size=8192 on the Run E2E step. |
||
|
|
5fa132205b
|
fix(deploy): use %%...%% for CDK user-data tokens (don't collide with @@ sed) (#21)
The systemd unit booted with a literal `@@OPENSWE_ENV@@` (fetch-config.sh got the token, not "dev" -> exit 2 -> crash-loop) because user-data.sh is double-templated: CDK substitutes @@tokens@@ AND user-data seds @@tokens@@ into the baked systemd/nginx files. CDK's `.replace(/@@OPENSWE_ENV@@/g, "dev")` clobbered the sed PATTERN (`s|@@OPENSWE_ENV@@|...|` -> `s|dev|...|`, a no-op), so the unit's token never got replaced. Same collision hit @@SERVER_NAME@@ (masked by nginx default_server). Fix: CDK tokens move to a DISTINCT delimiter %%...%% (rendered in app-service.ts); the @@...@@ tokens stay for the baked-template seds. No AMI rebuild (templates unchanged). Add a guard test asserting no unresolved %%CDK%% token survives in the synthesized user-data. Also add .github/scripts/** to build-artifacts paths so script-only changes trigger a publish. jest 20/20; tsc + shellcheck clean; rendered user-data: OPENSWE_ENV="dev", SERVER_NAME="openswe-dev.seahaven.com", @@ sed patterns preserved, 16872 B. |
||
|
|
9e2b215f08
|
fix(ci): avoid tar|grep -q SIGPIPE false-failure in package-artifacts (#20)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
`tar -tzf app.tar.gz | grep -qx` under set -o pipefail fails the pipeline when grep -q matches and exits early (SIGPIPEs tar -> 'write error' -> non-zero), a false 'missing agent/server.py'. List the archive once into a var, then grep the var. Same fix for the secret-guard pipe (which was also silently broken). |
||
|
|
2765b65acd
|
fix(ci): raise Node heap for the SPA build (vite OOM'd at ~2 GB) (#19)
The build-artifacts SPA build hit Node's default ~2 GB heap cap and aborted (JS heap OOM, exit 134) — the same memory-hungry vite build that needed an 8 GB swapfile on-box. The runner has ~16 GB, so set NODE_OPTIONS=--max-old-space-size=8192 on both build steps. |
||
|
|
404b3f6f75
|
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18)
* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)
Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.
deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
- the file provisioner failed uploading the templates dir ('scp: …: Is a
directory') — a trailing-slash contents-upload needs the dest dir to exist;
added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
dest trailing slash.
- the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
environment_vars never reached provision.sh (which runs under set -u and
aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.
infra:
- ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
discipline docs.
- app-service.ts: machineImage → bakedOpenSweArm64().
- open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
- cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
- README: Baked AMI + EBS-replacement-discipline section.
tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.
NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).
* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021
Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.
* fix(deploy): GitHub App + Slack required for prod only, not dev
Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.
Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).
* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)
Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.
Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
/opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
baked open-swe-base-arm64 AMI (folds in the held #16).
IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.
Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
target is healthy even before the first release); deploy.sh is base64-rendered
by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).
CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
ui/.output/public -> spa.tar.gz), package the Python source via git archive
(app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
push dev -> dev (auto); push main -> prod (env "prod" approval gate).
Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.
* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard
Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:
- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
not CommandInvocations[0], so a partial failure across the brief 2-instance
replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
(defense in depth over .gitignore; scoped to data extensions so *_credentials.py
source is not a false positive — verified against the real tree).
Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).
shellcheck/tsc/jest(16) clean; both stacks synth offline.
* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard
The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)
- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
GroupDescription + rule descriptions are pure ASCII, so this fails the build
instead of a deploy next time.
jest 18/18; tsc clean.
* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)
The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.
jest 18/18; tsc clean.
* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit
The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.
jest 19/19; minified deploy.sh passes bash -n + shellcheck.
|
||
|
|
70319cac0d
|
ci: path-filtered infra CI/CD + dual OIDC roles + prod approval gate (T18) (#15)
* ci: path-filtered infra CI/CD with dual OIDC roles + prod approval gate (T18)
Add the /infra half of the combined-repo pipeline (the Python agent keeps ci.yml):
- ci-infra.yml — PR check on infra/** : tsc + jest + cdk synth via the org
reusable ci-typescript-cdk.yaml (working-directory: infra).
- cd-infra.yml — push to dev/main on infra/** (or dispatch):
* job 'ci' (reusable) is the CI-green precondition (deploy needs: ci).
* deploy-dev (ref=dev, NO environment) → cdk deploy OpenSweDevStack,
assuming githubdeploy-open-swe-infra-dev (OIDC sub ref:refs/heads/dev). AUTO.
* deploy-prod (ref=main, environment: prod) → cdk deploy OpenSweProdStack,
assuming githubdeploy-open-swe-infra-prod (OIDC sub environment:prod). The
'prod' Environment's required reviewer is the manual-approval gate.
Deliberately self-contained (NOT the reusable cd-cdk.yaml) because that runs
'cdk deploy --all' — from a single-env push it would deploy the other env + the
shared IAM stack, breaking the per-env boundary. CD targets one stack per env;
the shared open-swe-iam stack is human-gated (T6), never deployed by CD.
Infra CI is enforced at the DEPLOY boundary (deploy jobs need ci), not as a
branch-protection required check — path-filtering a required check would deadlock
app-only PRs. Documented in infra/README.md along with the post-T6 prerequisites
(repo vars AWS_DEPLOY_ROLE_INFRA_{DEV,PROD}; a 'prod' Environment w/ reviewer).
Not active until the IAM roles are applied (T6) — assuming a nonexistent role
just fails closed. App-side CD (S3 artifact + SSM) is T19.
* fix(infra): commit jest.config.js (was ignored by *.js → infra CI used Babel)
The infra/.gitignore *.js rule (for compiled CDK output) silently swept up the
hand-authored jest.config.js, so it was never committed. Local jest passed (file
present in the working tree) but CI's fresh checkout lacked it → jest fell back to
the default Babel transform → 'Cannot use import statement outside a module' on the
TypeScript test. Surfaced now because T18 is the first workflow to run infra jest
in CI. Negate the ignore for this one file and commit it.
|
||
|
|
fcbdfb67aa
|
feat: open-swe dev/prod compute + ALB ingress (T12) (#14)
AppService construct wires the per-env EC2 box and its internet path. The
seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the
on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and
never owned/mutated; open-swe only ADDS its own resources.
Per env (open-swe-stack.ts → AppService):
- ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a,
in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination
(no RETAIN volume — replacement-tolerant; see ami-cache.ts).
userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh.
- Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress
via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress
(the imported, on-prem-owned SG is never mutated).
- Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays
loopback). Health check GET /healthz.
- Two rules on the imported :443 listener, both → the TG:
* webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/*
* site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api)
Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5)
or it would steal every webhook — first-match-by-ascending-priority.
- Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB.
- 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config).
Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed
medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large
GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on
/dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene);
XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed
critical/high.
Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked
open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean.
Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
|
||
|
|
aef6b26912
|
feat: Secrets Manager + SSM config store for open-swe (T11) (#10)
Create the per-env config surface the EC2 box reads at boot via
fetch-config.sh / seed_store.sh:
- ConfigStore construct (infra/lib/constructs/config-store.ts):
- 28 value-LESS Secrets Manager shells open-swe-<env>/<VAR>
(RemovalPolicy.RETAIN, no SecretString/generateSecretString — real
values are set out-of-band by put-config.sh, never in IaC/state).
- 8 IaC-managed SSM params /open-swe-<env>/<VAR> with real,
stable/derivable values (SANDBOX_TYPE, DEFAULT_REPO_OWNER/NAME,
ALLOWED_GITHUB_ORGS, DASHBOARD_*_URL/ORIGINS, LLM_MODEL_ID).
- OUT_OF_BAND_SSM documents the ~30 params CDK intentionally does NOT
own (operationally-variable / env-specific-unknown).
- Wire ConfigStore into OpenSwe<Env>Stack.
- KebabNamingAspect: exempt Secrets Manager + SSM names, which carry the
literal UPPER_SNAKE env-var segment (open-swe-dev/DASHBOARD_JWT_SECRET).
- deploy/seahaven/put-config.sh: out-of-band populator (placeholders only,
OPENSWE_PUT_<VAR> env indirection; no real values committed).
Synth-only; not deployed. Instance-role read grants on open-swe-<env>/*
already exist from T6 — no IAM/trust changes here.
|
||
|
|
434d0ad80c
|
feat(deploy): AWS-sourced fetch-config + seed_store + rotation docs (PR#7) (#8)
PR#7 of the AWS migration. deploy/seahaven/: fetch-config.sh materializes a service-user-owned 0600 tmpfs .env from Secrets Manager + SSM (fail-fast); seed_store.sh reseeds the in-memory store; ROTATION.md. Incorporates T5 /sh-security-review fixes: - seed_store no longer bash-sources the .env (closes the SH-INJ-001 RCE); uses a non-eval reader, jq --arg JSON bodies, and a loopback-pinned BASE. - fetch-config: .env owned by the openswe service user (app no longer runs as root); DEFAULT_REPO_OWNER hard-pinned; key-identifier validation + flat-namespace collision detection; dropped SSM --recursive. shellcheck + bash -n clean. |
||
|
|
ce9e5fb51b
|
feat(deploy): EC2 AMI recipe (Packer) + cloud-init/user-data (PR#2) (#7)
PR#2 of the AWS migration. deploy/ami/: Packer template (Ubuntu 24.04 arm64, uv+py3.12, nginx, awscli v2, CW agent; no swapfile), provisioning-only user-data (userDataCausesReplacement rationale), systemd unit + nginx + CW templates. Incorporates T5 /sh-security-review fixes: langgraph binds 127.0.0.1 (not 0.0.0.0); nginx is the sole ingress proxying only /dashboard/api/ + /webhooks/; ExecStartPre runs fetch-config as root (+) and passes the env arg; the app runs as the unprivileged openswe user reading an openswe-owned 0600 .env. packer validate clean. |