Commit graph

10 commits

Author SHA1 Message Date
Adam Moussa
a4ed19ba61
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)

Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.

- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
  region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
  FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module

* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8

The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.

Map profile effort to additional_model_request_fields:
  {thinking: {type: adaptive, display: summarized},
   output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.

* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids

Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
  (set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
  FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
  otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.

Surfaced by the cross-family review + verified against deploy/.

* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip

From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
  validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
  error code only, so the role ARN + account id in the raw botocore message never
  reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
  (Converse emits reasoning_content, not thinking) so the middleware is not a no-op
  on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)

* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids

Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
  us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
  each routed region (us-east-1/2, us-west-2). The model runs in the server process
  on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
  for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
  mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
  bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
  seed_store.sh's default via pick precedence, so the seed-script fix alone was
  insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
  ids to the Bedrock id (config.toml's model_id was an active, now-broken value).

AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.

* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)

Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
  (eval judge only — Bedrock builder/reviewer auth via the host IAM role).

REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
2026-06-29 15:57:19 -04:00
Ramon Nogueira
6117780982
feat(open-swe): let any org member post to a thread, with attribution (#1594)
* feat(dashboard): let any org member post to a thread, with attribution

Posting to an Agents chat thread from the web UI was restricted to the
thread owner. Open it to any authenticated org member (login is already
org-gated by OAuth) on both write paths — the queued follow-up
(send_dashboard_message) and the idle-thread run.start
(_enrich_run_start_command). Non-owner messages are prefixed with the
poster's verified GitHub login (@login:) so the agent and owner can tell
who sent them. Thread management (cancel/delete/resolve) stays owner-only,
and the UI now shows the composer to non-owners.

* fix(dashboard): keep non-run.start commands owner-only

Non-owner posting is allowed only via the attributed run.start path. Other
write commands (e.g. input.respond) carry unattributed user input, so the
commands proxy keeps them owner-only instead of readable-by-any-org-member.

* docs(e2e): drop per-test details from the E2E README
2026-06-23 11:06:53 -07:00
Johannes du Plessis
39a26e16b5
fix: optimize agent thread lists (#1570)
* fix: optimize agent thread lists

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: refresh missing thread run status

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-18 10:09:22 -07:00
Johannes du Plessis
c07434a221
fix: allow read-only cross-user access to agent threads via Open in Web links (#1568)
The thread detail endpoint already returned metadata for non-owners, but
the transcript hydration endpoints (state, stream/events, history, pr-diff)
all asserted ownership and 404-ed. This caused the UI to redirect non-owners
back to /agents when they clicked an "Open in Web" link shared in Slack.

Dashboard login is already gated by ALLOWED_GITHUB_ORGS, so any logged-in
user is a trusted org member. This commit:
- Adds _thread_is_readable / _assert_thread_readable helpers that grant
  read access to any surfaced-source thread for authenticated users
- Relaxes read endpoints (state, stream/events, history, pr-diff, SSE
  stream) to use readable checks instead of ownership checks
- Keeps write endpoints (send message, cancel, delete, resolve, run
  commands) owner-only
- Adds an isOwner field to the thread summary so the frontend can render
  a read-only mode (hides the prompt bar, resolve/delete buttons)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-18 09:10:09 -07:00
Johannes du Plessis
b19804536c
feat: handle images sent to non-vision models in Slack, Linear, and web UI (#1560)
* feat: handle images sent to non-vision models in Slack, Linear, and web UI

Add vision capability checks across all image input paths. When a user
sends images to a text-only model (e.g. GLM 5.2, DeepSeek V4 Pro), the
images are now skipped and a warning is injected into the prompt instead
of sending unsupported content to the model.

- Slack: resolve model at webhook time, skip image fetch + add warning
- Linear: same pattern as Slack
- Queued message middleware: read resolved model from thread metadata,
  strip images from queued payloads for text-only models
- Web UI: disable submit + show inline warning when images are attached
  to a non-vision model selection
- Shared: resolve_agent_model_id helper + vision_not_supported_warning

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: mock resolve_agent_model_id in Slack mention test

The test_process_slack_mention_queues_active_thread_message test was
missing a mock for the new resolve_agent_model_id call added to the
Slack webhook handler, causing a TypeError when image URLs triggered
the model resolution path.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: include vision warning in queued payload for text-only models

Update the prompt variable (not just content_blocks) before clearing
image_urls so the queued payload also carries the warning text when a
Slack/Linear follow-up arrives while the thread is busy.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 09:12:52 -07:00
Caroline di Vittorio
e8bb6b497b
feat: mark threads as resolved to hide from sidebar (#1500)
* feat: mark threads resolved to hide from sidebar [closes resolve-threads]

Add a resolved flag (thread metadata) so users can clear finished
threads from the sidebar without deleting them. Resolved threads move to
a collapsible "Resolved" group (capped at 20 with a "Show all" link) and
are fully searchable/filterable + paginated on a new /agents/threads
page driven by URL query params. Sending a new message auto-unresolves.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: refresh paginated agent thread lists

Invalidate all agent thread list queries when thread lifecycle events change sidebar data, keeping the new paginated sidebar cache in sync.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-06-12 12:01:42 -07:00
Johannes du Plessis
b8035a4758
fix: stream proxy preflight + optimistic thread staleTime (#1502)
* fix: preflight stream proxy before SSE starts, let optimistic thread survive refetch

* fix: seed sidebar thread details as stale so mark-viewed fetch still fires

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 13:07:24 -07:00
Johannes du Plessis
e0678e8c01
feat: track PR lifecycle state per thread for sidebar (#1492)
* feat: track PR lifecycle state per thread for sidebar

Persist a PR's draft/open/merged/closed state on the agent thread and keep
it in sync as PR webhooks fire, so the Agents UI can show per-thread PR
status the way Cursor does. open_pull_request now records the initial
draft/open state and pr_title; thread summaries expose diffStats; and the
PR webhook refreshes pr_state on close/reopen/draft/ready transitions.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: consolidate PR state mapping into shared derive_pr_state helper

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 12:21:59 -07:00
Christian Bromann
abf354bb05
feat(open-swe): stream agent chat via @langchain/react v2 protocol (#1475)
* feat(dashboard): stream agent chat via @langchain/react v2 protocol

Replace the bespoke SSE + React Query polling path with LangGraph’s
v2 event stream through credentialed dashboard proxies. Run starts go
through stream commands; mid-run follow-ups still queue via /messages.

* fix import path

* fix tests after rebase

* format

* PR feedback

* improved model fallback

* fix image handling

* embrace sdk

* cleanup

* cr

* more cleanup

* fix cors

* harden security

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 09:54:35 -07:00
Johannes du Plessis
5faf190954
fix: reject images for text-only models (#1439)
* fix: reject image uploads for text-only models

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: validate queued images against active model

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-06 13:09:31 -07:00