* fix(dashboard): managed-cloud OAuth hardening + admin user-mapping endpoint
Prepare the dashboard backend for the managed LangGraph Cloud + Vercel
runtime, where the API is HTTPS and cross-site from the UI.
- OAuth redirect_uri (#2): coerce a schemeless DASHBOARD_API_BASE_URL to
https:// in _api_base_url() so GitHub stops rejecting login with
"redirect_uri not associated with this application". _cookie_security()
now treats a schemeless (managed) value as Secure; SameSite=None too,
consistent with the coerced scheme.
- OAuth state cookie (#3): document that osw_oauth_state is host-only by
design (a Domain cookie is unsafe across *.vercel.app, a public suffix),
so login must always start on the stable alias to avoid "oauth state
mismatch". Operational contract; no behavioral change.
- Admin user mappings (#4): add POST /admin/user-mappings so an admin can
set the github_login -> work_email link from the dashboard instead of a
raw Store write. New "admin" MappingSource provenance value.
* fix(webapp): refresh user-mapping cache on GitHub webhook paths
On managed LangGraph Cloud the backend runs multiple replicas, so the
per-process GitHub<->work-email mapping cache can be stale on the replica
handling a webhook (a mapping created on another replica is invisible
until refresh). process_github_pr_comment and process_github_issue now
refresh the cache from the durable Store before resolving the author's
email, matching the existing Slack mention path (process_slack_mention).
* perf(webapp): defer deepagents import to speed custom-app cold start
The custom FastAPI app (agent.webapp:app, the langgraph.json http.app)
pulled deepagents -> langchain_anthropic -> anthropic into its import
graph via dashboard.routes, only to build skill/chat seed files. Defer
those create_file_data imports into the functions that use them. Removes
deepagents/langchain_anthropic/anthropic from app import entirely and
roughly halves module-import wall time (~0.6-0.8s -> ~0.35s warm; larger
cold-start saving since native anthropic init is skipped). Behavior
identical. (reviewer_diff already imports deepagents under TYPE_CHECKING.)
* feat(ui): set work_email user mappings from the admin dashboard
Add an "Add / update" form to the admin User mappings section and the
adminUpsertUserMapping API client method, wiring the new
POST /admin/user-mappings endpoint. Admins can now create or update a
github_login -> work_email mapping directly instead of waiting for the
user to self-connect Slack.
* docs: document managed LangGraph Cloud + Vercel deployment
- INSTALLATION §10: add the managed production env triad (LANGGRAPH_URL,
DASHBOARD_BASE_URL + DASHBOARD_API_BASE_URL with https://, empty
VITE_DASHBOARD_API_BASE_URL for same-origin), the stable-alias login
and vercel.json stable-deployment-URL requirements, multi-replica cache
note, plus redirect_uri-scheme and oauth-state-mismatch troubleshooting.
Refresh the langgraph.json snippet to all six graphs.
- README: reframe deployment around the managed migration; link the plan.
- deploy/MIGRATION.md: import the self-hosted -> managed migration plan.
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)
Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.
- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module
* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8
The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.
Map profile effort to additional_model_request_fields:
{thinking: {type: adaptive, display: summarized},
output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.
* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids
Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
(set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.
Surfaced by the cross-family review + verified against deploy/.
* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip
From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
error code only, so the role ARN + account id in the raw botocore message never
reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
(Converse emits reasoning_content, not thinking) so the middleware is not a no-op
on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)
* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids
Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
each routed region (us-east-1/2, us-west-2). The model runs in the server process
on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
seed_store.sh's default via pick precedence, so the seed-script fix alone was
insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
ids to the Bedrock id (config.toml's model_id was an active, now-broken value).
AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.
* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)
Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
(eval judge only — Bedrock builder/reviewer auth via the host IAM role).
REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
* feat: default Slack/dashboard/schedule PRs + commits to the app identity (#57)
Slack/dashboard/schedule runs now author PRs and run git/gh operations as the
GitHub App seahaven-openswe[bot] by default (matching GitHub-issue runs), so the
self-review 422 is impossible by construction rather than guarded in the prompt.
A profile flag author_prs_as_user restores per-user attribution.
- open_pull_request._resolve_pr_author_token + auth.resolve_github_token: default
to the installation token for these sources; per-user only when opted in.
- authorship: commit identity -> seahaven-openswe[bot] (numeric noreply;
accepted Vercel-resolution risk, documented inline).
- self-trigger safety: INTERNAL_BOT_LOGINS + webapp/reviewer_reconcile/reply
markers recognize seahaven-openswe[bot] (bot-authored events are now ours).
Supersedes the prompt-only guard in #58.
* fix: author commits as the app bot in the default path (SH-IDSPLIT-01)
Security review found the commit identity was NOT actually unified to the bot:
resolve_triggering_user_identity got a 403 from the installation token and fell
back to configurable['github_login'], so commits were still authored as the
triggering user (commit=user, push+PR=bot — a three-way split that missed the
stated goal). Now gate the triggering-user identity resolution on the same
default-bot decision as the token: slack/dashboard/schedule default to the app
bot identity unless author_prs_as_user is set.
* docs(security): record AUTHZ-SLACK-BOT-DEFAULT-001 as an accepted residual (#59)
Single-user deployment; bounded by App-on-pilot + ALLOWED_GITHUB_REPOS lock.
Revisit (add a per-user gate) before expanding users or the App installation.
* fix: enforce a replay window on Linear webhooks (AUTHZ-001)
verify_linear_signature accepted any correctly-signed body with no freshness
check, so a captured request could be replayed indefinitely. Parse the
signed webhookTimestamp (Unix ms) and reject requests outside a 60s window,
failing closed when the field is missing or malformed — mirroring the Slack
verifier.
* fix: stop leaking upstream auth-error bodies into user comments
get_github_token_for_user folded the raw upstream response text into the
error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the
full body server-side only and return a generic "GitHub auth failed (status
<code>)". Also document the accepted shared-installation-token blast radius on
the bot-token-only path (AUTHZ-003).
* fix: bind sandbox and token caches to repo to prevent thread-id collision
A PR head-branch name is attacker-controllable and get_thread_id_from_branch
derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01).
The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on
thread_id alone, and a cached sandbox was reused after only an echo-ping, so a
different repo's webhook could bind to another thread's sandbox or token.
Without changing the persistent thread-id scheme:
- Persist the bound repo (owner/name) in thread metadata on sandbox creation and
refuse to reuse a sandbox whose bound repo does not match the current event
(SandboxRepoMismatchError); the in-memory proxy also carries the binding.
- Bind the GitHub-token cache entries to their repo and evict on a cross-repo
read so a colliding thread_id cannot be served another repo's token.
- Thread repo through the reviewer and the webhook token resolvers.
* fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04)
The instance role and the GitHub deploy app role granted s3:ListBucket on the
whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts)
only ever lists under releases/, so add a StringLike s3:prefix=releases/*
condition. GetBucketLocation has no s3:prefix in its request context, so it
moves to its own unconditioned statement. Also document the accepted F-2
cross-env existence-oracle residual on BatchGetSecretValue.
* chore: suppress test-fixture credential false positive; document AUTHZ-002
Add a machine-level suppression for the fake Datadog key in the
test_team_credentials encryption-roundtrip fixture (CWE-798, not a real
credential). Clarify that the within-org thread-write path is intentional by
design (AUTHZ-002) — comment only, no behavior change.
* fix: casefold repo-binding keys to avoid spurious cross-repo mismatch
GitHub owner/name are case-insensitive. Casefold the owner/name key on both the
write (binding) and read (compare) sides — repo_cache_key and the metadata
bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate
same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2).
* fix: stop leaking upstream auth body in unexpected-result branch
The 2xx-but-missing-token/url branch echoed the parsed upstream response body
into the user-facing error. Return a generic message and log response_data
server-side only, mirroring the existing HTTPStatusError fix (Gap 4).
* fix: fail closed for unbound-legacy sandboxes and catch repo mismatch
Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no
recorded bound_repo (a pre-binding legacy thread, post-deploy) previously
reconnected-and-served the sandbox to the current repo, then rebound it. Now
fail closed: drop the stale id and recreate a fresh sandbox bound to this repo,
logging a reconnect-with-missing-binding event. A sandbox is never served to a
repo unless its binding is known and matches; new threads bind on first run
unchanged.
Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints,
log it for alarming, and surface a clean sanitized error instead of letting an
opaque deep-stack exception crash-loop the worker.
* chore: suppress test-fixture credential false positive in token-TTL tests
Add a machine-level suppression for the fake "ghp_secret" GitHub token used by
the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and
not a valid PAT; scoped to the unit test only.
Codify the box-only #4 customizations into Git so the AWS deployment
(which deploys from this repo) actually applies them — previously only
the retired sh-openswe box had them.
- prompt.py: branch names feature|bug|hotfix/<kebab> (optional <KEY->);
imperative PR titles with no conventional-commit type: prefix; PR body
Summary/Validation/Tests/Notes; handbook commit format. Rewrite the
collaboration template from an attribution MANDATE to a PROHIBITION —
no Co-authored-by bot trailer, no "Made by [Open SWE]" footer, no
agent/AI notes on any artifact.
- github_comments.py: add @seahaven-openswe (the deployed App slug) to
the mention triggers.
- authorship.py: remove the now-unused attribution helpers
(build_pr_attribution_footer, add_bot_coauthor_trailer,
add_pr_collaboration_note, PR_ATTRIBUTION_*). Keep OPEN_SWE_BOT_* —
server.py still uses them for the sandbox git identity.
- Flip the attribution unit tests to assert the no-attribution behavior;
drop tests for the removed helpers.
Commits stay authored as the triggering user for now — flipping
authorship to the bot account depends on the Vercel preview-deploy
constraint and is deferred to #11.
Refs: #4#11
Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi
* perf: cut review-chat time-to-first-token
The sandbox-less PR review chat paid several blocking network round-trips
before the first token on every message. Cache GitHub App installation
tokens in-process (per scope, until ~10m before expiry, above the proxy's
5m refresh window) so the chat graph factory and proxy stop re-minting one
each turn. Also drop the duplicate thread-metadata read in the commands
proxy and replace the heavy per-message get_review staleness check with a
single lightweight PR head-SHA lookup.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: keep review chat alive when reseed fails
Address review: a moved PR head now triggers _build_pr_context (and thus
get_review). For an existing chat, fall back to the last seeded context on
HTTPException instead of failing the command; fresh chats still surface the
error since they have no prior context.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: repo-scoped dynamic sandbox snapshots
Let admins build a per-repo sandbox image from a custom Dockerfile so runs
targeting that repo boot from a snapshot with its deps pre-baked. Snapshot
selection is purely additive: repos without a `ready` repo-scoped snapshot
always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID.
Backend adds a repo_snapshots store module (Dockerfile + build status keyed by
owner/name), threads the resolved repo through the LangSmith sandbox creation
path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a
throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The
UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor +
build status/logs).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden repo snapshot builds
Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins
cannot accidentally build a repo snapshot from a bare Python image that lacks
Open SWE's sandbox tools. Allow stale building records to be retried by tracking
build_started_at and treating old or missing timestamps as stale.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: document repo snapshot base image config
Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert
missing base-image configuration into a handled dashboard API error so admins see
a clear configuration message instead of an unhandled template-generation error.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode for read-only research and planning
Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce plan-mode read-only at tool layer and disable subagents
Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden plan-mode shell guard against wrapped mutations
Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow
- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline
Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff
Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.
- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).
Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.
* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright
The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.
Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.
* test(open-swe): record Playwright trace + video on every E2E run
Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
* feat(plan-mode): collaborative plan review with BlockNote + Yjs
When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.
- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
snapshots; plan content/status store; plan REST API (get/approve/reject,
owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
mounted under the agents shell, with a "Review plan" banner in the thread view
and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
flow, including cross-user comment sync and owner-only approval.
* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)
- plan-collab WS: authorize per-thread before joining a room (same read gate as
the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.
Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).
* style: ruff format plan_collab.py
* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS
- Slack "Approve & Implement" now verifies the clicking user is the plan
requester (owner, via the stored triggering_user_id) before implementing —
matching the dashboard API's owner-only approval. Non-owners are pointed to
Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
the REST require_same_origin CSRF defense.
* fix(plan-mode): enter plan mode only via the model + local mock dev harness
Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.
- enter_plan_mode returns a terminating ToolMessage, fixing the missing
ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
Alice/Bob mock users, and a GitHub login picker.
* docs(plan-mode): drop stale references to removed profile/team defaults
The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.
* feat(plan-mode): let any reviewer edit the plan, not just comment
Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.
* test(plan-mode): assert plan-mode entry via the tool's success message
plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.
---------
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: reviewer enforces AGENTS.md/CLAUDE.md repo rules as mandatory pass
The reviewer already fetched AGENTS.md but treated violations as optional
candidate findings. Now the reviewer runs a dedicated compliance pass that
checks every changed hunk against each rule in AGENTS.md (or CLAUDE.md as
fallback), treating violations as mandatory findings rather than style nits.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: oversized AGENTS.md returns None instead of falling back to CLAUDE.md
Only a 404 (file absent) triggers fallback to CLAUDE.md. Oversize,
HTTP errors, and unexpected status codes now return None immediately
so the reviewer does not enforce stale rules from a secondary file.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: validate LLM API keys on startup
* fix: correct relative import for options module
* refactor: move imports to top of file
* style: fix linting and formatting issues
* refactor: scope LLM validation to local dev and rename function
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: open-swe[bot] <johannes@langchain.dev>
* feat: add size caps for PR diff, fetch_url, Slack threads, pagination, message queue
Per-source byte/token caps with explicit truncation markers to prevent
unbounded payloads from blowing up LLM context/memory.
- reviewer_diff.py: cap PR diff at 200K chars with head+tail truncation
- fetch_url.py: cap markdownify output at 100K chars
- slack.py: cap thread message fetch at 500 messages
- github_comments.py: cap _fetch_paginated at 50 pages
- thread_ops.py: cap queued messages at 100 (drop oldest)
Closes OPE-51
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: compute diff line set from full diff, keep most recent Slack messages
Address PR review comments:
1. Truncated diffs rejected valid findings: fetch_pr_diff now returns the
full diff; truncate_diff is called separately in reviewer.py so the
line set used for add_finding/publish_review validation is computed
from the complete diff, not the truncated prompt text.
2. Slack cap dropped recent thread context: fetch_slack_thread_messages
now keeps the most recent SLACK_THREAD_MAX_MESSAGES messages (was
keeping the oldest). The tool surfaces a truncation marker in the
formatted output so the LLM knows the thread was truncated.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: shared GitHub HTTP helper with retries, rate-limit handling, and sane timeouts
Introduces agent/utils/github_http.py — a single place for GitHub API HTTP
calls with 30s/10s-connect timeouts (vs httpx's 5s default), exponential
backoff with jitter, Retry-After header support, and 429/secondary-rate-limit
detection. Migrates the reviewer publish path (reviewer_publish.py,
reviewer_diff.py, github_checks.py, github_ci.py) from one-shot
httpx.AsyncClient() calls to the shared helper.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: don't retry transport errors on non-idempotent GitHub writes
POST/DELETE/PATCH can create side effects server-side even when the client
gets a timeout or connection reset. Only retry transport errors for
idempotent methods (GET, HEAD, PUT, DELETE). 429/5xx status codes are still
retried for all methods since the server explicitly did not process the
request.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: don't retry 502/504 on non-idempotent GitHub writes
502 (bad gateway) and 504 (gateway timeout) are ambiguous — the upstream
may have processed the write before the gateway returned an error. Only
retry these for idempotent methods. 429 and 503 are still retried for all
methods since the server explicitly did not process the request.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: handle images sent to non-vision models in Slack, Linear, and web UI
Add vision capability checks across all image input paths. When a user
sends images to a text-only model (e.g. GLM 5.2, DeepSeek V4 Pro), the
images are now skipped and a warning is injected into the prompt instead
of sending unsupported content to the model.
- Slack: resolve model at webhook time, skip image fetch + add warning
- Linear: same pattern as Slack
- Queued message middleware: read resolved model from thread metadata,
strip images from queued payloads for text-only models
- Web UI: disable submit + show inline warning when images are attached
to a non-vision model selection
- Shared: resolve_agent_model_id helper + vision_not_supported_warning
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: mock resolve_agent_model_id in Slack mention test
The test_process_slack_mention_queues_active_thread_message test was
missing a mock for the new resolve_agent_model_id call added to the
Slack webhook handler, causing a TypeError when image URLs triggered
the model resolution path.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: include vision warning in queued payload for text-only models
Update the prompt variable (not just content_blocks) before clearing
image_urls so the queued payload also carries the warning text when a
Slack/Linear follow-up arrives while the thread is busy.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
The "Made by Open SWE" PR footer linked to the generic dashboard
homepage. Point it at the dashboard thread that generated the PR
(/agents/<thread_id>), falling back to the homepage when no thread
id is available.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: CI auto-fix and PR babysitting for agent PRs
Watch CI failures and review feedback on PRs Open SWE opened, then dispatch
confidence-gated fix runs on the originating agent thread. Adds CI webhook
ingestion (check_run/check_suite/workflow_run/status), a per-PR @open-swe
autofix on|off toggle, auto-response to review comments, and a polling
ci_monitor graph that also flags merge conflicts. Gated by the existing
autofix_mode/trigger_mode settings, the enabled-repos opt-in, base-branch and
human-commit skip rules, dedupe, and a per-PR attempt cap.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review feedback on CI auto-fix
- Security: gate the no-mention review-feedback path on author trust —
require a trusted author_association (OWNER/MEMBER/COLLABORATOR) plus a
GitHub write/maintain/admin permission check before dispatching a
write-capable agent run, preventing privilege escalation from
read/triage/outside reviewers.
- Auth: reuse the originating PR thread's source + login/email when
dispatching fix runs so the GitHub-token resolver authenticates them in
non-bot-token deployments (bespoke github_ci source failed to resolve).
- Docs: document the Commit statuses: Read-only permission required for the
Status webhook event.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: resolve trace URL project id by tracing project name
Graphs were split into separate LangSmith tracing projects
(open-swe-agent, open-swe-review) but the "View trace" link still used
a single fixed project-id env var pointing at the old combined project.
Resolve the project id from the tracing project name so agent and
reviewer links point at their respective projects, falling back to the
env var when resolution is unavailable.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* docs: document per-graph tracing projects for trace links
View trace links now resolve project IDs from the open-swe-agent /
open-swe-review project names. Document this so fresh deployments create
the right projects instead of relying on a single project ID.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
The top-level review comment's "Open in Web" link pointed at the
agent thread (/agents/{thread_id}). Point it at the dashboard review
detail page (/agents/reviews/{owner}/{repo}/{number}) instead.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: link Slack thread/Linear ticket in PRs and append ticket to title
When opening PRs, include any referenced Slack thread or Linear ticket
in the description and append the resolvable ticket number to the title.
Adds a Slack permalink to the webhook prompt so the agent has a
ready-to-link thread URL.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: use [closes <TICKET>] format in PR title suffix
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: deterministically append source refs to private-repo PRs only
Move the Slack/Linear source-reference linking out of the prompt and into
open_pull_request, gated to private repos so private Slack thread URLs and
Linear identifiers are never published to a public PR. Reverts the prompt
instruction and webhook permalink injection in favor of this server-side append.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
prepare_review_repo is best-effort, but the system prompt unconditionally
told the agent the repo is checked out at the PR head. On a reused sandbox,
a dirty worktree (or a transient fetch failure) makes 'git checkout <sha>'
fail; prep returns False, the old checkout stays in place, and the reviewer
confidently reads stale code — observed as 'I rechecked the current head'
replies quoting pre-push code.
- checkout with --force so leftover worktree state can't block it, verify
HEAD matches the requested sha, tolerate 'git fetch --all' failures
- when prep fails, the prompt now warns the checkout may be stale and tells
the agent to re-fetch/checkout (or fall back to API file contents) before
trusting local files
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: track PR lifecycle state per thread for sidebar
Persist a PR's draft/open/merged/closed state on the agent thread and keep
it in sync as PR webhooks fire, so the Agents UI can show per-thread PR
status the way Cursor does. open_pull_request now records the initial
draft/open state and pr_title; thread summaries expose diffStats; and the
PR webhook refreshes pr_state on close/reopen/draft/ready transitions.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor: consolidate PR state mapping into shared derive_pr_state helper
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: refresh sandbox GitHub proxy token before mid-run expiry
GitHub App installation tokens expire after exactly 1 hour. The LangSmith
sandbox proxy was configured once at run start with a snapshot of that
token, so runs longer than ~1h hit 401s on every gh/git call. Record the
proxy token's expiry per thread and add a before-model hook that
re-configures the proxy with a fresh token when it nears expiry.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preserve repo-scoped proxy token on mid-run refresh
Reviewer runs mint a repository-scoped installation token. Record the
repo scope per thread alongside the expiry so the before-model refresh
re-mints a token with the same scope instead of an installation-wide
token, avoiding privilege expansion on long reviewer runs.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: update passthrough stub for github_proxy_repositories param
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: keep review check visible on follow-up commits and stop using neutral
The "Open SWE Review" check completed as `neutral` whenever findings were
surfaced, which GitHub renders as a confusing "neutral check" group. Always
complete it as `success` (informational/non-blocking), matching Devin and
Corridor — the finding count stays in the title and findings post as comments.
Also create a fresh check run on the new head SHA in the push re-review path:
GitHub only shows check runs on a PR's current head, so the check vanished
after a follow-up push (and the stale id settled on an outdated commit).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: surface settled review check when push leaves diff unchanged
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: report Open SWE Review as a PR check run
Auto-review dispatch now creates an in-progress 'Open SWE Review' check run
on the PR head SHA; publish_review completes it (neutral with findings,
success when clean). An after-agent hook fails the check if the run dies
before publishing. Requires the GitHub App's Checks: Read & write permission;
all calls are best-effort so a missing permission never breaks reviews.
* fix: address review feedback on check-run settling
Keep review_check_run_id when the completion PATCH fails so a later
publish or the after-agent hook can retry instead of hanging the check;
count out-of-diff findings toward the check conclusion.
* fix: retry failed check completion with the real publish conclusion
A transient PATCH failure after a successful publish previously left the
check id for the after-agent hook, which settled it as 'failure'. Persist
the intended result as review_check_pending_result and have the hook
prefer it over the generic failure fallback.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: prep reviewer repo at init and load repo skills
Clone + checkout the PR head during reviewer agent init so SkillsMiddleware
can discover the repo's .agents/skills and .claude/skills from disk at its
one-shot scan, and so the LLM no longer narrates the clone mid-run.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: load reviewer skills from trusted base sha and fetch PR pull ref
Address review findings: skills are now extracted from the PR base sha via
git archive into a dir outside the checkout (prevents PR-authored SKILL.md
prompt injection), and repo prep fetches refs/pull/<n>/head with a strict
checkout so fork PRs fail loudly instead of silently reviewing the default
branch.
* fix: drop ref from skill-extraction log to satisfy CodeQL
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Pull the api-standards skill from the LangSmith Context Hub at reviewer
run start and inject it into the system prompt, gated on the PR adding or
modifying an API surface. Best-effort: failures fall back to no supplement.
Co-authored-by: GowriH-1 <218394553+GowriH-1@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Add conversations.info fetch so a repo:owner/name (or GitHub URL) token
in a Slack channel's topic/purpose pins the channel to a repo, slotting
in just below thread metadata in get_slack_repo_config.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat(reviewer): surface dashboard UI link on PR reviews [INF-0000]
Post a transient "review in progress" comment (with an "Open in Web"
dashboard link) when a reviewer run starts, then delete it once the
review lands. The published review body now carries the same
"Open in Web" link, so the link persists on the review itself.
The transient comment's id is tracked in reviewer thread metadata
(status_comment_id) so it can be deleted on completion.
* refactor(reviewer): inline dashboard URL helper, drop redundant future import [INF-0000]
The bot's numeric noreply (215916821+open-swe[bot]@users.noreply.github.com)
introduced in ae946d1a doesn't resolve to a GitHub account Vercel accepts,
breaking preview deploys on PRs in langchainplus. Revert OPEN_SWE_BOT_EMAIL
to open-swe@users.noreply.github.com.
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* feat: add scheduled web agents
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* fix: secure scheduled agent repositories
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* feat: rebuild scheduled agents as Automations tab
Scrap the inline ScheduledAgentsPanel and replace it with a dedicated
Automations tab: sidebar nav entry, list view with stat cards + empty
state, and a full editor (name, Active toggle, repo, scheduled trigger
picker, agent instructions + model).
* fix: clear collapsed-sidebar button on mobile in Automations
* fix: allow clearing automation repo on update
---------
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
Reviewer/push runs execute in a worker process, but the GitHub token was
cached by the webhook handler in the API server process — a different
process — so the worker's in-process cache was always cold. resolve_github_token
then failed (User not authenticated / Unknown source: github_push) and only the
app-token fallback kept reviews working, noisily.
The reviewer always acts as the GitHub App (open-swe[bot]), so resolve the
installation token directly at run start, scoped to the repo. This also bypasses
org SAML enforcement that blocks user OAuth tokens. Drop the now-dead
cross-process cache writes in the webhook reviewer-dispatch handlers, and stop
leave_failure_comment raising on the github_push source.
Open SWE Review now runs from automated PR triggers, so drop the old Slack/GitHub review keyword entrypoints and keep PR comments on the regular agent path.
* feat: add Slack Open in Web link
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: skip web link for Slack reviewer runs
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: accept include_dashboard_link kwarg in Slack reviewer test double
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <johannes@langchain.dev>
The Co-authored-by trailer and bot git identity used
open-swe@users.noreply.github.com, which resolves to the separate
open-swe *user* account rather than the open-swe[bot] GitHub App.
Switch OPEN_SWE_BOT_EMAIL to the bot's noreply address
(215916821+open-swe[bot]@users.noreply.github.com) so co-author credit
and the fallback author identity point at the bot.
Drive the prompt trailer and sandbox git config from the constant
instead of hardcoding the address.
* fix: scope public reviewer tokens
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor: simplify reviewer token wiring; fix push re-scope + red test
- Remove the redundant _check_or_recreate_sandbox_for_proxy /
_refresh_github_proxy_or_recreate_for_proxy wrappers and call the
underlying functions directly (they already default the token to None).
- process_github_push_event: re-scope the GitHub App token when the push
payload lacked repo privacy/id but PR metadata reveals a public repo, so
reviewer.py never proxies a full-installation token for a public PR.
- Clarify the two-token sequence in trigger_pr_review_from_ref.
- Fix pre-existing failing test test_proxy_refresh_failure_recreates_sandbox
and add coverage for _reviewer_token_for_repo + push-event scoping.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: deliver Slack account-link prompt as a visible threaded reply
Blocked Slack users got no prompt at all. Prod logs show chat.postEphemeral
returns ok, but ephemeral messages are silently dropped in Slack's assistant
threads (where Open SWE runs), so the user sees nothing. Post the prompt as a
normal threaded reply instead — the same channel the agent uses to reply.
* fix: deliver Slack auth-failure prompt as a visible threaded reply
leave_failure_comment() tried an ephemeral message first and only fell back
to a thread reply on failure. Ephemeral messages succeed (ok) but are dropped
in Slack's assistant threads, so the fallback never fired and the user saw no
auth-failure prompt. Post the visible threaded reply directly, matching the
account-link prompt fix.
* fix: prompt blocked Slack users with a generic, token-free dashboard link
Addresses the review findings that posting the per-user account-link token /
auth URL in a visible thread lets any channel member bind their GitHub account
to the triggering user's Slack identity.
Drop the per-user signed link entirely. Both the account-link prompt
(_post_account_link_prompt) and the runtime auth-failure prompt
(leave_failure_comment) now post a plain dashboard settings link
(build_settings_url) as a visible threaded reply. The user signs in with GitHub
from their own session and connects Slack via verified OIDC on the settings
page — no secret in the thread, nothing to hijack, and no DM machinery.
* feat: nudge first-time users to connect Slack from the dashboard home
Show a Connect Slack banner on the agents landing page whenever Slack OAuth is
enabled and the user hasn't linked Slack yet. A first-time user (no Slack
mapping) sees it immediately after signing in; it disappears once connected.
* feat: prompt first-time users to connect Slack via a dialog
Replace the inline Connect Slack card on the agents home with a modal dialog
(Base UI). It opens automatically once the mapping query resolves to
"not connected" and closes itself once Slack is linked; "Maybe later" dismisses
it for the session. No new dependency — uses the design system's Base UI.
* copy: frame Slack connect as resolving the user's GitHub account
Drop 'act/reply on your behalf' wording across the connect-Slack dialog, the
Slack thread prompts (blocked + auth-failure), and the settings description.
Connecting Slack lets Open SWE resolve the user's GitHub account when they tag
it in Slack.
Replace the double-attribution PR body footer (_Opened collaboratively by
{user} and open-swe._) with a single Cursor-style footer linking to the
project. Since PRs are now opened as the triggering user, the user no longer
needs to be named in the footer. The commit Co-authored-by trailer is kept.
Legacy footers are migrated on PR updates.
* feat: open Slack-triggered PRs as the triggering user
Route the Slack per-user GitHub token through the dashboard OAuth store
(the backend the self-service link prompt populates) and block runs that
lack a valid user token, prompting the user to (re-)link. Per-user OAuth
now wins over bot-token-only mode for mapped Slack/dashboard users.
Flip commit/PR authorship across all sources: the triggering user is the
commit author (via repo-local git identity using their resolvable GitHub
noreply email) and open-swe[bot] is the Co-authored-by collaborator.
* fix: address PR review — shell-escape commit identity, fix token cache impersonation
- Shell-escape the triggering user's name/email with shlex.quote before
embedding them in the repo-setup `git config` command, so a name like
O'Connor (or a crafted one) can't break or inject into the command.
- Stop consulting the shared thread-metadata token cache in
_resolve_dashboard_user_token. Slack thread ids are shared across the
conversation, so a cached token from a prior triggering user could be
returned for the current github_login. Always resolve by login from the
dashboard OAuth store instead.
* feat: dashboard self-service user mapping + UI cleanup
- Add session-scoped GET/PUT /dashboard/api/my-mapping so users can set their
own work email / Slack member ID (keyed by their GitHub login, source=self).
- Slack account-link prompt now redirects to Profile Settings after auth.
- Rename "My Settings" -> "Profile Settings" and "Cloud Agents" -> "Open SWE
Agent"; remove the Integrations tab/section (folded out, low value for now)
and redirect /integrations to Profile Settings.
- Add a "User mapping" section to Profile Settings (work email used by Slack
and Linear, optional Slack member ID).
- Make dashboard auth cookies scheme-aware: Secure;SameSite=None over HTTPS,
non-Secure;SameSite=Lax over http://localhost so local login works.
* feat: self-service Slack account linking via Sign in with Slack (OIDC)
Replace the spoofable manual work-email/Slack-ID form with a verified
"Sign in with Slack" flow so a logged-in GitHub user can only ever link
their own Slack identity.
- New agent/dashboard/slack_oauth.py: OIDC authorize URL, code exchange,
userInfo identity parse, optional workspace gate, configured check.
- routes.py: session-gated GET /slack/login and /slack/callback that upsert
the mapping from Slack-verified user_id + email (source=slack_oauth).
Remove the spoofable PUT /my-mapping; expose slack_oauth_enabled on /me.
- UI: drop the editable inputs; add a Connect Slack button + status to the
User mapping section.
Admin-managed mappings are unaffected and still resolve at trigger time.
* fix: include GitHub username in collaboration footer
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: clarify legacy footer replacement
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* chore: remove legacy mapping + admin profiles section, page user mappings
- Delete the hardcoded github_user_email_map.py and the one-time
POST /admin/user-mappings/import endpoint + Import legacy mapping UI
button (the Store is now the sole source of truth post-import).
- Remove the per-user profiles admin section and its GET/PUT
/admin/profiles endpoints + ProfileForm component; users still manage
their own profile via My Settings.
- Page the user mappings list: /admin/user-mappings now takes
page/page_size and returns {items,total,page,page_size}; the admin UI
shows 20 rows per page with Previous/Next controls.
* Address review: fix stale import-button doc + clamp mappings page on shrink
* Replace hardcoded GitHub-email map with Store-backed user mapping
Move the static GITHUB_USER_EMAIL_MAP to a Store-backed bidirectional
mapping (GitHub login <-> work email <-> optional Slack ID) with an
in-process cache, self-service onboarding, and admin management.
- agent/dashboard/user_mappings.py: Store CRUD + login/email/slack-id
indexes, sync cache readers for hot paths, async fallthrough, and a
bulk_import that preserves existing richer records.
- Migrate all read sites (auth.py, agent_overrides.py, authorship.py,
github_comments.py, webapp.py x2) off the dict.
- Unmapped Slack tags now run on the GitHub App installation token
(use_installation_token_fallback) and get an ephemeral "link your
GitHub account" prompt carrying the Slack id + email via a signed
account-link token threaded through the OAuth state.
- OAuth callback completes a self-service (org-gated) mapping from that
token, falling back to the verified GitHub email.
- Admin CRUD endpoints + one-time legacy import; dashboard UI section.
- Legacy dict retained only as the import payload (no longer read).
Tests: mapping store, account-link round-trip + completion, mapped vs
unmapped Slack flows; existing trust-gate tests updated to prime cache.
* Address review: cold-cache email resolution + stale alias de-indexing
- agent_overrides: add resolve_login_from_email_async that falls through to
the Store on a cold cache; use it at the async repo-resolution call sites
(Slack repo config, Linear comment, owner-metadata) so a mapped user still
resolves to their GitHub login + dashboard default_repo on a fresh worker.
- user_mappings.upsert_mapping: de-index the existing login before re-indexing
so a changed email/Slack id no longer leaves stale aliases resolving to the
login in-process.
- Tests for both fixes; update Slack repo-config test to patch the async resolver.
* Lock dashboard login to GitHub org members
Add an org-membership gate to the dashboard OAuth callback. After
resolving the GitHub login, enforce_org_login_gate(login) checks the
existing ALLOWED_GITHUB_ORGS allowlist before issuing a session.
- Reuses ALLOWED_GITHUB_ORGS (no new config knob) and
is_user_active_org_member (installation-token check, so no extra
OAuth scope and private memberships are visible).
- Fail-open when unset/blank so existing deployments keep working;
fail-closed on API errors.
- Gate runs before the session cookie/token is persisted.
Adds unit tests and documents the behavior in INSTALLATION.md.
* docs: document Organization Members permission required for org login gate