Commit graph

4 commits

Author SHA1 Message Date
Johannes du Plessis
3992d3ef5d
feat: repo-scoped dynamic sandbox snapshots (#1595)
* feat: repo-scoped dynamic sandbox snapshots

Let admins build a per-repo sandbox image from a custom Dockerfile so runs
targeting that repo boot from a snapshot with its deps pre-baked. Snapshot
selection is purely additive: repos without a `ready` repo-scoped snapshot
always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID.

Backend adds a repo_snapshots store module (Dockerfile + build status keyed by
owner/name), threads the resolved repo through the LangSmith sandbox creation
path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a
throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The
UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor +
build status/logs).

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden repo snapshot builds

Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins
cannot accidentally build a repo snapshot from a bare Python image that lacks
Open SWE's sandbox tools. Allow stale building records to be retried by tracking
build_started_at and treating old or missing timestamps as stale.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: document repo snapshot base image config

Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert
missing base-image configuration into a handled dashboard API error so admins see
a clear configuration message instead of an unhandled template-generation error.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
Johannes du Plessis
da20e57bcc
fix: refresh sandbox GitHub proxy token before mid-run expiry (#1496)
* fix: refresh sandbox GitHub proxy token before mid-run expiry

GitHub App installation tokens expire after exactly 1 hour. The LangSmith
sandbox proxy was configured once at run start with a snapshot of that
token, so runs longer than ~1h hit 401s on every gh/git call. Record the
proxy token's expiry per thread and add a before-model hook that
re-configures the proxy with a fresh token when it nears expiry.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve repo-scoped proxy token on mid-run refresh

Reviewer runs mint a repository-scoped installation token. Record the
repo scope per thread alongside the expiry so the before-model refresh
re-mints a token with the same scope instead of an installation-wide
token, avoiding privilege expansion on long reviewer runs.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: update passthrough stub for github_proxy_repositories param

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 10:59:21 -07:00
Johannes du Plessis
0a2e682364
fix: scope public reviewer tokens (#1389)
* fix: scope public reviewer tokens

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: simplify reviewer token wiring; fix push re-scope + red test

- Remove the redundant _check_or_recreate_sandbox_for_proxy /
  _refresh_github_proxy_or_recreate_for_proxy wrappers and call the
  underlying functions directly (they already default the token to None).
- process_github_push_event: re-scope the GitHub App token when the push
  payload lacked repo privacy/id but PR metadata reveals a public repo, so
  reviewer.py never proxies a full-installation token for a public PR.
- Clarify the two-token sequence in trigger_pr_review_from_ref.
- Fix pre-existing failing test test_proxy_refresh_failure_recreates_sandbox
  and add coverage for _reviewer_token_for_repo + push-event scoping.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-03 09:25:17 -07:00
Johannes du Plessis
4a55145bb1
feat: outcomes dataset + bootstrap/continual split via skills (#1365)
* fix: reset stale sandbox creation sentinel

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: treat SANDBOX_CREATING as a timestamped cross-process lock

Only reset the sentinel when proven stale (older than the creation
timeout); otherwise wait for the worker that holds the lock so a
concurrent run does not create a duplicate sandbox.

* feat(analyzer): outcomes dataset + bootstrap/continual split via skills

Rename the review_style_analyzer graph to `analyzer` and split it into two
modes, plus capture reviewer finding outcomes for continual learning.

- Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false
  positive), and GitHub/Slack thumbs findings into a single LangSmith dataset
  (openswe-reviewer-outcomes), keyed deterministically per finding+source.
  Emit points wired into update_finding, resolve_finding_thread, and the
  GitHub/Slack reaction handlers.
- Two playbooks delivered as deepagents skills (bootstrap-repo-analysis,
  continual-learning), served as virtual files via a CompositeBackend /skills/
  route + StateBackend (seeded into the run files channel at invoke time, never
  written to the sandbox). Mode is set by the launcher; continual runs fall
  back to the GitHub App installation token.
- Split launcher into start_bootstrap_analysis + start_continual_run; register
  a per-repo nightly continual-learning cron when bootstrap completes.
- New read_finding_outcomes tool feeds confirmed/dismissed findings back to the
  continual playbook.

Tests for outcome label mapping, skills helper, and cron idempotency.

* fix(analyzer): anchor continual cron runs to a real thread_id

The nightly continual-learning cron is threadless, and get_analyzer
early-returns an empty agent when configurable.thread_id is missing — so
every cron-launched run no-op'd before reading outcomes or saving a refined
prompt. Include the repo's deterministic analyzer thread_id in the continual
run configurable so the run executes; the threadless run carries no message
history, so nightly runs don't accumulate context.

* refactor(analyzer): move cron lifecycle calls out of the review-styles store

Drop the inline `analyzer_cron` imports from review_styles.py (added only to
dodge a circular import) by relocating the cron-trigger calls to the layer
above the store: registration to the save_review_style tool (after a prompt is
saved) and removal to the dashboard delete route. review_styles.py is now a
pure store again with top-level imports only.

* refactor: hoist reviewer_outcomes imports to module level

Move the two inline emit_finding_status_outcome imports introduced in this PR
(update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes
only depends on langsmith, so there is no circular import to avoid.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00