Commit graph

19 commits

Author SHA1 Message Date
Christian Bromann
abf354bb05
feat(open-swe): stream agent chat via @langchain/react v2 protocol (#1475)
* feat(dashboard): stream agent chat via @langchain/react v2 protocol

Replace the bespoke SSE + React Query polling path with LangGraph’s
v2 event stream through credentialed dashboard proxies. Run starts go
through stream commands; mid-run follow-ups still queue via /messages.

* fix import path

* fix tests after rebase

* format

* PR feedback

* improved model fallback

* fix image handling

* embrace sdk

* cleanup

* cr

* more cleanup

* fix cors

* harden security

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 09:54:35 -07:00
Johannes du Plessis
c7acf2ff27
Revert out-of-process usage-snapshot builder (#1473)
* Revert "fix: schedule usage snapshot runs with explicit empty input (#1470)"

This reverts commit 49aa638c58.

* Revert "feat: out-of-process usage-snapshot builder (Phase 1) (#1468)"

This reverts commit 798cc6edf1.
2026-06-09 13:19:11 -07:00
Johannes du Plessis
798cc6edf1
feat: out-of-process usage-snapshot builder (Phase 1) (#1468)
* feat: out-of-process usage-snapshot builder (Phase 1)

Move all usage-tab compute off the run-serving HTTP process (the #1434 bug
class). Read path is now a pure cache read with a typed computing placeholder
on cold miss; a dedicated usage_snapshot graph + global ~10min cron rebuild
every period's snapshot out-of-process, wrapped in asyncio.timeout and gated by
USAGE_SNAPSHOT_CRON_ENABLED. Lifespan makes one fire-and-forget scheduling call,
never a retained/looping task.

* fix: address PR review on usage-snapshot builder

- admin guard on /admin/usage/rebuild (was any logged-in user)
- cap lifespan loopback calls with asyncio.timeout(5) so a startup hang
  can't block boot
- memoize cron registration so steady-state requests skip the loopback
  check; reap duplicate crons from concurrent-replica races
- thread the computing flag through the usage payload too
2026-06-09 11:11:06 -07:00
Johannes du Plessis
c28b1641f8
fix: add 30-day thread ttl (#1430)
* fix: add two-week thread ttl

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: use 30-day thread ttl

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-05 11:47:15 -07:00
Johannes du Plessis
6895ddcedc
feat: add scheduled web agents (#1422)
* feat: add scheduled web agents

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* fix: secure scheduled agent repositories

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* feat: rebuild scheduled agents as Automations tab

Scrap the inline ScheduledAgentsPanel and replace it with a dedicated
Automations tab: sidebar nav entry, list view with stat cards + empty
state, and a full editor (name, Active toggle, repo, scheduled trigger
picker, agent instructions + model).

* fix: clear collapsed-sidebar button on mobile in Automations

* fix: allow clearing automation repo on update

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-05 02:20:24 +00:00
Johannes du Plessis
4a55145bb1
feat: outcomes dataset + bootstrap/continual split via skills (#1365)
* fix: reset stale sandbox creation sentinel

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: treat SANDBOX_CREATING as a timestamped cross-process lock

Only reset the sentinel when proven stale (older than the creation
timeout); otherwise wait for the worker that holds the lock so a
concurrent run does not create a duplicate sandbox.

* feat(analyzer): outcomes dataset + bootstrap/continual split via skills

Rename the review_style_analyzer graph to `analyzer` and split it into two
modes, plus capture reviewer finding outcomes for continual learning.

- Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false
  positive), and GitHub/Slack thumbs findings into a single LangSmith dataset
  (openswe-reviewer-outcomes), keyed deterministically per finding+source.
  Emit points wired into update_finding, resolve_finding_thread, and the
  GitHub/Slack reaction handlers.
- Two playbooks delivered as deepagents skills (bootstrap-repo-analysis,
  continual-learning), served as virtual files via a CompositeBackend /skills/
  route + StateBackend (seeded into the run files channel at invoke time, never
  written to the sandbox). Mode is set by the launcher; continual runs fall
  back to the GitHub App installation token.
- Split launcher into start_bootstrap_analysis + start_continual_run; register
  a per-repo nightly continual-learning cron when bootstrap completes.
- New read_finding_outcomes tool feeds confirmed/dismissed findings back to the
  continual playbook.

Tests for outcome label mapping, skills helper, and cron idempotency.

* fix(analyzer): anchor continual cron runs to a real thread_id

The nightly continual-learning cron is threadless, and get_analyzer
early-returns an empty agent when configurable.thread_id is missing — so
every cron-launched run no-op'd before reading outcomes or saving a refined
prompt. Include the repo's deterministic analyzer thread_id in the continual
run configurable so the run executes; the threadless run carries no message
history, so nightly runs don't accumulate context.

* refactor(analyzer): move cron lifecycle calls out of the review-styles store

Drop the inline `analyzer_cron` imports from review_styles.py (added only to
dodge a circular import) by relocating the cron-trigger calls to the layer
above the store: registration to the save_review_style tool (after a prompt is
saved) and removal to the dashboard delete route. review_styles.py is now a
pure store again with top-level imports only.

* refactor: hoist reviewer_outcomes imports to module level

Move the two inline emit_finding_status_outcome imports introduced in this PR
(update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes
only depends on langsmith, so there is no circular import to avoid.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
Johannes du Plessis
82852f9eda
feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312)
* feat: tune reviewer for precision — web/wiki tools + recalibrated prompt

Reviewer agent now has web_search, fetch_url, and http_request alongside the
finding tools, so it can verify library semantics and consult the DeepWiki
auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>)
before flagging cross-file or architectural concerns.

Prompt rewritten to push precision over recall:
- explicit severity ladder pushing reviews toward bimodal high/low instead of
  defaulting to medium
- ≤200-char description target (gold set averages ~186 chars; we were at ~436)
- mandatory docs / wiki / code lookup before flagging concurrency, security,
  or perf — the three categories that dominated false positives
- "do not flag" list covering compiler/linter-catchable nits, speculative
  claims without a concrete attacker/interleaving/scale, style preferences
  the codebase doesn't share, and test-quality nits on non-test diffs
- smart file-selection guidance for large PRs (deprioritize generated /
  vendored / pure-rename hunks)

Eval config switched to openai:gpt-5.5 + high reasoning effort for the next
benchmark run.

* trim prompt

* subagent prompting

* confidence ratings

* added medium

* enforce confidence threshold

* .

* reviewer: precision-tuned prompt + drop confidence gate

Rewrites the reviewer system prompt around a defensibility bar (anchor +
failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file
list (style nits, speculation, scope-policing, same-bug fan-out), and a
checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden
set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26%
style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new
prompt targets each class directly.

Confidence is still recorded on every finding for post-hoc calibration but
no longer gates publication — the audit showed the gate was a no-op (agent
self-rated 65% of findings "high" regardless), and the prompt's defensibility
bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD,
the confidence_threshold kwarg on filter_findings_for_publish, the
confidence_filtered score_mode, and the min_confidence kwarg on the eval
target's _extract_comments — all dead once the gate is gone.

Also removes the "informational" severity tier from the Severity enum,
SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for
FYI observations the dataset never rewards.

* benchmax

* adding google provider

* slight steering

* tuning

* more tuning

* fix

* cleanup

* reducing overfitting

* Add per-repo review style profiles and inject them into the reviewer.

Dashboard users can analyze historical PR review feedback per repository,
edit the resulting style guide, and have it loaded from LangGraph Store at
reviewer runtime (including Martian eval runs) keyed by owner/name.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix review style job errors leaking exception details to clients.

Return generic dashboard messages while logging full stack traces server-side.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-20 18:35:00 +00:00
Johannes du Plessis
ace71b0fd0
feat: add reviewer graph + eval target wiring (#1241)
* feat: add reviewer graph + eval target wiring

- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
  alongside the main `agent` graph. Reuses the same sandbox lifecycle,
  GH proxy auth, and middleware primitives from `agent.server`, but with
  a narrower tool set, a reviewer-specific system prompt, no
  commit/push, and the `task` (subagent) tool stripped via
  `_ToolExclusionMiddleware` so review stays in one context.

- New `github_comment` tool: agents call it once per issue with
  `(file, line, body, severity)` and the eval scores those calls
  against golden comments.

- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
  *not* on the reviewer's stack — that middleware exists to enforce the
  main agent's "always finalize via Slack/Linear/PR" contract, which
  the reviewer doesn't have. The main agent's behavior is unchanged.

- `evals/reviewer/target.py`: send PR info as a user message, extract
  every `github_comment` tool call (multiple expected per review) into
  the run output.

- `evals/reviewer/judge.py`: per-example evaluator now returns a list
  of metrics under `{"results": [...]}` so LangSmith averages each
  numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
  the UI. Dropped the broken `aggregate_pr` summary evaluator that
  reached for an attribute that doesn't exist on `RunTree`.

- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
  `client.list_examples(limit=N)` since `aevaluate` doesn't accept
  `max_examples`.

- Makefile: `dev` and `run` targets now use `uv run` so they work
  without an activated venv.

* resolve comments

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
Brace Sproul
bd52e5e09d
chore: Drop monorepo (#1029)
* chore: Drop monorepo

* cr
2026-03-06 16:10:34 -08:00
Brace Sproul
00e4eaff8e
fix: Add configurable headers to langgraph.json (#824) 2025-09-03 23:05:02 +00:00
open-swe[bot]
2f561796a5
feat: Implement Request Human Help Tool for Agent-Human Collaboration (#505)
* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* fix open-swe syntax errors

* fix render text input for human_request_help

* improve UI/UX post response (close input render human msg)

* remove comments

* move request_human_help to its own component file

* format

* update github comment and abstract getopensweurl util

* add username tag in request help comment, rm comments

* cr

* refactor

* cr

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Dylan Boudro <121908331+starmorph@users.noreply.github.com>
Co-authored-by: starmorph <dylan@starmorph.com>
Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2025-07-30 17:48:42 -07:00
Brace Sproul
ba0690caa0
feat: trigger runs from GH issues (#324)
* feat: trigger runs from GH issues

* cr

* working

* cr

* done

* cr

* cr

* cr

* cr

* diff label for dev vs prod

* cr

* cr

* cr
2025-06-30 16:43:09 -06:00
Brace Sproul
ff8a0f7167
[WIP] feat: New multi agent GitHub based workflow (#175)
* feat: New multi agent GitHub based workflow

* fix: read plan from gh issue in init issue node

* format

* cr

* fix how new planner threads are started, allow for creating new issues

* fix: read tasks from gh issue inside manager when taking any action

* cr

* update planner

* feat: port intrerrupt node to planner

* add interrupt routing in graph

* update programmer nodes and state

* read messages from issue in programmer

* lint fix

* tag issue when opening PR

* better pr title and body generation prompts

* start runs from planner interrupt and update issue with task plan

* fix lg server build issues

* cr

* fix cloning and pushing

* cr

* cr

* feat: v0 copy (#194)

* feat: v0 copy

* cr

* all pages 'working'

* cr

* cr

* cr

* cr
2025-06-18 11:03:39 -07:00
Brace Sproul
bf72e1faf7
feat: GitHub e2e auth (#106)
* feat: GitHub e2e auth

* cr

* reimplement proxy route

* fix github auth

* cr

* cr

* cr

* cr

* cr
2025-06-11 21:00:08 +00:00
Brace Sproul
8afe1e6803
feat: Update chat UI to be Open SWE specific (#76) (#91)
* init repo-selector

* add branch selector UI, extend GH hook w branch info

* add   reconnectOnMount: true resumeStream, remove repo description from selector, fix styling

* styling repo + branch

* udpate getBranch calls to include targetRepo arg

* format

* add search to repo + branch selectors

* add settings icon link to /github for dev convenience

* add default branch selection

* format

* minor cleanup

* init task list UI demo

* individual task component

* tasks real data

* task visual improvement

* persisting tasks, navigate to thread onClick

* task UI

* TaskSidebar, Navigate to tasks

* org tasks by thread

* group by thread

* improve sidebar, task-list, rm github modal on refresh

* collapse sidebar works when viewing a thread

* remove micro-tasks expansion from threads component, change taskId to taskIndex for persistence + simplicity

* init config sidebar ui

* open and close config sidebar

* add ref for getAllTasks + add graph_id metadata

* format

* create threadItem component

* format

* add threadUtils

* centralize + refactor types

* format

* improve thread status, realtime updates

* fix scroll issue

* remove temperature from config-sidebar

* Change model via sidebar

* fix flashing on updates

* init plan UI

* centralize types

* self CR: refactoring, cleanup, remove comments, add todos

* CR: clean comments + logs, fix repo selector bug

* fix: Update example env file in web (#77)

* disable auto scroll on pageload

* reposelector: use default branch not main master

* remove archived tab temp, default branch listed first in branchSelector dropdown

* auto select first repo

* improve status indicator

* refactor status indicator

* simplify status implementation to use Langgraph

* util moved from provider to thread-utils

* linting, improve task completion logic

* feat: Support followup requests (#60)

* feat: Support followup requests

* cr

* cr

* fix

* account for pr already exists

* remove task provider

* fix: Better prompting and context provided for codebase structure awareness (#85)

* fix: Better prompting and context provided for codebase structure awareness

* cr

* linting + skeleton loader

* fix: branch selector, add github api pagination

* small branch update

* update task completion status realtime

* threadItem stream remove flashing/polling

* feat: Create a shared package (#89)

* feat: Create a shared package

* formatting

* fix: drop passing target repo to getBranchName

* drop docs

* use shared package for config fields in config sidebar

* fix: styling and structure for repo/branch selectors

* fix: use shared types/objects for config fields

* default key for configs

* fix: improved styling for threads

* fix renering num tasks

* disable repo branch selector when chat started

* cr

* cr

* bump deps

* fix: rendering plan in interrupt

* cr

* cr

---------

Co-authored-by: Dylan Boudro <121908331+starmorph@users.noreply.github.com>
2025-06-09 00:29:40 +00:00
Brace Sproul
c6d867e1a7
feat: Monorepo (#22)
* feat: Monorepo

* cr

* cr
2025-05-26 13:02:51 -07:00
Brace Sproul
fd0d50be26
fix: Docs and scripts (#15)
* fix: Docs and scripts

* cr

* fix: better err handling

* cr
2025-05-23 16:43:14 -07:00
Brace Sproul
66f32b4afc
feat: Implement a planning subgraph (#12) 2025-05-22 19:05:29 -07:00
bracesproul
e2415aee3f init commit 2025-05-21 14:47:56 -07:00