feat: outcomes dataset + bootstrap/continual split via skills (#1365)
* fix: reset stale sandbox creation sentinel
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: treat SANDBOX_CREATING as a timestamped cross-process lock
Only reset the sentinel when proven stale (older than the creation
timeout); otherwise wait for the worker that holds the lock so a
concurrent run does not create a duplicate sandbox.
* feat(analyzer): outcomes dataset + bootstrap/continual split via skills
Rename the review_style_analyzer graph to `analyzer` and split it into two
modes, plus capture reviewer finding outcomes for continual learning.
- Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false
positive), and GitHub/Slack thumbs findings into a single LangSmith dataset
(openswe-reviewer-outcomes), keyed deterministically per finding+source.
Emit points wired into update_finding, resolve_finding_thread, and the
GitHub/Slack reaction handlers.
- Two playbooks delivered as deepagents skills (bootstrap-repo-analysis,
continual-learning), served as virtual files via a CompositeBackend /skills/
route + StateBackend (seeded into the run files channel at invoke time, never
written to the sandbox). Mode is set by the launcher; continual runs fall
back to the GitHub App installation token.
- Split launcher into start_bootstrap_analysis + start_continual_run; register
a per-repo nightly continual-learning cron when bootstrap completes.
- New read_finding_outcomes tool feeds confirmed/dismissed findings back to the
continual playbook.
Tests for outcome label mapping, skills helper, and cron idempotency.
* fix(analyzer): anchor continual cron runs to a real thread_id
The nightly continual-learning cron is threadless, and get_analyzer
early-returns an empty agent when configurable.thread_id is missing — so
every cron-launched run no-op'd before reading outcomes or saving a refined
prompt. Include the repo's deterministic analyzer thread_id in the continual
run configurable so the run executes; the threadless run carries no message
history, so nightly runs don't accumulate context.
* refactor(analyzer): move cron lifecycle calls out of the review-styles store
Drop the inline `analyzer_cron` imports from review_styles.py (added only to
dodge a circular import) by relocating the cron-trigger calls to the layer
above the store: registration to the save_review_style tool (after a prompt is
saved) and removal to the dashboard delete route. review_styles.py is now a
pure store again with top-level imports only.
* refactor: hoist reviewer_outcomes imports to module level
Move the two inline emit_finding_status_outcome imports introduced in this PR
(update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes
only depends on langsmith, so there is no circular import to avoid.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
|
|
|
"""Repo-bundled analyzer skills, served to the agent as virtual files.
|
|
|
|
|
|
|
|
|
|
The two analyzer playbooks live as ``SKILL.md`` files under ``agent/skills/``.
|
|
|
|
|
They are surfaced to the deepagents ``SkillsMiddleware`` via a ``StateBackend``
|
|
|
|
|
mounted at ``/skills/`` in a ``CompositeBackend`` — so the agent reads them with
|
|
|
|
|
``read_file`` without anything ever being written to the execution sandbox.
|
|
|
|
|
|
|
|
|
|
The ``files`` channel is seeded at invoke time (see the launchers). Because
|
|
|
|
|
``CompositeBackend`` strips the ``/skills/`` route prefix before delegating to the
|
|
|
|
|
``StateBackend``, the seeded keys are the *stripped* paths (e.g.
|
|
|
|
|
``/bootstrap-repo-analysis/SKILL.md``), while the agent and ``SkillsMiddleware``
|
|
|
|
|
address them under ``/skills/...``.
|
|
|
|
|
"""
|
|
|
|
|
|
|
|
|
|
from __future__ import annotations
|
|
|
|
|
|
|
|
|
|
from pathlib import Path
|
|
|
|
|
from typing import Any
|
|
|
|
|
|
|
|
|
|
SKILLS_DIR = Path(__file__).resolve().parent.parent / "skills"
|
|
|
|
|
SKILLS_ROUTE = "/skills/"
|
|
|
|
|
|
|
|
|
|
BOOTSTRAP_SKILL = "bootstrap-repo-analysis"
|
|
|
|
|
CONTINUAL_SKILL = "continual-learning"
|
|
|
|
|
|
|
|
|
|
ANALYZER_MODES = {"bootstrap": BOOTSTRAP_SKILL, "continual": CONTINUAL_SKILL}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def skill_path_for_mode(mode: str) -> str:
|
|
|
|
|
"""Return the agent-facing ``/skills/<name>/SKILL.md`` path for a run mode."""
|
|
|
|
|
skill = ANALYZER_MODES.get(mode, BOOTSTRAP_SKILL)
|
|
|
|
|
return f"{SKILLS_ROUTE}{skill}/SKILL.md"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def build_skill_files() -> dict[str, Any]:
|
|
|
|
|
"""Return ``{stripped_path: FileData}`` for every bundled analyzer skill.
|
|
|
|
|
|
|
|
|
|
Seed this into the run input's ``files`` so the ``/skills/`` StateBackend route
|
|
|
|
|
can serve them. Keys omit the ``/skills`` prefix (stripped by the composite
|
|
|
|
|
route); values are ``FileData`` v2 entries.
|
|
|
|
|
"""
|
feat: managed LangGraph Cloud + Vercel migration (PR2 — code fixes + docs) (#65)
* fix(dashboard): managed-cloud OAuth hardening + admin user-mapping endpoint
Prepare the dashboard backend for the managed LangGraph Cloud + Vercel
runtime, where the API is HTTPS and cross-site from the UI.
- OAuth redirect_uri (#2): coerce a schemeless DASHBOARD_API_BASE_URL to
https:// in _api_base_url() so GitHub stops rejecting login with
"redirect_uri not associated with this application". _cookie_security()
now treats a schemeless (managed) value as Secure; SameSite=None too,
consistent with the coerced scheme.
- OAuth state cookie (#3): document that osw_oauth_state is host-only by
design (a Domain cookie is unsafe across *.vercel.app, a public suffix),
so login must always start on the stable alias to avoid "oauth state
mismatch". Operational contract; no behavioral change.
- Admin user mappings (#4): add POST /admin/user-mappings so an admin can
set the github_login -> work_email link from the dashboard instead of a
raw Store write. New "admin" MappingSource provenance value.
* fix(webapp): refresh user-mapping cache on GitHub webhook paths
On managed LangGraph Cloud the backend runs multiple replicas, so the
per-process GitHub<->work-email mapping cache can be stale on the replica
handling a webhook (a mapping created on another replica is invisible
until refresh). process_github_pr_comment and process_github_issue now
refresh the cache from the durable Store before resolving the author's
email, matching the existing Slack mention path (process_slack_mention).
* perf(webapp): defer deepagents import to speed custom-app cold start
The custom FastAPI app (agent.webapp:app, the langgraph.json http.app)
pulled deepagents -> langchain_anthropic -> anthropic into its import
graph via dashboard.routes, only to build skill/chat seed files. Defer
those create_file_data imports into the functions that use them. Removes
deepagents/langchain_anthropic/anthropic from app import entirely and
roughly halves module-import wall time (~0.6-0.8s -> ~0.35s warm; larger
cold-start saving since native anthropic init is skipped). Behavior
identical. (reviewer_diff already imports deepagents under TYPE_CHECKING.)
* feat(ui): set work_email user mappings from the admin dashboard
Add an "Add / update" form to the admin User mappings section and the
adminUpsertUserMapping API client method, wiring the new
POST /admin/user-mappings endpoint. Admins can now create or update a
github_login -> work_email mapping directly instead of waiting for the
user to self-connect Slack.
* docs: document managed LangGraph Cloud + Vercel deployment
- INSTALLATION §10: add the managed production env triad (LANGGRAPH_URL,
DASHBOARD_BASE_URL + DASHBOARD_API_BASE_URL with https://, empty
VITE_DASHBOARD_API_BASE_URL for same-origin), the stable-alias login
and vercel.json stable-deployment-URL requirements, multi-replica cache
note, plus redirect_uri-scheme and oauth-state-mismatch troubleshooting.
Refresh the langgraph.json snippet to all six graphs.
- README: reframe deployment around the managed migration; link the plan.
- deploy/MIGRATION.md: import the self-hosted -> managed migration plan.
2026-06-29 19:58:38 -04:00
|
|
|
# Deferred import: deepagents (and its langchain_anthropic / anthropic
|
|
|
|
|
# transitive deps) is heavy (~0.7s) and is otherwise pulled into the custom
|
|
|
|
|
# FastAPI app's import chain via dashboard.routes, slowing cold start. It's
|
|
|
|
|
# only needed when a skill bundle is actually built (analyzer launch), so
|
|
|
|
|
# import it lazily here.
|
|
|
|
|
from deepagents.backends.utils import create_file_data
|
|
|
|
|
|
feat: outcomes dataset + bootstrap/continual split via skills (#1365)
* fix: reset stale sandbox creation sentinel
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: treat SANDBOX_CREATING as a timestamped cross-process lock
Only reset the sentinel when proven stale (older than the creation
timeout); otherwise wait for the worker that holds the lock so a
concurrent run does not create a duplicate sandbox.
* feat(analyzer): outcomes dataset + bootstrap/continual split via skills
Rename the review_style_analyzer graph to `analyzer` and split it into two
modes, plus capture reviewer finding outcomes for continual learning.
- Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false
positive), and GitHub/Slack thumbs findings into a single LangSmith dataset
(openswe-reviewer-outcomes), keyed deterministically per finding+source.
Emit points wired into update_finding, resolve_finding_thread, and the
GitHub/Slack reaction handlers.
- Two playbooks delivered as deepagents skills (bootstrap-repo-analysis,
continual-learning), served as virtual files via a CompositeBackend /skills/
route + StateBackend (seeded into the run files channel at invoke time, never
written to the sandbox). Mode is set by the launcher; continual runs fall
back to the GitHub App installation token.
- Split launcher into start_bootstrap_analysis + start_continual_run; register
a per-repo nightly continual-learning cron when bootstrap completes.
- New read_finding_outcomes tool feeds confirmed/dismissed findings back to the
continual playbook.
Tests for outcome label mapping, skills helper, and cron idempotency.
* fix(analyzer): anchor continual cron runs to a real thread_id
The nightly continual-learning cron is threadless, and get_analyzer
early-returns an empty agent when configurable.thread_id is missing — so
every cron-launched run no-op'd before reading outcomes or saving a refined
prompt. Include the repo's deterministic analyzer thread_id in the continual
run configurable so the run executes; the threadless run carries no message
history, so nightly runs don't accumulate context.
* refactor(analyzer): move cron lifecycle calls out of the review-styles store
Drop the inline `analyzer_cron` imports from review_styles.py (added only to
dodge a circular import) by relocating the cron-trigger calls to the layer
above the store: registration to the save_review_style tool (after a prompt is
saved) and removal to the dashboard delete route. review_styles.py is now a
pure store again with top-level imports only.
* refactor: hoist reviewer_outcomes imports to module level
Move the two inline emit_finding_status_outcome imports introduced in this PR
(update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes
only depends on langsmith, so there is no circular import to avoid.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
|
|
|
files: dict[str, Any] = {}
|
|
|
|
|
for skill in ANALYZER_MODES.values():
|
|
|
|
|
skill_md = SKILLS_DIR / skill / "SKILL.md"
|
|
|
|
|
text = skill_md.read_text(encoding="utf-8")
|
|
|
|
|
files[f"/{skill}/SKILL.md"] = create_file_data(text)
|
|
|
|
|
return files
|