open-swe/agent/tools/save_review_style.py
Johannes du Plessis 4a55145bb1
feat: outcomes dataset + bootstrap/continual split via skills (#1365)
* fix: reset stale sandbox creation sentinel

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: treat SANDBOX_CREATING as a timestamped cross-process lock

Only reset the sentinel when proven stale (older than the creation
timeout); otherwise wait for the worker that holds the lock so a
concurrent run does not create a duplicate sandbox.

* feat(analyzer): outcomes dataset + bootstrap/continual split via skills

Rename the review_style_analyzer graph to `analyzer` and split it into two
modes, plus capture reviewer finding outcomes for continual learning.

- Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false
  positive), and GitHub/Slack thumbs findings into a single LangSmith dataset
  (openswe-reviewer-outcomes), keyed deterministically per finding+source.
  Emit points wired into update_finding, resolve_finding_thread, and the
  GitHub/Slack reaction handlers.
- Two playbooks delivered as deepagents skills (bootstrap-repo-analysis,
  continual-learning), served as virtual files via a CompositeBackend /skills/
  route + StateBackend (seeded into the run files channel at invoke time, never
  written to the sandbox). Mode is set by the launcher; continual runs fall
  back to the GitHub App installation token.
- Split launcher into start_bootstrap_analysis + start_continual_run; register
  a per-repo nightly continual-learning cron when bootstrap completes.
- New read_finding_outcomes tool feeds confirmed/dismissed findings back to the
  continual playbook.

Tests for outcome label mapping, skills helper, and cron idempotency.

* fix(analyzer): anchor continual cron runs to a real thread_id

The nightly continual-learning cron is threadless, and get_analyzer
early-returns an empty agent when configurable.thread_id is missing — so
every cron-launched run no-op'd before reading outcomes or saving a refined
prompt. Include the repo's deterministic analyzer thread_id in the continual
run configurable so the run executes; the threadless run carries no message
history, so nightly runs don't accumulate context.

* refactor(analyzer): move cron lifecycle calls out of the review-styles store

Drop the inline `analyzer_cron` imports from review_styles.py (added only to
dodge a circular import) by relocating the cron-trigger calls to the layer
above the store: registration to the save_review_style tool (after a prompt is
saved) and removal to the dashboard delete route. review_styles.py is now a
pure store again with top-level imports only.

* refactor: hoist reviewer_outcomes imports to module level

Move the two inline emit_finding_status_outcome imports introduced in this PR
(update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes
only depends on langsmith, so there is no circular import to avoid.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00

71 lines
2.7 KiB
Python

"""Tool: persist synthesized per-repo review style prompt."""
from __future__ import annotations
import asyncio
import logging
from typing import Any
from langgraph.config import get_config
from ..dashboard.analyzer_cron import ensure_continual_cron
from ..dashboard.review_styles import mark_analysis_completed, mark_analysis_failed
logger = logging.getLogger(__name__)
async def _complete_and_register(full_name: str, **completed_kwargs: Any) -> dict[str, Any]:
"""Persist the prompt, then ensure the repo's nightly continual cron exists.
Cron registration is idempotent, so continual runs completing later don't
re-register it; it just guarantees a cron once a prompt first exists.
"""
record = await mark_analysis_completed(full_name, **completed_kwargs)
try:
await ensure_continual_cron(full_name)
except Exception:
logger.exception("Failed to ensure continual cron for %s", full_name)
return record
def save_review_style_prompt(
custom_prompt: str,
analysis_summary: str = "",
top_reviewers: str = "",
prs_sampled: int = 0,
reviews_sampled: int = 0,
) -> dict[str, Any]:
"""Save the synthesized repository-specific review style prompt.
Call this once at the end of style analysis with the final prompt text
that should be injected into the reviewer agent for this repository.
"""
config = get_config()
configurable = config.get("configurable") or {}
full_name = configurable.get("review_style_full_name")
if not isinstance(full_name, str) or "/" not in full_name:
return {"ok": False, "error": "review_style_full_name missing from config"}
reviewers_from_args = [r.strip() for r in top_reviewers.split(",") if r.strip()]
reviewers_from_config = configurable.get("review_style_top_reviewers") or []
merged_reviewers = reviewers_from_args or (
list(reviewers_from_config) if isinstance(reviewers_from_config, list) else []
)
prs_count = prs_sampled or int(configurable.get("review_style_prs_sampled") or 0)
reviews_count = reviews_sampled or int(configurable.get("review_style_reviews_sampled") or 0)
if not custom_prompt.strip():
asyncio.run(mark_analysis_failed(full_name, "custom_prompt was empty"))
return {"ok": False, "error": "custom_prompt cannot be empty"}
record = asyncio.run(
_complete_and_register(
full_name,
custom_prompt=custom_prompt.strip(),
analysis_summary=analysis_summary.strip(),
top_reviewers=merged_reviewers,
prs_sampled=prs_count,
reviews_sampled=reviews_count,
)
)
return {"ok": True, "full_name": full_name, "status": record.get("status")}