open-swe/agent/dashboard/analyzer_cron.py
Johannes du Plessis 4a55145bb1
feat: outcomes dataset + bootstrap/continual split via skills (#1365)
* fix: reset stale sandbox creation sentinel

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: treat SANDBOX_CREATING as a timestamped cross-process lock

Only reset the sentinel when proven stale (older than the creation
timeout); otherwise wait for the worker that holds the lock so a
concurrent run does not create a duplicate sandbox.

* feat(analyzer): outcomes dataset + bootstrap/continual split via skills

Rename the review_style_analyzer graph to `analyzer` and split it into two
modes, plus capture reviewer finding outcomes for continual learning.

- Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false
  positive), and GitHub/Slack thumbs findings into a single LangSmith dataset
  (openswe-reviewer-outcomes), keyed deterministically per finding+source.
  Emit points wired into update_finding, resolve_finding_thread, and the
  GitHub/Slack reaction handlers.
- Two playbooks delivered as deepagents skills (bootstrap-repo-analysis,
  continual-learning), served as virtual files via a CompositeBackend /skills/
  route + StateBackend (seeded into the run files channel at invoke time, never
  written to the sandbox). Mode is set by the launcher; continual runs fall
  back to the GitHub App installation token.
- Split launcher into start_bootstrap_analysis + start_continual_run; register
  a per-repo nightly continual-learning cron when bootstrap completes.
- New read_finding_outcomes tool feeds confirmed/dismissed findings back to the
  continual playbook.

Tests for outcome label mapping, skills helper, and cron idempotency.

* fix(analyzer): anchor continual cron runs to a real thread_id

The nightly continual-learning cron is threadless, and get_analyzer
early-returns an empty agent when configurable.thread_id is missing — so
every cron-launched run no-op'd before reading outcomes or saving a refined
prompt. Include the repo's deterministic analyzer thread_id in the continual
run configurable so the run executes; the threadless run carries no message
history, so nightly runs don't accumulate context.

* refactor(analyzer): move cron lifecycle calls out of the review-styles store

Drop the inline `analyzer_cron` imports from review_styles.py (added only to
dodge a circular import) by relocating the cron-trigger calls to the layer
above the store: registration to the save_review_style tool (after a prompt is
saved) and removal to the dashboard delete route. review_styles.py is now a
pure store again with top-level imports only.

* refactor: hoist reviewer_outcomes imports to module level

Move the two inline emit_finding_status_outcome imports introduced in this PR
(update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes
only depends on langsmith, so there is no circular import to avoid.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00

70 lines
2.6 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

"""Per-repo nightly continual-learning crons for the analyzer.
When a repo's bootstrap analysis completes we register one daily LangGraph cron
that fires a continual-learning run for that repo. Runs are threadless (a fresh
thread + sandbox each night) and authenticate via the GitHub App installation
token resolved inside ``get_analyzer`` (the cron carries no fresh user token).
"""
from __future__ import annotations
import hashlib
import logging
from .review_style_jobs import (
_client,
build_continual_run_configurable,
build_continual_run_input,
)
from .review_styles import get_review_style, update_review_style
logger = logging.getLogger(__name__)
_ASSISTANT_ID = "analyzer"
def _daily_schedule(full_name: str) -> str:
"""Daily cron expression, staggered per repo to avoid a thundering herd."""
digest = int(hashlib.sha256(full_name.encode()).hexdigest(), 16)
minute = digest % 60
hour = 5 + (digest // 60) % 4 # 05:00–08:59 UTC
return f"{minute} {hour} * * *"
async def ensure_continual_cron(full_name: str) -> str | None:
"""Idempotently register the per-repo nightly continual-learning cron."""
record = await get_review_style(full_name)
existing = record.get("continual_cron_id") if record else None
if isinstance(existing, str) and existing:
return existing
try:
cron = await _client().crons.create(
_ASSISTANT_ID,
schedule=_daily_schedule(full_name),
input=build_continual_run_input(full_name),
config={"configurable": build_continual_run_configurable(full_name)},
metadata={"kind": "analyzer_continual", "repo": full_name},
)
except Exception:
logger.exception("Failed to create continual cron for %s", full_name)
return None
cron_id = cron.get("cron_id") if isinstance(cron, dict) else getattr(cron, "cron_id", None)
if isinstance(cron_id, str) and cron_id:
await update_review_style(full_name, {"continual_cron_id": cron_id})
return cron_id
return None
async def remove_continual_cron(full_name: str) -> None:
"""Delete the per-repo continual-learning cron, if one is registered."""
record = await get_review_style(full_name)
cron_id = record.get("continual_cron_id") if record else None
if not (isinstance(cron_id, str) and cron_id):
return
try:
await _client().crons.delete(cron_id)
except Exception:
logger.debug("Could not delete continual cron %s for %s", cron_id, full_name, exc_info=True)
await update_review_style(full_name, {"continual_cron_id": None})