mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-10-02 10:53:17 +00:00
* fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
44 lines
2.4 KiB
Markdown
44 lines
2.4 KiB
Markdown
---
|
||
name: continual-learning
|
||
description: Nightly refinement of an existing per-repo review-style prompt using this reviewer's own finding outcomes. Read confirmed (resolved-by-commit / thumbs-up) and dismissed (thumbs-down) findings, promote the bug patterns the team actually fixes, demote the false-positive patterns, reconcile against the current prompt, and save the refined version. Use this once outcomes exist; use bootstrap-repo-analysis for a cold-start repo.
|
||
---
|
||
|
||
# Continual learning
|
||
|
||
You are **refining** the existing review-style prompt for the repository named in the
|
||
system prompt, using outcomes the reviewer has accrued since the last run. The goal is
|
||
to raise recall (catch more real bugs) without hurting precision (stop repeating
|
||
dismissed ones).
|
||
|
||
## 1. Read outcomes first
|
||
|
||
Call `read_finding_outcomes` once. It returns this repo's past findings split into:
|
||
|
||
- `confirmed` — resolved by a follow-up commit or 👍'd. These are **real** bug patterns
|
||
this team fixes. Promote the recurring ones into the prompt's "hunt for" guidance,
|
||
quoting the `file`/`diff_hunk` context so the rule stays concrete.
|
||
- `dismissed` — dismissed or 👎'd. These are **false-positive** patterns. Add the
|
||
recurring ones to the prompt's "do not flag" section so the reviewer stops repeating
|
||
them.
|
||
|
||
Look for repetition, not one-offs. A single dismissed finding is noise; the same class
|
||
dismissed several times is a rule.
|
||
|
||
## 2. Reconcile against the current prompt
|
||
|
||
The current `custom_prompt` is the starting point — you are editing it, not rewriting
|
||
from scratch. Read it (it is summarized for you / available via the dashboard record).
|
||
Keep what still holds, strengthen rules the outcomes confirm, and remove or soften rules
|
||
the outcomes contradict. Optionally do a **light** `gh` top-up
|
||
(`GH_TOKEN=dummy gh ...`) to confirm a pattern, but outcomes are the primary signal — do
|
||
not re-run a full PR crawl.
|
||
|
||
Stay aligned with the reviewer-agent themes in the system prompt.
|
||
|
||
## 3. Save
|
||
|
||
Call `save_review_style_prompt` once with the refined `custom_prompt` (400–1200 words),
|
||
an `analysis_summary` that names what changed this cycle (e.g. "promoted N-pattern after
|
||
3 confirmed fixes; dropped M-pattern after repeated dismissals"), and the
|
||
`top_reviewers` / counts you have. If outcomes were empty and nothing changed, say so in
|
||
`analysis_summary` and re-save the existing prompt unchanged rather than degrading it.
|