open-swe/agent/skills/continual-learning/SKILL.md
Johannes du Plessis 4a55145bb1
feat: outcomes dataset + bootstrap/continual split via skills (#1365)
* fix: reset stale sandbox creation sentinel

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: treat SANDBOX_CREATING as a timestamped cross-process lock

Only reset the sentinel when proven stale (older than the creation
timeout); otherwise wait for the worker that holds the lock so a
concurrent run does not create a duplicate sandbox.

* feat(analyzer): outcomes dataset + bootstrap/continual split via skills

Rename the review_style_analyzer graph to `analyzer` and split it into two
modes, plus capture reviewer finding outcomes for continual learning.

- Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false
  positive), and GitHub/Slack thumbs findings into a single LangSmith dataset
  (openswe-reviewer-outcomes), keyed deterministically per finding+source.
  Emit points wired into update_finding, resolve_finding_thread, and the
  GitHub/Slack reaction handlers.
- Two playbooks delivered as deepagents skills (bootstrap-repo-analysis,
  continual-learning), served as virtual files via a CompositeBackend /skills/
  route + StateBackend (seeded into the run files channel at invoke time, never
  written to the sandbox). Mode is set by the launcher; continual runs fall
  back to the GitHub App installation token.
- Split launcher into start_bootstrap_analysis + start_continual_run; register
  a per-repo nightly continual-learning cron when bootstrap completes.
- New read_finding_outcomes tool feeds confirmed/dismissed findings back to the
  continual playbook.

Tests for outcome label mapping, skills helper, and cron idempotency.

* fix(analyzer): anchor continual cron runs to a real thread_id

The nightly continual-learning cron is threadless, and get_analyzer
early-returns an empty agent when configurable.thread_id is missing — so
every cron-launched run no-op'd before reading outcomes or saving a refined
prompt. Include the repo's deterministic analyzer thread_id in the continual
run configurable so the run executes; the threadless run carries no message
history, so nightly runs don't accumulate context.

* refactor(analyzer): move cron lifecycle calls out of the review-styles store

Drop the inline `analyzer_cron` imports from review_styles.py (added only to
dodge a circular import) by relocating the cron-trigger calls to the layer
above the store: registration to the save_review_style tool (after a prompt is
saved) and removal to the dashboard delete route. review_styles.py is now a
pure store again with top-level imports only.

* refactor: hoist reviewer_outcomes imports to module level

Move the two inline emit_finding_status_outcome imports introduced in this PR
(update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes
only depends on langsmith, so there is no circular import to avoid.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00

2.4 KiB
Raw Blame History

name description
continual-learning Nightly refinement of an existing per-repo review-style prompt using this reviewer's own finding outcomes. Read confirmed (resolved-by-commit / thumbs-up) and dismissed (thumbs-down) findings, promote the bug patterns the team actually fixes, demote the false-positive patterns, reconcile against the current prompt, and save the refined version. Use this once outcomes exist; use bootstrap-repo-analysis for a cold-start repo.

Continual learning

You are refining the existing review-style prompt for the repository named in the system prompt, using outcomes the reviewer has accrued since the last run. The goal is to raise recall (catch more real bugs) without hurting precision (stop repeating dismissed ones).

1. Read outcomes first

Call read_finding_outcomes once. It returns this repo's past findings split into:

  • confirmed — resolved by a follow-up commit or 👍'd. These are real bug patterns this team fixes. Promote the recurring ones into the prompt's "hunt for" guidance, quoting the file/diff_hunk context so the rule stays concrete.
  • dismissed — dismissed or 👎'd. These are false-positive patterns. Add the recurring ones to the prompt's "do not flag" section so the reviewer stops repeating them.

Look for repetition, not one-offs. A single dismissed finding is noise; the same class dismissed several times is a rule.

2. Reconcile against the current prompt

The current custom_prompt is the starting point — you are editing it, not rewriting from scratch. Read it (it is summarized for you / available via the dashboard record). Keep what still holds, strengthen rules the outcomes confirm, and remove or soften rules the outcomes contradict. Optionally do a light gh top-up (GH_TOKEN=dummy gh ...) to confirm a pattern, but outcomes are the primary signal — do not re-run a full PR crawl.

Stay aligned with the reviewer-agent themes in the system prompt.

3. Save

Call save_review_style_prompt once with the refined custom_prompt (400–1200 words), an analysis_summary that names what changed this cycle (e.g. "promoted N-pattern after 3 confirmed fixes; dropped M-pattern after repeated dismissals"), and the top_reviewers / counts you have. If outcomes were empty and nothing changed, say so in analysis_summary and re-save the existing prompt unchanged rather than degrading it.