open-swe/agent/skills/continual-learning/SKILL.md
Johannes du Plessis 4a55145bb1
feat: outcomes dataset + bootstrap/continual split via skills (#1365)
* fix: reset stale sandbox creation sentinel

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: treat SANDBOX_CREATING as a timestamped cross-process lock

Only reset the sentinel when proven stale (older than the creation
timeout); otherwise wait for the worker that holds the lock so a
concurrent run does not create a duplicate sandbox.

* feat(analyzer): outcomes dataset + bootstrap/continual split via skills

Rename the review_style_analyzer graph to `analyzer` and split it into two
modes, plus capture reviewer finding outcomes for continual learning.

- Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false
  positive), and GitHub/Slack thumbs findings into a single LangSmith dataset
  (openswe-reviewer-outcomes), keyed deterministically per finding+source.
  Emit points wired into update_finding, resolve_finding_thread, and the
  GitHub/Slack reaction handlers.
- Two playbooks delivered as deepagents skills (bootstrap-repo-analysis,
  continual-learning), served as virtual files via a CompositeBackend /skills/
  route + StateBackend (seeded into the run files channel at invoke time, never
  written to the sandbox). Mode is set by the launcher; continual runs fall
  back to the GitHub App installation token.
- Split launcher into start_bootstrap_analysis + start_continual_run; register
  a per-repo nightly continual-learning cron when bootstrap completes.
- New read_finding_outcomes tool feeds confirmed/dismissed findings back to the
  continual playbook.

Tests for outcome label mapping, skills helper, and cron idempotency.

* fix(analyzer): anchor continual cron runs to a real thread_id

The nightly continual-learning cron is threadless, and get_analyzer
early-returns an empty agent when configurable.thread_id is missing — so
every cron-launched run no-op'd before reading outcomes or saving a refined
prompt. Include the repo's deterministic analyzer thread_id in the continual
run configurable so the run executes; the threadless run carries no message
history, so nightly runs don't accumulate context.

* refactor(analyzer): move cron lifecycle calls out of the review-styles store

Drop the inline `analyzer_cron` imports from review_styles.py (added only to
dodge a circular import) by relocating the cron-trigger calls to the layer
above the store: registration to the save_review_style tool (after a prompt is
saved) and removal to the dashboard delete route. review_styles.py is now a
pure store again with top-level imports only.

* refactor: hoist reviewer_outcomes imports to module level

Move the two inline emit_finding_status_outcome imports introduced in this PR
(update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes
only depends on langsmith, so there is no circular import to avoid.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00

44 lines
2.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
name: continual-learning
description: Nightly refinement of an existing per-repo review-style prompt using this reviewer's own finding outcomes. Read confirmed (resolved-by-commit / thumbs-up) and dismissed (thumbs-down) findings, promote the bug patterns the team actually fixes, demote the false-positive patterns, reconcile against the current prompt, and save the refined version. Use this once outcomes exist; use bootstrap-repo-analysis for a cold-start repo.
---
# Continual learning
You are **refining** the existing review-style prompt for the repository named in the
system prompt, using outcomes the reviewer has accrued since the last run. The goal is
to raise recall (catch more real bugs) without hurting precision (stop repeating
dismissed ones).
## 1. Read outcomes first
Call `read_finding_outcomes` once. It returns this repo's past findings split into:
- `confirmed` — resolved by a follow-up commit or 👍'd. These are **real** bug patterns
this team fixes. Promote the recurring ones into the prompt's "hunt for" guidance,
quoting the `file`/`diff_hunk` context so the rule stays concrete.
- `dismissed` — dismissed or 👎'd. These are **false-positive** patterns. Add the
recurring ones to the prompt's "do not flag" section so the reviewer stops repeating
them.
Look for repetition, not one-offs. A single dismissed finding is noise; the same class
dismissed several times is a rule.
## 2. Reconcile against the current prompt
The current `custom_prompt` is the starting point — you are editing it, not rewriting
from scratch. Read it (it is summarized for you / available via the dashboard record).
Keep what still holds, strengthen rules the outcomes confirm, and remove or soften rules
the outcomes contradict. Optionally do a **light** `gh` top-up
(`GH_TOKEN=dummy gh ...`) to confirm a pattern, but outcomes are the primary signal — do
not re-run a full PR crawl.
Stay aligned with the reviewer-agent themes in the system prompt.
## 3. Save
Call `save_review_style_prompt` once with the refined `custom_prompt` (400–1200 words),
an `analysis_summary` that names what changed this cycle (e.g. "promoted N-pattern after
3 confirmed fixes; dropped M-pattern after repeated dismissals"), and the
`top_reviewers` / counts you have. If outcomes were empty and nothing changed, say so in
`analysis_summary` and re-save the existing prompt unchanged rather than degrading it.