mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-10-03 22:33:19 +00:00
* fix: reset stale sandbox creation sentinel Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> * fix: treat SANDBOX_CREATING as a timestamped cross-process lock Only reset the sentinel when proven stale (older than the creation timeout); otherwise wait for the worker that holds the lock so a concurrent run does not create a duplicate sandbox. * feat(analyzer): outcomes dataset + bootstrap/continual split via skills Rename the review_style_analyzer graph to `analyzer` and split it into two modes, plus capture reviewer finding outcomes for continual learning. - Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false positive), and GitHub/Slack thumbs findings into a single LangSmith dataset (openswe-reviewer-outcomes), keyed deterministically per finding+source. Emit points wired into update_finding, resolve_finding_thread, and the GitHub/Slack reaction handlers. - Two playbooks delivered as deepagents skills (bootstrap-repo-analysis, continual-learning), served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run files channel at invoke time, never written to the sandbox). Mode is set by the launcher; continual runs fall back to the GitHub App installation token. - Split launcher into start_bootstrap_analysis + start_continual_run; register a per-repo nightly continual-learning cron when bootstrap completes. - New read_finding_outcomes tool feeds confirmed/dismissed findings back to the continual playbook. Tests for outcome label mapping, skills helper, and cron idempotency. * fix(analyzer): anchor continual cron runs to a real thread_id The nightly continual-learning cron is threadless, and get_analyzer early-returns an empty agent when configurable.thread_id is missing — so every cron-launched run no-op'd before reading outcomes or saving a refined prompt. Include the repo's deterministic analyzer thread_id in the continual run configurable so the run executes; the threadless run carries no message history, so nightly runs don't accumulate context. * refactor(analyzer): move cron lifecycle calls out of the review-styles store Drop the inline `analyzer_cron` imports from review_styles.py (added only to dodge a circular import) by relocating the cron-trigger calls to the layer above the store: registration to the save_review_style tool (after a prompt is saved) and removal to the dashboard delete route. review_styles.py is now a pure store again with top-level imports only. * refactor: hoist reviewer_outcomes imports to module level Move the two inline emit_finding_status_outcome imports introduced in this PR (update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes only depends on langsmith, so there is no circular import to avoid. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
61 lines
2.9 KiB
Markdown
61 lines
2.9 KiB
Markdown
---
|
||
name: bootstrap-repo-analysis
|
||
description: First-time analysis of a repository with no prior reviewer outcomes. Crawl historical merged-PR review feedback with the gh CLI (plus any preloaded samples), extract the team's review norms, and synthesize the initial per-repo review-style prompt. Use this for a cold-start repo; use continual-learning instead once the reviewer has accumulated finding outcomes.
|
||
---
|
||
|
||
# Bootstrap repo analysis
|
||
|
||
You are writing the **first** review-style prompt for the repository named in the
|
||
system prompt. There is no outcomes history yet, so your signal comes entirely from
|
||
the repo's own historical PR review feedback. Do not call `read_finding_outcomes` in
|
||
this mode — it will be empty.
|
||
|
||
Always invoke gh as: `GH_TOKEN=dummy gh <command>`.
|
||
|
||
## 1. Research (required)
|
||
|
||
Browse historical **merged** PR review feedback until you have catalogued at least
|
||
**8 substantive human** review comments (skip `[bot]` accounts and obvious automation
|
||
like codecov / dependabot). Useful commands:
|
||
|
||
```
|
||
GH_TOKEN=dummy gh pr list --repo <owner>/<repo> --state merged --limit 30
|
||
GH_TOKEN=dummy gh api repos/<owner>/<repo>/pulls/<PR_NUMBER>/reviews
|
||
GH_TOKEN=dummy gh api repos/<owner>/<repo>/pulls/<PR_NUMBER>/comments
|
||
GH_TOKEN=dummy gh api repos/<owner>/<repo>/issues/<PR_NUMBER>/comments
|
||
```
|
||
|
||
If the first batch is sparse, raise `--limit` or walk older PR numbers. The user
|
||
message may include **preloaded samples** — verify and extend them with `gh`, don't
|
||
just trust them.
|
||
|
||
Identify the top ~5 human reviewers by volume and note their phrasing, what severity
|
||
they assign, and what they routinely ignore.
|
||
|
||
## 2. Extract concrete, repo-specific patterns
|
||
|
||
The highest-value content is a **bug taxonomy tied to this repo's stack** — concrete
|
||
"hunt for X" rules a maintainer would catch on first read — plus a calibrated "do not
|
||
flag" list. Pair each pattern with the failure mode and, where you saw it, the kind of
|
||
diff that triggered it. Avoid generic advice that would apply to any repo.
|
||
|
||
Cover:
|
||
- What the team routinely flags vs. skips (paraphrased patterns, not invented quotes)
|
||
- Severity calibration tied to user-visible / runtime consequence
|
||
- Tone and test expectations
|
||
- Repo-specific conventions (frameworks, repository/data-access boundaries, naming)
|
||
- Anti-patterns the reviewers here deliberately avoid
|
||
|
||
Stay aligned with the reviewer-agent themes in the system prompt (high-signal,
|
||
diff-anchored defects — not nits).
|
||
|
||
## 3. Save
|
||
|
||
Only after real research, call `save_review_style_prompt` once with:
|
||
- `custom_prompt`: 400–1200 words teaching the reviewer this repo's norms.
|
||
- `analysis_summary`: 2–4 sentences for the dashboard.
|
||
- `top_reviewers` (comma-separated logins), `prs_sampled`, `reviews_sampled`.
|
||
|
||
Do **not** save a generic guide after one or two commands. Only after ~25+ merged PRs
|
||
with zero human feedback may you save a short, conservative guide — and say so in
|
||
`analysis_summary`.
|