mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-10-01 09:43:14 +00:00
* feat: tune reviewer for precision — web/wiki tools + recalibrated prompt Reviewer agent now has web_search, fetch_url, and http_request alongside the finding tools, so it can verify library semantics and consult the DeepWiki auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>) before flagging cross-file or architectural concerns. Prompt rewritten to push precision over recall: - explicit severity ladder pushing reviews toward bimodal high/low instead of defaulting to medium - ≤200-char description target (gold set averages ~186 chars; we were at ~436) - mandatory docs / wiki / code lookup before flagging concurrency, security, or perf — the three categories that dominated false positives - "do not flag" list covering compiler/linter-catchable nits, speculative claims without a concrete attacker/interleaving/scale, style preferences the codebase doesn't share, and test-quality nits on non-test diffs - smart file-selection guidance for large PRs (deprioritize generated / vendored / pure-rename hunks) Eval config switched to openai:gpt-5.5 + high reasoning effort for the next benchmark run. * trim prompt * subagent prompting * confidence ratings * added medium * enforce confidence threshold * . * reviewer: precision-tuned prompt + drop confidence gate Rewrites the reviewer system prompt around a defensibility bar (anchor + failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file list (style nits, speculation, scope-policing, same-bug fan-out), and a checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26% style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new prompt targets each class directly. Confidence is still recorded on every finding for post-hoc calibration but no longer gates publication — the audit showed the gate was a no-op (agent self-rated 65% of findings "high" regardless), and the prompt's defensibility bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD, the confidence_threshold kwarg on filter_findings_for_publish, the confidence_filtered score_mode, and the min_confidence kwarg on the eval target's _extract_comments — all dead once the gate is gone. Also removes the "informational" severity tier from the Severity enum, SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for FYI observations the dataset never rewards. * benchmax * adding google provider * slight steering * tuning * more tuning * fix * cleanup * reducing overfitting * Add per-repo review style profiles and inject them into the reviewer. Dashboard users can analyze historical PR review feedback per repository, edit the resulting style guide, and have it loaded from LangGraph Store at reviewer runtime (including Martian eval runs) keyed by owner/name. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix review style job errors leaking exception details to clients. Return generic dashboard messages while logging full stack traces server-side. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
160 lines
5.7 KiB
Python
160 lines
5.7 KiB
Python
"""Review style analyzer graph.
|
||
|
||
Uses the same sandbox + ``gh`` pattern as the reviewer agent. The dashboard
|
||
user's OAuth token is injected into the LangSmith GitHub proxy so ``gh`` works
|
||
on public repos even when the GitHub App is not installed on them.
|
||
"""
|
||
# ruff: noqa: E402
|
||
|
||
from __future__ import annotations
|
||
|
||
import asyncio
|
||
import logging
|
||
import os
|
||
import warnings
|
||
|
||
from langgraph.graph.state import RunnableConfig
|
||
from langgraph.pregel import Pregel
|
||
|
||
warnings.filterwarnings("ignore", module="langchain_core._api.deprecation")
|
||
warnings.filterwarnings("ignore", message=".*Pydantic V1.*", category=UserWarning)
|
||
|
||
from deepagents import create_deep_agent
|
||
from deepagents.backends.protocol import SandboxBackendProtocol
|
||
from langchain.agents.middleware import ModelCallLimitMiddleware
|
||
|
||
from .integrations.langsmith import _configure_github_proxy
|
||
from .middleware import SanitizeToolInputsMiddleware, ToolErrorMiddleware
|
||
from .review_style_guidance import REVIEWER_STYLE_THEMES
|
||
from .server import (
|
||
DEFAULT_LLM_MAX_TOKENS,
|
||
DEFAULT_LLM_MODEL_ID,
|
||
DEFAULT_RECURSION_LIMIT,
|
||
ensure_sandbox_for_thread,
|
||
graph_loaded_for_execution,
|
||
)
|
||
from .tools.save_review_style import save_review_style_prompt
|
||
from .utils.model import DEFAULT_LLM_REASONING, make_model, provider_model_kwargs
|
||
from .utils.sandbox_paths import aresolve_sandbox_work_dir
|
||
from .utils.sandbox_state import unwrap_sandbox_backend
|
||
|
||
logger = logging.getLogger(__name__)
|
||
|
||
STYLE_ANALYZER_MODEL_CALL_LIMIT = 80
|
||
|
||
STYLE_ANALYZER_PROMPT = """You are a code-review style analyst for `{repo_owner}/{repo_name}`.
|
||
|
||
Sandbox: `{working_dir}`. Use the shell (``execute``) to run GitHub commands.
|
||
|
||
**Always invoke gh as:** `GH_TOKEN=dummy gh <command>`
|
||
|
||
# How to research (required)
|
||
|
||
Browse historical **merged** PR review feedback until you have catalogued at least
|
||
**8 substantive human** review comments (not bots). Suggested commands:
|
||
|
||
```
|
||
GH_TOKEN=dummy gh pr list --repo {repo_owner}/{repo_name} --state merged --limit 30
|
||
GH_TOKEN=dummy gh api repos/{repo_owner}/{repo_name}/pulls/<PR_NUMBER>/reviews
|
||
GH_TOKEN=dummy gh api repos/{repo_owner}/{repo_name}/pulls/<PR_NUMBER>/comments
|
||
GH_TOKEN=dummy gh api repos/{repo_owner}/{repo_name}/issues/<PR_NUMBER>/comments
|
||
```
|
||
|
||
If the first batch is sparse, increase `--limit` or walk older PR numbers. Skip
|
||
`[bot]` accounts and obvious automation (codecov, dependabot, etc.).
|
||
|
||
Identify the top ~5 human reviewers by volume and note phrasing, severity, and
|
||
what they ignore.
|
||
|
||
# When you may call `save_review_style_prompt`
|
||
|
||
Only after real research. Your `custom_prompt` (400–1200 words) must teach our
|
||
reviewer agent this repo's norms:
|
||
|
||
- What the team routinely flags vs skips (paraphrased patterns, not invented quotes)
|
||
- Severity calibration
|
||
- Tone and test expectations
|
||
- Repo-specific conventions
|
||
- Anti-patterns reviewers here avoid
|
||
|
||
`analysis_summary`: 2–4 sentences for the dashboard.
|
||
Pass `prs_sampled`, `reviews_sampled`, and `top_reviewers` (comma-separated logins).
|
||
|
||
Do **not** save a generic guide after one or two commands. Only after ~25+ merged
|
||
PRs with zero human feedback may you save a short conservative guide and say so in
|
||
`analysis_summary`.
|
||
|
||
# Alignment with our reviewer agent
|
||
|
||
{reviewer_themes}
|
||
|
||
# Optional preloaded samples
|
||
|
||
The user message may include pre-collected samples — verify and extend with ``gh``.
|
||
"""
|
||
|
||
|
||
async def _configure_sandbox_github_proxy(
|
||
sandbox_backend: SandboxBackendProtocol,
|
||
github_token: str,
|
||
) -> None:
|
||
if os.getenv("SANDBOX_TYPE", "langsmith") != "langsmith":
|
||
return
|
||
backend = unwrap_sandbox_backend(sandbox_backend)
|
||
await asyncio.to_thread(_configure_github_proxy, backend.id, github_token)
|
||
|
||
|
||
async def get_review_style_analyzer(config: RunnableConfig) -> Pregel:
|
||
thread_id = config["configurable"].get("thread_id")
|
||
config["recursion_limit"] = DEFAULT_RECURSION_LIMIT
|
||
|
||
if thread_id is None or not graph_loaded_for_execution(config):
|
||
return create_deep_agent(system_prompt="", tools=[]).with_config(config)
|
||
|
||
sandbox_backend = await ensure_sandbox_for_thread(thread_id)
|
||
work_dir = await aresolve_sandbox_work_dir(sandbox_backend)
|
||
|
||
configurable = config["configurable"]
|
||
full_name = str(configurable.get("review_style_full_name") or "owner/repo")
|
||
owner, _, name = full_name.partition("/")
|
||
samples_text = str(configurable.get("review_style_samples_text") or "")
|
||
github_token = configurable.get("review_style_github_token")
|
||
if isinstance(github_token, str) and github_token:
|
||
await _configure_sandbox_github_proxy(sandbox_backend, github_token)
|
||
|
||
model_id = DEFAULT_LLM_MODEL_ID
|
||
model_kwargs = provider_model_kwargs(
|
||
model_id,
|
||
None,
|
||
max_tokens=DEFAULT_LLM_MAX_TOKENS,
|
||
openai_reasoning_default=DEFAULT_LLM_REASONING,
|
||
)
|
||
|
||
system_prompt = STYLE_ANALYZER_PROMPT.format(
|
||
repo_owner=owner or "<owner>",
|
||
repo_name=name or "<repo>",
|
||
working_dir=work_dir,
|
||
reviewer_themes=REVIEWER_STYLE_THEMES.strip(),
|
||
)
|
||
user_context = (
|
||
f"Repository: `{full_name}`\n\n"
|
||
f"{samples_text}\n\n"
|
||
"Research review style with `GH_TOKEN=dummy gh ...` via execute, then call "
|
||
"`save_review_style_prompt` once you have enough evidence."
|
||
)
|
||
system_prompt = f"{system_prompt}\n\n{user_context}"
|
||
|
||
return create_deep_agent(
|
||
model=make_model(model_id, **model_kwargs),
|
||
system_prompt=system_prompt,
|
||
tools=[save_review_style_prompt],
|
||
backend=sandbox_backend,
|
||
middleware=[
|
||
SanitizeToolInputsMiddleware(),
|
||
ModelCallLimitMiddleware(
|
||
run_limit=STYLE_ANALYZER_MODEL_CALL_LIMIT,
|
||
exit_behavior="end",
|
||
),
|
||
ToolErrorMiddleware(),
|
||
],
|
||
).with_config(config)
|