feat: open-swe dashboard for per-user profile config (#1302)
* feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints
Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering:
- GitHub App OAuth login → JWT cookie session (cross-domain ready)
- profile CRUD against LangGraph Store with model+effort validation
- admin gate via CONFIGURED_ADMINS
- /repos via /user/installations using the user's encrypted OAuth token
CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the
Vercel-hosted frontend can call the LangSmith deployment with credentials.
* feat: apply dashboard profile model/effort overrides in get_agent
Look up the triggering user's GitHub login from config (direct field or
GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store,
and apply default_model + reasoning_effort to make_model when both are
valid. Effort 'max' is captured on the profile but not yet wired through —
the OpenAI Reasoning Literal doesn't accept it.
* feat: ui/ TanStack Start dashboard for profile config
Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template,
base-ui primitives, Tailwind v4). Three routes:
- /login — Sign in with GitHub (links to /dashboard/api/auth/login)
- /profile — Edit default model, reasoning effort, default repo
- /admin — Admin-only: list users and edit other profiles
API client (src/lib/api.ts) uses credentials: include so the osw_session
cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL
points at the LangSmith deployment.
Effort options re-render when the model changes; 'max' on Opus 4.7 is
captured on the profile but ignored downstream until anthropic reasoning
is wired through make_model.
* feat: searchable Combobox for default repo picker
Replaces the Select with a base-ui Combobox so users can filter by typing,
the popup is wider than the trigger so full owner/repo names are readable,
and the list caps at max-h-80 to stay on screen.
* fix: address review comments + wire default_repo and Anthropic thinking
Security/correctness fixes from PR review:
* Open redirect: validate `redirect_to` in `/auth/login` against
`DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it
into the state JWT. Anything off-allowlist falls back to the dashboard
base URL. (PR #1302 r3250054386)
* Login CSRF: bind the OAuth `state` to the requesting browser. At
`/auth/login` we generate a fresh nonce, set it as a short-lived
HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and
embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback`
we require the cookie nonce to hash-match the state JWT's nonce_hash
(constant-time compare). (PR #1302 r3250054395)
* RMW race in profile vs token writes: split storage into two
namespaces — `["profiles"]` for user-editable settings and
`["oauth_tokens"]` for the encrypted GitHub token. Each upsert now
only writes its own namespace so an in-flight profile save can no
longer clobber a fresh token from a concurrent re-login (and vice
versa). (PR #1302 r3250054393)
* /repos pagination: follow `Link: rel="next"` for both
`/user/installations` and per-installation `/repositories` with
per_page=100, capped at 1000 items. (PR #1302 r3250054401)
Feature wires:
* default_repo: applied as a fallback in `get_slack_repo_config` (after
explicit-repo / thread metadata, before the env defaults) and in the
Linear webhook (after comment-body extraction, before team mapping).
Both paths resolve the triggering user's GitHub login via
GITHUB_USER_EMAIL_MAP and read the profile's default_repo.
* Anthropic "thinking" effort: `make_model` now accepts a `thinking`
kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max}
to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is
anthropic. OpenAI path still ignores "max" since the Literal doesn't
accept it.
2026-05-15 11:23:53 -07:00
|
|
|
"""FastAPI router for the dashboard backend."""
|
|
|
|
|
|
|
|
|
|
from __future__ import annotations
|
|
|
|
|
|
|
|
|
|
import hmac
|
|
|
|
|
import logging
|
|
|
|
|
import os
|
|
|
|
|
from typing import Any
|
|
|
|
|
|
|
|
|
|
import httpx
|
|
|
|
|
from fastapi import APIRouter, Depends, HTTPException, Request
|
|
|
|
|
from fastapi.responses import RedirectResponse, Response
|
|
|
|
|
|
|
|
|
|
from .admin import is_admin
|
|
|
|
|
from .oauth import (
|
|
|
|
|
COOKIE_NAME,
|
|
|
|
|
SESSION_TTL_SECONDS,
|
|
|
|
|
STATE_COOKIE_NAME,
|
|
|
|
|
STATE_TTL_SECONDS,
|
|
|
|
|
decode_state,
|
|
|
|
|
exchange_code,
|
|
|
|
|
fetch_github_user,
|
|
|
|
|
hash_state_nonce,
|
|
|
|
|
issue_session,
|
|
|
|
|
issue_state,
|
|
|
|
|
new_state_nonce,
|
|
|
|
|
require_session,
|
|
|
|
|
sanitize_redirect_to,
|
|
|
|
|
)
|
|
|
|
|
from .options import SUPPORTED_MODELS
|
|
|
|
|
from .profiles import (
|
|
|
|
|
ProfileUpdate,
|
|
|
|
|
get_access_token,
|
|
|
|
|
get_profile,
|
|
|
|
|
list_profiles,
|
|
|
|
|
upsert_access_token,
|
|
|
|
|
upsert_profile,
|
|
|
|
|
)
|
feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312)
* feat: tune reviewer for precision — web/wiki tools + recalibrated prompt
Reviewer agent now has web_search, fetch_url, and http_request alongside the
finding tools, so it can verify library semantics and consult the DeepWiki
auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>)
before flagging cross-file or architectural concerns.
Prompt rewritten to push precision over recall:
- explicit severity ladder pushing reviews toward bimodal high/low instead of
defaulting to medium
- ≤200-char description target (gold set averages ~186 chars; we were at ~436)
- mandatory docs / wiki / code lookup before flagging concurrency, security,
or perf — the three categories that dominated false positives
- "do not flag" list covering compiler/linter-catchable nits, speculative
claims without a concrete attacker/interleaving/scale, style preferences
the codebase doesn't share, and test-quality nits on non-test diffs
- smart file-selection guidance for large PRs (deprioritize generated /
vendored / pure-rename hunks)
Eval config switched to openai:gpt-5.5 + high reasoning effort for the next
benchmark run.
* trim prompt
* subagent prompting
* confidence ratings
* added medium
* enforce confidence threshold
* .
* reviewer: precision-tuned prompt + drop confidence gate
Rewrites the reviewer system prompt around a defensibility bar (anchor +
failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file
list (style nits, speculation, scope-policing, same-bug fan-out), and a
checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden
set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26%
style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new
prompt targets each class directly.
Confidence is still recorded on every finding for post-hoc calibration but
no longer gates publication — the audit showed the gate was a no-op (agent
self-rated 65% of findings "high" regardless), and the prompt's defensibility
bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD,
the confidence_threshold kwarg on filter_findings_for_publish, the
confidence_filtered score_mode, and the min_confidence kwarg on the eval
target's _extract_comments — all dead once the gate is gone.
Also removes the "informational" severity tier from the Severity enum,
SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for
FYI observations the dataset never rewards.
* benchmax
* adding google provider
* slight steering
* tuning
* more tuning
* fix
* cleanup
* reducing overfitting
* Add per-repo review style profiles and inject them into the reviewer.
Dashboard users can analyze historical PR review feedback per repository,
edit the resulting style guide, and have it loaded from LangGraph Store at
reviewer runtime (including Martian eval runs) keyed by owner/name.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix review style job errors leaking exception details to clients.
Return generic dashboard messages while logging full stack traces server-side.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-20 11:35:00 -07:00
|
|
|
from .review_style_jobs import start_review_style_analysis, sync_review_style_run_status
|
|
|
|
|
from .review_styles import (
|
|
|
|
|
ReviewStyleCreate,
|
|
|
|
|
ReviewStylePromptUpdate,
|
|
|
|
|
create_review_style,
|
|
|
|
|
get_review_style,
|
|
|
|
|
list_review_styles,
|
|
|
|
|
normalize_repo_full_name,
|
|
|
|
|
set_custom_prompt,
|
|
|
|
|
)
|
feat: open-swe dashboard for per-user profile config (#1302)
* feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints
Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering:
- GitHub App OAuth login → JWT cookie session (cross-domain ready)
- profile CRUD against LangGraph Store with model+effort validation
- admin gate via CONFIGURED_ADMINS
- /repos via /user/installations using the user's encrypted OAuth token
CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the
Vercel-hosted frontend can call the LangSmith deployment with credentials.
* feat: apply dashboard profile model/effort overrides in get_agent
Look up the triggering user's GitHub login from config (direct field or
GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store,
and apply default_model + reasoning_effort to make_model when both are
valid. Effort 'max' is captured on the profile but not yet wired through —
the OpenAI Reasoning Literal doesn't accept it.
* feat: ui/ TanStack Start dashboard for profile config
Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template,
base-ui primitives, Tailwind v4). Three routes:
- /login — Sign in with GitHub (links to /dashboard/api/auth/login)
- /profile — Edit default model, reasoning effort, default repo
- /admin — Admin-only: list users and edit other profiles
API client (src/lib/api.ts) uses credentials: include so the osw_session
cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL
points at the LangSmith deployment.
Effort options re-render when the model changes; 'max' on Opus 4.7 is
captured on the profile but ignored downstream until anthropic reasoning
is wired through make_model.
* feat: searchable Combobox for default repo picker
Replaces the Select with a base-ui Combobox so users can filter by typing,
the popup is wider than the trigger so full owner/repo names are readable,
and the list caps at max-h-80 to stay on screen.
* fix: address review comments + wire default_repo and Anthropic thinking
Security/correctness fixes from PR review:
* Open redirect: validate `redirect_to` in `/auth/login` against
`DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it
into the state JWT. Anything off-allowlist falls back to the dashboard
base URL. (PR #1302 r3250054386)
* Login CSRF: bind the OAuth `state` to the requesting browser. At
`/auth/login` we generate a fresh nonce, set it as a short-lived
HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and
embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback`
we require the cookie nonce to hash-match the state JWT's nonce_hash
(constant-time compare). (PR #1302 r3250054395)
* RMW race in profile vs token writes: split storage into two
namespaces — `["profiles"]` for user-editable settings and
`["oauth_tokens"]` for the encrypted GitHub token. Each upsert now
only writes its own namespace so an in-flight profile save can no
longer clobber a fresh token from a concurrent re-login (and vice
versa). (PR #1302 r3250054393)
* /repos pagination: follow `Link: rel="next"` for both
`/user/installations` and per-installation `/repositories` with
per_page=100, capped at 1000 items. (PR #1302 r3250054401)
Feature wires:
* default_repo: applied as a fallback in `get_slack_repo_config` (after
explicit-repo / thread metadata, before the env defaults) and in the
Linear webhook (after comment-body extraction, before team mapping).
Both paths resolve the triggering user's GitHub login via
GITHUB_USER_EMAIL_MAP and read the profile's default_repo.
* Anthropic "thinking" effort: `make_model` now accepts a `thinking`
kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max}
to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is
anthropic. OpenAI path still ignores "max" since the Literal doesn't
accept it.
2026-05-15 11:23:53 -07:00
|
|
|
|
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
|
|
|
|
|
|
router = APIRouter(prefix="/dashboard/api", tags=["dashboard"])
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _require_admin(session: dict[str, Any]) -> dict[str, Any]:
|
|
|
|
|
if not is_admin(session.get("email")):
|
|
|
|
|
raise HTTPException(403, "admin only")
|
|
|
|
|
return session
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
_SESSION_DEP = Depends(require_session)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _admin_session(session: dict[str, Any] = _SESSION_DEP) -> dict[str, Any]:
|
|
|
|
|
return _require_admin(session)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
_ADMIN_DEP = Depends(_admin_session)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _api_base_url() -> str:
|
|
|
|
|
v = os.environ.get("DASHBOARD_API_BASE_URL", "").rstrip("/")
|
|
|
|
|
if not v:
|
|
|
|
|
raise HTTPException(500, "DASHBOARD_API_BASE_URL not configured")
|
|
|
|
|
return v
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _frontend_base_url() -> str:
|
|
|
|
|
v = os.environ.get("DASHBOARD_BASE_URL", "").rstrip("/")
|
|
|
|
|
if not v:
|
|
|
|
|
raise HTTPException(500, "DASHBOARD_BASE_URL not configured")
|
|
|
|
|
return v
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _set_session_cookie(response: Response, jwt_token: str) -> None:
|
|
|
|
|
response.set_cookie(
|
|
|
|
|
key=COOKIE_NAME,
|
|
|
|
|
value=jwt_token,
|
|
|
|
|
max_age=SESSION_TTL_SECONDS,
|
|
|
|
|
httponly=True,
|
|
|
|
|
secure=True,
|
|
|
|
|
samesite="none",
|
|
|
|
|
path="/",
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _set_state_cookie(response: Response, nonce: str) -> None:
|
|
|
|
|
# SameSite=Lax so GitHub's top-level redirect back to /auth/callback
|
|
|
|
|
# still presents this cookie; the cookie is single-purpose and lives
|
|
|
|
|
# only for the duration of one OAuth round-trip.
|
|
|
|
|
response.set_cookie(
|
|
|
|
|
key=STATE_COOKIE_NAME,
|
|
|
|
|
value=nonce,
|
|
|
|
|
max_age=STATE_TTL_SECONDS,
|
|
|
|
|
httponly=True,
|
|
|
|
|
secure=True,
|
|
|
|
|
samesite="lax",
|
|
|
|
|
path="/dashboard/api/auth",
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _clear_state_cookie(response: Response) -> None:
|
|
|
|
|
response.delete_cookie(
|
|
|
|
|
STATE_COOKIE_NAME, path="/dashboard/api/auth", samesite="lax", secure=True
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.get("/auth/login")
|
|
|
|
|
async def auth_login(request: Request, redirect_to: str | None = None) -> RedirectResponse:
|
|
|
|
|
client_id = os.environ.get("GITHUB_APP_CLIENT_ID", "")
|
|
|
|
|
if not client_id:
|
|
|
|
|
raise HTTPException(500, "GITHUB_APP_CLIENT_ID not configured")
|
|
|
|
|
safe_redirect = sanitize_redirect_to(redirect_to) or _frontend_base_url()
|
|
|
|
|
|
|
|
|
|
nonce = new_state_nonce()
|
|
|
|
|
state = issue_state(redirect_to=safe_redirect, nonce_hash=hash_state_nonce(nonce))
|
|
|
|
|
redirect_uri = f"{_api_base_url()}/dashboard/api/auth/callback"
|
|
|
|
|
url = (
|
|
|
|
|
"https://github.com/login/oauth/authorize"
|
|
|
|
|
f"?client_id={client_id}"
|
|
|
|
|
f"&redirect_uri={redirect_uri}"
|
|
|
|
|
f"&state={state}"
|
|
|
|
|
)
|
|
|
|
|
response = RedirectResponse(url, status_code=302)
|
|
|
|
|
_set_state_cookie(response, nonce)
|
|
|
|
|
return response
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.get("/auth/callback")
|
|
|
|
|
async def auth_callback(request: Request, code: str, state: str) -> RedirectResponse:
|
|
|
|
|
state_payload = decode_state(state)
|
|
|
|
|
state_nonce_hash = state_payload.get("nonce_hash")
|
|
|
|
|
cookie_nonce = request.cookies.get(STATE_COOKIE_NAME)
|
|
|
|
|
if (
|
|
|
|
|
not isinstance(state_nonce_hash, str)
|
|
|
|
|
or not cookie_nonce
|
|
|
|
|
or not hmac.compare_digest(hash_state_nonce(cookie_nonce), state_nonce_hash)
|
|
|
|
|
):
|
|
|
|
|
# Either the cookie went missing (different browser, expired,
|
|
|
|
|
# cookies blocked) or the state was issued for a different session.
|
|
|
|
|
raise HTTPException(400, "oauth state mismatch — please retry login")
|
|
|
|
|
|
|
|
|
|
redirect_to = sanitize_redirect_to(state_payload.get("redirect_to")) or _frontend_base_url()
|
|
|
|
|
|
|
|
|
|
access_token = await exchange_code(code)
|
|
|
|
|
user, email = await fetch_github_user(access_token)
|
|
|
|
|
login = user.get("login")
|
|
|
|
|
if not login:
|
|
|
|
|
raise HTTPException(400, "could not resolve GitHub login")
|
|
|
|
|
|
|
|
|
|
await upsert_access_token(login, email or "", access_token)
|
|
|
|
|
|
|
|
|
|
session_jwt = issue_session(login=login, email=email, avatar_url=user.get("avatar_url"))
|
|
|
|
|
response = RedirectResponse(redirect_to, status_code=302)
|
|
|
|
|
_set_session_cookie(response, session_jwt)
|
|
|
|
|
_clear_state_cookie(response)
|
|
|
|
|
return response
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.post("/auth/logout")
|
|
|
|
|
async def auth_logout() -> Response:
|
|
|
|
|
response = Response(status_code=204)
|
|
|
|
|
response.delete_cookie(COOKIE_NAME, path="/", samesite="none", secure=True)
|
|
|
|
|
return response
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.get("/me")
|
|
|
|
|
async def me(session: dict[str, Any] = _SESSION_DEP) -> dict[str, Any]:
|
|
|
|
|
return {
|
|
|
|
|
"login": session["sub"],
|
|
|
|
|
"email": session.get("email"),
|
|
|
|
|
"avatar_url": session.get("avatar_url"),
|
|
|
|
|
"is_admin": is_admin(session.get("email")),
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.get("/options")
|
|
|
|
|
async def options() -> dict[str, Any]:
|
|
|
|
|
return {"models": SUPPORTED_MODELS}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.get("/profile")
|
|
|
|
|
async def get_my_profile(
|
|
|
|
|
session: dict[str, Any] = _SESSION_DEP,
|
|
|
|
|
) -> dict[str, Any]:
|
|
|
|
|
profile = await get_profile(session["sub"])
|
|
|
|
|
return profile or {}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.put("/profile")
|
|
|
|
|
async def put_my_profile(
|
|
|
|
|
update: ProfileUpdate,
|
|
|
|
|
session: dict[str, Any] = _SESSION_DEP,
|
|
|
|
|
) -> dict[str, Any]:
|
|
|
|
|
update.validate_pairing()
|
|
|
|
|
return await upsert_profile(session["sub"], session.get("email") or "", update)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.get("/admin/profiles")
|
|
|
|
|
async def admin_list_profiles(
|
|
|
|
|
_admin: dict[str, Any] = _ADMIN_DEP,
|
|
|
|
|
) -> list[dict[str, Any]]:
|
|
|
|
|
return await list_profiles()
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
class AdminProfileUpdate(ProfileUpdate):
|
|
|
|
|
email: str | None = None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.put("/admin/profiles/{login}")
|
|
|
|
|
async def admin_put_profile(
|
|
|
|
|
login: str,
|
|
|
|
|
update: AdminProfileUpdate,
|
|
|
|
|
_admin: dict[str, Any] = _ADMIN_DEP,
|
|
|
|
|
) -> dict[str, Any]:
|
|
|
|
|
update.validate_pairing()
|
|
|
|
|
existing = await get_profile(login) or {}
|
|
|
|
|
email = update.email or existing.get("email") or ""
|
|
|
|
|
base = ProfileUpdate(
|
|
|
|
|
default_model=update.default_model,
|
|
|
|
|
reasoning_effort=update.reasoning_effort,
|
|
|
|
|
default_repo=update.default_repo,
|
|
|
|
|
)
|
|
|
|
|
return await upsert_profile(login, email, base)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _next_link_url(link_header: str | None) -> str | None:
|
|
|
|
|
if not link_header:
|
|
|
|
|
return None
|
|
|
|
|
# GitHub Link header is comma-separated: '<url>; rel="next", <url>; rel="last"'
|
|
|
|
|
for part in link_header.split(","):
|
|
|
|
|
segments = [s.strip() for s in part.split(";")]
|
|
|
|
|
if len(segments) >= 2 and 'rel="next"' in segments[1] and segments[0].startswith("<"):
|
|
|
|
|
return segments[0][1:-1]
|
|
|
|
|
return None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
async def _paginate(
|
|
|
|
|
client: httpx.AsyncClient,
|
|
|
|
|
url: str,
|
|
|
|
|
*,
|
|
|
|
|
headers: dict[str, str],
|
|
|
|
|
items_key: str | None,
|
|
|
|
|
cap: int = 1000,
|
|
|
|
|
) -> list[dict[str, Any]]:
|
|
|
|
|
"""Follow ``Link: rel="next"`` until exhausted (or cap reached).
|
|
|
|
|
|
|
|
|
|
``items_key`` is the JSON key holding the list when the endpoint returns
|
|
|
|
|
a wrapper object (e.g. ``/user/installations`` returns
|
|
|
|
|
``{"total_count": N, "installations": [...]}``). When ``None`` the
|
|
|
|
|
response body itself is treated as the list.
|
|
|
|
|
"""
|
|
|
|
|
out: list[dict[str, Any]] = []
|
|
|
|
|
next_url: str | None = url
|
|
|
|
|
first = True
|
|
|
|
|
while next_url and len(out) < cap:
|
|
|
|
|
params = {"per_page": "100"} if first else None
|
|
|
|
|
r = await client.get(next_url, headers=headers, params=params)
|
|
|
|
|
if r.status_code == 401:
|
|
|
|
|
raise HTTPException(401, "github token expired, re-login required")
|
|
|
|
|
r.raise_for_status()
|
|
|
|
|
body = r.json()
|
|
|
|
|
page = body.get(items_key, []) if items_key else body
|
|
|
|
|
if isinstance(page, list):
|
|
|
|
|
out.extend(page)
|
|
|
|
|
next_url = _next_link_url(r.headers.get("Link"))
|
|
|
|
|
first = False
|
|
|
|
|
return out
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.get("/repos")
|
|
|
|
|
async def list_repos(
|
|
|
|
|
session: dict[str, Any] = _SESSION_DEP,
|
|
|
|
|
) -> dict[str, Any]:
|
|
|
|
|
"""List repos where open-swe is installed and the user has access.
|
|
|
|
|
|
|
|
|
|
Paginates both ``/user/installations`` and per-installation
|
|
|
|
|
``/user/installations/{id}/repositories`` so users with multiple
|
|
|
|
|
installations or >30 accessible repos get the complete set.
|
|
|
|
|
"""
|
|
|
|
|
token = await get_access_token(session["sub"])
|
|
|
|
|
if not token:
|
|
|
|
|
raise HTTPException(401, "github token unavailable, re-login required")
|
|
|
|
|
headers = {
|
|
|
|
|
"Authorization": f"Bearer {token}",
|
|
|
|
|
"Accept": "application/vnd.github+json",
|
|
|
|
|
"X-GitHub-Api-Version": "2022-11-28",
|
|
|
|
|
}
|
|
|
|
|
async with httpx.AsyncClient() as client:
|
|
|
|
|
installations = await _paginate(
|
|
|
|
|
client,
|
|
|
|
|
"https://api.github.com/user/installations",
|
|
|
|
|
headers=headers,
|
|
|
|
|
items_key="installations",
|
|
|
|
|
)
|
|
|
|
|
repositories: list[dict[str, Any]] = []
|
|
|
|
|
for inst in installations:
|
|
|
|
|
inst_id = inst.get("id")
|
|
|
|
|
if inst_id is None:
|
|
|
|
|
continue
|
|
|
|
|
try:
|
|
|
|
|
repos = await _paginate(
|
|
|
|
|
client,
|
|
|
|
|
f"https://api.github.com/user/installations/{inst_id}/repositories",
|
|
|
|
|
headers=headers,
|
|
|
|
|
items_key="repositories",
|
|
|
|
|
)
|
|
|
|
|
except HTTPException:
|
|
|
|
|
raise
|
|
|
|
|
except httpx.HTTPStatusError:
|
|
|
|
|
continue
|
|
|
|
|
repositories.extend(repos)
|
|
|
|
|
return {
|
|
|
|
|
"installations": [
|
|
|
|
|
{
|
|
|
|
|
"id": i.get("id"),
|
|
|
|
|
"account": (i.get("account") or {}).get("login"),
|
|
|
|
|
"account_type": (i.get("account") or {}).get("type"),
|
|
|
|
|
}
|
|
|
|
|
for i in installations
|
|
|
|
|
],
|
|
|
|
|
"repositories": [
|
|
|
|
|
{"full_name": r.get("full_name"), "private": r.get("private", False)}
|
|
|
|
|
for r in repositories
|
|
|
|
|
if r.get("full_name")
|
|
|
|
|
],
|
|
|
|
|
}
|
feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312)
* feat: tune reviewer for precision — web/wiki tools + recalibrated prompt
Reviewer agent now has web_search, fetch_url, and http_request alongside the
finding tools, so it can verify library semantics and consult the DeepWiki
auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>)
before flagging cross-file or architectural concerns.
Prompt rewritten to push precision over recall:
- explicit severity ladder pushing reviews toward bimodal high/low instead of
defaulting to medium
- ≤200-char description target (gold set averages ~186 chars; we were at ~436)
- mandatory docs / wiki / code lookup before flagging concurrency, security,
or perf — the three categories that dominated false positives
- "do not flag" list covering compiler/linter-catchable nits, speculative
claims without a concrete attacker/interleaving/scale, style preferences
the codebase doesn't share, and test-quality nits on non-test diffs
- smart file-selection guidance for large PRs (deprioritize generated /
vendored / pure-rename hunks)
Eval config switched to openai:gpt-5.5 + high reasoning effort for the next
benchmark run.
* trim prompt
* subagent prompting
* confidence ratings
* added medium
* enforce confidence threshold
* .
* reviewer: precision-tuned prompt + drop confidence gate
Rewrites the reviewer system prompt around a defensibility bar (anchor +
failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file
list (style nits, speculation, scope-policing, same-bug fan-out), and a
checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden
set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26%
style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new
prompt targets each class directly.
Confidence is still recorded on every finding for post-hoc calibration but
no longer gates publication — the audit showed the gate was a no-op (agent
self-rated 65% of findings "high" regardless), and the prompt's defensibility
bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD,
the confidence_threshold kwarg on filter_findings_for_publish, the
confidence_filtered score_mode, and the min_confidence kwarg on the eval
target's _extract_comments — all dead once the gate is gone.
Also removes the "informational" severity tier from the Severity enum,
SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for
FYI observations the dataset never rewards.
* benchmax
* adding google provider
* slight steering
* tuning
* more tuning
* fix
* cleanup
* reducing overfitting
* Add per-repo review style profiles and inject them into the reviewer.
Dashboard users can analyze historical PR review feedback per repository,
edit the resulting style guide, and have it loaded from LangGraph Store at
reviewer runtime (including Martian eval runs) keyed by owner/name.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix review style job errors leaking exception details to clients.
Return generic dashboard messages while logging full stack traces server-side.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-20 11:35:00 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
async def _assert_repo_available_for_style_analysis(full_name: str, token: str) -> None:
|
|
|
|
|
"""Ensure the repo exists and is readable for style learning.
|
|
|
|
|
|
|
|
|
|
Public repositories are allowed without the GitHub App installed on them.
|
|
|
|
|
Private repositories require the authenticated user to have read access.
|
|
|
|
|
"""
|
|
|
|
|
full_name = normalize_repo_full_name(full_name)
|
|
|
|
|
headers = {
|
|
|
|
|
"Authorization": f"Bearer {token}",
|
|
|
|
|
"Accept": "application/vnd.github+json",
|
|
|
|
|
"X-GitHub-Api-Version": "2022-11-28",
|
|
|
|
|
}
|
|
|
|
|
owner, name = full_name.split("/", 1)
|
|
|
|
|
async with httpx.AsyncClient() as client:
|
|
|
|
|
r = await client.get(
|
|
|
|
|
f"https://api.github.com/repos/{owner}/{name}",
|
|
|
|
|
headers=headers,
|
|
|
|
|
)
|
|
|
|
|
if r.status_code == 404:
|
|
|
|
|
raise HTTPException(404, "repository not found")
|
|
|
|
|
if r.status_code == 403:
|
|
|
|
|
raise HTTPException(403, "no access to this private repository")
|
|
|
|
|
if r.status_code != 200:
|
|
|
|
|
raise HTTPException(502, f"github API error ({r.status_code})")
|
|
|
|
|
body = r.json()
|
|
|
|
|
if body.get("private") is not True:
|
|
|
|
|
return
|
|
|
|
|
# Private repo: 200 from GitHub implies the user's token can read it.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.get("/review-styles")
|
|
|
|
|
async def api_list_review_styles(
|
|
|
|
|
session: dict[str, Any] = _SESSION_DEP,
|
|
|
|
|
) -> list[dict[str, Any]]:
|
|
|
|
|
records = await list_review_styles()
|
|
|
|
|
out: list[dict[str, Any]] = []
|
|
|
|
|
for record in records:
|
|
|
|
|
if record.get("status") == "running":
|
|
|
|
|
synced = await sync_review_style_run_status(record["full_name"])
|
|
|
|
|
out.append(synced)
|
|
|
|
|
else:
|
|
|
|
|
out.append(record)
|
|
|
|
|
return out
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.post("/review-styles")
|
|
|
|
|
async def api_create_review_style(
|
|
|
|
|
body: ReviewStyleCreate,
|
|
|
|
|
session: dict[str, Any] = _SESSION_DEP,
|
|
|
|
|
) -> dict[str, Any]:
|
|
|
|
|
token = await get_access_token(session["sub"])
|
|
|
|
|
if not token:
|
|
|
|
|
raise HTTPException(401, "github token unavailable, re-login required")
|
|
|
|
|
await _assert_repo_available_for_style_analysis(body.full_name, token)
|
|
|
|
|
return await create_review_style(body.full_name, session["sub"])
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.get("/review-styles/{full_name:path}")
|
|
|
|
|
async def api_get_review_style(
|
|
|
|
|
full_name: str,
|
|
|
|
|
session: dict[str, Any] = _SESSION_DEP,
|
|
|
|
|
) -> dict[str, Any]:
|
|
|
|
|
full_name = normalize_repo_full_name(full_name)
|
|
|
|
|
record = await get_review_style(full_name)
|
|
|
|
|
if not record:
|
|
|
|
|
raise HTTPException(404, "review style not found")
|
|
|
|
|
if record.get("status") == "running":
|
|
|
|
|
record = await sync_review_style_run_status(full_name)
|
|
|
|
|
return record
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.put("/review-styles/{full_name:path}")
|
|
|
|
|
async def api_update_review_style_prompt(
|
|
|
|
|
full_name: str,
|
|
|
|
|
body: ReviewStylePromptUpdate,
|
|
|
|
|
session: dict[str, Any] = _SESSION_DEP,
|
|
|
|
|
) -> dict[str, Any]:
|
|
|
|
|
full_name = normalize_repo_full_name(full_name)
|
|
|
|
|
record = await get_review_style(full_name)
|
|
|
|
|
if not record:
|
|
|
|
|
raise HTTPException(404, "review style not found")
|
|
|
|
|
return await set_custom_prompt(full_name, body.custom_prompt)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@router.post("/review-styles/{full_name:path}/analyze")
|
|
|
|
|
async def api_analyze_review_style(
|
|
|
|
|
full_name: str,
|
|
|
|
|
session: dict[str, Any] = _SESSION_DEP,
|
|
|
|
|
) -> dict[str, Any]:
|
|
|
|
|
full_name = normalize_repo_full_name(full_name)
|
|
|
|
|
token = await get_access_token(session["sub"])
|
|
|
|
|
if not token:
|
|
|
|
|
raise HTTPException(401, "github token unavailable, re-login required")
|
|
|
|
|
await _assert_repo_available_for_style_analysis(full_name, token)
|
|
|
|
|
record = await get_review_style(full_name)
|
|
|
|
|
if not record:
|
|
|
|
|
record = await create_review_style(full_name, session["sub"])
|
|
|
|
|
if record.get("status") == "running":
|
|
|
|
|
raise HTTPException(409, "analysis already running")
|
|
|
|
|
return await start_review_style_analysis(
|
|
|
|
|
full_name,
|
|
|
|
|
github_token=token,
|
|
|
|
|
created_by=session["sub"],
|
|
|
|
|
)
|