open-swe/agent/review_style_analyzer.py
Johannes du Plessis e87085139b
feat: add Agents chat UI for cloud threads (#1323)
* feat(ui): add Agents chat UI ported from open-swe-app

Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(dashboard): wire Agents UI to LangGraph thread APIs

Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(dashboard): single agent reply per turn in Agents UI

Use UUID thread IDs LangGraph accepts, skip confirming_completion for
dashboard threads, and merge adapter agent messages so duplicate bubbles
do not render.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(ui): polish Agents UI with floating prompt and layout cleanup

Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar
from open-swe-app, and refine chat layout so messages scroll behind the input.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent): patch deepagents reducer for None messages on checkpoint replay

LangGraph thread state could 500 when cancelled runs left messages as None.
Apply the reducer guard before graph import, fall back to metadata in the
dashboard API, and adjust Agents prompt bar layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(ui): unify sidebar user menu and clean up Agents UI navigation

Extract SidebarUserMenu so the dashboard and Agents sidebars render the
same profile button, drop the redundant Agents nav row in favor of the
existing Back to Agents link, add the open-swe logo header to the Agents
sidebar, flatten the New Agent button, and cap the home screen run list
to keep the prompt input in view.

* feat(ui): resizable/collapsible sidebar shared across dashboard and Agents

Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars
share a persisted width (default 260px, drag to resize, 200-420 range)
and a collapse toggle that hides the panel and surfaces a floating
reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on-
hover thread delete control in the Agents sidebar.

* feat(ui): instant user message and busy indicator on Agents transition

Stash submitted prompts in sessionStorage, pre-populate the new thread
detail cache, and merge pending prompts into the rendered message list
so the Agents page renders the user bubble plus the existing thinking
spinner immediately instead of flashing a skeleton and "Agent is
starting" while the run boots.

* feat(ui): token-stream agent replies in the Agents thread view

Opt the LangGraph runs into messages-tuple streaming and forward those
events through the existing SSE channel. The frontend now applies
AIMessageChunk deltas directly to the cached thread (cancelling any
in-flight refetch first so optimistic tokens are not clobbered) and
keeps positional pending prompts so the user bubble stays in the right
place while the agent streams its reply.

* fix(dashboard): await threads.join_stream before iterating

threads.join_stream is async def returning an AsyncIterator, so it must
be awaited before async for. The SSE endpoint was raising
TypeError: 'async for' requires an object with __aiter__ method, got
coroutine on every connection.

* fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns

Setting stream_mode=["values","messages-tuple","updates"] on
runs.create forces langchain_anthropic into streaming, and on the
second model call (after tool execution) its serialized thinking
blocks come back malformed, so Anthropic rejects the request with
'messages.1.content.0.thinking.thinking: Field required'. Revert to
the default stream_mode so claude-opus thinking + tool use runs to
completion. The frontend keeps the messages-event handler in place
as a no-op fallback for when streaming is re-enabled.

* feat(agents): per-thread model picker wired through to the run

Add optional model_id/effort to the create-thread and send-message
request bodies, forward them as agent_model_id/agent_effort in the
LangGraph run configurable, and record the resolved choice in thread
metadata so the UI can show the model the run is actually using.
get_agent now picks the per-thread override last (highest priority over
team default + profile override) and falls back gracefully when it is
absent or unsupported.

The frontend prompt bar becomes a controlled component fed by a
shared useModelOptions hook (options + profile -> defaultSelection).
AgentsHome seeds the picker from the user's profile default; the
thread view seeds from the thread's recorded model/effort and lets
each follow-up retarget the run.

* refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar

Drop the absolute-positioned send button, restore the original
px-4 py-3.5 min-h-[106px] flex-col container, and move the model
picker into a mt-auto pt-2 footer row so the placeholder text and
the model selector share the same horizontal padding.

* chore: fix lint/format CI failures

Remove unused imports and reformat two files flagged by ruff.

* fix(tests): stop messages-reducer patch tests from polluting the suite

Restore agent modules after reducer patch tests and import LangSmithSandbox
from agent.server in proxy refresh tests so isinstance checks stay valid.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 18:15:59 +00:00

164 lines
5.8 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

"""Review style analyzer graph.
Uses the same sandbox + ``gh`` pattern as the reviewer agent. The dashboard
user's OAuth token is injected into the LangSmith GitHub proxy so ``gh`` works
on public repos even when the GitHub App is not installed on them.
"""
# ruff: noqa: E402
from __future__ import annotations
import asyncio
import logging
import os
import warnings
from langgraph.graph.state import RunnableConfig
from langgraph.pregel import Pregel
warnings.filterwarnings("ignore", module="langchain_core._api.deprecation")
warnings.filterwarnings("ignore", message=".*Pydantic V1.*", category=UserWarning)
from ._patch_messages_reducer import _apply as _apply_messages_reducer_patch
_apply_messages_reducer_patch()
from deepagents import create_deep_agent
from deepagents.backends.protocol import SandboxBackendProtocol
from langchain.agents.middleware import ModelCallLimitMiddleware
from .integrations.langsmith import _configure_github_proxy
from .middleware import SanitizeToolInputsMiddleware, ToolErrorMiddleware
from .review_style_guidance import REVIEWER_STYLE_THEMES
from .server import (
DEFAULT_LLM_MAX_TOKENS,
DEFAULT_LLM_MODEL_ID,
DEFAULT_RECURSION_LIMIT,
ensure_sandbox_for_thread,
graph_loaded_for_execution,
)
from .tools.save_review_style import save_review_style_prompt
from .utils.model import DEFAULT_LLM_REASONING, make_model, provider_model_kwargs
from .utils.sandbox_paths import aresolve_sandbox_work_dir
from .utils.sandbox_state import unwrap_sandbox_backend
logger = logging.getLogger(__name__)
STYLE_ANALYZER_MODEL_CALL_LIMIT = 80
STYLE_ANALYZER_PROMPT = """You are a code-review style analyst for `{repo_owner}/{repo_name}`.
Sandbox: `{working_dir}`. Use the shell (``execute``) to run GitHub commands.
**Always invoke gh as:** `GH_TOKEN=dummy gh <command>`
# How to research (required)
Browse historical **merged** PR review feedback until you have catalogued at least
**8 substantive human** review comments (not bots). Suggested commands:
```
GH_TOKEN=dummy gh pr list --repo {repo_owner}/{repo_name} --state merged --limit 30
GH_TOKEN=dummy gh api repos/{repo_owner}/{repo_name}/pulls/<PR_NUMBER>/reviews
GH_TOKEN=dummy gh api repos/{repo_owner}/{repo_name}/pulls/<PR_NUMBER>/comments
GH_TOKEN=dummy gh api repos/{repo_owner}/{repo_name}/issues/<PR_NUMBER>/comments
```
If the first batch is sparse, increase `--limit` or walk older PR numbers. Skip
`[bot]` accounts and obvious automation (codecov, dependabot, etc.).
Identify the top ~5 human reviewers by volume and note phrasing, severity, and
what they ignore.
# When you may call `save_review_style_prompt`
Only after real research. Your `custom_prompt` (400–1200 words) must teach our
reviewer agent this repo's norms:
- What the team routinely flags vs skips (paraphrased patterns, not invented quotes)
- Severity calibration
- Tone and test expectations
- Repo-specific conventions
- Anti-patterns reviewers here avoid
`analysis_summary`: 2–4 sentences for the dashboard.
Pass `prs_sampled`, `reviews_sampled`, and `top_reviewers` (comma-separated logins).
Do **not** save a generic guide after one or two commands. Only after ~25+ merged
PRs with zero human feedback may you save a short conservative guide and say so in
`analysis_summary`.
# Alignment with our reviewer agent
{reviewer_themes}
# Optional preloaded samples
The user message may include pre-collected samples — verify and extend with ``gh``.
"""
async def _configure_sandbox_github_proxy(
sandbox_backend: SandboxBackendProtocol,
github_token: str,
) -> None:
if os.getenv("SANDBOX_TYPE", "langsmith") != "langsmith":
return
backend = unwrap_sandbox_backend(sandbox_backend)
await asyncio.to_thread(_configure_github_proxy, backend.id, github_token)
async def get_review_style_analyzer(config: RunnableConfig) -> Pregel:
thread_id = config["configurable"].get("thread_id")
config["recursion_limit"] = DEFAULT_RECURSION_LIMIT
if thread_id is None or not graph_loaded_for_execution(config):
return create_deep_agent(system_prompt="", tools=[]).with_config(config)
sandbox_backend = await ensure_sandbox_for_thread(thread_id)
work_dir = await aresolve_sandbox_work_dir(sandbox_backend)
configurable = config["configurable"]
full_name = str(configurable.get("review_style_full_name") or "owner/repo")
owner, _, name = full_name.partition("/")
samples_text = str(configurable.get("review_style_samples_text") or "")
github_token = configurable.get("review_style_github_token")
if isinstance(github_token, str) and github_token:
await _configure_sandbox_github_proxy(sandbox_backend, github_token)
model_id = DEFAULT_LLM_MODEL_ID
model_kwargs = provider_model_kwargs(
model_id,
None,
max_tokens=DEFAULT_LLM_MAX_TOKENS,
openai_reasoning_default=DEFAULT_LLM_REASONING,
)
system_prompt = STYLE_ANALYZER_PROMPT.format(
repo_owner=owner or "<owner>",
repo_name=name or "<repo>",
working_dir=work_dir,
reviewer_themes=REVIEWER_STYLE_THEMES.strip(),
)
user_context = (
f"Repository: `{full_name}`\n\n"
f"{samples_text}\n\n"
"Research review style with `GH_TOKEN=dummy gh ...` via execute, then call "
"`save_review_style_prompt` once you have enough evidence."
)
system_prompt = f"{system_prompt}\n\n{user_context}"
return create_deep_agent(
model=make_model(model_id, **model_kwargs),
system_prompt=system_prompt,
tools=[save_review_style_prompt],
backend=sandbox_backend,
middleware=[
SanitizeToolInputsMiddleware(),
ModelCallLimitMiddleware(
run_limit=STYLE_ANALYZER_MODEL_CALL_LIMIT,
exit_behavior="end",
),
ToolErrorMiddleware(),
],
).with_config(config)