open-swe/tests
Johannes du Plessis 40162a6d9e
fix: inject existing PR review threads into reviewer context (#1331)
* fix(reviewer): inject existing PR review threads into reviewer context

The reviewer agent was filing the same inline comment on every re-review
because it only saw findings recorded on its own thread metadata — not
the live PR review-thread state on GitHub. When a previous finding was
still open (code unchanged, or a human reply explained it), the agent
rediscovered the same defect on the next push and called `add_finding`
again, producing duplicate comments.

This change fetches the PR's review threads (across all reviewers, with
replies and isResolved status) via GraphQL and renders them into the
first-review and re-review contexts as a "Pre-existing PR review
threads" block. The system prompt now lists overlap with that block as
a hard "Do NOT file" rule, and treats threads addressed by a human
reply as resolved.

This also gives the reviewer comment-awareness on its very first run
on a PR, so it skips findings already raised by another reviewer or
bot.

* fix(reviewer): wrap PR review threads in untrusted-data XML block

Addresses the reviewer comment on this PR
(https://github.com/langchain-ai/open-swe/pull/1331#discussion_r3295497533):
PR review comment bodies are attacker-controlled (anyone who can comment
on the PR can put anything in them), and they were being concatenated
into the reviewer's system prompt with instruction-priority.

Switches the existing-threads section from a Markdown block to an XML
data block:

  <pr_review_threads>
    <thread location="path:line" status="open">
      <comment author="open-swe[bot]">
        <body>...</body>
      </comment>
      <comment author="romain-priour-lc">
        <body>We added defaults in the template</body>
      </comment>
    </thread>
  </pr_review_threads>

The system prompt now explicitly names the wrapper, tells the agent that
everything inside it is untrusted data from the PR (not instructions),
and that prompt-injection payloads inside bodies must be disregarded.
We keep the bodies so the agent can actually read engineer replies —
that's the whole point of comment-awareness — but they're delimited as
data, not concatenated as prose. Modern frontier models are well-trained
to honor this contract.

Additional defenses:
- Author logins are validated against the GitHub username grammar; any
  unexpected value is rendered as "unknown" so the `author` attribute
  can't smuggle freeform text.
- Literal closing tags (`</body>`, `</pr_review_threads>`, etc.) in
  bodies are neutered so a body can't break out of its wrapper.
- Body length is capped at 4000 chars per comment to bound the prompt.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-24 16:36:50 -07:00
..
middleware fix: keep sandbox backend stable across recovery (#1294)w 2026-05-11 16:03:38 -07:00
conftest.py feat: restructure Open SWE Review tab + wire create_prs (#1319) 2026-05-21 09:17:07 -07:00
test_anthropic_effort.py feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
test_auth_sources.py chore: Drop monorepo (#1029) 2026-03-06 16:10:34 -08:00
test_dashboard_message_adapter.py feat: add Agents chat UI for cloud threads (#1323) 2026-05-22 18:15:59 +00:00
test_daytona_integration.py fix(daytona): make sandbox snapshot configurable (#1220) 2026-05-01 22:51:34 +00:00
test_encryption.py feat: support TOKEN_ENCRYPTION_KEY rotation via MultiFernet [closes AB-2323] (#1275) 2026-05-08 14:29:23 -07:00
test_ensure_no_empty_msg.py feat: add Agents chat UI for cloud threads (#1323) 2026-05-22 18:15:59 +00:00
test_github_comment_prompts.py feat: give agent self-awareness of its own repo (langchain-ai/open-swe) (#1266) 2026-05-08 13:13:35 -07:00
test_github_issue_webhook.py feat: restructure Open SWE Review tab + wire create_prs (#1319) 2026-05-21 09:17:07 -07:00
test_github_oauth_refresh.py fix: review style prompts UX, stale runs, and OAuth refresh (#1321) 2026-05-21 17:44:04 +00:00
test_github_token_ttl.py feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322] (#1280) 2026-05-08 22:57:01 +00:00
test_google_model.py feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
test_http_security.py fix: harden http_request SSRF guard against DNS rebinding [closes AB-2321] (#1277) 2026-05-08 22:06:25 +00:00
test_langsmith_sandbox_config.py feat: add idle TTL and delete-after-stop sandbox lifecycle controls (#1265) 2026-05-08 00:30:14 -04:00
test_messages_reducer_patch.py feat: add Agents chat UI for cloud threads (#1323) 2026-05-22 18:15:59 +00:00
test_model_fallback_middleware.py feat: cross-provider model fallback on transient errors (#1281) 2026-05-08 15:35:13 -07:00
test_multimodal.py chore: Drop monorepo (#1029) 2026-03-06 16:10:34 -08:00
test_normalize_repo.py feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
test_notify_step_limit_middleware.py fix: notify users via Slack when agent hits model call step limit (#1204) 2026-05-01 14:24:25 -07:00
test_pr_ready_auto_review.py feat: auto-review PRs on opened / ready-for-review (#1325) 2026-05-22 13:36:14 -07:00
test_proxy_auth.py feat: add Agents chat UI for cloud threads (#1323) 2026-05-22 18:15:59 +00:00
test_public_repo_org_gate.py feat: gate @open-swe mentions on public repos to org members (#1273) 2026-05-08 11:38:29 -07:00
test_recent_comments.py chore: Drop monorepo (#1029) 2026-03-06 16:10:34 -08:00
test_refresh_slack_status_middleware.py feat: add optional Slack Assistants API typing status indicator (#1269) 2026-05-08 10:21:55 -07:00
test_repo_extraction.py feat: support allowed repos in addition to allowed orgs for webhook filtering (#1092) 2026-05-08 13:26:30 -07:00
test_review_style_collector.py feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
test_review_style_sync.py fix: review style prompts UX, stale runs, and OAuth refresh (#1321) 2026-05-21 17:44:04 +00:00
test_review_styles_store.py feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
test_reviewer.py fix: inject existing PR review threads into reviewer context (#1331) 2026-05-24 16:36:50 -07:00
test_reviewer_diff.py feat: implement reviewer findings, publish_review, and watch mode (#1253) 2026-05-07 14:48:43 -07:00
test_reviewer_eval_run.py feat: Adds ability to run evals against deployment (#1311) 2026-05-18 15:47:13 -07:00
test_reviewer_eval_target.py feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
test_reviewer_findings.py feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
test_reviewer_publish.py fix: inject existing PR review threads into reviewer context (#1331) 2026-05-24 16:36:50 -07:00
test_reviewer_tools.py fix: remove line collapsing (#1330) 2026-05-23 00:43:48 +00:00
test_reviewer_watch.py feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322] (#1280) 2026-05-08 22:57:01 +00:00
test_sandbox_paths.py fix: better custom backend support (#1071) 2026-03-17 11:55:36 -07:00
test_sanitize_tool_inputs.py fix: coerce malformed integer strings in read_file offset/limit params (#1216) 2026-05-01 14:29:48 -07:00
test_slack_assistants_status.py chore: remove slack assistants api feature flag (#1295) 2026-05-12 00:41:51 +00:00
test_slack_context.py fix: stop inferring Slack repos from message text (#1306) 2026-05-15 22:36:44 +00:00
test_slack_feedback.py feat: add Slack reaction feedback to LangSmith (#1231) 2026-05-08 14:24:12 -07:00