open-swe/tests/e2e/patches.py
seahaven-openswe[bot] 0546085672
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
feat: port durable dispatch hardening and startup latency improvements (#160)
* feat: port plan-review & workflow-approval UX (#135)

Port six upstream commits onto dev:

- c03a6be7 (already ported): keep plan guidance high-level
- 546042a4: add workflow approval UI with diff preview, approval URLs,
  web review links, and polling for approval status during active runs
- 216cf181: remove workflow token elevation; approved pushes pass
  through directly without proxy token rewriting
- 3dbc0282: preserve plan redirects after login by accepting relative
  same-origin redirect_to values and rejecting blocked paths
- bb104d93: submit plan comments with cmd+enter
- 90cb6caa: terse Slack replies, shared content via save_plan outside
  plan mode (PLAN_STATUS_SHARED), reject shared-content mutations

Refs: #135

* feat: port durable dispatch hardening and startup latency improvements

Port five upstream PRs onto dev:

- #1621 / #1658: durable dispatch with loopback webhook defense,
  create_durable_run helper, _config_with_prepare_run_id, degradation
  to None for relative/loopback completion webhook URLs
- #1696: run-level completion webhook deduplication (replace
  claim-then-post with post-then-flag per run_id), DeferredErrorModel
  for graph-factory resilience, ToolRetryMiddleware for task subagents,
  TimeoutWrapupMiddleware for all three graphs
- #1697: lazy-load __init__.py for agent.middleware, agent.tools,
  agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search,
  agent.webapp in request_pr_review, deepagents in sandbox.py); add
  ttl_cache.py with stale-while-revalidate for tool loaders

Refs: #137

* fix: restore login page render and clear CI lint/format

The plan-review port removed the authRedirectUrl import from login.tsx
but left its call site, crashing the login page at runtime (blank page,
no 'Sign in to open-swe'). Pass the relative path straight to loginUrl,
matching the plan route and the backend relative-redirect handling.

Also drop an unused os import in the guard test and reformat
workflow_push_guard.py to satisfy ruff.

* fix: restore RepairOrphaned middleware export and repoint model fake to deferred_model boundary

* fix: restore RepairOrphanedToolCallsMiddleware, fix E2E model-fake patch, drop dead ttl_cache

- Re-add RepairOrphanedToolCallsMiddleware to the lazy middleware __init__
  (_MIDDLEWARE_MODULES, __all__, TYPE_CHECKING) so agent.reviewer can import it.
- Reroute E2E model patching to deferred_model.make_model so make_model_or_defer
  (used by all three graph factories) returns the scripted fake instead of
  building a real model with fake credentials.
- Drop unused agent/utils/ttl_cache.py — no agent module imports it.
- Fix import ordering in agent/reviewer.py and agent/analyzer.py (ruff I001).
- Format tests/test_dispatch.py.

* fix: claim-then-post run-level failure dedup; stop permanent suppression

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-09 17:11:25 -04:00

88 lines
3.4 KiB
Python

"""Boundary monkeypatches: fake the LLM and the external SaaS endpoints.
Everything patched here is an *external boundary*, not agent logic:
- the LLM (model factory) -> scripted fake
- GitHub App token mint + GitHub REST base URL -> dummy token + fake GitHub
- Slack API base URL -> fake Slack
- the api.github.com/user identity lookup -> offline (falls back to config)
Applied at import of both the graph entrypoint and the HTTP harness (same dev
process), so it runs before the first run regardless of import order. Idempotent.
"""
from __future__ import annotations
import logging
import os
import e2e_env # noqa: F401 (sets env before any agent import)
logger = logging.getLogger(__name__)
_applied = False
def apply() -> None:
global _applied
if _applied:
return
import importlib
from agent.utils import auth, authorship
from agent.utils import slack as slack_utils
# NB: ``from agent.tools import open_pull_request`` returns the re-exported
# *function* (the tools package __init__ shadows the submodule), so patch the
# actual module object by name instead.
opr = importlib.import_module("agent.tools.open_pull_request")
from e2e_env import FAKE_GITHUB_API, FAKE_SLACK_API
# The LLM is the only agent-internal piece we fake, and only by default.
# Set E2E_REAL_LLM=1 to drive the harness (mock Slack/GitHub, real agent)
# with a real model — useful for manually exercising plan review etc. The
# provider key (e.g. ANTHROPIC_API_KEY) must be in the environment.
if os.environ.get("E2E_REAL_LLM"):
logger.warning("E2E_REAL_LLM set — using the real model factory, not the scripted fake")
else:
from fake_llm import FakeScriptedChatModel, build_script
def _fake_make_model(model_id: str, **kwargs: object): # noqa: ARG001
return FakeScriptedChatModel(script=build_script())
# ``make_model_or_defer`` (used by server/reviewer/analyzer) resolves
# ``make_model`` via the module global at call time, so patch it there.
import agent.utils.deferred_model as deferred_model
deferred_model.make_model = _fake_make_model
async def _dummy_install_token_with_expiry() -> tuple[str, str | None]:
return "dummy-installation-token", None
async def _dummy_install_token() -> str:
return "dummy-installation-token"
auth.get_github_app_installation_token_with_expiry = _dummy_install_token_with_expiry
opr.get_github_app_installation_token = _dummy_install_token
# Point the real PR/Slack code at the in-process fakes.
opr.GITHUB_API = FAKE_GITHUB_API
slack_utils.SLACK_API_BASE_URL = FAKE_SLACK_API
# Keep the triggering-user identity lookup offline; the real fallback to
# config-derived identity (Slack name/email) still runs.
authorship._identity_from_github_token = lambda _token: None # noqa: SLF001
# OAuth-token store is an external credential boundary. Stub it so a web
# follow-up (dashboard run.start) and PR-as-user resolution have a token;
# the real ownership/authorization checks still run.
from agent.dashboard import profiles, thread_api
async def _dummy_user_token(login: str, **_kwargs: object) -> str: # noqa: ARG001
return "dummy-user-oauth-token"
profiles.get_valid_access_token = _dummy_user_token
thread_api.get_valid_access_token = _dummy_user_token
_applied = True