open-swe/agent/middleware/__init__.py
seahaven-openswe[bot] 0651de2ebf
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
feat: surface attributed PR creation failures (#180)
* feat: surface attributed PR creation failures

Port upstream #1659: adds PullRequestCreationGuardMiddleware that
blocks shell fallbacks (gh pr create, gh api /pulls, curl) when
open_pull_request fails, keeping failures visible. Also adds preflight
branch/repo visibility checks in open_pull_request with structured
failure payloads, and updates the prompt to forbid PR creation
fallbacks.

Refs: #134

* fix: fall back to core GitHub App scope when optional grants missing (#1701)

* fix: fall back to core GitHub App scope when optional grants missing

Proxy-token minting requested workflows:write and actions:read in the
permission set used for every sandbox. GitHub 422s a token request that
asks for a permission the installation hasn't granted, so any install
without workflows:write failed to mint a token and every run died in
before-agent setup with "GitHub App installation token is unavailable".

_resolve_proxy_token now walks a permission ladder (full -> +workflows ->
core) and returns the first scope that mints, recording the granted scope
so hourly proxy refreshes stay consistent. A missing optional grant now
degrades to the install-time core scope instead of failing the run;
workflow-file HITL pushes still require workflows:write and fail at push
time when it is absent.

* refactor: flatten proxy-token ladder loop with continue

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit f53caff1aa24a7b29d851b267aa3bdfe62c1e935)

Sea Haven fork deviation: upstream #1701 folds workflows:write into the
standing BASE/RUNTIME scope. This fork deliberately keeps workflows:write
OUT of the standing permission ladder (RUNTIME = core + actions:read;
LADDER = (RUNTIME, CORE)) so the sandbox proxy token cannot push
.github/workflows/* during normal operation. workflows:write is minted only
transiently by WorkflowPushGuardMiddleware for an approved HITL push and
dropped on restore, preserving token scope as a backstop for the workflow-
push approval control. Security-reviewed (agentic fan-out + GPT-4.1 cross
review); the standing-scope-carries-workflows:write bypass was blocked.

* fix(open-swe): harden proxy-token restore and mint error handling

Two low-severity follow-ups from the security review of the #1701 port.

Restore the recorded baseline scope after a workflow-push elevation instead
of a hardcoded RUNTIME. An install granted workflows:write but not actions:read
resolves its standing token to core; hardcoding RUNTIME on restore requested the
ungranted actions:read, 422'd, and fired a false "SECURITY: failed to downscope"
error on every approved workflow push before the core fallback recovered. The
guard now captures the run's recorded scope before elevating (via the new
get_recorded_proxy_permissions) and restores exactly that, falling back to the
guaranteed core scope only when the baseline restore fails.

Classify installation-token mint failures. get_github_app_installation_token_
with_expiry now treats HTTP 422 (a permission the installation hasn't granted)
as the ladder's expected descend signal and keeps it at debug, while a non-422
failure (network/5xx/timeout) is surfaced at WARNING even when errors are
otherwise suppressed — so a transient blip no longer silently downscopes a whole
run under a debug-only trace. The reduced-scope warning no longer asserts a
missing grant as the sole cause.

* chore(triage): mark upstream #1701 landed on this branch

Ported via PR #181 as Option A (workflows:write kept out of the standing
proxy-token scope). Regenerated triage.md from triage.jsonl.

* fix: restructure PR creation to POST-first with diagnose-on-failure

Move preflight checks from an authoritative gate (before POST) to a
diagnostic run after POST failure. This avoids false-positive failures
when a just-pushed head branch is momentarily invisible to GitHub ref
endpoints, and eliminates 2-3 extra serial API round-trips on the happy
path.

Also drop unused _PR_CREATED_FALSE indirection and add a docstring to
pr_creation_guard acknowledging the fail-open detection design.

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-13 14:45:19 -04:00

101 lines
4.3 KiB
Python

import sys
from types import ModuleType
from typing import TYPE_CHECKING, Any
_MIDDLEWARE_MODULES = {
"check_message_queue_before_model": ".check_message_queue",
"ensure_no_empty_msg": ".ensure_no_empty_msg",
"ExcludeToolsMiddleware": ".exclude_tools",
"ModelFallbackMiddleware": ".model_fallback",
"notify_step_limit_reached": ".notify_step_limit",
"PlanModeMiddleware": ".plan_mode",
"PullRequestCreationGuardMiddleware": ".pr_creation_guard",
"refresh_github_proxy_before_model": ".refresh_github_proxy",
"RepairOrphanedToolCallsMiddleware": ".repair_orphaned_tool_calls",
"SlackAssistantStatusMiddleware": ".refresh_slack_status",
"SandboxCircuitBreakerMiddleware": ".sandbox_circuit_breaker",
"SanitizeFireworksMessagesMiddleware": ".sanitize_fireworks_messages",
"SanitizeOpenAIResponsesMiddleware": ".sanitize_openai_responses",
"SanitizeThinkingBlocksMiddleware": ".sanitize_thinking_blocks",
"SanitizeToolInputsMiddleware": ".sanitize_tool_inputs",
"settle_review_check_on_exit": ".settle_review_check",
"SubdirAgentsReadMiddleware": ".subdir_agents",
"task_on_failure": ".task_retry",
"task_retry_on": ".task_retry",
"TimeoutWrapupMiddleware": ".timeout_wrapup",
"ToolArtifactMiddleware": ".tool_artifact",
"ToolErrorMiddleware": ".tool_error_handler",
"WorkflowPushGuardMiddleware": ".workflow_push_guard",
}
__all__ = [
"ExcludeToolsMiddleware",
"ModelFallbackMiddleware",
"PlanModeMiddleware",
"PullRequestCreationGuardMiddleware",
"RepairOrphanedToolCallsMiddleware",
"SanitizeFireworksMessagesMiddleware",
"SanitizeOpenAIResponsesMiddleware",
"SanitizeThinkingBlocksMiddleware",
"SanitizeToolInputsMiddleware",
"SubdirAgentsReadMiddleware",
"ToolArtifactMiddleware",
"ToolErrorMiddleware",
"TimeoutWrapupMiddleware",
"WorkflowPushGuardMiddleware",
"SandboxCircuitBreakerMiddleware",
"SlackAssistantStatusMiddleware",
"check_message_queue_before_model",
"ensure_no_empty_msg",
"notify_step_limit_reached",
"refresh_github_proxy_before_model",
"settle_review_check_on_exit",
"task_on_failure",
"task_retry_on",
]
if TYPE_CHECKING:
from .check_message_queue import check_message_queue_before_model
from .ensure_no_empty_msg import ensure_no_empty_msg
from .exclude_tools import ExcludeToolsMiddleware
from .model_fallback import ModelFallbackMiddleware
from .notify_step_limit import notify_step_limit_reached
from .plan_mode import PlanModeMiddleware
from .pr_creation_guard import PullRequestCreationGuardMiddleware
from .refresh_github_proxy import refresh_github_proxy_before_model
from .refresh_slack_status import SlackAssistantStatusMiddleware
from .repair_orphaned_tool_calls import RepairOrphanedToolCallsMiddleware
from .sandbox_circuit_breaker import SandboxCircuitBreakerMiddleware
from .sanitize_fireworks_messages import SanitizeFireworksMessagesMiddleware
from .sanitize_openai_responses import SanitizeOpenAIResponsesMiddleware
from .sanitize_thinking_blocks import SanitizeThinkingBlocksMiddleware
from .sanitize_tool_inputs import SanitizeToolInputsMiddleware
from .settle_review_check import settle_review_check_on_exit
from .subdir_agents import SubdirAgentsReadMiddleware
from .task_retry import task_on_failure, task_retry_on
from .timeout_wrapup import TimeoutWrapupMiddleware
from .tool_artifact import ToolArtifactMiddleware
from .tool_error_handler import ToolErrorMiddleware
from .workflow_push_guard import WorkflowPushGuardMiddleware
def _load_export(name: str) -> Any:
module_name = _MIDDLEWARE_MODULES.get(name)
if module_name is None:
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
from importlib import import_module
value = getattr(import_module(module_name, __name__), name)
globals()[name] = value
return value
class _LazyMiddlewareModule(ModuleType):
def __getattribute__(self, name: str) -> Any:
module_map = ModuleType.__getattribute__(self, "__dict__").get("_MIDDLEWARE_MODULES", {})
if name in module_map:
return _load_export(name)
return ModuleType.__getattribute__(self, name)
sys.modules[__name__].__class__ = _LazyMiddlewareModule