open-swe/tests/e2e/fakes.py
seahaven-openswe[bot] 0651de2ebf
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
feat: surface attributed PR creation failures (#180)
* feat: surface attributed PR creation failures

Port upstream #1659: adds PullRequestCreationGuardMiddleware that
blocks shell fallbacks (gh pr create, gh api /pulls, curl) when
open_pull_request fails, keeping failures visible. Also adds preflight
branch/repo visibility checks in open_pull_request with structured
failure payloads, and updates the prompt to forbid PR creation
fallbacks.

Refs: #134

* fix: fall back to core GitHub App scope when optional grants missing (#1701)

* fix: fall back to core GitHub App scope when optional grants missing

Proxy-token minting requested workflows:write and actions:read in the
permission set used for every sandbox. GitHub 422s a token request that
asks for a permission the installation hasn't granted, so any install
without workflows:write failed to mint a token and every run died in
before-agent setup with "GitHub App installation token is unavailable".

_resolve_proxy_token now walks a permission ladder (full -> +workflows ->
core) and returns the first scope that mints, recording the granted scope
so hourly proxy refreshes stay consistent. A missing optional grant now
degrades to the install-time core scope instead of failing the run;
workflow-file HITL pushes still require workflows:write and fail at push
time when it is absent.

* refactor: flatten proxy-token ladder loop with continue

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit f53caff1aa24a7b29d851b267aa3bdfe62c1e935)

Sea Haven fork deviation: upstream #1701 folds workflows:write into the
standing BASE/RUNTIME scope. This fork deliberately keeps workflows:write
OUT of the standing permission ladder (RUNTIME = core + actions:read;
LADDER = (RUNTIME, CORE)) so the sandbox proxy token cannot push
.github/workflows/* during normal operation. workflows:write is minted only
transiently by WorkflowPushGuardMiddleware for an approved HITL push and
dropped on restore, preserving token scope as a backstop for the workflow-
push approval control. Security-reviewed (agentic fan-out + GPT-4.1 cross
review); the standing-scope-carries-workflows:write bypass was blocked.

* fix(open-swe): harden proxy-token restore and mint error handling

Two low-severity follow-ups from the security review of the #1701 port.

Restore the recorded baseline scope after a workflow-push elevation instead
of a hardcoded RUNTIME. An install granted workflows:write but not actions:read
resolves its standing token to core; hardcoding RUNTIME on restore requested the
ungranted actions:read, 422'd, and fired a false "SECURITY: failed to downscope"
error on every approved workflow push before the core fallback recovered. The
guard now captures the run's recorded scope before elevating (via the new
get_recorded_proxy_permissions) and restores exactly that, falling back to the
guaranteed core scope only when the baseline restore fails.

Classify installation-token mint failures. get_github_app_installation_token_
with_expiry now treats HTTP 422 (a permission the installation hasn't granted)
as the ladder's expected descend signal and keeps it at debug, while a non-422
failure (network/5xx/timeout) is surfaced at WARNING even when errors are
otherwise suppressed — so a transient blip no longer silently downscopes a whole
run under a debug-only trace. The reduced-scope warning no longer asserts a
missing grant as the sole cause.

* chore(triage): mark upstream #1701 landed on this branch

Ported via PR #181 as Option A (workflows:write kept out of the standing
proxy-token scope). Regenerated triage.md from triage.jsonl.

* fix: restructure PR creation to POST-first with diagnose-on-failure

Move preflight checks from an authoritative gate (before POST) to a
diagnostic run after POST failure. This avoids false-positive failures
when a just-pushed head branch is momentarily invisible to GitHub ref
endpoints, and eliminates 2-3 extra serial API round-trips on the happy
path.

Also drop unused _PR_CREATED_FALSE indirection and add a docstring to
pr_creation_guard acknowledging the fail-open detection design.

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-13 14:45:19 -04:00

170 lines
5.1 KiB
Python

"""In-memory state + git plumbing behind the fake GitHub and fake Slack.
These stores are the single source of truth that both the real agent code
(via the faked HTTP endpoints) and the mock UIs read from — so what Playwright
sees in the UI is exactly what the agent produced.
"""
from __future__ import annotations
import shutil
import subprocess
import time
from pathlib import Path
from typing import Any
from e2e_env import BARE_REMOTE, BASE_BRANCH, OWNER, REPO
# --- Slack -----------------------------------------------------------------
# (channel, thread_ts) -> list of {user, text, ts, blocks, is_bot}
SLACK_MESSAGES: dict[tuple[str, str], list[dict[str, Any]]] = {}
_slack_seq = [1]
def next_slack_ts() -> str:
_slack_seq[0] += 1
return f"1700000000.{_slack_seq[0]:06d}"
_thread_seq = [0]
def new_thread_ts() -> str:
"""A globally-unique thread ts so every send maps to a fresh LangGraph thread
(the in-mem store persists across restarts, so reused ids would carry state).
Not reset by reset(), so back-to-back tests never collide."""
_thread_seq[0] += 1
return f"{int(time.time())}.{_thread_seq[0]:06d}"
def add_slack_message(
channel: str, thread_ts: str, *, user: str, text: str, blocks: Any = None, is_bot: bool = False
) -> str:
ts = next_slack_ts()
actual_thread_ts = thread_ts or ts
SLACK_MESSAGES.setdefault((channel, actual_thread_ts), []).append(
{
"user": user,
"text": text,
"ts": ts,
"thread_ts": actual_thread_ts,
"blocks": blocks,
"is_bot": is_bot,
}
)
return ts
def slack_thread(channel: str, thread_ts: str) -> list[dict[str, Any]]:
return SLACK_MESSAGES.get((channel, thread_ts), [])
def slack_messages(channel: str) -> list[dict[str, Any]]:
messages: list[dict[str, Any]] = []
for (message_channel, _thread_ts), thread_messages in SLACK_MESSAGES.items():
if message_channel == channel:
messages.extend(thread_messages)
return sorted(messages, key=lambda message: message["ts"])
# --- GitHub ----------------------------------------------------------------
PULLS: list[dict[str, Any]] = []
_pr_seq = [0]
def _git(*args: str, cwd: Path | None = None) -> str:
result = subprocess.run(
["git", *args],
cwd=str(cwd) if cwd else None,
capture_output=True,
text=True,
check=True,
)
return result.stdout
def seed_bare_remote() -> None:
"""Create a fresh bare repo (the fake GitHub remote) with one commit on main."""
if BARE_REMOTE.exists():
shutil.rmtree(BARE_REMOTE)
seed_work = BARE_REMOTE.parent / f"seed-{OWNER}-{REPO}"
if seed_work.exists():
shutil.rmtree(seed_work)
seed_work.mkdir(parents=True)
ident = ["-c", "user.email=seed@example.com", "-c", "user.name=Seed"]
_git("init", "-b", BASE_BRANCH, str(seed_work))
(seed_work / "README.md").write_text("# demo\n\nA tiny demo repo.\n")
_git("add", "-A", cwd=seed_work)
_git(*ident, "commit", "-m", "Initial commit", cwd=seed_work)
_git("init", "--bare", "-b", BASE_BRANCH, str(BARE_REMOTE))
_git("remote", "add", "origin", str(BARE_REMOTE), cwd=seed_work)
_git("push", "origin", BASE_BRANCH, cwd=seed_work)
shutil.rmtree(seed_work)
def _diff_files(base: str, head: str) -> list[dict[str, Any]]:
"""Compute changed files for a PR from the pushed branch in the bare remote."""
try:
out = _git("--git-dir", str(BARE_REMOTE), "diff", "--numstat", base, head)
except subprocess.CalledProcessError:
return []
files = []
for line in out.splitlines():
parts = line.split("\t")
if len(parts) == 3:
adds, dels, name = parts
files.append(
{
"filename": name,
"additions": int(adds) if adds.isdigit() else 0,
"deletions": int(dels) if dels.isdigit() else 0,
}
)
return files
def branch_exists(branch: str) -> bool:
"""Check whether a branch exists in the bare remote (the fake GitHub)."""
try:
_git("--git-dir", str(BARE_REMOTE), "rev-parse", "--verify", f"refs/heads/{branch}")
return True
except subprocess.CalledProcessError:
return False
def create_pull(
owner: str, repo: str, *, head: str, base: str, title: str, body: str, draft: bool
) -> dict[str, Any]:
_pr_seq[0] += 1
number = _pr_seq[0]
files = _diff_files(base, head)
pr = {
"number": number,
"owner": owner,
"repo": repo,
"head": head,
"base": base,
"title": title,
"body": body,
"draft": draft,
"state": "open",
"merged": False,
"author": "open-swe[bot]",
"files": files,
"additions": sum(f["additions"] for f in files),
"deletions": sum(f["deletions"] for f in files),
}
PULLS.append(pr)
return pr
def find_pull(number: int) -> dict[str, Any] | None:
return next((p for p in PULLS if p["number"] == number), None)
def reset() -> None:
SLACK_MESSAGES.clear()
PULLS.clear()
_pr_seq[0] = 0
seed_bare_remote()