* feat: port plan-review & workflow-approval UX (#135)
Port six upstream commits onto dev:
- c03a6be7 (already ported): keep plan guidance high-level
- 546042a4: add workflow approval UI with diff preview, approval URLs,
web review links, and polling for approval status during active runs
- 216cf181: remove workflow token elevation; approved pushes pass
through directly without proxy token rewriting
- 3dbc0282: preserve plan redirects after login by accepting relative
same-origin redirect_to values and rejecting blocked paths
- bb104d93: submit plan comments with cmd+enter
- 90cb6caa: terse Slack replies, shared content via save_plan outside
plan mode (PLAN_STATUS_SHARED), reject shared-content mutations
Refs: #135
* feat: port durable dispatch hardening and startup latency improvements
Port five upstream PRs onto dev:
- #1621 / #1658: durable dispatch with loopback webhook defense,
create_durable_run helper, _config_with_prepare_run_id, degradation
to None for relative/loopback completion webhook URLs
- #1696: run-level completion webhook deduplication (replace
claim-then-post with post-then-flag per run_id), DeferredErrorModel
for graph-factory resilience, ToolRetryMiddleware for task subagents,
TimeoutWrapupMiddleware for all three graphs
- #1697: lazy-load __init__.py for agent.middleware, agent.tools,
agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search,
agent.webapp in request_pr_review, deepagents in sandbox.py); add
ttl_cache.py with stale-while-revalidate for tool loaders
Refs: #137
* fix: restore login page render and clear CI lint/format
The plan-review port removed the authRedirectUrl import from login.tsx
but left its call site, crashing the login page at runtime (blank page,
no 'Sign in to open-swe'). Pass the relative path straight to loginUrl,
matching the plan route and the backend relative-redirect handling.
Also drop an unused os import in the guard test and reformat
workflow_push_guard.py to satisfy ruff.
* feat: port model-fallback resilience from upstream (#1694, #1695)
- Add httpx.TransportError to the transient-exception set so
incomplete chunked reads on streamed responses trigger a
fallback instead of cancelling the run (#1694 / c9f6dd86).
- Rewrite fallback to alternate primary/fallback with exponential
backoff instead of a single failover, so the agent survives
multi-minute gateway outages spanning both providers (#1695 /
c9a9a7cd).
- Default backoff schedule (0, 5, 15, 30, 45) reaches past the
gateway's ~30s recovery window; jittered ±25%.
- On exhaustion, surface a terminal AIMessage explaining the
outage instead of crashing — progress is checkpointed so the
user can retrigger to continue.
- Preserve fork conventions: Bedrock ClientError retryability
check, sync wrap_model_call (using time.sleep instead of
asyncio.sleep), and existing access-error surfacing for
Anthropic, OpenAI, and Bedrock (botocore) provider errors.
- Update triage ledger (c9f6dd86, c9a9a7cd → landed) and
re-render triage.md.
Refs: #139
* fix: align workflow-push-guard tests with dev's transient-elevation impl
The dev merge auto-combined dev's elevation tests with the stale passthrough
tests inherited from the durable-dispatch branch; the passthrough tests
contradict dev's restored _run_with_workflow_token impl. Take dev's test file.
* fix: drop dead ttl_cache module; make fallback backoff jitter two-sided
ttl_cache.py was re-introduced via the dev merge but dev/#160 deliberately
removed it as dead code (no agent importer). Remove it to match dev.
Also make _jittered_delay symmetric (±25%) to match its docstring.
* chore(upstream-sync): triage 4 new upstream commits (#1708-#1713)
Synced ledger to upstream/main (71e3b818). New rows all deferred:
- #1708 add GPT-5.6 OpenAI models (FLAG-HUMAN: fork picker is Bedrock/Fireworks-only)
- #1709 stale admin model defaults after upgrades
- #1710 bump langchain-fireworks 1.4.4
- #1713 align reviewer eval with published findings
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>