open-swe/agent/utils/multimodal.py
Adam Moussa b3b0274403
feat: Re-land deferred upstream features on modular webhooks (#80) (#128)
* fix(webhooks): fall back to vision model for Slack/Linear image threads

Re-land upstream #1626 onto the modular webhook structure. When a
Slack mention or Linear issue carries images but the resolved model is
text-only, fall back to a vision-capable model instead of dropping the
images. Re-points default_vision_model_pair at the fork's image-capable
models (Opus 4.8 default, else any supports_images model) rather than
upstream's openai:/anthropic: provider filter.

Refs #80, upstream #1626

* fix(slack): persist trace_message_ts so web-handoff updates the trace reply

Re-land upstream #1630 onto the modular structure. The first-mention
store_slack_run_mapping call did not pass trace_message_ts, so it was
never persisted (nothing to preserve from on first mention) and
_notify_slack_web_handoff always skipped the trace-reply update on web
handoff. Pass it through and cover it with a test.

Refs #80, upstream #1630

* feat(slack): include channel context in Slack prompts

Re-land upstream #1633 onto the modular structure. Fetch cached Slack
channel metadata once per event (_get_slack_channel_context) and thread
it through the docs-plz gate, repo resolution, and process_slack_mention
so prompts carry the channel name and a clearly-marked untrusted
channel description. Avoids duplicate conversations.info calls.

Refs #80, upstream #1633

* feat(tools): add slack_start_new_thread breakout tool

Re-land upstream #1638 onto the modular structure. Adds the
slack_start_new_thread tool (posts a top-level Slack message and
dispatches a fresh agent run for a broken-out task via the durable
dispatch_agent_run contract), wires it into the agent tool list and
tools/__init__, adds prompt guidance, and excludes it from plan mode so
it can't bypass the approval flow. Tool imports only live modules.

Refs #80, upstream #1638

* feat(plan): notify Slack on plan approval

Re-land upstream #1632 onto the modular structure. When a plan is
approved via the dashboard approve endpoint, post a thread reply to the
originating Slack thread noting the comment count and approver, after
the follow-up run is dispatched. Slack post failures never break
approval. Adapted to the fork's approve_plan (no plan_markdown read).

Refs #80, upstream #1632

* feat(plan): publish plans from sandbox files

Re-land upstream #1635 onto the modular structure, completing the
partially-ported change so dev is internally consistent. save_plan now
takes a plan_file_path, reads the agent-authored Markdown file from
/workspace/plans/ (validating extension/location/UTF-8/size) and
publishes it, instead of taking a plan_markdown string. Removes
write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can
author the plan file, updates enter_plan_mode/reject_plan guidance and
the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev).

Refs #80, upstream #1635

* fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs

INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop
revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/
Linear image URL could 302-redirect the fetch to an internal host / cloud
metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot
is_url_safe check. Route image fetches through the same per-hop resolve+pin+
revalidate loop the http_request tool uses, lifted into url_safety as the shared
request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token
on redirect so it can't be replayed to a redirect target.

SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at
DEBUG; multimodal logged them at INFO on every fetch. Log host-only.

Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened
its reach by no longer dropping images for text-only models. Fixing on the base
branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression
tests (redirect-to-internal blocked; auth stripped on redirect).
2026-07-08 18:32:43 -04:00

116 lines
4.2 KiB
Python

"""Utilities for building multimodal content blocks."""
from __future__ import annotations
import base64
import logging
import mimetypes
import os
import re
from typing import Any
from urllib.parse import urlparse
import httpx
from langchain_core.messages.content import create_image_block
from .url_safety import request_with_safe_redirects
logger = logging.getLogger(__name__)
IMAGE_MARKDOWN_RE = re.compile(r"!\[[^\]]*\]\((https?://[^\s)]+)\)")
IMAGE_URL_RE = re.compile(
r"(https?://[^\s)]+\.(?:png|jpe?g|gif|webp|bmp|tiff)(?:\?[^\s)]+)?)",
re.IGNORECASE,
)
def extract_image_urls(text: str) -> list[str]:
"""Extract image URLs from markdown image syntax and direct image links."""
if not text:
return []
urls: list[str] = []
urls.extend(IMAGE_MARKDOWN_RE.findall(text))
urls.extend(IMAGE_URL_RE.findall(text))
deduped = dedupe_urls(urls)
if deduped:
logger.debug("Extracted %d image URL(s)", len(deduped))
return deduped
def vision_not_supported_warning(model_id: str, image_count: int) -> str:
"""Build a prompt-visible warning when images are sent to a text-only model."""
return (
f"\n\n**Note:** {image_count} image(s) were attached but the current model "
f"({model_id}) does not support image input. The images were not included. "
"Please switch to a vision-enabled model to process images."
)
async def fetch_image_block(
image_url: str,
client: httpx.AsyncClient,
) -> dict[str, Any] | None:
"""Fetch image bytes and build an image content block.
The fetch validates and pins every redirect hop (SSRF guard) and drops any
Authorization header once the URL redirects, so a per-host token is never
replayed to a redirect target the caller never chose to authenticate to.
URLs are logged host-only — a signed image URL can carry a bearer token.
"""
host = (urlparse(image_url).hostname or "").lower()
try:
headers = None
if host == "uploads.linear.app" or host.endswith(".uploads.linear.app"):
linear_api_key = os.environ.get("LINEAR_API_KEY", "")
if linear_api_key:
headers = {"Authorization": linear_api_key}
else:
logger.warning(
"LINEAR_API_KEY not set; cannot authenticate image fetch for %s", host
)
elif host == "files.slack.com" or host.endswith(".files.slack.com"):
slack_bot_token = os.environ.get("SLACK_BOT_TOKEN", "")
if slack_bot_token:
headers = {"Authorization": f"Bearer {slack_bot_token}"}
else:
logger.warning(
"SLACK_BOT_TOKEN not set; cannot authenticate image fetch for %s", host
)
response, blocked = await request_with_safe_redirects(
client, "GET", image_url, headers=headers, strip_auth_on_redirect=True
)
if blocked is not None:
_, reason = blocked
logger.warning("Refusing to fetch image (SSRF guard) from %s: %s", host, reason)
return None
response.raise_for_status()
content_type = response.headers.get("Content-Type", "").split(";")[0].strip()
if not content_type:
guessed, _ = mimetypes.guess_type(image_url)
if not guessed:
logger.warning("Could not determine content type from %s; skipping image", host)
return None
content_type = guessed
supported_types = {"image/jpeg", "image/png", "image/gif", "image/webp"}
if content_type not in supported_types:
logger.warning(
"Unsupported content type '%s' from %s; skipping image", content_type, host
)
return None
encoded = base64.b64encode(response.content).decode("ascii")
logger.info(
"Fetched image from %s (%s, %d bytes)", host, content_type, len(response.content)
)
return create_image_block(base64=encoded, mime_type=content_type)
except Exception:
logger.exception("Failed to fetch image from %s", host)
return None
def dedupe_urls(urls: list[str]) -> list[str]:
return list(dict.fromkeys(urls))