open-swe/agent/webhooks/jira.py

231 lines
9.5 KiB
Python
Raw Permalink Normal View History

feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"""Jira webhook handler — mirrors ``agent/webhooks/linear.py`` for Jira issues.
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
Helpers and constants stay in common.py; they are accessed through the module
object (``common.X``) so tests that monkeypatch them keep working.
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"""
from typing import Any
from urllib.parse import urlparse
import httpx
from langchain_core.messages.content import create_text_block
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
from . import common
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
async def process_jira_issue( # noqa: PLR0912, PLR0915
issue_data: dict[str, Any], repo_config: dict[str, str]
) -> None:
"""Process a Jira issue comment by creating a new LangGraph thread and run.
Args:
issue_data: The Jira issue data from the webhook (basic info + the
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
triggering comment; see ``jira_routes.jira_webhook`` for the shape).
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
repo_config: The repo configuration with owner and name.
"""
issue_key = issue_data.get("key", "")
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
common.logger.info(
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"Processing Jira issue %s for repo %s/%s",
issue_key,
repo_config.get("owner"),
repo_config.get("name"),
)
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
thread_id = common.generate_thread_id_from_jira_issue(issue_key)
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
full_issue = await common.fetch_jira_issue_details(issue_key)
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
if not full_issue:
full_issue = {}
# Actor email for token attribution: restricted to the comment author only,
# resolved once (by account_id) in the webhook handler and carried through
# here — mirrors the Linear handler's restriction to the comment author,
# so a PR is never opened as a non-actor.
comment_author = issue_data.get("comment_author") or {}
actor_email = comment_author.get("email")
user_name = comment_author.get("name") or None
user_email = actor_email
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
common.logger.info("User email for issue %s: %s", issue_key, user_email)
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
title = full_issue.get("title") or "No title"
description = full_issue.get("description") or "No description"
image_urls: list[str] = []
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
description_image_urls = common.extract_image_urls(description)
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
if description_image_urls:
image_urls.extend(description_image_urls)
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
common.logger.debug(
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"Found %d image URL(s) in issue description",
len(description_image_urls),
)
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
raw_comments = await common.fetch_jira_issue_comments(issue_key)
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
comments = [{**comment, "createdAt": comment.get("created", "")} for comment in raw_comments]
comments_text = ""
triggering_comment = issue_data.get("triggering_comment", "")
triggering_comment_id = issue_data.get("triggering_comment_id", "")
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
bot_message_prefixes = common._GITHUB_BOT_MESSAGE_PREFIXES
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
comment_ids: set[str] = set()
comment_id_to_index: dict[str, int] = {}
if comments:
for i, comment in enumerate(comments):
comment_id = comment.get("id", "")
if comment_id:
comment_ids.add(comment_id)
comment_id_to_index[comment_id] = i
relevant_comments = []
trigger_index = None
if triggering_comment_id:
trigger_index = comment_id_to_index.get(triggering_comment_id)
if trigger_index is not None:
relevant_comments = comments[trigger_index:]
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
common.logger.debug(
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"Using triggering comment index %d to build relevant comments",
trigger_index,
)
else:
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
relevant_comments = common.get_recent_comments(comments, bot_message_prefixes)
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
if relevant_comments:
comments_text = "\n\n## Comments:\n"
for comment in relevant_comments:
author = (comment.get("author") or {}).get("name") or "User"
body = comment.get("body", "")
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
body_image_urls = common.extract_image_urls(body)
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
if body_image_urls:
image_urls.extend(body_image_urls)
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
common.logger.debug(
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"Found %d image URL(s) in comment by %s",
len(body_image_urls),
author,
)
if any(body.startswith(prefix) for prefix in bot_message_prefixes):
continue
comments_text += f"\n**{author}:** {body}\n"
if triggering_comment and triggering_comment_id not in comment_ids:
if not comments_text:
comments_text = "\n\n## Comments:\n"
trigger_author = comment_author.get("name") or "Unknown"
trigger_body = triggering_comment
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
trigger_image_urls = common.extract_image_urls(trigger_body)
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
if trigger_image_urls:
image_urls.extend(trigger_image_urls)
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
common.logger.debug(
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"Found %d image URL(s) in triggering comment by %s",
len(trigger_image_urls),
trigger_author,
)
comments_text += f"\n**{trigger_author}:** {trigger_body}\n"
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
common.logger.debug(
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"Appended triggering comment %s not present in issue comments list",
triggering_comment_id or "<missing-id>",
)
project_key = issue_data.get("project_key", "") or full_issue.get("project_key", "")
issue_number = issue_key.split("-", 1)[1] if "-" in issue_key else ""
issue_url = full_issue.get("url", "")
triggered_by_line = f"## Triggered by: {user_name}\n\n" if user_name else ""
tag_instruction = (
f"When calling jira_comment, tag @{user_name} if you are asking them a question, need their input, or are notifying them of something important (e.g. a completed PR). For simple answers, tagging is not required."
if user_name
else ""
)
prompt = (
f"Please work on the following issue:\n\n"
f"## Repository: {repo_config.get('owner')}/{repo_config.get('name')}\n\n"
f"## Title: {title}\n\n"
f"{triggered_by_line}"
f"## Jira Ticket: {issue_key}\n\n"
f"## Description:\n{description}\n"
f"{comments_text}\n\n"
f"Please analyze this issue and implement the necessary changes. "
f"When you're done, commit and push your changes. {tag_instruction}"
)
content_blocks: list[dict[str, Any]] = [create_text_block(prompt)]
# Resolve the GitHub login from the actor's Jira email via the same
# user-mapping store Slack/Linear use, so PRs open *as the triggering user*
# and the thread is tagged for the dashboard. Restricted to the comment
# author so token attribution never falls back to reporter/assignee.
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
mapped_login = await common.resolve_login_from_email_async(actor_email) if actor_email else None
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
# Only attribute to an *active* user mapping; a pending/unconfirmed mapping
# must never drive PR authorship or token resolution.
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
if mapped_login and not common.is_login_mapped(mapped_login):
common.logger.info(
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"Jira actor login %s is not an active mapping; running unattributed", mapped_login
)
mapped_login = None
image_model_override: tuple[str, str] | None = None
if image_urls:
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
image_urls = common.dedupe_urls(image_urls)
resolved_model_id = await common.resolve_agent_model_id(mapped_login)
if not common.model_supports_images(resolved_model_id):
fallback_model_id, fallback_effort = common.default_vision_model_pair()
common.logger.info(
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"Using vision fallback model %s for %d Jira image(s); configured model %s "
"does not support images",
fallback_model_id,
len(image_urls),
resolved_model_id,
)
resolved_model_id = fallback_model_id
image_model_override = (fallback_model_id, fallback_effort)
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
common.logger.info("Preparing %d image(s) for multimodal content", len(image_urls))
common.logger.debug("Image hosts: %s", [urlparse(u).hostname for u in image_urls])
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
async with httpx.AsyncClient(timeout=common.DEFAULT_HTTP_TIMEOUT) as client:
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
for image_url in image_urls:
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
image_block = await common.fetch_image_block(image_url, client)
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
if image_block:
content_blocks.append(image_block)
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
common.logger.info("Built %d content block(s) for prompt", len(content_blocks))
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
configurable: dict[str, Any] = {
"repo": repo_config,
"jira_issue": {
"key": issue_key,
"url": issue_url,
"project_key": project_key,
"issue_number": issue_number,
"title": title,
"triggering_user_name": user_name or "",
},
"user_email": user_email,
"source": "jira",
}
if mapped_login:
configurable["github_login"] = mapped_login
if image_model_override:
configurable["agent_model_id"] = image_model_override[0]
configurable["agent_effort"] = image_model_override[1]
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
await common.upsert_agent_thread_owner_metadata(
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
thread_id,
source="jira",
repo_config=repo_config,
github_login=mapped_login or "",
user_email=user_email or "",
title=title or issue_key or "Jira issue",
source_context={"jira_issue": configurable["jira_issue"]},
)
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
run = await common.dispatch_agent_run(
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
thread_id,
content_blocks,
configurable,
source="jira",
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
metadata=common._AGENT_VERSION_METADATA,
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
)
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
common.logger.info(
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"LangGraph run dispatched for thread %s (run=%s)",
thread_id,
run.get("run_id") if isinstance(run, dict) else None,
)
refactor: split webapp.py into api/ + per-source webhook routes Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved decisions 1-2): split the 2,590-line agent/webapp.py monolith into agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py (composition), agent/api/health.py (/health + /webhooks/run-complete), and per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian Connect lifecycle + descriptor routes (/connect/*) fold into confluence_routes.py; webapp.py becomes the upstream-shaped compatibility shim (from .api.app import app). langgraph.json http.app stays agent.webapp:app via the shim. Fork content, upstream layout: linear/slack route files verified content-identical to upstream 8356eb34 and taken verbatim; github_routes is upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are fork-only, transformed to the same common.X / service.X module-attribute style. All signature verification (GitHub HMAC, Slack, Linear timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo binding, _is_repo_auto_review_enabled gates, and public-repo org gate move unchanged. Handlers rewired from webapp.X to common.X; test monkeypatch sites across 26 files + conftest.py + e2e/harness.py retargeted to webhook_common/handler/route modules per upstream's pattern. Residual agent.webapp importers: only the shim, langgraph.json http.app, Makefile uvicorn target, and docs (doc-path updates land in C7). Gates: ruff check + format, pytest --co, full unit (1637 passed), full Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
2026-07-17 14:30:05 -04:00
await common.post_jira_trace_comment(issue_key, thread_id)