feat: Jira + Confluence integration (tools + triggers) (#182)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run

* feat(open-swe): add Jira tool plane (Phase 1)

Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear
tools:

- utils/jira.py: service-account REST client (Basic auth) with get/
  create/update issue, comments, list projects, trace comment; issue and
  comment bodies normalized to markdown.
- utils/adf.py: minimal ADF <-> markdown conversion (read paths convert
  Jira ADF to markdown; agent comments convert prose to ADF).
- tools/jira_{comment,get_issue,get_issue_comments,create_issue,
  update_issue,list_projects}.py wired into the tool registry and the
  main agent tool list.
- tests/test_jira_utils.py: ADF conversion + mocked-transport util tests.

Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env
returns a clean error, so this is safe to land dark. Trigger plane,
prompt guidance, and config plumbing follow in Phase 2.

* feat(open-swe): add Confluence tool plane (Phase 3)

Curated Confluence Cloud REST toolset for the agent, mirroring the Jira
tools:

- utils/confluence.py: service-account REST client (Basic auth) with
  get/create/update page, add comment, CQL search. Page bodies are XHTML
  storage format (not ADF), with minimal storage<->text converters;
  update_page reads the current version and bumps it, as Confluence
  requires.
- tools/confluence_{get_page,create_page,update_page,comment,search}.py
  registered in the tool registry.
- tests/test_confluence_utils.py: converter + mocked-transport tests
  including the version-bump path.

Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN;
unset env returns a clean error. Activation in the agent tool list lands
with the Phase 2 server.py wiring.

* feat(open-swe): add Jira trigger plane (Phase 2)

Make an @openswe comment on a Jira issue spawn an agent run, mirroring
the Linear trigger plane:

- webhooks/jira.py: process_jira_issue clones process_linear_issue —
  deterministic thread id, full-issue fetch, actor accountId->email
  attribution feeding resolve_login_from_email_async (PRs open as the
  human), multimodal image handling, source="jira" + jira_issue config.
- webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time
  X-Automation-Webhook-Token check, fails closed), repo-resolution
  cascade, get_repo_config_from_jira_mapping.
- utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder
  entry — real project->repo mappings still needed).
- utils/jira.py: get_user_email (accountId -> email) for attribution.
- completion.py: source=="jira" failure-reply branch.
- prompt.py: Jira-triggered notify guidance + Refs:/branch key from
  {jira_project_key}-{jira_issue_number}.
- server.py: read jira_issue config + pass jira key to the system
  prompt; also activates the Phase 3 Confluence tools in the agent list.

Jira Automation lacks native webhook HMAC signing, so trust is a shared
secret header (decision D2); replay protection is weaker than Linear's
HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the
outstanding gate/hardening before push.

* fix(open-swe): harden Jira webhook trust (sh-security-review)

Resolves findings from the Phase 2 security review (detector fan-out +
proof-or-kill verifier). The unsigned Jira Automation webhook body was
trusted for identity, comment content, repo routing, and issue
existence; a JIRA_WEBHOOK_SECRET holder could forge those fields.

- Corroborate against the real Jira record: the webhook body is now only
  a pointer (issue_key + required comment_id). The triggering comment's
  author and text are re-fetched server-side via get_comment/fetch_jira_
  comment, and identity, the @openswe check, prompt text, and project
  key are derived from that authoritative record — never payload author/
  body fields. An uncorroborated comment is rejected. (closes the
  account-id impersonation, unsigned-body prompt injection, and
  fabricated-issue findings)
- Validate issue_key against the Jira key format and percent-encode all
  untrusted path segments (_seg) so a crafted key can't traverse to a
  different Jira REST endpoint or inject query params. (closes the path-
  traversal / query-injection findings)
- Route source=="jira" through the bot-token-default / author_prs_as_
  user opt-in path in resolve_github_token, matching Linear, instead of
  unconditionally resolving a per-user OAuth token from a payload email.
- Gate attribution on an active user mapping (is_login_mapped) so a
  pending/unconfirmed mapping can't drive PR authorship.

Adds regression tests: server-corroboration wins over payload, malformed
issue_key rejected, uncorroborated comment rejected, path-segment
encoding, project-key derivation, active-mapping gate.

Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/
REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body +
timestamp on the Automation payload to close the residual replay gap.

* harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist

Folds the two deployment-hardening items from the Phase 2 security review
into code (all opt-in / default-off, so existing and upstream deployments
are unaffected):

- JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must
  carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by
  JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by
  verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear
  HMAC+freshness model). Closes the static-token model's replay/forgery
  gap when enabled.
- JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's
  direct client IP (verify_jira_source_ip). Documented as direct-peer
  only; behind a proxy/LB, allowlist Atlassian's ranges at that layer.
- REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail
  CLOSED instead of the back-compat allow-all, plus a startup fail-open
  warning. Applies to all channels for consistency.

Documents all new vars (and a Jira section) in .env.example. Adds tests
for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and
the fail-closed allowlist.

* feat(open-swe): Confluence Atlassian Connect trigger (Phase 4)

Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian
Connect app. Designed and adversarially verified with the ultracode
workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus
skeptics on the implemented crypto).

- utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's
  official test vector), PyJWT HS256 webhook verifier with alg-pinning,
  issuer binding, and qsh-verified-last ordering; RS256 signed-install
  lifecycle verifier against Atlassian's published keys; installation
  store keyed by clientKey with the sharedSecret encrypted at rest
  (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned).
- webhooks/confluence.py: install/uninstall lifecycle + comment handler.
  The JWT-signed webhook body is only a pointer; the comment's real
  author/text/container are re-fetched server-side via the Basic-auth
  service account (Phase-2 corroboration lesson), with active-only login
  attribution and the repo allowlist.
- utils/confluence.py: get_comment / get_user_email (path-encoded).
- webapp.py: GET /connect/atlassian-connect.json (served dynamically),
  POST /connect/{installed,uninstalled,webhook/comment-created}, the
  space->repo resolver, thread-id, and fetch helpers.
- completion.py: source=="confluence" failure-reply branch.

Security: the sh-security-review verify pass confirmed one HIGH — the
symmetric signed-install=false first-install was trust-on-first-use gated
only by the public Confluence hostname (webhook-auth bypass). Fixed by
switching to signed-install=true + RS256 verification of lifecycle
callbacks, which cryptographically authenticates the first install. All
other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall
DoS/corroboration/injection) were defeated; residuals are deployment
config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover
bodies; comment-trigger prompt injection, shared with all sources).

New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN,
CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets
require the durable Postgres LangGraph store in prod.

Outstanding before push: /sh-security-review on the real diff and the
GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory.

* docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs

- prompt.py: Confluence-triggered runs notify via confluence_comment on
  the triggering page; add Confluence to the shared-base source list.
- CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian
  triggers (Jira Automation shared-secret webhook; Confluence Connect app
  with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the
  server-side corroboration + encrypted install store.

Phase 5 also verified the trigger surface end-to-end against a running
uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed
without valid auth) and recorded the integration in project memory.

* fix(open-swe): resolve /sh-security-review findings on the Atlassian surface

Formal sh-security-review (detector fan-out + verifier) over the Phase-4
Connect surface (esp. the new RS256 signed-install code, unseen by the
earlier adversarial verify) and the Phase-2 opt-in hardening.

CRITICAL — cross-tenant install (origin validation, CWE-346): signed-
install proves the caller is *an* Atlassian tenant, not *ours*, and the
descriptor is served publicly, so any attacker could install the app on
their own Confluence site and drive agent runs against our allowlisted
repos. The baseUrl body field is attacker-controlled and cannot bind the
tenant; only the signature-verified clientKey (JWT iss) can. Added a
MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in
process_install after signature+iss verification.

HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment
ids are per-instance, so generate_thread_id_from_confluence_comment now
salts the hash with the verified clientKey (plumbed from the webhook JWT
iss) to prevent thread hijack across tenants.

HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page
interpolated page_id into the REST path unencoded (update_page on a
mutating PUT with no params= backstop). Now _seg()-encoded, matching the
rest of the module.

MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no
bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID
guard mirroring the Linear botActor / Jira comment_author_is_bot checks.

LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as
opt-in defense-in-depth (the clientKey allowlist is the real gate).

Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp,
kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption,
constant-time comparisons, and the Phase-2 hardening. New regression
tests for each fix; full suite green (1602).

* harden(open-swe): GPT-4.1 cross-family review follow-ups

Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no
critical/high issues and confirmed the auth boundary is fail-closed and
correct. Two low-cost defense-in-depth items applied:

- Validate the signed-install JWT 'kid' against a strict charset before
  the public-key fetch, so a malformed kid fails fast with no network
  call (on top of the existing fixed host + percent-encoding).
- Make JWT nbf verification explicit (verify_nbf) on both the RS256
  lifecycle and HS256 webhook decodes.

Other suggestions triaged as already-handled (aud cross-app replay is
blocked by the per-tenant iss->secret lookup; documented static-token/IP/
baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra
(Fernet rotation via MultiFernet; rate limiting at the gateway).

* docs(open-swe): document Jira + Confluence in installation & customization guides

- INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret,
  service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect
  app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note,
  CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars +
  REQUIRE_REPO_ALLOWLIST.
- CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction
  note covers all four sources.
- AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth).
- README.md: invocation section, tools table, and overview line.
This commit is contained in:
Adam Moussa 2026-07-13 19:45:54 -04:00 • committed by GitHub
parent 2fb1122630
commit 1f80500cf3
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
39 changed files with 4045 additions and 23 deletions

View file

@ -52,6 +52,10 @@ ALLOWED_GITHUB_ORGS="" # e.g. "my-org,my-other-org"
# Slack mentions are not rejected from regex-inferred repository text; repository access is bounded by GitHub App installation permissions.
# Leave both empty to allow all repos.
ALLOWED_GITHUB_REPOS="" # e.g. "some-user/their-repo,another-org/specific-repo"
# When "true", an empty allowlist fails CLOSED (allow nothing) instead of the
# back-compat allow-all. Set this once the allowlist above is populated so a
# forged/misconfigured trigger can't steer the agent at an arbitrary repo.
REQUIRE_REPO_ALLOWLIST="" # "true" to fail closed when the allowlist is empty
# === Default Repository ===
# Used across all triggers when no repo is specified.
@ -84,6 +88,47 @@ LANGGRAPH_URL="http://localhost:2024"
LINEAR_API_KEY="" # From step 5
LINEAR_WEBHOOK_SECRET="" # From step 5
# === Jira (if using Jira trigger / tools) ===
JIRA_BASE_URL="" # e.g. "https://your-site.atlassian.net"
JIRA_SERVICE_EMAIL="" # Service-account email for Basic auth
JIRA_API_TOKEN="" # Service-account API token
# Shared secret the Jira Automation rule sends in X-Automation-Webhook-Token.
JIRA_WEBHOOK_SECRET="" # Generate with: openssl rand -hex 32
# Opt-in stronger trust: when "true", the Automation rule must also send
# X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by
# JIRA_WEBHOOK_SECRET) and a fresh `timestamp` (Unix ms) in the body. Closes the
# static-token model's replay/forgery gap. Leave empty to keep token-only.
JIRA_WEBHOOK_REQUIRE_SIGNATURE="" # "true" to require HMAC body signature + timestamp
# Optional CIDR allowlist for the webhook's DIRECT client IP (comma-separated).
# Only meaningful when the app terminates connections directly; behind a proxy
# or LB, allowlist Atlassian's egress ranges at that layer instead.
JIRA_WEBHOOK_IP_ALLOWLIST="" # e.g. "1.2.3.0/24,5.6.7.8/32"
# === Confluence (if using Confluence tools / trigger) ===
# Basic-auth service account for the Confluence REST tools + webhook corroboration.
CONFLUENCE_BASE_URL="" # e.g. "https://your-site.atlassian.net"
CONFLUENCE_EMAIL="" # Service-account email
CONFLUENCE_API_TOKEN="" # Service-account API token
# Atlassian Connect app (the Confluence @openswe comment trigger). The descriptor
# is served at GET /connect/atlassian-connect.json with this baseUrl (public
# origin, EMPTY context path so descriptor path == request path == qsh path).
CONNECT_BASE_URL="" # e.g. "https://openswe.seahavenind.com"
# REQUIRED tenant binding: comma/space-separated clientKey(s) allowed to install.
# signed-install proves the caller is *an* Atlassian tenant, not *ours*, and the
# descriptor is public — so without this any tenant could install the app and
# trigger runs. Empty => reject ALL installs (fail closed). Bootstrap: attempt an
# install, read the rejected clientKey from the logs, add it here, re-install.
CONNECT_EXPECTED_CLIENT_KEYS="" # e.g. "a1b2c3d4-....-jira-confluence"
# Optional defense-in-depth on the (untrusted) install baseUrl host; enforced
# only when set. The clientKey allowlist above is the real tenant gate.
CONNECT_EXPECTED_BASE_URL="" # e.g. "your-site.atlassian.net"
# Optional: the app's own Confluence service-account accountId. When set,
# comments it authored are ignored (self-trigger loop guard).
CONFLUENCE_BOT_ACCOUNT_ID=""
# NOTE: install sharedSecrets are stored encrypted via TOKEN_ENCRYPTION_KEY (below)
# in the LangGraph store — that MUST be the durable Postgres store in prod, or
# installs are lost on restart and every webhook 401s until reinstall.
# === Slack (if using Slack trigger) ===
SLACK_BOT_TOKEN="" # From step 5
SLACK_BOT_USER_ID=""

View file

@ -4,7 +4,7 @@ This file provides guidance to Coding Agents when working with code in this repo
## Project
Open SWE is an open-source coding-agent framework built on **LangGraph** + **Deep Agents** (`deepagents.create_deep_agent`). It runs as a LangGraph app: each thread spawns its own isolated cloud sandbox, and the agent is invoked from Slack, Linear, or GitHub (PR comments, plus auto-review on opened / ready-for-review).
Open SWE is an open-source coding-agent framework built on **LangGraph** + **Deep Agents** (`deepagents.create_deep_agent`). It runs as a LangGraph app: each thread spawns its own isolated cloud sandbox, and the agent is invoked from Slack, Linear, Jira, Confluence, or GitHub (PR comments, plus auto-review on opened / ready-for-review).
A separate **reviewer** graph runs read-only code reviews on PRs, and a **review-style analyzer** graph learns per-repo review style from historical PRs.
@ -27,7 +27,7 @@ make format # ruff format + ruff check --fix
| Graph | Entrypoint | Purpose |
|---|---|---|
| `agent` | `agent.server:get_agent` | Main coding agent (Slack/Linear/GitHub-triggered). |
| `agent` | `agent.server:get_agent` | Main coding agent (Slack/Linear/Jira/Confluence/GitHub-triggered). |
| `reviewer` | `agent.reviewer:get_reviewer_agent` | Read-only PR reviewer. Findings model + `publish_review`. |
| `analyzer` | `agent.analyzer:get_analyzer` | Learns per-repo reviewer style from historical PRs and this reviewer's own finding outcomes. |
| `ci_monitor` | `agent.ci_monitor:get_ci_monitor` | Polling fallback for CI auto-fix: each tick sweeps open agent-authored PRs for failing checks / merge conflicts via `agent.ci_autofix.sweep_open_prs`. |
@ -43,7 +43,7 @@ CI auto-fix ("PR babysitting") lives in `agent/ci_autofix.py`: when a CI check f
- **`agent/server.py` → `get_agent(config)`** — main graph factory. Called per-thread. Resolves the GitHub token, gets-or-creates the sandbox for the thread, resolves the team/profile/per-thread model + effort, then constructs a fresh `create_deep_agent(...)` with the curated tool list and middleware stack. The agent itself is stateless — all per-thread state lives in the sandbox + thread metadata.
- **`agent/reviewer.py` → `get_reviewer_agent(config)`** — reviewer graph factory. Shares `ensure_sandbox_for_thread` with the main agent but wires a reviewer-only toolset (`add_finding`, `update_finding`, `list_findings`, `publish_review`, `web_search`, `fetch_url`, `http_request`) and a different system prompt that pins the single-evolving-findings model and the diff-anchored bar for filing a finding. Read-only: no commit/push/PR-opening tools.
- **`agent/analyzer.py` → `get_analyzer(config)`** — small graph that emits a per-repo style prompt via the `save_review_style_prompt` tool, consumed by the reviewer as a "repository-specific review style" appendix. It runs in one of two modes (`analyzer_mode` in `configurable`): **bootstrap** (cold-start: crawl historical PR reviews) and **continual** (nightly: refine using this reviewer's own finding outcomes via `read_finding_outcomes`). Each mode's procedure lives in a deepagents **skill** (`agent/skills/bootstrap-repo-analysis/`, `agent/skills/continual-learning/`) served as virtual files via a `CompositeBackend` `/skills/` route + `StateBackend` (seeded into the run's `files` channel by the launcher — never written to the sandbox). Launchers and the per-repo nightly cron live in `agent/dashboard/review_style_jobs.py` and `agent/dashboard/analyzer_cron.py`; the cron is registered when bootstrap completes.
- **`agent/webapp.py`** — custom FastAPI routes mounted alongside the LangGraph server. Webhooks land here (GitHub, Linear, Slack). Each webhook resolves a deterministic `thread_id` (so follow-up messages route to the same agent run) and triggers/streams a run via the `langgraph_sdk` client. Also auto-reviews PRs on `opened` / `ready_for_review` events when the repo+author opt in.
- **`agent/webapp.py`** — custom FastAPI routes mounted alongside the LangGraph server. Webhooks land here (GitHub, Linear, Slack, Jira, and the Confluence Atlassian Connect `/connect/*` routes). Each webhook resolves a deterministic `thread_id` (so follow-up messages route to the same agent run) and triggers/streams a run via the `langgraph_sdk` client. Also auto-reviews PRs on `opened` / `ready_for_review` events when the repo+author opt in. The Atlassian triggers (`agent/webhooks/{jira,confluence}.py`, `agent/utils/atlassian_connect.py`) verify webhook trust — a Jira Automation shared secret / optional HMAC, and a Confluence Connect HS256-JWT + `qsh` with RS256 signed-install — then re-fetch the triggering comment server-side before deriving identity.
- **`agent/dashboard/`** — `router` mounted under the FastAPI app at startup (`app.include_router(dashboard_router)`). Owns GitHub OAuth, per-user profiles, admin endpoints, team defaults, enabled-repo lists, review-style management, and the Agents chat thread API used by the UI in `ui/`.
### Sandbox lifecycle (the tricky part)
@ -86,7 +86,7 @@ There is intentionally no after-agent safety net that opens a PR for the agent.
All tools live in `agent/tools/` and are flat-imported via `agent/tools/__init__.py`. The set is intentionally small and curated — see README "Tools — Curated, Not Accumulated".
Wired into `get_agent`:
`http_request`, `fetch_url`, `web_search`, `linear_comment`, `linear_create_issue`, `linear_delete_issue`, `linear_get_issue`, `linear_get_issue_comments`, `linear_list_teams`, `linear_update_issue`, `request_pr_review`, `schedule_thread_wakeup`, `slack_add_reaction`, `slack_read_thread_messages`, `slack_thread_reply`.
`http_request`, `fetch_url`, `web_search`, `linear_comment`, `linear_create_issue`, `linear_delete_issue`, `linear_get_issue`, `linear_get_issue_comments`, `linear_list_teams`, `linear_update_issue`, `jira_comment`, `jira_create_issue`, `jira_get_issue`, `jira_get_issue_comments`, `jira_list_projects`, `jira_update_issue`, `confluence_get_page`, `confluence_create_page`, `confluence_update_page`, `confluence_comment`, `confluence_search`, `request_pr_review`, `schedule_thread_wakeup`, `slack_add_reaction`, `slack_read_thread_messages`, `slack_thread_reply`.
Reviewer-only tools (in `agent/reviewer.py`): `add_finding`, `update_finding`, `list_findings`, `publish_review`. The review-style analyzer uses `save_review_style` (exported as `save_review_style_prompt`).
@ -105,7 +105,7 @@ Supported model IDs and per-model effort/reasoning rules live in `agent/dashboar
### Auth
- **GitHub**: dual-mode. User OAuth tokens are encrypted at rest in the dashboard OAuth store and cached only in process during a run (`utils/auth.py:resolve_github_token`, `utils/github_token.py`). When no user token is available, falls back to a GitHub App installation token (`utils/github_app.py`). The installation token is also what configures the LangSmith sandbox's GitHub proxy.
- **Webhooks**: GitHub signatures verified in `utils/github_comments.py:verify_github_signature`; Slack/Linear handled in their respective utils.
- **Webhooks**: GitHub signatures verified in `utils/github_comments.py:verify_github_signature`; Slack/Linear handled in their respective utils; Jira via a shared-secret header (`verify_jira_secret`, optional HMAC/timestamp); Confluence via Atlassian Connect JWT + `qsh` and RS256 signed-install (`utils/atlassian_connect.py`), with install secrets stored encrypted.
- **Dashboard / UI**: GitHub OAuth login lives in `agent/dashboard/oauth.py` and `routes.py` (`/auth/login`, `/auth/callback`, `/auth/logout`, `/me`).
### Thread-id derivation

View file

@ -4,7 +4,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
## Project
Open SWE is an open-source coding-agent framework built on **LangGraph** + **Deep Agents** (`deepagents.create_deep_agent`). It runs as a LangGraph app: each thread spawns its own isolated cloud sandbox, and the agent is invoked from Slack, Linear, or GitHub (PR comments, plus auto-review on opened / ready-for-review).
Open SWE is an open-source coding-agent framework built on **LangGraph** + **Deep Agents** (`deepagents.create_deep_agent`). It runs as a LangGraph app: each thread spawns its own isolated cloud sandbox, and the agent is invoked from Slack, Linear, Jira, Confluence, or GitHub (PR comments, plus auto-review on opened / ready-for-review).
A separate **reviewer** graph runs read-only code reviews on PRs, and a **review-style analyzer** graph learns per-repo review style from historical PRs.
@ -27,7 +27,7 @@ make format # ruff format + ruff check --fix
| Graph | Entrypoint | Purpose |
|---|---|---|
| `agent` | `agent.server:traced_agent` (wraps `get_agent`) | Main coding agent (Slack/Linear/GitHub-triggered). |
| `agent` | `agent.server:traced_agent` (wraps `get_agent`) | Main coding agent (Slack/Linear/Jira/Confluence/GitHub-triggered). |
| `reviewer` | `agent.reviewer:traced_reviewer_agent` (wraps `get_reviewer_agent`) | Read-only PR reviewer. Findings model + `publish_review`. |
| `analyzer` | `agent.analyzer:traced_analyzer` (wraps `get_analyzer`) | Learns per-repo reviewer style from historical PRs and this reviewer's own finding outcomes. |
@ -40,7 +40,8 @@ The FastAPI app is `agent.webapp:app`.
- **`agent/server.py` → `get_agent(config)`** — main graph factory. Called per-thread. Resolves the GitHub token, gets-or-creates the sandbox for the thread, resolves the team/profile/per-thread model + effort, then constructs a fresh `create_deep_agent(...)` with the curated tool list and middleware stack. The agent itself is stateless — all per-thread state lives in the sandbox + thread metadata.
- **`agent/reviewer.py` → `get_reviewer_agent(config)`** — reviewer graph factory. Shares `ensure_sandbox_for_thread` with the main agent but wires a reviewer-only toolset (`add_finding`, `update_finding`, `list_findings`, `publish_review`, `web_search`, `fetch_url`, `http_request`) and a different system prompt that pins the single-evolving-findings model and the diff-anchored bar for filing a finding. Read-only: no commit/push/PR-opening tools.
- **`agent/analyzer.py` → `get_analyzer(config)`** — small graph that emits a per-repo style prompt via the `save_review_style_prompt` tool, consumed by the reviewer as a "repository-specific review style" appendix. It runs in one of two modes (`analyzer_mode` in `configurable`): **bootstrap** (cold-start: crawl historical PR reviews) and **continual** (nightly: refine using this reviewer's own finding outcomes via `read_finding_outcomes`). Each mode's procedure lives in a deepagents **skill** (`agent/skills/bootstrap-repo-analysis/`, `agent/skills/continual-learning/`) served as virtual files via a `CompositeBackend` `/skills/` route + `StateBackend` (seeded into the run's `files` channel by the launcher — never written to the sandbox). Launchers and the per-repo nightly cron live in `agent/dashboard/review_style_jobs.py` and `agent/dashboard/analyzer_cron.py`; the cron is registered when bootstrap completes.
- **`agent/webapp.py`** — thin FastAPI routing layer mounted alongside the LangGraph server. Defines the webhook routes (GitHub, Linear, Slack) plus `/webhooks/run-complete`, and keeps the shared helpers/constants; the per-source handlers live in **`agent/webhooks/{github,slack,linear}.py`** (re-exported from `webapp` so existing call sites and tests keep working). Each webhook resolves a deterministic `thread_id` (so follow-up messages route to the same agent run) and triggers a run through the single durable dispatch contract in **`agent/dispatch.py`** (`dispatch_agent_run`: `multitask_strategy="interrupt"` + `durability="sync"` + completion webhook); `agent/completion.py` posts a failure reply if a run dies, and `agent/reconcile.py` (a `scheduler`-graph sweep) catches stragglers. The GitHub handler also auto-reviews PRs on `opened` / `ready_for_review` and drives the CI auto-fix flow (`agent/ci_autofix.py`).
- **`agent/webapp.py`** — thin FastAPI routing layer mounted alongside the LangGraph server. Defines the webhook routes (GitHub, Linear, Slack, Jira, Confluence Connect `/connect/*`) plus `/webhooks/run-complete`, and keeps the shared helpers/constants; the per-source handlers live in **`agent/webhooks/{github,slack,linear,jira,confluence}.py`** (re-exported from `webapp` so existing call sites and tests keep working). Each webhook resolves a deterministic `thread_id` (so follow-up messages route to the same agent run) and triggers a run through the single durable dispatch contract in **`agent/dispatch.py`** (`dispatch_agent_run`: `multitask_strategy="interrupt"` + `durability="sync"` + completion webhook); `agent/completion.py` posts a failure reply if a run dies, and `agent/reconcile.py` (a `scheduler`-graph sweep) catches stragglers. The GitHub handler also auto-reviews PRs on `opened` / `ready_for_review` and drives the CI auto-fix flow (`agent/ci_autofix.py`).
- **Atlassian triggers.** The Jira trigger is a Jira **Automation** rule POSTing to `/webhooks/jira` with a shared-secret header (`verify_jira_secret`; optional HMAC-body+timestamp via `JIRA_WEBHOOK_REQUIRE_SIGNATURE`), since Jira Cloud has no native webhook signing. The Confluence trigger is a private **Atlassian Connect app** (`agent/utils/atlassian_connect.py`): the `comment_created` webhook is HS256-JWT-verified against the per-tenant stored `sharedSecret` with a hand-rolled `qsh` (query-string-hash) check; `signed-install` is on, so install/uninstall lifecycle callbacks are RS256-verified against Atlassian's published keys (no trust-on-first-use). Install secrets are stored **encrypted** in the LangGraph store, keyed by `clientKey`. Both Atlassian webhook bodies are treated as pointers only — the triggering comment's real author/text is re-fetched server-side via the Basic-auth service account before anything security-relevant is derived, and attribution is gated on an active user mapping (mirrors the GitHub-login / token-attribution flow). Descriptor served at `GET /connect/atlassian-connect.json`.
- **`agent/dashboard/`** — `router` mounted under the FastAPI app at startup (`app.include_router(dashboard_router)`). Owns GitHub OAuth, per-user profiles, admin endpoints, team defaults, enabled-repo lists, review-style management, and the Agents chat thread API used by the UI in `ui/`.
### Sandbox lifecycle (the tricky part)
@ -81,7 +82,9 @@ There is intentionally no after-agent safety net that opens a PR for the agent.
All tools live in `agent/tools/` and are flat-imported via `agent/tools/__init__.py`. The set is intentionally small and curated — see README "Tools — Curated, Not Accumulated".
Wired into `get_agent`:
`http_request`, `fetch_url`, `web_search`, `linear_comment`, `linear_create_issue`, `linear_delete_issue`, `linear_get_issue`, `linear_get_issue_comments`, `linear_list_teams`, `linear_update_issue`, `request_pr_review`, `schedule_thread_wakeup`, `slack_add_reaction`, `slack_read_thread_messages`, `slack_thread_reply`.
`http_request`, `fetch_url`, `web_search`, `linear_comment`, `linear_create_issue`, `linear_delete_issue`, `linear_get_issue`, `linear_get_issue_comments`, `linear_list_teams`, `linear_update_issue`, `jira_comment`, `jira_create_issue`, `jira_get_issue`, `jira_get_issue_comments`, `jira_list_projects`, `jira_update_issue`, `confluence_get_page`, `confluence_create_page`, `confluence_update_page`, `confluence_comment`, `confluence_search`, `request_pr_review`, `schedule_thread_wakeup`, `slack_add_reaction`, `slack_read_thread_messages`, `slack_thread_reply`.
Jira uses a service-account REST client (`agent/utils/jira.py`, Basic auth) with ADF↔markdown conversion (`agent/utils/adf.py`); Confluence likewise (`agent/utils/confluence.py`, XHTML storage-format). Both are dark-safe: unset env returns a clean error.
Reviewer-only tools (in `agent/reviewer.py`): `add_finding`, `update_finding`, `list_findings`, `publish_review`. The review-style analyzer uses `save_review_style` (exported as `save_review_style_prompt`).

View file

@ -223,6 +223,8 @@ Open SWE ships with a small set of custom tools on top of the built-in Deep Agen
| `fetch_url` | `agent/tools/fetch_url.py` | Fetch web pages as markdown |
| `http_request` | `agent/tools/http_request.py` | HTTP API calls |
| `linear_comment` | `agent/tools/linear_comment.py` | Post comments on Linear tickets |
| `jira_comment`, `jira_get_issue`, … | `agent/tools/jira_*.py` | Read/comment/create/update Jira issues (`agent/utils/jira.py`) |
| `confluence_get_page`, `confluence_update_page`, … | `agent/tools/confluence_*.py` | Read/write Confluence pages + comments (`agent/utils/confluence.py`) |
| `slack_thread_reply` | `agent/tools/slack_thread_reply.py` | Reply in Slack threads |
### Adding a tool
@ -322,7 +324,7 @@ These are used as the fallback when:
### Repository extraction from messages
Both Slack and Linear support specifying a target repo directly in the message or comment text. The shared utility `extract_repo_from_text()` in `agent/utils/repo.py` handles parsing these formats:
Slack, Linear, Jira, and Confluence all support specifying a target repo directly in the message or comment text. The shared utility `extract_repo_from_text()` in `agent/utils/repo.py` handles parsing these formats:
- `repo:owner/name` — explicit org and repo
- `repo owner/name` — space syntax (same result)

View file

@ -36,7 +36,7 @@ You'll need the ngrok URL in subsequent steps when configuring webhooks, so star
ngrok http 2024 --url https://some-url-you-configure.ngrok.dev
```
You don't need to pass the `--url` flag, however doing so will use the same subdomain each time you startup the server. Without this, you'll need to update the webhook URL in GitHub, Slack and Linear every time you restart your server for local development.
You don't need to pass the `--url` flag, however doing so will use the same subdomain each time you startup the server. Without this, you'll need to update the webhook URL in GitHub, Slack, Linear, Jira, and Confluence (Connect `CONNECT_BASE_URL`) every time you restart your server for local development.
Copy the HTTPS URL you set, or if you didn't pass `--url`, the one ngrok gives you. You'll paste this into the webhook settings in steps 3 and 5.
@ -224,7 +224,7 @@ REPO_SNAPSHOT_BASE_IMAGE="<your-docker-hub>/<name-of-your-image>"
## 5. Set up triggers
Open SWE can be triggered from GitHub, Linear, and/or Slack. **Configure whichever surfaces your team uses — you don't need all of them.**
Open SWE can be triggered from GitHub, Linear, Jira, Confluence, and/or Slack. **Configure whichever surfaces your team uses — you don't need all of them.**
### GitHub
@ -248,7 +248,7 @@ ALLOWED_GITHUB_ORGS="langchain-ai,anthropics"
ALLOWED_GITHUB_REPOS="some-user/their-repo,another-org/specific-repo"
```
A GitHub or Linear webhook is accepted if the resolved repo's org is in `ALLOWED_GITHUB_ORGS` **or** the `owner/repo` is in `ALLOWED_GITHUB_REPOS`. If both are empty, all repos are allowed. Slack mentions are not rejected from regex-inferred repository text; repository access is bounded by the GitHub App installation permissions.
A GitHub, Linear, Jira, or Confluence webhook is accepted if the resolved repo's org is in `ALLOWED_GITHUB_ORGS` **or** the `owner/repo` is in `ALLOWED_GITHUB_REPOS`. If both are empty, all repos are allowed by default — set `REQUIRE_REPO_ALLOWLIST=true` to fail closed instead (recommended in production). Slack mentions are not rejected from regex-inferred repository text; repository access is bounded by the GitHub App installation permissions.
`ALLOWED_GITHUB_ORGS` also gates **dashboard login**: when set, only GitHub accounts that are active members of one of the listed organizations can complete the OAuth login and receive a session. Membership is verified server-side with the GitHub App installation token (so private memberships are visible and no extra OAuth scope is required), and the check fails closed on any API error. When `ALLOWED_GITHUB_ORGS` is empty, dashboard login is open to any GitHub account (the prior behavior).
@ -402,6 +402,71 @@ The dashboard can let a user link their Slack identity to their GitHub login via
If `SLACK_CLIENT_ID`/`SLACK_CLIENT_SECRET` are unset, the "Sign in with Slack" link is simply disabled; the rest of Slack triggering still works.
### Jira (optional)
Open SWE listens for Jira issue comments that mention `@openswe`. It also exposes Jira tools (`jira_get_issue`, `jira_comment`, …) to the agent.
**Set up the service account** (used for the tools and to re-fetch the triggering comment server-side):
1. Create/choose a Jira service account and generate an API token at [id.atlassian.com/manage-profile/security/api-tokens](https://id.atlassian.com/manage-profile/security/api-tokens).
2. Set `JIRA_BASE_URL` (e.g. `https://your-site.atlassian.net`), `JIRA_SERVICE_EMAIL`, and `JIRA_API_TOKEN`.
**Create the trigger.** Jira Cloud has no natively-signed outgoing webhook, so Open SWE is fronted by a Jira **Automation** rule with a shared-secret header:
1. Generate a secret with `openssl rand -hex 32` and save it as `JIRA_WEBHOOK_SECRET`.
2. In **Project settings → Automation → Create rule**:
- **Trigger**: *Issue commented*.
- **Action**: *Send web request* → **URL** `https://<your-ngrok-url>/webhooks/jira`, **Method** POST, **Header** `X-Automation-Webhook-Token: <JIRA_WEBHOOK_SECRET>`, and a **custom JSON body** built from smart values:
```json
{
"issue_key": "{{issue.key}}",
"comment_id": "{{comment.id}}",
"comment_author_is_bot": false
}
```
`issue_key` and `comment_id` are the only fields trusted from the body, and only as a pointer — the triggering comment's real author and text are re-fetched from Jira server-side before anything security-relevant is derived.
**Optional stronger trust:** set `JIRA_WEBHOOK_REQUIRE_SIGNATURE=true` to also require an `X-Openswe-Signature` (hex HMAC-SHA256 of the raw body keyed by `JIRA_WEBHOOK_SECRET`) plus a fresh `timestamp` in the body (closes the static-token replay gap), and/or `JIRA_WEBHOOK_IP_ALLOWLIST` to restrict the source IP.
**Configure project-to-repo mapping** in `agent/utils/jira_project_repo_map.py`:
```python
JIRA_PROJECT_TO_REPO = {
"PROJ": {"owner": "my-org", "name": "my-repo"},
}
```
As with Linear, a user can override per-comment with `repo:owner/name` in the `@openswe` comment; the mapping is the fallback.
### Confluence (optional)
Open SWE listens for Confluence page comments that mention `@openswe`, via a private **Atlassian Connect** app, and exposes Confluence tools (`confluence_get_page`, `confluence_update_page`, …).
**Set up the service account** (tools + server-side comment re-fetch): set `CONFLUENCE_BASE_URL`, `CONFLUENCE_EMAIL`, `CONFLUENCE_API_TOKEN` (same Atlassian API-token flow as Jira).
**Prerequisites for the Connect app:**
- `CONNECT_BASE_URL` — the app's public origin with an **empty context path** (e.g. `https://<your-ngrok-url>`); the descriptor `baseUrl` and the JWT `aud` are derived from it.
- `TOKEN_ENCRYPTION_KEY` — install `sharedSecret`s are encrypted at rest with it (see §6). **In production the LangGraph store must be the durable Postgres store**, or installations are lost on restart and every webhook 401s until reinstall.
**Install the app:**
1. In Confluence, go to **Settings → Apps → Manage apps**, enable **Development mode**, then **Upload app** and give it the descriptor URL: `https://<your-ngrok-url>/connect/atlassian-connect.json`.
2. **Tenant binding (required).** The descriptor is public, so `CONNECT_EXPECTED_CLIENT_KEYS` is a fail-closed allowlist of the Confluence tenant(s) allowed to install — without it, *any* Atlassian tenant could install the app and trigger runs. Bootstrap it: the first install is **rejected** and the backend logs the `clientKey`; add that value to `CONNECT_EXPECTED_CLIENT_KEYS` (comma/space-separated) and re-install.
3. (Optional) `CONNECT_EXPECTED_BASE_URL` pins the install baseUrl host (defense-in-depth), and `CONFLUENCE_BOT_ACCOUNT_ID` suppresses self-triggering on the app's own comments.
Lifecycle callbacks (install/uninstall) are RS256-verified against Atlassian's published keys, and the `comment_created` webhook is HS256-JWT-verified against the stored per-tenant secret.
**Configure space-to-repo mapping** in `agent/utils/confluence_space_repo_map.py`:
```python
CONFLUENCE_SPACE_TO_REPO = {
"IT": {"owner": "my-org", "name": "my-repo"},
}
```
## 6. Environment variables
Create a `.env` file in the project root. Below is the full list — only fill in the sections relevant to the triggers you configured.
@ -488,6 +553,27 @@ LANGGRAPH_URL="http://localhost:2024"
LINEAR_API_KEY="" # From step 5
LINEAR_WEBHOOK_SECRET="" # From step 5
# === Jira (if using Jira tools / trigger) ===
JIRA_BASE_URL="" # e.g. "https://your-site.atlassian.net"
JIRA_SERVICE_EMAIL="" # Service-account email (Basic auth)
JIRA_API_TOKEN="" # Service-account API token
JIRA_WEBHOOK_SECRET="" # Shared secret on the Automation rule's X-Automation-Webhook-Token header
JIRA_WEBHOOK_REQUIRE_SIGNATURE="" # "true" to also require an HMAC body signature + fresh timestamp
JIRA_WEBHOOK_IP_ALLOWLIST="" # optional CIDR allowlist for the direct client IP
# === Confluence (if using Confluence tools / trigger) ===
CONFLUENCE_BASE_URL="" # e.g. "https://your-site.atlassian.net"
CONFLUENCE_EMAIL="" # Service-account email
CONFLUENCE_API_TOKEN="" # Service-account API token
CONNECT_BASE_URL="" # Public origin of the Connect app (empty context path), e.g. "https://<ngrok>"
CONNECT_EXPECTED_CLIENT_KEYS="" # REQUIRED tenant allowlist — clientKey(s) allowed to install (fail closed; bootstrap from logs)
CONNECT_EXPECTED_BASE_URL="" # optional defense-in-depth: pin the install baseUrl host
CONFLUENCE_BOT_ACCOUNT_ID="" # optional: the app's own accountId, to suppress self-triggering
# NOTE: TOKEN_ENCRYPTION_KEY (below) is required to store Connect install secrets, and prod must use the durable Postgres store.
# === Fail-closed repo allowlist (optional, recommended in prod) ===
REQUIRE_REPO_ALLOWLIST="" # "true" => empty ALLOWED_GITHUB_ORGS/REPOS rejects all repos instead of allowing all
# === Slack (if using Slack trigger) ===
SLACK_BOT_TOKEN="" # From step 5
SLACK_BOT_USER_ID=""

View file

@ -25,7 +25,7 @@
Elite engineering orgs like Stripe, Ramp, and Coinbase are building their own internal coding agents — Slackbots, CLIs, and web apps that meet engineers where they already work. These agents are connected to internal systems with the right context, permissioning, and safety boundaries to operate with minimal human oversight.
Open SWE is the open-source version of this pattern. Built on [LangGraph](https://langchain-ai.github.io/langgraph/) and [Deep Agents](https://github.com/langchain-ai/deepagents), it gives you the same architecture those companies built internally: cloud sandboxes, Slack and Linear invocation, subagent orchestration, and automatic PR creation — ready to customize for your own codebase and workflows.
Open SWE is the open-source version of this pattern. Built on [LangGraph](https://langchain-ai.github.io/langgraph/) and [Deep Agents](https://github.com/langchain-ai/deepagents), it gives you the same architecture those companies built internally: cloud sandboxes, Slack / Linear / Jira / Confluence / GitHub invocation, subagent orchestration, and automatic PR creation — ready to customize for your own codebase and workflows.
> [!NOTE]
> Read the **announcement blog post [here](https://blog.langchain.com/open-swe-an-open-source-framework-for-internal-coding-agents/)**
@ -72,6 +72,8 @@ Stripe's key insight: *tool curation matters more than tool quantity.* Open SWE
| `fetch_url` | Fetch web pages as markdown |
| `http_request` | API calls (GET, POST, etc.) |
| `linear_comment` | Post updates to Linear tickets |
| `jira_*` | Read/comment/create/update Jira issues |
| `confluence_*` | Read/write Confluence pages + comments |
| `slack_add_reaction` | React to Slack messages |
| `slack_thread_reply` | Reply in Slack threads |
@ -100,14 +102,18 @@ Open SWE's orchestration has two layers:
- **`notify_step_limit_reached`** — After-agent hook that posts a Slack reply when the agent hits the model-call limit, so users get a clear signal instead of silence.
- **`ToolErrorMiddleware`** — Catches and handles tool errors gracefully.
### 6. Invocation — Slack, Linear, and GitHub
### 6. Invocation — Slack, Linear, Jira, Confluence, and GitHub
All three companies in the article converge on **Slack as the primary invocation surface**. Open SWE does the same:
- **Slack** — Mention the bot in any thread. Supports `repo:owner/name` syntax to specify which repo to work on. The agent replies in-thread with status updates and PR links.
- **Linear** — Comment `@openswe` on any issue. The agent reacts with 👀 to acknowledge, reads the full issue context, and posts results back as comments.
- **Jira** — Comment `@openswe` on any issue (fronted by a Jira Automation rule → `/webhooks/jira`). The agent reads the issue and posts results back as a comment.
- **Confluence** — Comment `@openswe` on a page. A private Atlassian Connect app delivers the `comment_created` event; the agent acts and replies on the page.
- **GitHub** — Tag `@openswe` in PR comments on agent-created PRs to have it address review feedback and push fixes to the same branch.
See **[INSTALLATION.md](./INSTALLATION.md) §5** for per-surface trigger setup.
Each invocation creates a deterministic thread ID, so follow-up messages on the same issue or thread route to the same running agent.
**Trigger tags (Sea Haven fork):** a mention is a case-insensitive substring match on the comment body — `@openswe`, `@open-swe`, `@openswe-dev`, or `@seahaven-openswe` (the deployed App slug). GitHub won't linkify `@seahaven-openswe` (App `[bot]` accounts aren't user-mentionable), but the text still fires a run.

View file

@ -23,9 +23,11 @@ import os
from collections.abc import Awaitable, Callable
from typing import Any
from .utils.confluence import add_comment as add_confluence_comment
from .utils.dashboard_links import dashboard_thread_url
from .utils.github_app import get_github_app_installation_token
from .utils.github_comments import post_github_comment
from .utils.jira import comment_on_issue as comment_on_jira_issue
from .utils.linear import comment_on_linear_issue
from .utils.slack import post_slack_thread_reply
from .utils.thread_ops import langgraph_client
@ -124,6 +126,25 @@ async def _post_failure_reply(
return await comment_on_linear_issue(issue_id, text)
return False
if source == "jira":
jira_issue = ctx.get("jira_issue")
if isinstance(jira_issue, dict):
issue_key = jira_issue.get("key")
if issue_key:
await claim()
return await comment_on_jira_issue(issue_key, text)
return False
if source == "confluence":
confluence = ctx.get("confluence")
if isinstance(confluence, dict):
page_id = confluence.get("page_id")
if page_id:
await claim()
result = await add_confluence_comment(page_id, text)
return bool(result.get("success"))
return False
if source in ("github", "github_issue"):
repo_config = metadata.get("repo")
number = ctx.get("pr_number")

View file

@ -58,7 +58,7 @@ def _load_default_prompt() -> str:
# deepagents' generic base prompt so there is a single Open SWE voice. The
# per-thread, main-agent-specific prompt (working dir, repo setup, PR workflow,
# source-channel reply) is layered in front of this via `construct_system_prompt`.
OPEN_SWE_SHARED_BASE = """You are **Open SWE**, an open-source agent built on LangGraph and Deep Agents, operating in a remote, git-backed Linux sandbox invoked from Slack, Linear, or GitHub.
OPEN_SWE_SHARED_BASE = """You are **Open SWE**, an open-source agent built on LangGraph and Deep Agents, operating in a remote, git-backed Linux sandbox invoked from Slack, Linear, Jira, Confluence, or GitHub.
### Core Behavior
@ -266,7 +266,7 @@ Steps, in order:
```
feat: add retry logic for transient upstream failures [<KEY>]
```
With no resolvable key, drop the suffix: `feat: add retry logic for transient upstream failures`. Resolve the key from the Linear-triggered run when present (`{linear_project_id}-{linear_issue_number}`), or from a Linear ticket referenced in the Slack thread / task context.
With no resolvable key, drop the suffix: `feat: add retry logic for transient upstream failures`. Resolve the key from the Linear-triggered run when present (`{linear_project_id}-{linear_issue_number}`), from the Jira-triggered run when present (`{jira_project_key}-{jira_issue_number}`), or from a Linear/Jira ticket referenced in the Slack thread / task context.
**PR Body** — use this structure. Omit a section only when it would be empty:
```
@ -315,6 +315,8 @@ Steps, in order:
4. **Notify the source** immediately after pushing and, when applicable, PR creation/update succeeds. Include a brief summary plus the PR link or branch URL:
- Linear-triggered: use `linear_comment` with an `@mention` of the user who triggered the task
- Jira-triggered: use `jira_comment` with an `@mention` of the user who triggered the task
- Confluence-triggered: use `confluence_comment` on the triggering page (the page id is in the task context)
- Slack-triggered: use `slack_thread_reply`
- GitHub-triggered: use `GH_TOKEN=dummy gh issue comment` or `GH_TOKEN=dummy gh pr comment`
- If the task was not triggered from a known source channel (no Slack thread, no Linear ticket, no GitHub issue context), skip the notification step.
@ -421,6 +423,8 @@ def construct_system_prompt(
working_dir: str,
linear_project_id: str = "",
linear_issue_number: str = "",
jira_project_key: str = "",
jira_issue_number: str = "",
triggering_user_identity: CollaboratorIdentity | None = None,
create_prs: bool = False,
default_repo: dict[str, str] | None = None,
@ -450,6 +454,8 @@ def construct_system_prompt(
working_dir=working_dir,
linear_project_id=linear_project_id or "<PROJECT_ID>",
linear_issue_number=linear_issue_number or "<ISSUE_NUMBER>",
jira_project_key=jira_project_key or "<JIRA_PROJECT_KEY>",
jira_issue_number=jira_issue_number or "<ISSUE_NUMBER>",
plan_review_url=plan_url or "(the dashboard plan-review page)",
plan_mode_section=(
PLAN_MODE_SECTION.format(plan_url=plan_url or "(plan-review link unavailable)")

View file

@ -85,9 +85,20 @@ from .middleware import (
)
from .prompt import construct_system_prompt
from .tools import (
confluence_comment,
confluence_create_page,
confluence_get_page,
confluence_search,
confluence_update_page,
enter_plan_mode,
fetch_url,
http_request,
jira_comment,
jira_create_issue,
jira_get_issue,
jira_get_issue_comments,
jira_list_projects,
jira_update_issue,
linear_comment,
linear_create_issue,
linear_delete_issue,
@ -816,6 +827,10 @@ async def get_agent(config: RunnableConfig) -> Pregel:
linear_project_id = linear_issue.get("linear_project_id", "")
linear_issue_number = linear_issue.get("linear_issue_number", "")
jira_issue = config["configurable"].get("jira_issue", {})
jira_project_key = jira_issue.get("project_key", "")
jira_issue_number = jira_issue.get("issue_number", "")
work_dir = await aresolve_sandbox_work_dir(sandbox_backend)
def backend_factory(_runtime: object, _thread_id: str = thread_id) -> SandboxBackendProtocol:
@ -980,6 +995,8 @@ async def get_agent(config: RunnableConfig) -> Pregel:
working_dir=work_dir,
linear_project_id=linear_project_id,
linear_issue_number=linear_issue_number,
jira_project_key=jira_project_key,
jira_issue_number=jira_issue_number,
triggering_user_identity=triggering_user_identity,
create_prs=always_create_prs,
default_repo=prompt_default_repo,
@ -1003,6 +1020,17 @@ async def get_agent(config: RunnableConfig) -> Pregel:
linear_get_issue_comments,
linear_list_teams,
linear_update_issue,
jira_comment,
jira_create_issue,
jira_get_issue,
jira_get_issue_comments,
jira_list_projects,
jira_update_issue,
confluence_get_page,
confluence_create_page,
confluence_update_page,
confluence_comment,
confluence_search,
open_pull_request,
request_pr_review,
report_platform_issue,

View file

@ -4,9 +4,20 @@ from typing import TYPE_CHECKING, Any
_TOOL_MODULES = {
"add_finding": ".add_finding",
"confluence_comment": ".confluence_comment",
"confluence_create_page": ".confluence_create_page",
"confluence_get_page": ".confluence_get_page",
"confluence_search": ".confluence_search",
"confluence_update_page": ".confluence_update_page",
"enter_plan_mode": ".enter_plan_mode",
"fetch_url": ".fetch_url",
"http_request": ".http_request",
"jira_comment": ".jira_comment",
"jira_create_issue": ".jira_create_issue",
"jira_get_issue": ".jira_get_issue",
"jira_get_issue_comments": ".jira_get_issue_comments",
"jira_list_projects": ".jira_list_projects",
"jira_update_issue": ".jira_update_issue",
"linear_comment": ".linear_comment",
"linear_create_issue": ".linear_create_issue",
"linear_delete_issue": ".linear_delete_issue",
@ -36,9 +47,20 @@ _TOOL_MODULES = {
__all__ = [
"add_finding",
"confluence_comment",
"confluence_create_page",
"confluence_get_page",
"confluence_search",
"confluence_update_page",
"enter_plan_mode",
"fetch_url",
"http_request",
"jira_comment",
"jira_create_issue",
"jira_get_issue",
"jira_get_issue_comments",
"jira_list_projects",
"jira_update_issue",
"linear_comment",
"linear_create_issue",
"linear_delete_issue",
@ -68,9 +90,20 @@ __all__ = [
if TYPE_CHECKING:
from .add_finding import add_finding
from .confluence_comment import confluence_comment
from .confluence_create_page import confluence_create_page
from .confluence_get_page import confluence_get_page
from .confluence_search import confluence_search
from .confluence_update_page import confluence_update_page
from .enter_plan_mode import enter_plan_mode
from .fetch_url import fetch_url
from .http_request import http_request
from .jira_comment import jira_comment
from .jira_create_issue import jira_create_issue
from .jira_get_issue import jira_get_issue
from .jira_get_issue_comments import jira_get_issue_comments
from .jira_list_projects import jira_list_projects
from .jira_update_issue import jira_update_issue
from .linear_comment import linear_comment
from .linear_create_issue import linear_create_issue
from .linear_delete_issue import linear_delete_issue

View file

@ -0,0 +1,16 @@
from typing import Any
from ..utils.confluence import add_comment
async def confluence_comment(page_id: str, comment_body: str) -> dict[str, Any]:
"""Post a comment to a Confluence page.
Args:
page_id: The Confluence page id to comment on.
comment_body: Plain-text comment text to post.
Returns:
Dictionary with 'success' (bool) key.
"""
return await add_comment(page_id, comment_body)

View file

@ -0,0 +1,23 @@
from typing import Any
from ..utils.confluence import create_page
async def confluence_create_page(
space_key: str,
title: str,
body: str,
parent_id: str | None = None,
) -> dict[str, Any]:
"""Create a new Confluence page.
Args:
space_key: The Confluence space key to create the page in (e.g. IT).
title: The page title.
body: Plain-text page body (blank-line-separated blocks become paragraphs).
parent_id: Optional parent page id to nest this page under.
Returns:
Dictionary with 'success' (bool) and 'page' (id, title, url), or 'error'.
"""
return await create_page(space_key, title, body, parent_id=parent_id)

View file

@ -0,0 +1,15 @@
from typing import Any
from ..utils.confluence import get_page
async def confluence_get_page(page_id: str) -> dict[str, Any]:
"""Get a Confluence page by its id.
Args:
page_id: The Confluence page id.
Returns:
Dictionary with 'page' containing the normalized page (plain-text body).
"""
return await get_page(page_id)

View file

@ -0,0 +1,15 @@
from typing import Any
from ..utils.confluence import search
async def confluence_search(cql: str) -> dict[str, Any]:
"""Search Confluence content using CQL (Confluence Query Language).
Args:
cql: A CQL query string (e.g. 'space = "IT" AND title ~ "Architecture"').
Returns:
Dictionary with 'results' (list of {id, title, type, url}).
"""
return await search(cql)

View file

@ -0,0 +1,25 @@
from typing import Any
from ..utils.confluence import update_page
async def confluence_update_page(
page_id: str,
title: str | None = None,
body: str | None = None,
) -> dict[str, Any]:
"""Update an existing Confluence page.
Use this tool to keep architecture/documentation pages in sync with code
changes. Confluence versions every edit, so this reads the current page
first and bumps the version number automatically.
Args:
page_id: The Confluence page id to update.
title: Optional new title (existing title is reused if not provided).
body: Optional new plain-text page body (replaces the existing body).
Returns:
Dictionary with 'success' (bool) and 'page' (id, title, url), or 'error'.
"""
return await update_page(page_id, title=title, body=body)

View file

@ -0,0 +1,25 @@
from typing import Any
from ..utils.jira import comment_on_issue
async def jira_comment(comment_body: str, issue_key: str) -> dict[str, Any]:
"""Post a comment to a Jira issue.
Use this tool to communicate progress and completion to stakeholders on Jira.
**When to use:**
- After opening/updating a draft PR, post a comment on the Jira issue to let
stakeholders know the task is complete and include the PR link. For example:
"I've completed the implementation and opened a PR: <pr_url>"
- When answering a question or sharing an update (no code changes needed).
Args:
comment_body: Markdown-formatted comment text to post to the Jira issue.
issue_key: The Jira issue key to post the comment to (e.g. PROJ-123).
Returns:
Dictionary with 'success' (bool) key.
"""
success = await comment_on_issue(issue_key, comment_body)
return {"success": success}

View file

@ -0,0 +1,34 @@
from typing import Any
from ..utils.jira import create_issue
async def jira_create_issue(
project_key: str,
summary: str,
description: str | None = None,
issue_type: str = "Task",
priority: str | None = None,
labels: list[str] | None = None,
) -> dict[str, Any]:
"""Create a new Jira issue.
Args:
project_key: The Jira project key to create the issue in (e.g. PROJ).
summary: The issue title/summary.
description: Optional markdown-formatted issue description.
issue_type: The issue type name (default "Task").
priority: Optional priority name (e.g. "High").
labels: Optional list of label strings.
Returns:
Dictionary with 'success' (bool) and 'issue' (key, id, url).
"""
return await create_issue(
project_key,
summary,
description=description,
issue_type=issue_type,
priority=priority,
labels=labels,
)

View file

@ -0,0 +1,15 @@
from typing import Any
from ..utils.jira import get_issue
async def jira_get_issue(issue_key: str) -> dict[str, Any]:
"""Get a Jira issue by its key (e.g. PROJ-123).
Args:
issue_key: The Jira issue key (e.g. PROJ-123).
Returns:
Dictionary with 'issue' containing the normalized issue (markdown body).
"""
return await get_issue(issue_key)

View file

@ -0,0 +1,15 @@
from typing import Any
from ..utils.jira import get_issue_comments
async def jira_get_issue_comments(issue_key: str) -> dict[str, Any]:
"""Get comments for a Jira issue.
Args:
issue_key: The Jira issue key (e.g. PROJ-123).
Returns:
Dictionary with 'comments' (list) each containing a markdown body.
"""
return await get_issue_comments(issue_key)

View file

@ -0,0 +1,12 @@
from typing import Any
from ..utils.jira import list_projects
async def jira_list_projects() -> dict[str, Any]:
"""List Jira projects visible to the agent's service account.
Returns:
Dictionary with 'projects' (list of {key, name, id}).
"""
return await list_projects()

View file

@ -0,0 +1,31 @@
from typing import Any
from ..utils.jira import update_issue
async def jira_update_issue(
issue_key: str,
summary: str | None = None,
description: str | None = None,
priority: str | None = None,
labels: list[str] | None = None,
) -> dict[str, Any]:
"""Update an existing Jira issue.
Args:
issue_key: The Jira issue key to update (e.g. PROJ-123).
summary: Optional new summary/title.
description: Optional new markdown-formatted description.
priority: Optional priority name (e.g. "High").
labels: Optional list of label strings (replaces existing labels).
Returns:
Dictionary with 'success' (bool) and 'issue' (key, url), or 'error'.
"""
return await update_issue(
issue_key,
summary=summary,
description=description,
priority=priority,
labels=labels,
)

117
agent/utils/adf.py Normal file
View file

@ -0,0 +1,117 @@
"""Minimal Atlassian Document Format (ADF) <-> markdown conversion.
Jira Cloud v3 issue descriptions and comments are ADF JSON, not markdown. These
helpers convert the node types Atlassian actually emits in descriptions/comments
so bodies read as markdown in prompts, and produce valid ADF for agent-authored
comments (prose). ``markdown_to_adf`` is intentionally minimal: agent comments
are plain prose, so each blank-line-separated block becomes one paragraph.
"""
from __future__ import annotations
from typing import Any
_MARK_WRAP = {
"strong": "**",
"em": "_",
"code": "`",
"strikethrough": "~~",
}
def _apply_marks(text: str, marks: list[dict[str, Any]]) -> str:
for mark in marks:
mtype = mark.get("type")
if mtype == "link":
href = (mark.get("attrs") or {}).get("href", "")
text = f"[{text}]({href})"
elif mtype in _MARK_WRAP:
wrap = _MARK_WRAP[mtype]
text = f"{wrap}{text}{wrap}"
return text
def _render_nodes(nodes: list[dict[str, Any]]) -> str:
return "".join(_render_node(n) for n in nodes)
def _render_node(node: dict[str, Any]) -> str: # noqa: PLR0911, PLR0912
ntype = node.get("type")
content = node.get("content", []) or []
attrs = node.get("attrs", {}) or {}
if ntype == "text":
return _apply_marks(node.get("text", ""), node.get("marks", []) or [])
if ntype == "hardBreak":
return "\n"
if ntype == "paragraph":
return _render_nodes(content) + "\n\n"
if ntype == "heading":
level = min(int(attrs.get("level", 1)), 6)
return f"{'#' * level} {_render_nodes(content)}\n\n"
if ntype == "blockquote":
inner = _render_nodes(content).strip()
return "".join(f"> {line}\n" for line in inner.splitlines()) + "\n"
if ntype == "codeBlock":
lang = attrs.get("language", "")
return f"```{lang}\n{_render_nodes(content)}\n```\n\n"
if ntype == "rule":
return "---\n\n"
if ntype == "bulletList":
return (
"".join(f"- {_render_nodes(li.get('content', [])).strip()}\n" for li in content) + "\n"
)
if ntype == "orderedList":
out = []
for i, li in enumerate(content, start=1):
out.append(f"{i}. {_render_nodes(li.get('content', [])).strip()}\n")
return "".join(out) + "\n"
if ntype == "listItem":
return _render_nodes(content)
if ntype in ("mediaSingle", "mediaGroup"):
return _render_nodes(content)
if ntype == "media":
alt = attrs.get("alt") or attrs.get("id", "media")
url = attrs.get("url", "")
return f"![{alt}]({url})\n\n" if url else f"[media: {alt}]\n\n"
if ntype == "inlineCard":
return (attrs.get("url", "")) or ""
if ntype == "mention":
return attrs.get("text", "") or ""
if ntype == "emoji":
return attrs.get("text", "") or attrs.get("shortName", "") or ""
# Unknown/container node: recurse into content.
return _render_nodes(content)
def adf_to_markdown(adf: Any) -> str:
"""Convert an ADF document (or None) to a markdown string."""
if not adf or not isinstance(adf, dict):
return ""
return _render_nodes(adf.get("content", []) or []).strip()
def markdown_to_adf(text: str) -> dict[str, Any]:
"""Convert plain markdown/prose to a minimal ADF document.
Blank-line-separated blocks become paragraphs; single newlines within a
block become hardBreaks. Inline markdown is left as literal text.
"""
text = text or ""
blocks = text.split("\n\n")
paragraphs: list[dict[str, Any]] = []
for block in blocks:
if not block.strip():
continue
lines = block.split("\n")
para_content: list[dict[str, Any]] = []
for idx, line in enumerate(lines):
if idx > 0:
para_content.append({"type": "hardBreak"})
if line:
para_content.append({"type": "text", "text": line})
paragraphs.append({"type": "paragraph", "content": para_content})
if not paragraphs:
paragraphs = [{"type": "paragraph", "content": []}]
return {"type": "doc", "version": 1, "content": paragraphs}

View file

@ -0,0 +1,395 @@
"""Atlassian Connect (Confluence) trust surface: qsh, JWT verify, secret store.
Private Connect app, symmetric (HS256) shared-secret flow. Atlassian mints a
``sharedSecret`` per install and signs each request with it; we verify the
signature, the ``exp``, and the ``qsh`` (query-string-hash) claim that binds a
token to one exact method+path+query (defeating cross-endpoint replay). The
shared secret is stored encrypted-at-rest in the LangGraph store, keyed by the
tenant ``clientKey``.
``qsh`` is Connect-specific (no JWT library provides it) so it is hand-rolled
here from the Atlassian spec and pinned to the official test vector in
``tests/test_atlassian_connect.py``. A subtle bug here is a silent auth bypass.
"""
from __future__ import annotations
import hashlib
import hmac
import logging
import os
import re
import time
from typing import Any
from urllib.parse import parse_qs, quote
import httpx
import jwt
from langgraph_sdk import get_client
from ..encryption import decrypt_token, encrypt_token
from .http import DEFAULT_HTTP_TIMEOUT
logger = logging.getLogger(__name__)
LANGGRAPH_URL = os.environ.get("LANGGRAPH_URL") or os.environ.get(
"LANGGRAPH_URL_PROD", "http://localhost:2024"
)
# The app's public origin — must equal the descriptor baseUrl and the `aud` in
# Atlassian's signed-install lifecycle JWTs.
CONNECT_BASE_URL = os.environ.get("CONNECT_BASE_URL", "").rstrip("/")
# Atlassian's CDN of public keys for signed-install (asymmetric) lifecycle JWTs.
_CONNECT_INSTALL_KEYS_BASE = "https://connect-install-keys.atlassian.com"
# Tenant binding (REQUIRED): the signature-verified clientKey(s) — i.e. the JWT
# `iss` — we accept installs from. signed-install proves the caller is *an*
# Atlassian tenant, not *ours*, and the descriptor is served publicly, so
# without this any tenant could install the app and drive runs. Empty => reject
# ALL installs (fail closed). The install `baseUrl` is an untrusted body field
# and is NOT a valid binding; only the signed `iss` is. Bootstrap: attempt an
# install, read the rejected clientKey from the logs, add it here, re-install.
CONNECT_EXPECTED_CLIENT_KEYS: frozenset[str] = frozenset(
key.strip()
for key in os.environ.get("CONNECT_EXPECTED_CLIENT_KEYS", "").replace(",", " ").split()
if key.strip()
)
# Optional defense-in-depth on the (untrusted) install baseUrl host. Enforced
# only when set — the clientKey allowlist above is the real tenant gate.
CONNECT_EXPECTED_BASE_URL_HOSTS: frozenset[str] = frozenset(
host.strip().lower()
for host in os.environ.get("CONNECT_EXPECTED_BASE_URL", "").replace(",", " ").split()
if host.strip()
)
_JWT_LEEWAY_SECONDS = 10
_INSTALL_NS = ("atlassian_connect", "installations")
# --- qsh (query-string hash) -----------------------------------------------
def _encode(value: Any) -> str:
"""RFC-3986 encode a single component (Atlassian encodeRfc3986: space->%20)."""
return quote(str(value), safe="")
def _canonical_uri(path: str) -> str:
if not path:
return "/"
if len(path) > 1 and path.endswith("/"):
path = path[:-1]
return path.replace("&", "%26")
def _canonical_query(query: str) -> str:
if not query:
return ""
parsed = parse_qs(query, keep_blank_values=True)
parsed.pop("jwt", None)
pairs = [
f"{_encode(key)}={','.join(sorted(_encode(v) for v in values))}"
for key, values in parsed.items()
]
pairs.sort()
return "&".join(pairs)
def canonical_request(method: str, path: str, query: str = "") -> str:
return "&".join([method.upper(), _canonical_uri(path), _canonical_query(query)])
def compute_qsh(method: str, path: str, query: str = "") -> str:
"""SHA-256 hex of the canonical `METHOD&path&query` request string."""
return hashlib.sha256(canonical_request(method, path, query).encode()).hexdigest()
# --- token extraction + JWT verification -----------------------------------
def extract_connect_token(request: Any) -> str | None:
"""Pull the Connect JWT from `Authorization: JWT <t>` or the `?jwt=` param."""
header = request.headers.get("Authorization") or request.headers.get("authorization") or ""
if header[:4].upper() == "JWT ":
token = header[4:].strip()
if token:
return token
query_token = request.query_params.get("jwt") if hasattr(request, "query_params") else None
return query_token or None
def verify_connect_jwt(
request: Any,
*,
shared_secret: str | None = None,
expected_client_key: str | None = None,
qsh_required: bool = True,
) -> dict[str, Any] | None:
"""Verify a Connect JWT. Returns the claims on success, None on any failure.
``shared_secret`` is passed explicitly by the lifecycle routes (the stored
secret); the webhook resolves it from the token's ``iss`` via the store.
``qsh_required`` is True for the webhook and False (verify-if-present) for
lifecycle callbacks, whose auth strength comes from the signature + issuer
binding rather than qsh. Every failure path returns None with no side effect.
"""
token = extract_connect_token(request)
if not token:
logger.warning("Connect JWT missing — rejecting")
return None
# Fail-fast alg pin before any lookup: blocks alg=none and RS256/ES256
# algorithm-confusion against a symmetric secret.
try:
alg = jwt.get_unverified_header(token).get("alg")
except jwt.PyJWTError:
logger.warning("Connect JWT header undecodable — rejecting")
return None
if alg != "HS256":
logger.warning("Connect JWT alg %r is not HS256 — rejecting", alg)
return None
# Read the (untrusted) issuer to resolve the secret; trust nothing yet.
try:
unverified = jwt.decode(token, options={"verify_signature": False})
except jwt.PyJWTError:
logger.warning("Connect JWT undecodable — rejecting")
return None
issuer = unverified.get("iss")
if not issuer:
logger.warning("Connect JWT missing iss — rejecting")
return None
# The caller always supplies the secret: lifecycle passes the stored secret;
# the webhook resolves it from `iss` via get_shared_secret and passes it.
# A missing secret fails closed (never verify against nothing / a default).
if not shared_secret:
logger.warning("No shared secret provided for Connect JWT — rejecting")
return None
try:
claims = jwt.decode(
token,
shared_secret,
algorithms=["HS256"],
options={
"require": ["exp", "iss"],
"verify_signature": True,
"verify_exp": True,
"verify_nbf": True,
},
leeway=_JWT_LEEWAY_SECONDS,
)
except jwt.PyJWTError as exc:
logger.warning("Connect JWT signature/claims invalid: %s", exc.__class__.__name__)
return None
if expected_client_key is not None and claims.get("iss") != expected_client_key:
logger.warning("Connect JWT iss does not match expected client key — rejecting")
return None
# qsh verified LAST — it is a signed claim, only trustworthy post-signature.
qsh_claim = claims.get("qsh")
if qsh_claim == "context-qsh":
logger.warning("Connect JWT carries context-qsh (iframe token) — rejecting")
return None
if qsh_claim is None:
if qsh_required:
logger.warning("Connect JWT missing required qsh — rejecting")
return None
else:
expected_qsh = compute_qsh(request.method, request.url.path, request.url.query)
if not hmac.compare_digest(expected_qsh, qsh_claim):
logger.warning("Connect JWT qsh mismatch — rejecting")
return None
return claims
_KID_RE = re.compile(r"^[A-Za-z0-9._-]+$")
async def _fetch_atlassian_public_key(kid: str) -> str | None:
"""Fetch Atlassian's PEM public key for a signed-install ``kid``.
``kid`` is validated to a strict charset before use (defense-in-depth on top
of the fixed host + percent-encoding) so a malformed kid fails fast without
a network call and can never influence the request path.
"""
if not _KID_RE.match(kid):
logger.warning("Rejecting Connect install: malformed kid")
return None
async with httpx.AsyncClient(timeout=DEFAULT_HTTP_TIMEOUT) as client:
try:
response = await client.get(f"{_CONNECT_INSTALL_KEYS_BASE}/{quote(kid, safe='')}")
response.raise_for_status()
return response.text
except Exception as exc: # noqa: BLE001
logger.warning("Failed to fetch Atlassian install public key: %s", exc)
return None
async def verify_asymmetric_install_jwt(
request: Any, *, expected_client_key: str | None = None
) -> dict[str, Any] | None:
"""Verify a signed-install (RS256) lifecycle JWT against Atlassian's keys.
signed-install cryptographically authenticates the install/uninstall
callbacks — including the FIRST install — so first install is not
trust-on-first-use. The token is RS256-signed by Atlassian; its ``kid``
selects a published public key, and its ``aud`` must equal this app's
baseUrl (blocking tokens minted for another Connect app). qsh is
verify-if-present on lifecycle. Returns claims on success, None otherwise.
"""
if not CONNECT_BASE_URL:
logger.warning("CONNECT_BASE_URL unset — cannot verify signed-install aud; rejecting")
return None
token = extract_connect_token(request)
if not token:
return None
try:
header = jwt.get_unverified_header(token)
except jwt.PyJWTError:
return None
if header.get("alg") != "RS256":
logger.warning("Signed-install JWT alg %r is not RS256 — rejecting", header.get("alg"))
return None
kid = header.get("kid")
if not kid:
return None
public_key = await _fetch_atlassian_public_key(kid)
if not public_key:
return None
try:
claims = jwt.decode(
token,
public_key,
algorithms=["RS256"],
audience=CONNECT_BASE_URL,
options={
"require": ["exp", "iss", "aud"],
"verify_signature": True,
"verify_exp": True,
"verify_nbf": True,
"verify_aud": True,
},
leeway=_JWT_LEEWAY_SECONDS,
)
except jwt.PyJWTError as exc:
logger.warning("Signed-install JWT invalid: %s", exc.__class__.__name__)
return None
if expected_client_key is not None and claims.get("iss") != expected_client_key:
logger.warning("Signed-install JWT iss does not match clientKey — rejecting")
return None
qsh_claim = claims.get("qsh")
if qsh_claim == "context-qsh":
return None
if qsh_claim is not None:
expected_qsh = compute_qsh(request.method, request.url.path, request.url.query)
if not hmac.compare_digest(expected_qsh, qsh_claim):
logger.warning("Signed-install JWT qsh mismatch — rejecting")
return None
return claims
async def verify_connect_webhook(request: Any) -> dict[str, Any] | None:
"""Resolve the shared secret from the token's issuer, then verify (qsh required).
The async counterpart to verify_connect_jwt for the webhook route, where the
secret must be looked up from the store by the (untrusted-until-verified)
issuer. Returns claims on success, None on any failure.
"""
token = extract_connect_token(request)
if not token:
return None
try:
issuer = jwt.decode(token, options={"verify_signature": False}).get("iss")
except jwt.PyJWTError:
return None
if not issuer:
return None
secret = await get_shared_secret(issuer)
if not secret:
logger.warning("No installation for Connect issuer — rejecting webhook")
return None
return verify_connect_jwt(
request, shared_secret=secret, expected_client_key=issuer, qsh_required=True
)
# --- encrypted installation store ------------------------------------------
def _client() -> Any:
return get_client(url=LANGGRAPH_URL)
async def get_installation(client_key: str) -> dict[str, Any] | None:
"""Return the stored installation record (secret decrypted), or None."""
item = await _client().store.get_item(_INSTALL_NS, client_key)
if not item:
return None
value = item.get("value")
if not isinstance(value, dict):
return None
record = dict(value)
try:
record["shared_secret"] = decrypt_token(record["shared_secret_enc"])
except Exception: # noqa: BLE001 — fail closed on decrypt/missing-key failure
logger.warning("Failed to decrypt stored Connect shared secret for %s", client_key)
return None
return record
async def get_shared_secret(client_key: str) -> str | None:
record = await get_installation(client_key)
return (record or {}).get("shared_secret") or None
async def put_installation(
client_key: str,
shared_secret: str,
base_url: str,
product_type: str,
*,
first_install: bool,
) -> None:
now_ms = int(time.time() * 1000)
installed_at = now_ms
if not first_install:
existing = await get_installation(client_key)
installed_at = (existing or {}).get("installed_at_ms", now_ms)
record = {
"client_key": client_key,
"shared_secret_enc": encrypt_token(shared_secret),
"base_url": base_url,
"product_type": product_type,
"installed_at_ms": installed_at,
"updated_at_ms": now_ms,
}
await _client().store.put_item(_INSTALL_NS, client_key, record)
async def delete_installation(client_key: str) -> None:
await _client().store.delete_item(_INSTALL_NS, client_key)
def client_key_allowed(client_key: str) -> bool:
"""Whether a signature-verified clientKey (JWT iss) is an accepted tenant.
Fail closed: an empty allowlist accepts no installs (the mandatory tenant
binding — see CONNECT_EXPECTED_CLIENT_KEYS).
"""
return bool(client_key) and client_key in CONNECT_EXPECTED_CLIENT_KEYS
def base_url_host_allowed(base_url: str) -> bool:
"""Whether an install baseUrl's host is in the optional CONNECT_EXPECTED_BASE_URL allowlist.
Defense-in-depth only (the baseUrl is an untrusted body field); the real
tenant gate is client_key_allowed. Returns False when the allowlist is empty.
"""
if not CONNECT_EXPECTED_BASE_URL_HOSTS:
return False
from urllib.parse import urlparse
host = (urlparse(base_url).hostname or "").lower()
return bool(host) and host in CONNECT_EXPECTED_BASE_URL_HOSTS

View file

@ -450,7 +450,7 @@ async def resolve_github_token(config: RunnableConfig, thread_id: str) -> tuple[
# restores the per-user OAuth token so the run is attributed to the triggering
# user.
if (
source in ("slack", "linear", "dashboard", "schedule")
source in ("slack", "linear", "jira", "dashboard", "schedule")
and isinstance(github_login, str)
and github_login.strip()
):

269
agent/utils/confluence.py Normal file
View file

@ -0,0 +1,269 @@
"""Confluence Cloud REST API utilities.
Mirrors ``agent/utils/jira.py`` but talks to Confluence Cloud REST v1 (the
``/wiki/rest/api`` base path) with a single service-account (Basic auth over
``email:api_token``). Confluence page/comment bodies are XHTML "storage
format," not Atlassian Document Format, so this module has its own tiny
storage <-> text converters instead of importing ``agent/utils/adf.py``.
"""
from __future__ import annotations
import base64
import os
import re
from html import unescape
from typing import Any
from urllib.parse import quote
import httpx
from .http import DEFAULT_HTTP_TIMEOUT
def _seg(value: str) -> str:
"""Percent-encode a single untrusted URL path segment."""
return quote(value, safe="")
CONFLUENCE_BASE_URL = os.environ.get("CONFLUENCE_BASE_URL", "").rstrip("/")
CONFLUENCE_EMAIL = os.environ.get("CONFLUENCE_EMAIL", "")
CONFLUENCE_API_TOKEN = os.environ.get("CONFLUENCE_API_TOKEN", "")
_PAGE_EXPAND = "body.storage,version,space"
def _headers() -> dict[str, str]:
token = base64.b64encode(f"{CONFLUENCE_EMAIL}:{CONFLUENCE_API_TOKEN}".encode()).decode()
return {
"Authorization": f"Basic {token}",
"Content-Type": "application/json",
"Accept": "application/json",
}
async def _request(
method: str,
path: str,
*,
json: dict[str, Any] | None = None,
params: dict[str, Any] | None = None,
) -> dict[str, Any]:
"""Execute a REST request against the Confluence Cloud API."""
if not (CONFLUENCE_BASE_URL and CONFLUENCE_EMAIL and CONFLUENCE_API_TOKEN):
return {
"error": "CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN are not set"
}
async with httpx.AsyncClient(timeout=DEFAULT_HTTP_TIMEOUT) as client:
try:
response = await client.request(
method,
f"{CONFLUENCE_BASE_URL}/wiki/rest/api{path}",
headers=_headers(),
json=json,
params=params,
)
response.raise_for_status()
return response.json() if response.content else {}
except Exception as e: # noqa: BLE001
return {"error": str(e)}
def _page_url(webui_path: str | None) -> str:
if not webui_path or not CONFLUENCE_BASE_URL:
return ""
return f"{CONFLUENCE_BASE_URL}/wiki{webui_path}"
def text_to_storage(text: str) -> str:
"""Wrap blank-line-separated blocks of plain text in ``<p>`` tags.
Minimal converter for agent-authored prose; not a general HTML sanitizer.
"""
blocks = [b.strip() for b in text.split("\n\n")]
escaped = (
b.replace("&", "&amp;").replace("<", "&lt;").replace(">", "&gt;") for b in blocks if b
)
return "".join(f"<p>{b}</p>" for b in escaped)
_TAG_RE = re.compile(r"<[^>]+>")
def storage_to_text(xhtml: str) -> str:
"""Strip XHTML storage-format markup down to plain text."""
if not xhtml:
return ""
text = _TAG_RE.sub("", xhtml)
return unescape(text).strip()
def _normalize_page(raw: dict[str, Any]) -> dict[str, Any]:
body = raw.get("body", {}) or {}
storage = body.get("storage", {}) or {}
version = raw.get("version", {}) or {}
space = raw.get("space", {}) or {}
links = raw.get("_links", {}) or {}
return {
"id": raw.get("id", ""),
"title": raw.get("title", ""),
"body": storage_to_text(storage.get("value", "")),
"version": version.get("number", 0),
"space_key": space.get("key", ""),
"url": _page_url(links.get("webui")),
}
async def get_page(page_id: str) -> dict[str, Any]:
"""Get a Confluence page by id, with its body normalized to plain text."""
result = await _request("GET", f"/content/{_seg(page_id)}", params={"expand": _PAGE_EXPAND})
if "error" in result:
return result
return {"page": _normalize_page(result)}
async def create_page(
space_key: str,
title: str,
body: str,
parent_id: str | None = None,
) -> dict[str, Any]:
"""Create a new Confluence page in the given space."""
payload: dict[str, Any] = {
"type": "page",
"title": title,
"space": {"key": space_key},
"body": {"storage": {"value": text_to_storage(body), "representation": "storage"}},
}
if parent_id is not None:
payload["ancestors"] = [{"id": parent_id}]
result = await _request("POST", "/content", json=payload)
if "error" in result:
return result
links = result.get("_links", {}) or {}
return {
"success": bool(result.get("id")),
"page": {
"id": result.get("id", ""),
"title": result.get("title", ""),
"url": _page_url(links.get("webui")),
},
}
async def update_page(
page_id: str,
title: str | None = None,
body: str | None = None,
) -> dict[str, Any]:
"""Update an existing Confluence page.
Confluence requires the next version number on every update, so this
first reads the current page to learn ``version.number`` and the
existing title (title is a required field on the PUT even when unchanged).
"""
current = await get_page(page_id)
if "error" in current:
return current
page = current["page"]
new_title = title if title is not None else page["title"]
payload: dict[str, Any] = {
"type": "page",
"title": new_title,
"version": {"number": page["version"] + 1},
}
if body is not None:
payload["body"] = {"storage": {"value": text_to_storage(body), "representation": "storage"}}
result = await _request("PUT", f"/content/{_seg(page_id)}", json=payload)
if "error" in result:
return result
links = result.get("_links", {}) or {}
return {
"success": bool(result.get("id")),
"page": {
"id": result.get("id", ""),
"title": result.get("title", ""),
"url": _page_url(links.get("webui")),
},
}
async def add_comment(page_id: str, body: str) -> dict[str, Any]:
"""Add a comment to a Confluence page."""
payload = {
"type": "comment",
"container": {"id": page_id, "type": "page"},
"body": {"storage": {"value": text_to_storage(body), "representation": "storage"}},
}
result = await _request("POST", "/content", json=payload)
if "error" in result:
return result
return {"success": bool(result.get("id")), "id": result.get("id", "")}
def _normalize_search_result(raw: dict[str, Any]) -> dict[str, Any]:
links = raw.get("_links", {}) or {}
return {
"id": raw.get("id", ""),
"title": raw.get("title", ""),
"type": raw.get("type", ""),
"url": _page_url(links.get("webui")),
}
async def search(cql: str) -> dict[str, Any]:
"""Search Confluence content using CQL."""
result = await _request("GET", "/content/search", params={"cql": cql})
if "error" in result:
return result
results = result.get("results", [])
return {"results": [_normalize_search_result(r) for r in results]}
_COMMENT_EXPAND = "body.storage,history.createdBy,ancestors,space"
def _normalize_comment(raw: dict[str, Any]) -> dict[str, Any]:
body = (raw.get("body") or {}).get("storage") or {}
author = ((raw.get("history") or {}).get("createdBy")) or {}
ancestors = raw.get("ancestors") or []
space = raw.get("space") or {}
# A comment's container page is its nearest ancestor.
page_id = ancestors[-1].get("id", "") if ancestors else ""
return {
"id": raw.get("id", ""),
"body": storage_to_text(body.get("value", "")),
"author": {
"account_id": author.get("accountId"),
"name": author.get("displayName"),
"email": author.get("email"),
},
"page_id": page_id,
"space_key": space.get("key", ""),
}
async def get_comment(comment_id: str) -> dict[str, Any]:
"""Fetch a Confluence comment by id — the authoritative record for a webhook.
Connect webhook bodies are only a pointer; the comment's real author, text,
and container are read here via the Basic-auth service account.
"""
result = await _request(
"GET", f"/content/{_seg(comment_id)}", params={"expand": _COMMENT_EXPAND}
)
if "error" in result:
return result
return {"comment": _normalize_comment(result)}
async def get_user_email(account_id: str) -> str | None:
"""Look up a Confluence user's email by accountId (webhooks carry accountId)."""
result = await _request("GET", "/user", params={"accountId": account_id})
if "error" in result:
return None
return result.get("email")

View file

@ -0,0 +1,5 @@
from typing import Any
# Maps a Confluence space key to the repo a comment-triggered run targets.
# Real entries are deployment config; empty falls through to the team default.
CONFLUENCE_SPACE_TO_REPO: dict[str, dict[str, Any] | dict[str, str]] = {}

268
agent/utils/jira.py Normal file
View file

@ -0,0 +1,268 @@
"""Jira Cloud REST API utilities.
Mirrors ``agent/utils/linear.py`` but talks to Jira Cloud REST v3 with a single
service-account (Basic auth over ``email:api_token``). Issue/comment bodies are
Atlassian Document Format (ADF), so read paths convert ADF -> markdown and write
paths convert markdown -> ADF via ``agent/utils/adf.py``.
"""
from __future__ import annotations
import base64
import logging
import os
import re
from typing import Any
from urllib.parse import quote
import httpx
from agent.utils.langsmith import get_langsmith_trace_url
from .adf import adf_to_markdown, markdown_to_adf
from .http import DEFAULT_HTTP_TIMEOUT
logger = logging.getLogger(__name__)
# Jira issue keys are `<PROJECT>-<number>` (e.g. PROJ-123). Untrusted webhook
# input is interpolated into REST paths, so keys are validated against this and
# path segments are percent-encoded to prevent traversal / query injection.
_ISSUE_KEY_RE = re.compile(r"^[A-Za-z][A-Za-z0-9]*-\d+$")
def is_valid_issue_key(issue_key: str) -> bool:
"""Whether ``issue_key`` matches the Jira `<PROJECT>-<number>` format."""
return bool(issue_key) and bool(_ISSUE_KEY_RE.match(issue_key))
def _seg(value: str) -> str:
"""Percent-encode a single untrusted URL path segment (no '/' passthrough)."""
return quote(value, safe="")
JIRA_BASE_URL = os.environ.get("JIRA_BASE_URL", "").rstrip("/") # https://seahaven.atlassian.net
JIRA_EMAIL = os.environ.get("JIRA_SERVICE_EMAIL", "")
JIRA_API_TOKEN = os.environ.get("JIRA_API_TOKEN", "")
_ISSUE_FIELDS = (
"summary,description,status,assignee,reporter,priority,labels,"
"project,issuetype,created,updated,comment"
)
def _headers() -> dict[str, str]:
token = base64.b64encode(f"{JIRA_EMAIL}:{JIRA_API_TOKEN}".encode()).decode()
return {
"Authorization": f"Basic {token}",
"Content-Type": "application/json",
"Accept": "application/json",
}
async def _request(
method: str,
path: str,
*,
json: dict[str, Any] | None = None,
params: dict[str, Any] | None = None,
) -> dict[str, Any]:
"""Execute a REST request against the Jira Cloud v3 API."""
if not (JIRA_BASE_URL and JIRA_EMAIL and JIRA_API_TOKEN):
return {"error": "JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN are not set"}
async with httpx.AsyncClient(timeout=DEFAULT_HTTP_TIMEOUT) as client:
try:
response = await client.request(
method,
f"{JIRA_BASE_URL}/rest/api/3{path}",
headers=_headers(),
json=json,
params=params,
)
response.raise_for_status()
return response.json() if response.content else {}
except Exception as e: # noqa: BLE001
return {"error": str(e)}
def _issue_url(issue_key: str) -> str:
return f"{JIRA_BASE_URL}/browse/{_seg(issue_key)}" if JIRA_BASE_URL else ""
def _normalize_issue(raw: dict[str, Any]) -> dict[str, Any]:
"""Flatten a raw Jira issue into an agent-friendly dict with markdown bodies."""
fields = raw.get("fields", {}) or {}
project = fields.get("project") or {}
status = fields.get("status") or {}
assignee = fields.get("assignee") or {}
priority = fields.get("priority") or {}
issue_type = fields.get("issuetype") or {}
return {
"key": raw.get("key", ""),
"id": raw.get("id", ""),
"title": fields.get("summary", ""),
"description": adf_to_markdown(fields.get("description")),
"status": status.get("name", ""),
"assignee": {
"name": assignee.get("displayName"),
"email": assignee.get("emailAddress"),
"account_id": assignee.get("accountId"),
}
if assignee
else None,
"priority": priority.get("name", ""),
"labels": fields.get("labels", []),
"project_key": project.get("key", ""),
"project_name": project.get("name", ""),
"issue_type": issue_type.get("name", ""),
"created": fields.get("created", ""),
"updated": fields.get("updated", ""),
"url": _issue_url(raw.get("key", "")),
}
def _normalize_comment(raw: dict[str, Any]) -> dict[str, Any]:
author = raw.get("author") or {}
return {
"id": raw.get("id", ""),
"body": adf_to_markdown(raw.get("body")),
"created": raw.get("created", ""),
"updated": raw.get("updated", ""),
"author": {
"name": author.get("displayName"),
"email": author.get("emailAddress"),
"account_id": author.get("accountId"),
},
}
async def comment_on_issue(issue_key: str, comment_body: str) -> bool:
"""Add a comment (markdown) to a Jira issue. Returns True on success."""
result = await _request(
"POST",
f"/issue/{_seg(issue_key)}/comment",
json={"body": markdown_to_adf(comment_body)},
)
return bool(result.get("id")) and "error" not in result
async def post_jira_trace_comment(issue_key: str, thread_id: str) -> None:
"""Post a trace URL comment on a Jira issue."""
trace_url = get_langsmith_trace_url(thread_id)
body = f"On it! [View trace]({trace_url})" if trace_url else "On it!"
await comment_on_issue(issue_key, body)
async def get_user_email(account_id: str) -> str | None:
"""Look up a Jira user's email by accountId (webhooks only carry accountId)."""
result = await _request("GET", "/user", params={"accountId": account_id})
if "error" in result:
return None
return result.get("emailAddress")
async def get_issue(issue_key: str) -> dict[str, Any]:
"""Get a Jira issue by its key (e.g. PROJ-123)."""
result = await _request("GET", f"/issue/{_seg(issue_key)}", params={"fields": _ISSUE_FIELDS})
if "error" in result:
return result
return {"issue": _normalize_issue(result)}
async def get_issue_comments(issue_key: str) -> dict[str, Any]:
"""Get comments for a Jira issue (newest ordering as returned by Jira)."""
result = await _request("GET", f"/issue/{_seg(issue_key)}/comment")
if "error" in result:
return result
comments = result.get("comments", [])
return {"comments": [_normalize_comment(c) for c in comments]}
async def get_comment(issue_key: str, comment_id: str) -> dict[str, Any]:
"""Fetch a single comment by id — the authoritative record for a webhook.
Webhook payloads are unsigned, so the triggering comment's real author and
body must be read from Jira server-side (matched by comment_id) rather than
trusted from the payload.
"""
result = await _request("GET", f"/issue/{_seg(issue_key)}/comment/{_seg(comment_id)}")
if "error" in result:
return result
return {"comment": _normalize_comment(result)}
async def create_issue(
project_key: str,
summary: str,
description: str | None = None,
issue_type: str = "Task",
assignee_account_id: str | None = None,
priority: str | None = None,
labels: list[str] | None = None,
) -> dict[str, Any]:
"""Create a new Jira issue."""
fields: dict[str, Any] = {
"project": {"key": project_key},
"summary": summary,
"issuetype": {"name": issue_type},
}
if description is not None:
fields["description"] = markdown_to_adf(description)
if assignee_account_id is not None:
fields["assignee"] = {"accountId": assignee_account_id}
if priority is not None:
fields["priority"] = {"name": priority}
if labels is not None:
fields["labels"] = labels
result = await _request("POST", "/issue", json={"fields": fields})
if "error" in result:
return result
key = result.get("key", "")
return {
"success": bool(key),
"issue": {"key": key, "id": result.get("id", ""), "url": _issue_url(key)},
}
async def update_issue(
issue_key: str,
summary: str | None = None,
description: str | None = None,
assignee_account_id: str | None = None,
priority: str | None = None,
labels: list[str] | None = None,
) -> dict[str, Any]:
"""Update an existing Jira issue."""
fields: dict[str, Any] = {}
if summary is not None:
fields["summary"] = summary
if description is not None:
fields["description"] = markdown_to_adf(description)
if assignee_account_id is not None:
fields["assignee"] = {"accountId": assignee_account_id}
if priority is not None:
fields["priority"] = {"name": priority}
if labels is not None:
fields["labels"] = labels
if not fields:
return {"error": "No fields to update"}
# A 204 (empty body) is success; _request returns {} in that case.
result = await _request("PUT", f"/issue/{_seg(issue_key)}", json={"fields": fields})
if "error" in result:
return result
return {"success": True, "issue": {"key": issue_key, "url": _issue_url(issue_key)}}
async def list_projects() -> dict[str, Any]:
"""List projects visible to the service account."""
result = await _request("GET", "/project/search")
if "error" in result:
return result
projects = [
{"key": p.get("key", ""), "name": p.get("name", ""), "id": p.get("id", "")}
for p in result.get("values", [])
]
return {"projects": projects}

View file

@ -0,0 +1,10 @@
"""Static Jira project -> GitHub repo mapping.
Mirrors ``agent/utils/linear_team_repo_map.py``, but flat: Jira has no
team/project two-level split the way Linear does, so this is keyed directly by
project key.
"""
JIRA_PROJECT_TO_REPO: dict[str, dict[str, str]] = {
"OS": {"owner": "langchain-ai", "name": "open-swe"},
}

View file

@ -2,6 +2,7 @@
import hashlib
import hmac
import ipaddress
import json
import logging
import os
@ -13,7 +14,7 @@ from typing import Any
from urllib.parse import parse_qs, quote
import httpx
from fastapi import BackgroundTasks, FastAPI, HTTPException, Request
from fastapi import BackgroundTasks, FastAPI, HTTPException, Request, Response
from fastapi.middleware.cors import CORSMiddleware
from langgraph_sdk import get_client
from langgraph_sdk.client import LangGraphClient
@ -39,6 +40,7 @@ from .dashboard.team_settings import (
)
from .dashboard.user_mappings import (
email_for_login, # noqa: F401
is_login_mapped, # noqa: F401
login_for_email, # noqa: F401
login_for_slack_id, # noqa: F401
)
@ -58,11 +60,16 @@ from .reviewer_findings import (
)
from .reviewer_publish import fetch_pr_review_threads, post_review_started_comment # noqa: F401
from .reviewer_reconcile import reconcile_findings_with_review_threads # noqa: F401
from .utils.atlassian_connect import verify_connect_webhook
from .utils.auth import (
is_bot_token_only_mode,
resolve_github_token_from_email,
)
from .utils.comments import get_recent_comments # noqa: F401
from .utils.confluence import get_comment as get_confluence_comment
from .utils.confluence import get_page as get_confluence_page
from .utils.confluence import get_user_email as get_confluence_user_email # noqa: F401
from .utils.confluence_space_repo_map import CONFLUENCE_SPACE_TO_REPO
from .utils.dashboard_links import dashboard_thread_url # noqa: F401
from .utils.github_app import (
get_github_app_installation_token, # noqa: F401
@ -90,6 +97,13 @@ from .utils.github_token import (
invalidate_cached_github_token,
)
from .utils.http import DEFAULT_HTTP_TIMEOUT
from .utils.jira import get_comment as get_jira_comment
from .utils.jira import get_issue as get_jira_issue
from .utils.jira import get_issue_comments as get_jira_issue_comments
from .utils.jira import get_user_email as get_jira_user_email
from .utils.jira import is_valid_issue_key as is_valid_jira_issue_key
from .utils.jira import post_jira_trace_comment # noqa: F401
from .utils.jira_project_repo_map import JIRA_PROJECT_TO_REPO
from .utils.linear import post_linear_trace_comment # noqa: F401
from .utils.linear_team_repo_map import LINEAR_TEAM_TO_REPO
from .utils.multimodal import (
@ -188,7 +202,30 @@ app.include_router(plan_router)
app.include_router(workflow_approval_router)
LINEAR_WEBHOOK_SECRET = os.environ.get("LINEAR_WEBHOOK_SECRET", "")
JIRA_WEBHOOK_SECRET = os.environ.get("JIRA_WEBHOOK_SECRET", "")
# Opt-in stronger trust for the Jira webhook: when true, the Automation payload
# must carry a valid HMAC-SHA256 body signature (X-Openswe-Signature) plus a
# fresh `timestamp`, closing the replay/forgery gap of the static-token model.
JIRA_WEBHOOK_REQUIRE_SIGNATURE = os.environ.get(
"JIRA_WEBHOOK_REQUIRE_SIGNATURE", ""
).strip().lower() in (
"1",
"true",
"yes",
)
JIRA_WEBHOOK_MAX_AGE_SECONDS = 300
# Opt-in CIDR allowlist for the Jira webhook's direct client IP. Empty = off.
# Only meaningful when the app terminates connections directly; behind a proxy
# or load balancer, allowlist Atlassian's published egress ranges at that layer
# instead (this checks the immediate peer, not X-Forwarded-For).
JIRA_WEBHOOK_IP_ALLOWLIST: tuple[str, ...] = tuple(
cidr.strip()
for cidr in os.environ.get("JIRA_WEBHOOK_IP_ALLOWLIST", "").split(",")
if cidr.strip()
)
GITHUB_WEBHOOK_SECRET = os.environ.get("GITHUB_WEBHOOK_SECRET", "")
# Public origin the Atlassian Connect descriptor advertises (empty context path).
CONNECT_BASE_URL = os.environ.get("CONNECT_BASE_URL", "").rstrip("/")
SLACK_SIGNING_SECRET = os.environ.get("SLACK_SIGNING_SECRET", "")
SLACK_BOT_USER_ID = os.environ.get("SLACK_BOT_USER_ID", "")
SLACK_BOT_USERNAME = os.environ.get("SLACK_BOT_USERNAME", "")
@ -225,6 +262,21 @@ ALLOWED_GITHUB_REPOS: frozenset[str] = frozenset(
for repo in os.environ.get("ALLOWED_GITHUB_REPOS", "").split(",")
if repo.strip()
)
# When true, an empty allowlist is treated as "allow nothing" (fail closed)
# rather than "allow all" (the back-compat default). Set this once ALLOWED_
# GITHUB_ORGS/REPOS are configured to prevent a forged/misconfigured trigger
# from steering the agent at an arbitrary repo.
REQUIRE_REPO_ALLOWLIST = os.environ.get("REQUIRE_REPO_ALLOWLIST", "").strip().lower() in (
"1",
"true",
"yes",
)
if not ALLOWED_GITHUB_ORGS and not ALLOWED_GITHUB_REPOS and not REQUIRE_REPO_ALLOWLIST:
logger.warning(
"No repo allowlist configured (ALLOWED_GITHUB_ORGS/ALLOWED_GITHUB_REPOS empty) and "
"REQUIRE_REPO_ALLOWLIST is off — all repos are permitted (fail-open). Configure the "
"allowlist and set REQUIRE_REPO_ALLOWLIST=true to fail closed."
)
LINEAR_API_KEY = os.environ.get("LINEAR_API_KEY", "")
@ -264,6 +316,26 @@ def get_repo_config_from_team_mapping(
return fallback
def get_repo_config_from_jira_mapping(project_key: str) -> dict[str, str]:
"""Look up repository configuration from JIRA_PROJECT_TO_REPO mapping.
Flat lookup (no team/project split, unlike Linear): Jira issues carry a
single project key.
"""
fallback = {"owner": DEFAULT_REPO_OWNER, "name": DEFAULT_REPO_NAME} if DEFAULT_REPO_NAME else {}
if not project_key:
return fallback
return JIRA_PROJECT_TO_REPO.get(project_key, fallback)
def get_repo_config_from_confluence_mapping(space_key: str) -> dict[str, str]:
"""Look up repository configuration from CONFLUENCE_SPACE_TO_REPO mapping."""
fallback = {"owner": DEFAULT_REPO_OWNER, "name": DEFAULT_REPO_NAME} if DEFAULT_REPO_NAME else {}
if not space_key:
return fallback
return CONFLUENCE_SPACE_TO_REPO.get(space_key, fallback)
async def react_to_linear_comment(comment_id: str, emoji: str = "👀") -> bool:
"""Add an emoji reaction to a Linear comment.
@ -375,6 +447,64 @@ async def fetch_linear_issue_details(issue_id: str) -> dict[str, Any] | None:
return None
async def fetch_jira_issue_details(issue_key: str) -> dict[str, Any] | None:
"""Fetch full issue details from Jira (title/description/etc.).
Thin wrapper over ``agent.utils.jira.get_issue``, mirroring
``fetch_linear_issue_details``. Returns None on error so callers can fall
back to the (thinner) webhook-supplied issue data.
"""
result = await get_jira_issue(issue_key)
if "error" in result:
logger.warning("Failed to fetch Jira issue %s: %s", issue_key, result["error"])
return None
return result.get("issue")
async def fetch_jira_issue_comments(issue_key: str) -> list[dict[str, Any]]:
"""Fetch normalized comments for a Jira issue, or [] on error."""
result = await get_jira_issue_comments(issue_key)
if "error" in result:
logger.warning("Failed to fetch Jira comments for %s: %s", issue_key, result["error"])
return []
return result.get("comments", [])
async def fetch_confluence_comment(comment_id: str) -> dict[str, Any] | None:
"""Fetch the authoritative Confluence comment (author + body + container)."""
result = await get_confluence_comment(comment_id)
if "error" in result:
logger.warning("Failed to fetch Confluence comment %s: %s", comment_id, result["error"])
return None
return result.get("comment")
async def fetch_confluence_page(page_id: str) -> dict[str, Any] | None:
"""Fetch a Confluence page (title/url/etc.) for prompt context, or None."""
result = await get_confluence_page(page_id)
if "error" in result:
logger.warning("Failed to fetch Confluence page %s: %s", page_id, result["error"])
return None
return result.get("page")
async def fetch_jira_comment(issue_key: str, comment_id: str) -> dict[str, Any] | None:
"""Fetch the authoritative triggering comment (author + body) from Jira.
Webhook payloads are unsigned, so the trigger's real author and text are
read server-side (matched by comment_id) rather than trusted from the body.
Returns None when the comment can't be fetched (nonexistent / unreadable),
which the webhook treats as a hard reject.
"""
result = await get_jira_comment(issue_key, comment_id)
if "error" in result:
logger.warning(
"Failed to fetch Jira comment %s on %s: %s", comment_id, issue_key, result["error"]
)
return None
return result.get("comment")
def generate_thread_id_from_issue(issue_id: str) -> str:
"""Generate a deterministic thread ID from a Linear issue ID.
@ -391,6 +521,37 @@ def generate_thread_id_from_issue(issue_id: str) -> str:
)
def generate_thread_id_from_jira_issue(issue_key: str) -> str:
"""Generate a deterministic thread ID from a Jira issue key.
Args:
issue_key: The Jira issue key (e.g. PROJ-123)
Returns:
A UUID-formatted thread ID derived from the issue key
"""
hash_bytes = hashlib.sha256(f"jira-issue:{issue_key}".encode()).hexdigest()
return (
f"{hash_bytes[:8]}-{hash_bytes[8:12]}-{hash_bytes[12:16]}-"
f"{hash_bytes[16:20]}-{hash_bytes[20:32]}"
)
def generate_thread_id_from_confluence_comment(client_key: str, comment_id: str) -> str:
"""Deterministic thread id from tenant clientKey + comment id.
Confluence comment ids are per-instance (not globally unique), so the
verified clientKey salts the hash to prevent cross-tenant thread collisions.
"""
hash_bytes = hashlib.sha256(
f"confluence-comment:{client_key}:{comment_id}".encode()
).hexdigest()
return (
f"{hash_bytes[:8]}-{hash_bytes[8:12]}-{hash_bytes[12:16]}-"
f"{hash_bytes[16:20]}-{hash_bytes[20:32]}"
)
def generate_thread_id_from_github_issue(issue_id: str) -> str:
"""Generate a deterministic thread ID from a GitHub issue ID."""
hash_bytes = hashlib.sha256(f"github-issue:{issue_id}".encode()).hexdigest()
@ -468,12 +629,14 @@ async def _is_docs_plz_slack_channel(
def _is_repo_allowed(repo_config: dict[str, str]) -> bool:
"""Check if the repo is in the allowlist.
Returns True if no allowlist is configured (both ALLOWED_GITHUB_ORGS and
ALLOWED_GITHUB_REPOS are empty), or if the repo owner is in
ALLOWED_GITHUB_ORGS, or if owner/name is in ALLOWED_GITHUB_REPOS.
When no allowlist is configured (both ALLOWED_GITHUB_ORGS and
ALLOWED_GITHUB_REPOS empty), returns True (allow-all, back-compat) unless
REQUIRE_REPO_ALLOWLIST is set, in which case it fails closed. Otherwise
allows the repo when its owner is in ALLOWED_GITHUB_ORGS or owner/name is in
ALLOWED_GITHUB_REPOS.
"""
if not ALLOWED_GITHUB_ORGS and not ALLOWED_GITHUB_REPOS:
return True
return not REQUIRE_REPO_ALLOWLIST
owner = repo_config.get("owner", "").lower()
name = repo_config.get("name", "").lower()
if ALLOWED_GITHUB_ORGS and owner in ALLOWED_GITHUB_ORGS:
@ -895,6 +1058,101 @@ def verify_linear_signature(body: bytes, signature: str, secret: str) -> bool:
return _linear_timestamp_is_fresh(body)
def verify_jira_secret(headers: Any) -> bool:
"""Verify the shared-secret header on a Jira Automation webhook.
Jira Cloud Automation "Send web request" actions aren't HMAC-body-signed
like Linear's webhooks — the rule can only attach static headers. So this
is a constant-time comparison of the ``X-Automation-Webhook-Token`` header
against ``JIRA_WEBHOOK_SECRET`` (configured on the Automation rule's
outgoing webhook action to match this deployment's secret). Fails closed
when the secret is unset.
"""
secret = JIRA_WEBHOOK_SECRET
if not secret:
logger.warning("JIRA_WEBHOOK_SECRET is not configured — rejecting webhook request")
return False
token = headers.get("X-Automation-Webhook-Token", "") or ""
if not token:
return False
return hmac.compare_digest(token, secret)
def _jira_timestamp_is_fresh(body: bytes) -> bool:
"""Reject replays: the payload's ``timestamp`` (Unix ms) must be recent."""
try:
ts_ms = json.loads(body)["timestamp"]
except (json.JSONDecodeError, KeyError, TypeError):
logger.warning("Jira webhook missing/invalid timestamp — rejecting")
return False
if not isinstance(ts_ms, (int, float)) or isinstance(ts_ms, bool):
logger.warning("Jira webhook timestamp is not numeric — rejecting")
return False
now_ms = datetime.now(UTC).timestamp() * 1000
if abs(now_ms - ts_ms) > JIRA_WEBHOOK_MAX_AGE_SECONDS * 1000:
logger.warning("Jira webhook timestamp outside freshness window — rejecting")
return False
return True
def verify_jira_signature(body: bytes, headers: Any) -> bool:
"""Optionally verify an HMAC body signature + fresh timestamp (opt-in).
A no-op returning True unless ``JIRA_WEBHOOK_REQUIRE_SIGNATURE`` is set, so
the default static-token deployments are unaffected. When enabled, the
Automation rule must send ``X-Openswe-Signature`` = hex HMAC-SHA256 of the
raw body keyed by ``JIRA_WEBHOOK_SECRET``, plus a fresh ``timestamp`` field
in the body — binding the request to its exact content and a time window,
which the static token alone cannot. Fails closed.
"""
if not JIRA_WEBHOOK_REQUIRE_SIGNATURE:
return True
secret = JIRA_WEBHOOK_SECRET
if not secret:
logger.warning("JIRA_WEBHOOK_SECRET is not configured — rejecting signed webhook")
return False
signature = headers.get("X-Openswe-Signature", "") or ""
if not signature:
logger.warning("Jira webhook signature required but missing — rejecting")
return False
expected = hmac.new(secret.encode("utf-8"), body, hashlib.sha256).hexdigest()
if not hmac.compare_digest(expected, signature):
logger.warning("Jira webhook signature mismatch — rejecting")
return False
return _jira_timestamp_is_fresh(body)
def verify_jira_source_ip(request: Request) -> bool:
"""Optionally require the direct client IP to fall in an allowlisted CIDR.
A no-op returning True unless ``JIRA_WEBHOOK_IP_ALLOWLIST`` is set. Checks
the immediate peer (``request.client.host``), not ``X-Forwarded-For`` — so
it is only meaningful when the app terminates connections directly. Behind a
proxy/load balancer, allowlist Atlassian's egress ranges at that layer.
"""
if not JIRA_WEBHOOK_IP_ALLOWLIST:
return True
client = request.client
if client is None:
logger.warning("Jira webhook has no client address — rejecting (IP allowlist on)")
return False
try:
peer = ipaddress.ip_address(client.host)
except ValueError:
logger.warning("Jira webhook client host %r is not a valid IP — rejecting", client.host)
return False
for cidr in JIRA_WEBHOOK_IP_ALLOWLIST:
try:
if peer in ipaddress.ip_network(cidr, strict=False):
return True
except ValueError:
logger.warning("Ignoring malformed JIRA_WEBHOOK_IP_ALLOWLIST entry %r", cidr)
logger.warning(
"Jira webhook client %s not in JIRA_WEBHOOK_IP_ALLOWLIST — rejecting", client.host
)
return False
@app.post("/webhooks/linear")
async def linear_webhook( # noqa: PLR0911, PLR0912, PLR0915
request: Request, background_tasks: BackgroundTasks
@ -1053,6 +1311,248 @@ async def linear_webhook_verify() -> dict[str, str]:
return {"status": "ok", "message": "Linear webhook endpoint is active"}
@app.post("/webhooks/jira")
async def jira_webhook( # noqa: PLR0911, PLR0912
request: Request, background_tasks: BackgroundTasks
) -> dict[str, str]:
"""Handle Jira Automation webhooks.
Triggers a new LangGraph run when a comment mentioning ``@openswe`` is
added to an issue. Unlike Linear, Jira Cloud has no native outgoing-webhook
signing, so this is fronted by a Jira **Automation** rule (trigger:
"Issue commented") with a "Send web request" action posting a custom JSON
body to this route, carrying the shared-secret token in
``X-Automation-Webhook-Token``.
Expected payload (the Automation rule's custom JSON body, built from smart
values)::
{
"issue_key": "PROJ-123",
"comment_id": "10050",
"comment_author_is_bot": false
}
``issue_key`` (validated against the Jira key format) and ``comment_id`` are
**required** — they are the only fields trusted from the unsigned body, and
only as a pointer. The triggering comment's real author and text are then
re-fetched from Jira server-side (``fetch_jira_comment``) and everything
security-relevant (identity/attribution, the ``@openswe`` trigger check, the
prompt text, repo routing) is derived from that authoritative record, never
from payload-supplied author/body fields. ``comment_author_is_bot`` is an
optional cheap early-out only. A comment that cannot be corroborated
server-side is rejected.
"""
logger.info("Received Jira webhook")
if not verify_jira_source_ip(request):
raise HTTPException(status_code=403, detail="Source IP not allowed")
if not verify_jira_secret(request.headers):
logger.warning("Invalid Jira webhook token")
raise HTTPException(status_code=401, detail="Invalid token")
body = await request.body()
if not verify_jira_signature(body, request.headers):
raise HTTPException(status_code=401, detail="Invalid signature")
try:
payload = json.loads(body)
except json.JSONDecodeError:
logger.exception("Failed to parse Jira webhook JSON")
return {"status": "error", "message": "Invalid JSON"}
# Cheap early-out on the (untrusted) payload before any Jira API call.
if payload.get("comment_author_is_bot"):
logger.debug("Ignoring webhook: comment is from a bot")
return {"status": "ignored", "reason": "Comment is from a bot"}
issue_key = payload.get("issue_key", "") or ""
if not is_valid_jira_issue_key(issue_key):
logger.debug("Ignoring webhook: missing or malformed issue key")
return {"status": "ignored", "reason": "Missing or malformed issue key"}
comment_id = payload.get("comment_id", "") or ""
if not comment_id:
logger.debug("Ignoring webhook: no comment id to corroborate")
return {"status": "ignored", "reason": "No comment id in payload"}
# Corroborate against the real Jira record. The webhook body is unsigned, so
# the triggering comment's author and text are read server-side (matched by
# comment_id) rather than trusted from the payload — this is what prevents a
# secret-holder from spoofing the author (to hijack another user's token) or
# injecting arbitrary agent instructions. A comment that can't be fetched
# (nonexistent issue/comment or a forged event) is rejected.
server_comment = await fetch_jira_comment(issue_key, comment_id)
if not server_comment:
logger.warning(
"Rejecting Jira webhook: comment %s on %s could not be corroborated",
comment_id,
issue_key,
)
return {"status": "ignored", "reason": "Triggering comment not found"}
author = server_comment.get("author") or {}
account_id = author.get("account_id") or ""
display_name = author.get("name") or ""
comment_body = server_comment.get("body") or ""
for prefix in _GITHUB_BOT_MESSAGE_PREFIXES:
if comment_body.startswith(prefix):
logger.debug("Ignoring webhook: comment is our own bot message")
return {"status": "ignored", "reason": "Comment is our own bot message"}
if "@openswe" not in comment_body.lower():
logger.debug("Ignoring webhook: comment doesn't mention @openswe")
return {"status": "ignored", "reason": "Comment doesn't mention @openswe"}
# Derive the project key from the (validated, corroborated) issue key rather
# than trusting the payload's project_key for repo routing.
project_key = issue_key.split("-", 1)[0]
actor_email = await get_jira_user_email(account_id) if account_id else None
repo_config = extract_repo_from_text(comment_body, default_owner=DEFAULT_REPO_OWNER)
if repo_config:
logger.debug(
"Using repo from comment body: %s/%s",
repo_config["owner"],
repo_config["name"],
)
else:
try:
profile_repo = await get_profile_default_repo(
await resolve_login_from_email_async(actor_email) if actor_email else None
)
except Exception: # noqa: BLE001
logger.exception("Failed to apply dashboard default_repo for Jira user")
profile_repo = None
if profile_repo:
logger.info(
"Applying dashboard default_repo for Jira user %s: %s/%s",
account_id,
profile_repo["owner"],
profile_repo["name"],
)
repo_config = profile_repo
if not repo_config:
repo_config = get_repo_config_from_jira_mapping(project_key)
if not repo_config:
repo_config = await get_team_default_repo()
if not repo_config:
return {"status": "ignored", "reason": "No default repository configured"}
if not _is_repo_allowed(repo_config):
logger.warning(
"Rejecting Jira webhook: repo '%s/%s' not in allowlist",
repo_config.get("owner"),
repo_config.get("name"),
)
return {"status": "ignored", "reason": "Repository not in allowlist"}
issue_data = {
"key": issue_key,
"project_key": project_key,
"triggering_comment": comment_body,
"triggering_comment_id": comment_id,
"comment_author": {
"account_id": account_id,
"email": actor_email,
"name": display_name,
},
}
logger.info(
"Accepted webhook for issue '%s', scheduling background task",
issue_key,
)
background_tasks.add_task(process_jira_issue, issue_data, repo_config)
return {
"status": "accepted",
"message": f"Processing issue '{issue_key}' for repo "
f"{repo_config['owner']}/{repo_config['name']}",
}
@app.get("/webhooks/jira")
async def jira_webhook_verify() -> dict[str, str]:
"""Verify endpoint for Jira webhook setup."""
return {"status": "ok", "message": "Jira webhook endpoint is active"}
# --- Atlassian Connect (Confluence trigger) --------------------------------
@app.get("/connect/atlassian-connect.json")
async def connect_descriptor() -> dict[str, Any]:
"""Serve the Atlassian Connect app descriptor (baseUrl from CONNECT_BASE_URL).
signed-install is true: Atlassian asymmetrically (RS256) signs the lifecycle
callbacks, so install/uninstall are cryptographically authenticated against
Atlassian's published keys (no trust-on-first-use). The comment_created
webhook stays symmetric (HS256 against the stored per-tenant sharedSecret).
"""
return {
"key": "sea-haven-open-swe-confluence",
"name": "Open SWE",
"description": "Triggers Open SWE runs from Confluence comments mentioning @openswe.",
"baseUrl": CONNECT_BASE_URL,
"vendor": {"name": "Sea Haven Industries", "url": "https://seahavenind.com"},
"authentication": {"type": "jwt"},
"apiMigrations": {"signed-install": True, "gdpr": True},
"lifecycle": {"installed": "/connect/installed", "uninstalled": "/connect/uninstalled"},
"scopes": ["READ"],
"modules": {
"webhooks": [{"event": "comment_created", "url": "/connect/webhook/comment-created"}]
},
}
@app.post("/connect/installed")
async def connect_installed(request: Request) -> Response:
"""Connect install lifecycle: trust-on-first-use (host-gated), verify re-install."""
try:
body = await request.json()
except Exception: # noqa: BLE001
raise HTTPException(status_code=400, detail="Invalid JSON") from None
code, detail = await process_install(request, body)
if code >= 400:
raise HTTPException(status_code=code, detail=detail)
return Response(status_code=code)
@app.post("/connect/uninstalled")
async def connect_uninstalled(request: Request) -> Response:
"""Connect uninstall lifecycle: verify against the stored secret before deleting."""
try:
body = await request.json()
except Exception: # noqa: BLE001
raise HTTPException(status_code=400, detail="Invalid JSON") from None
code, detail = await process_uninstall(request, body)
if code >= 400:
raise HTTPException(status_code=code, detail=detail)
return Response(status_code=code)
@app.post("/connect/webhook/comment-created")
async def connect_comment_created(
request: Request, background_tasks: BackgroundTasks
) -> dict[str, str]:
"""JWT-verified Confluence comment_created trigger."""
claims = await verify_connect_webhook(request)
if claims is None:
raise HTTPException(status_code=401, detail="Invalid Connect JWT")
try:
payload = await request.json()
except Exception: # noqa: BLE001
return {"status": "error", "message": "Invalid JSON"}
background_tasks.add_task(process_confluence_comment, payload, claims.get("iss", ""))
return {"status": "accepted"}
@app.post("/webhooks/slack")
async def slack_webhook(request: Request, background_tasks: BackgroundTasks) -> dict[str, str]:
"""Handle Slack Event API webhooks for app mentions."""
@ -2068,6 +2568,11 @@ async def github_webhook(request: Request, background_tasks: BackgroundTasks) ->
# ---- Webhook handlers (moved to agent/webhooks/, re-exported here) ----
# Re-exported so the @app routes above and the test suite (which references
# webapp.process_github_issue, webapp.build_github_issue_prompt, etc.) keep working.
from .webhooks.confluence import ( # noqa: E402,F401
process_confluence_comment,
process_install,
process_uninstall,
)
from .webhooks.github import ( # noqa: E402,F401
_dispatch_first_review_from_pr_payload,
_is_actionable_review_payload,
@ -2088,5 +2593,6 @@ from .webhooks.github import ( # noqa: E402,F401
process_github_review_finding_reply,
trigger_pr_review_from_ref,
)
from .webhooks.jira import process_jira_issue # noqa: E402,F401
from .webhooks.linear import process_linear_issue # noqa: E402,F401
from .webhooks.slack import process_slack_mention # noqa: E402,F401

View file

@ -0,0 +1,215 @@
"""Confluence Atlassian Connect webhook: lifecycle + comment-created handler.
Lifecycle (installed/uninstalled) implements the verify-before-overwrite guard;
the comment handler mirrors ``webhooks/jira.py:process_jira_issue`` with the
Phase-2 lesson baked in: the JWT-signed webhook body is only a pointer, so the
triggering comment's real author, text, and container are re-fetched server-side
via the Basic-auth service account before anything security-relevant is derived.
"""
import os
from typing import Any
from langchain_core.messages.content import create_text_block
from agent import webapp
from agent.utils import atlassian_connect as ac
# The Connect app's own Confluence service-account accountId. When set, comments
# authored by it are ignored (self-trigger loop guard, like the Linear botActor
# / Jira comment_author_is_bot early-outs).
CONFLUENCE_BOT_ACCOUNT_ID = os.environ.get("CONFLUENCE_BOT_ACCOUNT_ID", "")
async def process_install(request: Any, body: dict[str, Any]) -> tuple[int, str]:
"""Handle POST /connect/installed. Returns (status_code, detail).
With signed-install, Atlassian RS256-signs every install callback (including
the first), so both first-install and re-install are verified against
Atlassian's published keys — there is no trust-on-first-use, and the
re-install path cannot be used to rotate our secret without a valid
Atlassian signature.
"""
client_key = body.get("clientKey") or ""
shared_secret = body.get("sharedSecret") or ""
base_url = body.get("baseUrl", "") or ""
product_type = body.get("productType", "") or ""
if not client_key or not shared_secret:
return 400, "Missing clientKey/sharedSecret"
claims = await ac.verify_asymmetric_install_jwt(request, expected_client_key=client_key)
if claims is None:
webapp.logger.warning("Rejecting Connect install for %s: signature unverified", client_key)
return 401, "Install verification failed"
# Mandatory tenant binding: signed-install proves the caller is *an*
# Atlassian tenant, not *ours*, so only installs from an allowlisted
# (signature-verified) clientKey are accepted. To bootstrap, add the
# clientKey logged here to CONNECT_EXPECTED_CLIENT_KEYS and re-install.
if not ac.client_key_allowed(client_key):
webapp.logger.warning(
"Rejecting Connect install: clientKey %s not in CONNECT_EXPECTED_CLIENT_KEYS",
client_key,
)
return 403, "clientKey not allowed"
# Defense-in-depth (only enforced when configured): the callback's baseUrl
# host must be our Confluence site.
if ac.CONNECT_EXPECTED_BASE_URL_HOSTS and not ac.base_url_host_allowed(base_url):
webapp.logger.warning("Rejecting Connect install: baseUrl %s not allowed", base_url)
return 403, "baseUrl host not allowed"
existing = await ac.get_installation(client_key)
await ac.put_installation(
client_key, shared_secret, base_url, product_type, first_install=existing is None
)
webapp.logger.info(
"Connect %s verified and stored for %s",
"first-install" if existing is None else "re-install",
client_key,
)
return 204, ""
async def process_uninstall(request: Any, body: dict[str, Any]) -> tuple[int, str]:
"""Handle POST /connect/uninstalled. Returns (status_code, detail)."""
client_key = body.get("clientKey") or ""
if not client_key:
return 400, "Missing clientKey"
claims = await ac.verify_asymmetric_install_jwt(request, expected_client_key=client_key)
if claims is None:
webapp.logger.warning("Rejecting Connect uninstall for %s: unverified", client_key)
return 401, "Uninstall verification failed"
existing = await ac.get_installation(client_key)
if existing is None:
return 204, "" # idempotent
await ac.delete_installation(client_key)
webapp.logger.info("Connect uninstall verified for %s", client_key)
return 204, ""
def _extract_comment_id(payload: dict[str, Any]) -> str:
comment = payload.get("comment")
if isinstance(comment, dict) and comment.get("id"):
return str(comment["id"])
if payload.get("commentId"):
return str(payload["commentId"])
content = payload.get("content")
if isinstance(content, dict) and content.get("id"):
return str(content["id"])
return ""
async def process_confluence_comment(payload: dict[str, Any], client_key: str = "") -> None:
"""Corroborate a comment_created event server-side and dispatch a run."""
comment_id = _extract_comment_id(payload)
if not comment_id:
webapp.logger.debug("Ignoring Confluence webhook: no comment id in payload")
return
server_comment = await webapp.fetch_confluence_comment(comment_id)
if not server_comment:
webapp.logger.warning(
"Rejecting Confluence webhook: comment %s could not be corroborated", comment_id
)
return
author = server_comment.get("author") or {}
account_id = author.get("account_id") or ""
display_name = author.get("name") or ""
body_text = server_comment.get("body") or ""
page_id = server_comment.get("page_id") or ""
space_key = server_comment.get("space_key") or ""
# Self-trigger loop guard: ignore the app's own comments (its confluence_comment
# replies can echo "@openswe" and otherwise re-trigger).
if CONFLUENCE_BOT_ACCOUNT_ID and account_id == CONFLUENCE_BOT_ACCOUNT_ID:
webapp.logger.debug("Ignoring Confluence webhook: comment authored by the bot account")
return
for prefix in webapp._GITHUB_BOT_MESSAGE_PREFIXES:
if body_text.startswith(prefix):
webapp.logger.debug("Ignoring Confluence webhook: comment is our own bot message")
return
if "@openswe" not in body_text.lower():
webapp.logger.debug("Ignoring Confluence webhook: comment doesn't mention @openswe")
return
actor_email = await webapp.get_confluence_user_email(account_id) if account_id else None
repo_config = webapp.extract_repo_from_text(body_text, default_owner=webapp.DEFAULT_REPO_OWNER)
if not repo_config:
repo_config = webapp.get_repo_config_from_confluence_mapping(space_key)
if not repo_config:
repo_config = await webapp.get_team_default_repo()
if not repo_config:
webapp.logger.info("Ignoring Confluence webhook: no repo resolved for space %s", space_key)
return
if not webapp._is_repo_allowed(repo_config):
webapp.logger.warning(
"Rejecting Confluence webhook: repo '%s/%s' not in allowlist",
repo_config.get("owner"),
repo_config.get("name"),
)
return
mapped_login = await webapp.resolve_login_from_email_async(actor_email) if actor_email else None
if mapped_login and not webapp.is_login_mapped(mapped_login):
webapp.logger.info(
"Confluence actor login %s is not an active mapping; running unattributed", mapped_login
)
mapped_login = None
thread_id = webapp.generate_thread_id_from_confluence_comment(client_key, comment_id)
page = await webapp.fetch_confluence_page(page_id) if page_id else None
page_title = (page or {}).get("title", "") or "Confluence page"
page_url = (page or {}).get("url", "")
triggered_by = f"## Triggered by: {display_name}\n\n" if display_name else ""
prompt = (
f"Please act on the following Confluence comment:\n\n"
f"## Repository: {repo_config.get('owner')}/{repo_config.get('name')}\n\n"
f"## Confluence page: {page_title} ({space_key}) - Page ID: {page_id}\n\n"
f"{triggered_by}"
f"## Comment:\n{body_text}\n\n"
f"Please analyze this and implement the necessary changes. When you're done, commit and "
f"push your changes."
)
content_blocks: list[dict[str, Any]] = [create_text_block(prompt)]
configurable: dict[str, Any] = {
"repo": repo_config,
"confluence": {
"comment_id": comment_id,
"page_id": page_id,
"space_key": space_key,
"url": page_url,
"triggering_user_name": display_name or "",
},
"user_email": actor_email,
"source": "confluence",
}
if mapped_login:
configurable["github_login"] = mapped_login
await webapp.upsert_agent_thread_owner_metadata(
thread_id,
source="confluence",
repo_config=repo_config,
github_login=mapped_login or "",
user_email=actor_email or "",
title=page_title,
source_context={"confluence": configurable["confluence"]},
)
run = await webapp.dispatch_agent_run(
thread_id,
content_blocks,
configurable,
source="confluence",
metadata=webapp._AGENT_VERSION_METADATA,
)
webapp.logger.info(
"LangGraph run dispatched for Confluence thread %s (run=%s)",
thread_id,
run.get("run_id") if isinstance(run, dict) else None,
)

230
agent/webhooks/jira.py Normal file
View file

@ -0,0 +1,230 @@
"""Jira webhook handler — mirrors ``agent/webhooks/linear.py`` for Jira issues.
Helpers and constants stay in webapp.py; they are accessed through the module
object (``webapp.X``) so tests that monkeypatch them keep working.
"""
from typing import Any
from urllib.parse import urlparse
import httpx
from langchain_core.messages.content import create_text_block
from agent import webapp
async def process_jira_issue( # noqa: PLR0912, PLR0915
issue_data: dict[str, Any], repo_config: dict[str, str]
) -> None:
"""Process a Jira issue comment by creating a new LangGraph thread and run.
Args:
issue_data: The Jira issue data from the webhook (basic info + the
triggering comment; see ``webapp.jira_webhook`` for the shape).
repo_config: The repo configuration with owner and name.
"""
issue_key = issue_data.get("key", "")
webapp.logger.info(
"Processing Jira issue %s for repo %s/%s",
issue_key,
repo_config.get("owner"),
repo_config.get("name"),
)
thread_id = webapp.generate_thread_id_from_jira_issue(issue_key)
full_issue = await webapp.fetch_jira_issue_details(issue_key)
if not full_issue:
full_issue = {}
# Actor email for token attribution: restricted to the comment author only,
# resolved once (by account_id) in the webhook handler and carried through
# here — mirrors the Linear handler's restriction to the comment author,
# so a PR is never opened as a non-actor.
comment_author = issue_data.get("comment_author") or {}
actor_email = comment_author.get("email")
user_name = comment_author.get("name") or None
user_email = actor_email
webapp.logger.info("User email for issue %s: %s", issue_key, user_email)
title = full_issue.get("title") or "No title"
description = full_issue.get("description") or "No description"
image_urls: list[str] = []
description_image_urls = webapp.extract_image_urls(description)
if description_image_urls:
image_urls.extend(description_image_urls)
webapp.logger.debug(
"Found %d image URL(s) in issue description",
len(description_image_urls),
)
raw_comments = await webapp.fetch_jira_issue_comments(issue_key)
comments = [{**comment, "createdAt": comment.get("created", "")} for comment in raw_comments]
comments_text = ""
triggering_comment = issue_data.get("triggering_comment", "")
triggering_comment_id = issue_data.get("triggering_comment_id", "")
bot_message_prefixes = webapp._GITHUB_BOT_MESSAGE_PREFIXES
comment_ids: set[str] = set()
comment_id_to_index: dict[str, int] = {}
if comments:
for i, comment in enumerate(comments):
comment_id = comment.get("id", "")
if comment_id:
comment_ids.add(comment_id)
comment_id_to_index[comment_id] = i
relevant_comments = []
trigger_index = None
if triggering_comment_id:
trigger_index = comment_id_to_index.get(triggering_comment_id)
if trigger_index is not None:
relevant_comments = comments[trigger_index:]
webapp.logger.debug(
"Using triggering comment index %d to build relevant comments",
trigger_index,
)
else:
relevant_comments = webapp.get_recent_comments(comments, bot_message_prefixes)
if relevant_comments:
comments_text = "\n\n## Comments:\n"
for comment in relevant_comments:
author = (comment.get("author") or {}).get("name") or "User"
body = comment.get("body", "")
body_image_urls = webapp.extract_image_urls(body)
if body_image_urls:
image_urls.extend(body_image_urls)
webapp.logger.debug(
"Found %d image URL(s) in comment by %s",
len(body_image_urls),
author,
)
if any(body.startswith(prefix) for prefix in bot_message_prefixes):
continue
comments_text += f"\n**{author}:** {body}\n"
if triggering_comment and triggering_comment_id not in comment_ids:
if not comments_text:
comments_text = "\n\n## Comments:\n"
trigger_author = comment_author.get("name") or "Unknown"
trigger_body = triggering_comment
trigger_image_urls = webapp.extract_image_urls(trigger_body)
if trigger_image_urls:
image_urls.extend(trigger_image_urls)
webapp.logger.debug(
"Found %d image URL(s) in triggering comment by %s",
len(trigger_image_urls),
trigger_author,
)
comments_text += f"\n**{trigger_author}:** {trigger_body}\n"
webapp.logger.debug(
"Appended triggering comment %s not present in issue comments list",
triggering_comment_id or "<missing-id>",
)
project_key = issue_data.get("project_key", "") or full_issue.get("project_key", "")
issue_number = issue_key.split("-", 1)[1] if "-" in issue_key else ""
issue_url = full_issue.get("url", "")
triggered_by_line = f"## Triggered by: {user_name}\n\n" if user_name else ""
tag_instruction = (
f"When calling jira_comment, tag @{user_name} if you are asking them a question, need their input, or are notifying them of something important (e.g. a completed PR). For simple answers, tagging is not required."
if user_name
else ""
)
prompt = (
f"Please work on the following issue:\n\n"
f"## Repository: {repo_config.get('owner')}/{repo_config.get('name')}\n\n"
f"## Title: {title}\n\n"
f"{triggered_by_line}"
f"## Jira Ticket: {issue_key}\n\n"
f"## Description:\n{description}\n"
f"{comments_text}\n\n"
f"Please analyze this issue and implement the necessary changes. "
f"When you're done, commit and push your changes. {tag_instruction}"
)
content_blocks: list[dict[str, Any]] = [create_text_block(prompt)]
# Resolve the GitHub login from the actor's Jira email via the same
# user-mapping store Slack/Linear use, so PRs open *as the triggering user*
# and the thread is tagged for the dashboard. Restricted to the comment
# author so token attribution never falls back to reporter/assignee.
mapped_login = await webapp.resolve_login_from_email_async(actor_email) if actor_email else None
# Only attribute to an *active* user mapping; a pending/unconfirmed mapping
# must never drive PR authorship or token resolution.
if mapped_login and not webapp.is_login_mapped(mapped_login):
webapp.logger.info(
"Jira actor login %s is not an active mapping; running unattributed", mapped_login
)
mapped_login = None
image_model_override: tuple[str, str] | None = None
if image_urls:
image_urls = webapp.dedupe_urls(image_urls)
resolved_model_id = await webapp.resolve_agent_model_id(mapped_login)
if not webapp.model_supports_images(resolved_model_id):
fallback_model_id, fallback_effort = webapp.default_vision_model_pair()
webapp.logger.info(
"Using vision fallback model %s for %d Jira image(s); configured model %s "
"does not support images",
fallback_model_id,
len(image_urls),
resolved_model_id,
)
resolved_model_id = fallback_model_id
image_model_override = (fallback_model_id, fallback_effort)
webapp.logger.info("Preparing %d image(s) for multimodal content", len(image_urls))
webapp.logger.debug("Image hosts: %s", [urlparse(u).hostname for u in image_urls])
async with httpx.AsyncClient(timeout=webapp.DEFAULT_HTTP_TIMEOUT) as client:
for image_url in image_urls:
image_block = await webapp.fetch_image_block(image_url, client)
if image_block:
content_blocks.append(image_block)
webapp.logger.info("Built %d content block(s) for prompt", len(content_blocks))
configurable: dict[str, Any] = {
"repo": repo_config,
"jira_issue": {
"key": issue_key,
"url": issue_url,
"project_key": project_key,
"issue_number": issue_number,
"title": title,
"triggering_user_name": user_name or "",
},
"user_email": user_email,
"source": "jira",
}
if mapped_login:
configurable["github_login"] = mapped_login
if image_model_override:
configurable["agent_model_id"] = image_model_override[0]
configurable["agent_effort"] = image_model_override[1]
await webapp.upsert_agent_thread_owner_metadata(
thread_id,
source="jira",
repo_config=repo_config,
github_login=mapped_login or "",
user_email=user_email or "",
title=title or issue_key or "Jira issue",
source_context={"jira_issue": configurable["jira_issue"]},
)
run = await webapp.dispatch_agent_run(
thread_id,
content_blocks,
configurable,
source="jira",
metadata=webapp._AGENT_VERSION_METADATA,
)
webapp.logger.info(
"LangGraph run dispatched for thread %s (run=%s)",
thread_id,
run.get("run_id") if isinstance(run, dict) else None,
)
await webapp.post_jira_trace_comment(issue_key, thread_id)

View file

@ -0,0 +1,305 @@
"""Atlassian Connect qsh vectors + JWT verification (auth boundary).
qsh correctness is a silent-auth-bypass surface, so the official Atlassian test
vector and the three endpoint vectors are pinned here. The JWT tests exercise
alg-pinning, signature, exp, issuer binding, and qsh binding.
"""
from __future__ import annotations
import time
from types import SimpleNamespace
import jwt
import pytest
from agent.utils import atlassian_connect as ac
_SECRET = "connect-shared-secret"
def _request(method: str, path: str, query: str = "", *, token: str | None = None) -> object:
headers = {"Authorization": f"JWT {token}"} if token else {}
return SimpleNamespace(
method=method,
url=SimpleNamespace(path=path, query=query),
headers=headers,
query_params={},
)
def _make_token(
*,
secret: str = _SECRET,
iss: str = "tenant-1",
alg: str = "HS256",
exp_delta: int = 180,
qsh: str | None = "auto",
method: str = "POST",
path: str = "/connect/webhook/comment-created",
query: str = "",
drop_exp: bool = False,
) -> str:
claims: dict = {"iss": iss}
if not drop_exp:
claims["exp"] = int(time.time()) + exp_delta
if qsh == "auto":
claims["qsh"] = ac.compute_qsh(method, path, query)
elif qsh is not None:
claims["qsh"] = qsh
return jwt.encode(claims, secret, algorithm=alg)
# --- qsh vectors -----------------------------------------------------------
def test_official_atlassian_qsh_vector() -> None:
canon = ac.canonical_request(
"GET",
"/path/to/service",
"zee_last=param&repeated=parameter 1&first=param&repeated=parameter 2",
)
assert canon == (
"GET&/path/to/service&first=param&repeated=parameter%201,parameter%202&zee_last=param"
)
@pytest.mark.parametrize(
("path", "expected"),
[
("/connect/installed", "72c0a77bd4d709a202e9b2561ed003fdb400318f7a1cfabe47576d1e1d5b5dd7"),
(
"/connect/uninstalled",
"ef0c0673ed4cf59a823d82cdc5c397c8643d79db724ce7d567342ea15e02acfe",
),
(
"/connect/webhook/comment-created",
"72e058a8906e894732718ec80dbbbf073640341b9ac6ed7cd86f887f23a00b4d",
),
],
)
def test_endpoint_qsh_vectors(path: str, expected: str) -> None:
assert ac.compute_qsh("POST", path, "") == expected
def test_qsh_drops_jwt_param_and_encodes_space_not_plus() -> None:
# jwt param is excluded; space must be %20 (never +).
with_jwt = ac.compute_qsh("GET", "/x", "a=b c&jwt=zzz")
without = ac.compute_qsh("GET", "/x", "a=b c")
assert with_jwt == without
assert "%20" in ac.canonical_request("GET", "/x", "a=b c")
assert "+" not in ac.canonical_request("GET", "/x", "a=b c")
# --- JWT verification ------------------------------------------------------
def test_valid_token_accepted() -> None:
token = _make_token()
req = _request("POST", "/connect/webhook/comment-created", token=token)
claims = ac.verify_connect_jwt(req, shared_secret=_SECRET)
assert claims is not None
assert claims["iss"] == "tenant-1"
def test_missing_token_rejected() -> None:
req = _request("POST", "/connect/webhook/comment-created")
assert ac.verify_connect_jwt(req, shared_secret=_SECRET) is None
def test_alg_none_rejected() -> None:
token = jwt.encode({"iss": "t", "exp": int(time.time()) + 60}, "", algorithm="none")
req = _request("POST", "/connect/webhook/comment-created", token=token)
assert ac.verify_connect_jwt(req, shared_secret=_SECRET) is None
def test_wrong_secret_rejected() -> None:
token = _make_token(secret="attacker-secret")
req = _request("POST", "/connect/webhook/comment-created", token=token)
assert ac.verify_connect_jwt(req, shared_secret=_SECRET) is None
def test_expired_token_rejected() -> None:
token = _make_token(exp_delta=-3600)
req = _request("POST", "/connect/webhook/comment-created", token=token)
assert ac.verify_connect_jwt(req, shared_secret=_SECRET) is None
def test_missing_secret_fails_closed() -> None:
token = _make_token()
req = _request("POST", "/connect/webhook/comment-created", token=token)
assert ac.verify_connect_jwt(req, shared_secret=None) is None
def test_issuer_binding_mismatch_rejected() -> None:
token = _make_token(iss="tenant-1")
req = _request("POST", "/connect/webhook/comment-created", token=token)
assert ac.verify_connect_jwt(req, shared_secret=_SECRET, expected_client_key="tenant-2") is None
def test_qsh_mismatch_rejected_cross_endpoint_replay() -> None:
# Token signed with the qsh for /installed, replayed at the webhook endpoint.
token = _make_token(path="/connect/installed")
req = _request("POST", "/connect/webhook/comment-created", token=token)
assert ac.verify_connect_jwt(req, shared_secret=_SECRET) is None
def test_missing_qsh_rejected_when_required() -> None:
token = _make_token(qsh=None)
req = _request("POST", "/connect/webhook/comment-created", token=token)
assert ac.verify_connect_jwt(req, shared_secret=_SECRET, qsh_required=True) is None
def test_missing_qsh_allowed_on_lifecycle_when_not_required() -> None:
token = _make_token(qsh=None)
req = _request("POST", "/connect/installed", token=token)
assert ac.verify_connect_jwt(req, shared_secret=_SECRET, qsh_required=False) is not None
def test_context_qsh_rejected() -> None:
token = _make_token(qsh="context-qsh")
req = _request("POST", "/connect/webhook/comment-created", token=token)
assert ac.verify_connect_jwt(req, shared_secret=_SECRET, qsh_required=False) is None
def test_missing_exp_rejected() -> None:
token = _make_token(drop_exp=True)
req = _request("POST", "/connect/webhook/comment-created", token=token)
assert ac.verify_connect_jwt(req, shared_secret=_SECRET) is None
# --- baseUrl host allowlist (first-install gate) ---------------------------
def test_base_url_allowlist(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(
ac, "CONNECT_EXPECTED_BASE_URL_HOSTS", frozenset({"seahaven.atlassian.net"})
)
assert ac.base_url_host_allowed("https://seahaven.atlassian.net/wiki") is True
assert ac.base_url_host_allowed("https://evil.example.com/wiki") is False
def test_base_url_allowlist_empty_fails_closed(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(ac, "CONNECT_EXPECTED_BASE_URL_HOSTS", frozenset())
assert ac.base_url_host_allowed("https://seahaven.atlassian.net/wiki") is False
# --- signed-install (asymmetric RS256) lifecycle verification ---------------
_AUD = "https://openswe.example.com"
def _rsa_keypair() -> tuple[str, str]:
from cryptography.hazmat.primitives import serialization
from cryptography.hazmat.primitives.asymmetric import rsa
key = rsa.generate_private_key(public_exponent=65537, key_size=2048)
priv = key.private_bytes(
serialization.Encoding.PEM,
serialization.PrivateFormat.PKCS8,
serialization.NoEncryption(),
).decode()
pub = (
key.public_key()
.public_bytes(serialization.Encoding.PEM, serialization.PublicFormat.SubjectPublicKeyInfo)
.decode()
)
return priv, pub
def _install_token(priv_pem: str, *, iss="tenant-1", aud=_AUD, exp_delta=180) -> str:
return jwt.encode(
{"iss": iss, "aud": aud, "exp": int(time.time()) + exp_delta},
priv_pem,
algorithm="RS256",
headers={"kid": "install-key-1"},
)
def _install_req(token: str) -> object:
return SimpleNamespace(
method="POST",
url=SimpleNamespace(path="/connect/installed", query=""),
headers={"Authorization": f"JWT {token}"},
query_params={},
)
def _run_install_verify(token, pub_pem, monkeypatch, *, expected_client_key=None):
import asyncio
from unittest.mock import AsyncMock, patch
monkeypatch.setattr(ac, "CONNECT_BASE_URL", _AUD)
with patch.object(ac, "_fetch_atlassian_public_key", new=AsyncMock(return_value=pub_pem)):
return asyncio.run(
ac.verify_asymmetric_install_jwt(
_install_req(token), expected_client_key=expected_client_key
)
)
def test_signed_install_valid_accepted(monkeypatch: pytest.MonkeyPatch) -> None:
priv, pub = _rsa_keypair()
claims = _run_install_verify(_install_token(priv), pub, monkeypatch)
assert claims is not None and claims["iss"] == "tenant-1"
def test_signed_install_hs256_rejected(monkeypatch: pytest.MonkeyPatch) -> None:
# An HS256 token (symmetric alg-confusion) must not pass asymmetric verify.
_priv, pub = _rsa_keypair()
hs = jwt.encode({"iss": "t", "aud": _AUD, "exp": int(time.time()) + 60}, "x", algorithm="HS256")
assert _run_install_verify(hs, pub, monkeypatch) is None
def test_signed_install_wrong_audience_rejected(monkeypatch: pytest.MonkeyPatch) -> None:
priv, pub = _rsa_keypair()
token = _install_token(priv, aud="https://some-other-app.example.com")
assert _run_install_verify(token, pub, monkeypatch) is None
def test_signed_install_wrong_key_rejected(monkeypatch: pytest.MonkeyPatch) -> None:
priv, _pub = _rsa_keypair()
_priv2, pub2 = _rsa_keypair()
assert _run_install_verify(_install_token(priv), pub2, monkeypatch) is None
def test_signed_install_expired_rejected(monkeypatch: pytest.MonkeyPatch) -> None:
priv, pub = _rsa_keypair()
assert _run_install_verify(_install_token(priv, exp_delta=-3600), pub, monkeypatch) is None
def test_signed_install_issuer_binding(monkeypatch: pytest.MonkeyPatch) -> None:
priv, pub = _rsa_keypair()
token = _install_token(priv, iss="tenant-1")
assert _run_install_verify(token, pub, monkeypatch, expected_client_key="tenant-2") is None
def test_signed_install_key_fetch_failure_rejected(monkeypatch: pytest.MonkeyPatch) -> None:
priv, _pub = _rsa_keypair()
assert _run_install_verify(_install_token(priv), None, monkeypatch) is None
async def test_fetch_public_key_rejects_malformed_kid() -> None:
# Defense-in-depth: a kid with path/URL chars is rejected before any fetch.
for bad in ["../../evil", "a/b", "http://evil.com", "a b", ""]:
assert await ac._fetch_atlassian_public_key(bad) is None
def test_client_key_allowed_fails_closed(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(ac, "CONNECT_EXPECTED_CLIENT_KEYS", frozenset())
assert ac.client_key_allowed("anything") is False
monkeypatch.setattr(ac, "CONNECT_EXPECTED_CLIENT_KEYS", frozenset({"ours"}))
assert ac.client_key_allowed("ours") is True
assert ac.client_key_allowed("attacker") is False
assert ac.client_key_allowed("") is False
def test_signed_install_no_base_url_rejected(monkeypatch: pytest.MonkeyPatch) -> None:
import asyncio
from unittest.mock import AsyncMock, patch
priv, pub = _rsa_keypair()
monkeypatch.setattr(ac, "CONNECT_BASE_URL", "")
with patch.object(ac, "_fetch_atlassian_public_key", new=AsyncMock(return_value=pub)):
result = asyncio.run(ac.verify_asymmetric_install_jwt(_install_req(_install_token(priv))))
assert result is None

View file

@ -138,6 +138,22 @@ async def test_linear_source_comments_on_issue(monkeypatch: pytest.MonkeyPatch)
assert comment.await_args.args[0] == "iss_1"
@pytest.mark.asyncio
async def test_jira_source_comments_on_issue(monkeypatch: pytest.MonkeyPatch) -> None:
client = _FakeClient({"source": "jira", "source_context": {"jira_issue": {"key": "PROJ-42"}}})
monkeypatch.setattr(completion, "langgraph_client", lambda: client)
comment = AsyncMock(return_value=True)
monkeypatch.setattr(completion, "comment_on_jira_issue", comment)
result = await completion.handle_run_completion(
{"thread_id": "t1", "run_id": "run-1", "status": "timeout"}
)
assert result["status"] == "ok"
comment.assert_awaited_once()
assert comment.await_args.args[0] == "PROJ-42"
@pytest.mark.asyncio
async def test_missing_thread_id_is_ignored() -> None:
result = await completion.handle_run_completion({"run_id": "run-1", "status": "error"})

View file

@ -0,0 +1,211 @@
"""Unit tests for the Confluence REST utilities and storage-format conversion."""
from __future__ import annotations
from typing import Any
import pytest
from agent.utils import confluence
# --- storage-format conversion ---------------------------------------------
def test_text_to_storage_wraps_blocks_in_paragraphs() -> None:
storage = confluence.text_to_storage("first block\n\nsecond block")
assert storage == "<p>first block</p><p>second block</p>"
def test_text_to_storage_escapes_html() -> None:
storage = confluence.text_to_storage("a < b & c > d")
assert storage == "<p>a &lt; b &amp; c &gt; d</p>"
def test_text_to_storage_empty_is_empty() -> None:
assert confluence.text_to_storage("") == ""
def test_storage_to_text_strips_tags() -> None:
assert confluence.storage_to_text("<p>hello <strong>world</strong></p>") == "hello world"
def test_storage_to_text_handles_none_and_empty() -> None:
assert confluence.storage_to_text("") == ""
def test_storage_to_text_unescapes_entities() -> None:
assert confluence.storage_to_text("<p>a &lt; b &amp; c</p>") == "a < b & c"
# --- Confluence REST utilities (mocked transport) ---------------------------
@pytest.fixture
def _confluence_env(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(confluence, "CONFLUENCE_BASE_URL", "https://seahaven.atlassian.net")
monkeypatch.setattr(confluence, "CONFLUENCE_EMAIL", "bot@seahavenind.com")
monkeypatch.setattr(confluence, "CONFLUENCE_API_TOKEN", "token")
def _mock_request(
monkeypatch: pytest.MonkeyPatch, responses: dict[str, Any] | list[dict[str, Any]]
) -> list[dict[str, Any]]:
calls: list[dict[str, Any]] = []
queue = responses if isinstance(responses, list) else None
async def fake_request(method: str, path: str, *, json=None, params=None):
calls.append({"method": method, "path": path, "json": json, "params": params})
if queue is not None:
return queue[len(calls) - 1]
return responses
monkeypatch.setattr(confluence, "_request", fake_request)
return calls
async def test_get_page_normalizes_fields(
_confluence_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
raw = {
"id": "123",
"title": "Architecture Map",
"body": {"storage": {"value": "<p>System overview</p>"}},
"version": {"number": 3},
"space": {"key": "IT"},
"_links": {"webui": "/spaces/IT/pages/123/Architecture+Map"},
}
_mock_request(monkeypatch, raw)
result = await confluence.get_page("123")
page = result["page"]
assert page["id"] == "123"
assert page["title"] == "Architecture Map"
assert page["body"] == "System overview"
assert page["version"] == 3
assert page["space_key"] == "IT"
assert page["url"] == (
"https://seahaven.atlassian.net/wiki/spaces/IT/pages/123/Architecture+Map"
)
async def test_create_page_builds_payload(
_confluence_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
calls = _mock_request(
monkeypatch,
{"id": "456", "title": "New Page", "_links": {"webui": "/spaces/IT/pages/456/New+Page"}},
)
result = await confluence.create_page("IT", "New Page", "hello world", parent_id="100")
assert result["success"] is True
assert result["page"]["id"] == "456"
sent = calls[0]["json"]
assert sent["type"] == "page"
assert sent["space"] == {"key": "IT"}
assert sent["title"] == "New Page"
assert sent["body"]["storage"]["value"] == "<p>hello world</p>"
assert sent["body"]["storage"]["representation"] == "storage"
assert sent["ancestors"] == [{"id": "100"}]
async def test_update_page_reads_current_then_increments_version(
_confluence_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
get_response = {
"id": "123",
"title": "Old Title",
"body": {"storage": {"value": "<p>old</p>"}},
"version": {"number": 5},
"space": {"key": "IT"},
"_links": {"webui": "/spaces/IT/pages/123/Old+Title"},
}
put_response = {
"id": "123",
"title": "Old Title",
"_links": {"webui": "/spaces/IT/pages/123/Old+Title"},
}
calls = _mock_request(monkeypatch, [get_response, put_response])
result = await confluence.update_page("123", body="new body")
assert result["success"] is True
assert calls[0]["method"] == "GET"
assert calls[1]["method"] == "PUT"
assert calls[1]["path"] == "/content/123"
sent = calls[1]["json"]
assert sent["version"]["number"] == 6
assert sent["title"] == "Old Title"
assert sent["body"]["storage"]["value"] == "<p>new body</p>"
async def test_add_comment_builds_container(
_confluence_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
calls = _mock_request(monkeypatch, {"id": "999"})
result = await confluence.add_comment("123", "great work")
assert result["success"] is True
sent = calls[0]["json"]
assert sent["type"] == "comment"
assert sent["container"] == {"id": "123", "type": "page"}
assert sent["body"]["storage"]["value"] == "<p>great work</p>"
async def test_search_normalizes_results(
_confluence_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
_mock_request(
monkeypatch,
{
"results": [
{
"id": "123",
"title": "Architecture Map",
"type": "page",
"_links": {"webui": "/spaces/IT/pages/123/Architecture+Map"},
}
]
},
)
result = await confluence.search('space = "IT"')
assert result["results"][0]["id"] == "123"
assert result["results"][0]["title"] == "Architecture Map"
assert result["results"][0]["type"] == "page"
async def test_request_without_env_returns_error(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(confluence, "CONFLUENCE_BASE_URL", "")
monkeypatch.setattr(confluence, "CONFLUENCE_API_TOKEN", "")
result = await confluence._request("GET", "/content/123")
assert "error" in result
async def test_get_page_propagates_error(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(confluence, "CONFLUENCE_BASE_URL", "")
monkeypatch.setattr(confluence, "CONFLUENCE_API_TOKEN", "")
result = await confluence.get_page("123")
assert "error" in result
async def test_get_page_encodes_page_id(
_confluence_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
calls = _mock_request(monkeypatch, {"id": "x", "title": "t"})
await confluence.get_page("../../admin/foo")
assert calls[0]["path"] == "/content/..%2F..%2Fadmin%2Ffoo"
async def test_update_page_encodes_page_id(
_confluence_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
get_resp = {
"id": "1",
"title": "T",
"body": {"storage": {"value": ""}},
"version": {"number": 1},
"space": {"key": "IT"},
"_links": {"webui": "/x"},
}
calls = _mock_request(
monkeypatch, [get_resp, {"id": "1", "title": "T", "_links": {"webui": "/x"}}]
)
await confluence.update_page("1?status=trashed", body="new")
# Both the internal get_page and the PUT must encode the id (no raw query).
assert calls[0]["path"] == "/content/1%3Fstatus%3Dtrashed"
assert calls[1]["method"] == "PUT"
assert calls[1]["path"] == "/content/1%3Fstatus%3Dtrashed"

View file

@ -0,0 +1,285 @@
"""Confluence Connect lifecycle overwrite guard, webhook corroboration, descriptor.
The lifecycle tests mock the JWT verifier (verified separately in
test_atlassian_connect.py) to isolate the accept/verify/store/reject logic —
especially that a failed re-install/uninstall leaves the stored secret intact.
"""
from __future__ import annotations
import asyncio
from types import SimpleNamespace
from typing import Any
from unittest.mock import AsyncMock, patch
from agent import webapp
from agent.utils import atlassian_connect as ac
from agent.webhooks import confluence as cf
def _req() -> object:
return SimpleNamespace(
method="POST",
url=SimpleNamespace(path="/connect/installed", query=""),
headers={},
query_params={},
)
def _install(fn, body: dict[str, Any], *, existing, verify_ok: bool = True):
puts: list[dict] = []
dels: list[str] = []
async def fake_get(_ck):
return existing
async def fake_put(client_key, shared_secret, base_url, product_type, *, first_install):
puts.append({"client_key": client_key, "secret": shared_secret, "first": first_install})
async def fake_del(client_key):
dels.append(client_key)
with (
patch.object(ac, "get_installation", new=AsyncMock(side_effect=fake_get)),
patch.object(ac, "put_installation", new=AsyncMock(side_effect=fake_put)),
patch.object(ac, "delete_installation", new=AsyncMock(side_effect=fake_del)),
patch.object(ac, "CONNECT_EXPECTED_CLIENT_KEYS", frozenset({"t"})),
patch.object(
ac,
"verify_asymmetric_install_jwt",
new=AsyncMock(return_value=({"iss": "t"} if verify_ok else None)),
),
):
code, _detail = asyncio.run(fn(_req(), body))
return code, puts, dels
# --- first install (trust-on-first-use, host-gated) ------------------------
def test_first_install_stores_secret() -> None:
code, puts, _ = _install(
cf.process_install,
{"clientKey": "t", "sharedSecret": "s1", "baseUrl": "https://x.atlassian.net"},
existing=None,
)
assert code == 204
assert puts == [{"client_key": "t", "secret": "s1", "first": True}]
def test_first_install_missing_secret_400() -> None:
code, puts, _ = _install(cf.process_install, {"clientKey": "t"}, existing=None)
assert code == 400
assert puts == []
def test_install_bad_signature_rejected() -> None:
# Even a first install now requires a valid Atlassian signature (no TOFU).
code, puts, _ = _install(
cf.process_install,
{"clientKey": "t", "sharedSecret": "s", "baseUrl": "https://x.atlassian.net"},
existing=None,
verify_ok=False,
)
assert code == 401
assert puts == []
def test_install_rejected_when_client_key_not_allowed() -> None:
# CONF-01: a valid Atlassian signature from a NON-allowlisted tenant (any
# attacker who installs the public descriptor on their own site) is rejected.
puts: list = []
async def fake_put(*a, **k):
puts.append(a)
with (
patch.object(ac, "get_installation", new=AsyncMock(return_value=None)),
patch.object(
ac, "verify_asymmetric_install_jwt", new=AsyncMock(return_value={"iss": "attacker"})
),
patch.object(ac, "CONNECT_EXPECTED_CLIENT_KEYS", frozenset({"our-tenant"})),
patch.object(ac, "put_installation", new=AsyncMock(side_effect=fake_put)),
):
code, _ = asyncio.run(
cf.process_install(
_req(),
{
"clientKey": "attacker",
"sharedSecret": "s",
"baseUrl": "https://attacker.atlassian.net",
},
)
)
assert code == 403
assert puts == []
def test_install_bad_host_rejected_when_allowlist_configured() -> None:
# Defense-in-depth host check (only enforced when CONNECT_EXPECTED_BASE_URL set).
puts: list = []
async def fake_put(*a, **k):
puts.append(a)
with (
patch.object(ac, "get_installation", new=AsyncMock(return_value=None)),
patch.object(ac, "verify_asymmetric_install_jwt", new=AsyncMock(return_value={"iss": "t"})),
patch.object(ac, "CONNECT_EXPECTED_CLIENT_KEYS", frozenset({"t"})),
patch.object(ac, "CONNECT_EXPECTED_BASE_URL_HOSTS", frozenset({"x.atlassian.net"})),
patch.object(ac, "base_url_host_allowed", return_value=False),
patch.object(ac, "put_installation", new=AsyncMock(side_effect=fake_put)),
):
code, _ = asyncio.run(
cf.process_install(
_req(), {"clientKey": "t", "sharedSecret": "s", "baseUrl": "https://evil.com"}
)
)
assert code == 403
assert puts == []
# --- re-install overwrite guard (the security-critical path) ---------------
_EXISTING = {"client_key": "t", "shared_secret": "stored-secret"}
def test_reinstall_bad_jwt_preserves_stored_secret() -> None:
code, puts, _ = _install(
cf.process_install,
{"clientKey": "t", "sharedSecret": "attacker", "baseUrl": "https://x.atlassian.net"},
existing=_EXISTING,
verify_ok=False,
)
assert code == 401
assert puts == [] # stored secret NOT overwritten
def test_reinstall_valid_jwt_overwrites() -> None:
code, puts, _ = _install(
cf.process_install,
{"clientKey": "t", "sharedSecret": "rotated", "baseUrl": "https://x.atlassian.net"},
existing=_EXISTING,
verify_ok=True,
)
assert code == 204
assert puts == [{"client_key": "t", "secret": "rotated", "first": False}]
# --- uninstall guard -------------------------------------------------------
def test_uninstall_bad_jwt_keeps_record() -> None:
code, _puts, dels = _install(
cf.process_uninstall, {"clientKey": "t"}, existing=_EXISTING, verify_ok=False
)
assert code == 401
assert dels == []
def test_uninstall_valid_jwt_deletes() -> None:
code, _puts, dels = _install(
cf.process_uninstall, {"clientKey": "t"}, existing=_EXISTING, verify_ok=True
)
assert code == 204
assert dels == ["t"]
def test_uninstall_no_record_idempotent() -> None:
code, _puts, dels = _install(cf.process_uninstall, {"clientKey": "t"}, existing=None)
assert code == 204
assert dels == []
# --- webhook corroboration (identity/body from server, not payload) --------
def _run_comment(payload: dict, server_comment: dict | None, *, active: set[str] | None = None):
captured: dict = {}
active = {"jane"} if active is None else active
async def fake_dispatch(
thread_id, content, configurable, *, source, metadata=None, client=None
):
captured["configurable"] = configurable
captured["source"] = source
return {"run_id": "r1"}
with (
patch.object(
webapp, "fetch_confluence_comment", new=AsyncMock(return_value=server_comment)
),
patch.object(
webapp, "fetch_confluence_page", new=AsyncMock(return_value={"title": "P", "url": "u"})
),
patch.object(webapp, "get_confluence_user_email", new=AsyncMock(return_value="jane@x.com")),
patch.object(webapp, "resolve_login_from_email_async", new=AsyncMock(return_value="jane")),
patch.object(webapp, "is_login_mapped", side_effect=lambda login: login in active),
patch.object(
webapp,
"get_repo_config_from_confluence_mapping",
return_value={"owner": "o", "name": "n"},
),
patch.object(webapp, "_is_repo_allowed", return_value=True),
patch.object(webapp, "generate_thread_id_from_confluence_comment", return_value="th-1"),
patch.object(
webapp, "upsert_agent_thread_owner_metadata", new=AsyncMock(return_value=None)
),
patch.object(webapp, "dispatch_agent_run", side_effect=fake_dispatch),
):
asyncio.run(cf.process_confluence_comment(payload))
return captured
def _server_comment(
*, account_id="real", name="Real", body="@openswe fix it", space="IT", page="99"
):
return {
"author": {"account_id": account_id, "name": name},
"body": body,
"page_id": page,
"space_key": space,
}
def test_webhook_uses_server_comment_not_payload() -> None:
payload = {"comment": {"id": "555"}, "userAccountId": "victim", "body": "benign"}
cap = _run_comment(payload, _server_comment())
assert cap["source"] == "confluence"
conf = cap["configurable"]["confluence"]
assert conf["comment_id"] == "555"
assert conf["space_key"] == "IT"
assert cap["configurable"]["github_login"] == "jane"
def test_webhook_uncorroborated_comment_dropped() -> None:
cap = _run_comment({"comment": {"id": "555"}}, None)
assert cap == {} # no dispatch
def test_webhook_without_mention_dropped() -> None:
cap = _run_comment({"comment": {"id": "555"}}, _server_comment(body="just a normal comment"))
assert cap == {}
def test_webhook_pending_mapping_unattributed() -> None:
cap = _run_comment({"comment": {"id": "555"}}, _server_comment(), active=set())
assert "github_login" not in cap["configurable"]
def test_webhook_bot_own_comment_dropped() -> None:
# CONF-02: a comment authored by the app's own account is ignored (no loop).
with patch.object(cf, "CONFLUENCE_BOT_ACCOUNT_ID", "bot-acct"):
cap = _run_comment({"comment": {"id": "555"}}, _server_comment(account_id="bot-acct"))
assert cap == {}
# --- descriptor ------------------------------------------------------------
def test_descriptor_signed_install_true_and_read_scope() -> None:
desc = asyncio.run(webapp.connect_descriptor())
assert desc["apiMigrations"]["signed-install"] is True
assert desc["scopes"] == ["READ"]
assert desc["authentication"]["type"] == "jwt"
assert desc["modules"]["webhooks"][0]["event"] == "comment_created"

274
tests/test_jira_utils.py Normal file
View file

@ -0,0 +1,274 @@
"""Unit tests for the Jira REST utilities and ADF conversion."""
from __future__ import annotations
from typing import Any
import pytest
from agent.utils import adf, jira
# --- ADF conversion --------------------------------------------------------
def test_adf_to_markdown_handles_none_and_empty() -> None:
assert adf.adf_to_markdown(None) == ""
assert adf.adf_to_markdown({}) == ""
assert adf.adf_to_markdown("not a dict") == ""
def test_adf_to_markdown_paragraphs_marks_and_links() -> None:
doc = {
"type": "doc",
"version": 1,
"content": [
{
"type": "paragraph",
"content": [
{"type": "text", "text": "Hello "},
{"type": "text", "text": "world", "marks": [{"type": "strong"}]},
],
},
{
"type": "paragraph",
"content": [
{
"type": "text",
"text": "a link",
"marks": [{"type": "link", "attrs": {"href": "https://x.com"}}],
}
],
},
],
}
md = adf.adf_to_markdown(doc)
assert "Hello **world**" in md
assert "[a link](https://x.com)" in md
def test_adf_to_markdown_bullet_and_code() -> None:
doc = {
"type": "doc",
"content": [
{
"type": "bulletList",
"content": [
{
"type": "listItem",
"content": [
{"type": "paragraph", "content": [{"type": "text", "text": "one"}]}
],
},
{
"type": "listItem",
"content": [
{"type": "paragraph", "content": [{"type": "text", "text": "two"}]}
],
},
],
},
{
"type": "codeBlock",
"attrs": {"language": "python"},
"content": [{"type": "text", "text": "print(1)"}],
},
],
}
md = adf.adf_to_markdown(doc)
assert "- one" in md
assert "- two" in md
assert "```python" in md
assert "print(1)" in md
def test_markdown_to_adf_structure() -> None:
doc = adf.markdown_to_adf("first block\n\nsecond block")
assert doc["type"] == "doc"
assert doc["version"] == 1
assert len(doc["content"]) == 2
assert doc["content"][0]["content"][0]["text"] == "first block"
def test_markdown_to_adf_empty_is_valid_doc() -> None:
doc = adf.markdown_to_adf("")
assert doc["type"] == "doc"
assert doc["content"] == [{"type": "paragraph", "content": []}]
def test_markdown_to_adf_multiline_block_uses_hardbreaks() -> None:
doc = adf.markdown_to_adf("line one\nline two")
para = doc["content"][0]["content"]
assert {"type": "hardBreak"} in para
# --- Jira REST utilities (mocked transport) --------------------------------
@pytest.fixture
def _jira_env(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(jira, "JIRA_BASE_URL", "https://seahaven.atlassian.net")
monkeypatch.setattr(jira, "JIRA_EMAIL", "bot@seahavenind.com")
monkeypatch.setattr(jira, "JIRA_API_TOKEN", "token")
def _mock_request(
monkeypatch: pytest.MonkeyPatch, response: dict[str, Any]
) -> list[dict[str, Any]]:
calls: list[dict[str, Any]] = []
async def fake_request(method: str, path: str, *, json=None, params=None):
calls.append({"method": method, "path": path, "json": json, "params": params})
return response
monkeypatch.setattr(jira, "_request", fake_request)
return calls
async def test_get_issue_normalizes_fields(
_jira_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
raw = {
"key": "PROJ-123",
"id": "10001",
"fields": {
"summary": "Fix the bug",
"description": {
"type": "doc",
"content": [
{"type": "paragraph", "content": [{"type": "text", "text": "details"}]}
],
},
"status": {"name": "In Progress"},
"assignee": {"displayName": "Ada", "emailAddress": "ada@x.com", "accountId": "acc1"},
"priority": {"name": "High"},
"labels": ["backend"],
"project": {"key": "PROJ", "name": "Project"},
"issuetype": {"name": "Bug"},
},
}
_mock_request(monkeypatch, raw)
result = await jira.get_issue("PROJ-123")
issue = result["issue"]
assert issue["key"] == "PROJ-123"
assert issue["title"] == "Fix the bug"
assert issue["description"] == "details"
assert issue["assignee"]["email"] == "ada@x.com"
assert issue["project_key"] == "PROJ"
assert issue["url"] == "https://seahaven.atlassian.net/browse/PROJ-123"
async def test_comment_on_issue_success(_jira_env: None, monkeypatch: pytest.MonkeyPatch) -> None:
calls = _mock_request(monkeypatch, {"id": "5001"})
ok = await jira.comment_on_issue("PROJ-1", "done, see PR")
assert ok is True
assert calls[0]["method"] == "POST"
assert calls[0]["path"] == "/issue/PROJ-1/comment"
# Body must be ADF, not raw markdown.
assert calls[0]["json"]["body"]["type"] == "doc"
async def test_comment_on_issue_error_returns_false(
_jira_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
_mock_request(monkeypatch, {"error": "boom"})
assert await jira.comment_on_issue("PROJ-1", "x") is False
async def test_get_issue_comments_normalizes(
_jira_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
_mock_request(
monkeypatch,
{
"comments": [
{
"id": "1",
"author": {"displayName": "Ada", "emailAddress": "ada@x.com", "accountId": "a"},
"body": {
"type": "doc",
"content": [
{"type": "paragraph", "content": [{"type": "text", "text": "hi"}]}
],
},
}
]
},
)
result = await jira.get_issue_comments("PROJ-1")
assert result["comments"][0]["body"] == "hi"
assert result["comments"][0]["author"]["email"] == "ada@x.com"
async def test_create_issue_builds_fields(_jira_env: None, monkeypatch: pytest.MonkeyPatch) -> None:
calls = _mock_request(monkeypatch, {"key": "PROJ-9", "id": "999"})
result = await jira.create_issue("PROJ", "New thing", description="body", priority="High")
assert result["success"] is True
assert result["issue"]["key"] == "PROJ-9"
sent = calls[0]["json"]["fields"]
assert sent["project"] == {"key": "PROJ"}
assert sent["summary"] == "New thing"
assert sent["description"]["type"] == "doc"
assert sent["priority"] == {"name": "High"}
async def test_update_issue_no_fields_errors(_jira_env: None) -> None:
result = await jira.update_issue("PROJ-1")
assert result["error"] == "No fields to update"
async def test_request_without_env_returns_error(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(jira, "JIRA_BASE_URL", "")
monkeypatch.setattr(jira, "JIRA_API_TOKEN", "")
result = await jira._request("GET", "/issue/PROJ-1")
assert "error" in result
# --- issue_key validation + path-segment encoding (INJ hardening) ----------
def test_is_valid_issue_key() -> None:
assert jira.is_valid_issue_key("PROJ-123")
assert jira.is_valid_issue_key("OS-1")
assert not jira.is_valid_issue_key("")
assert not jira.is_valid_issue_key("../../../../rest/api/2/permissions")
assert not jira.is_valid_issue_key("PROJ-1?expand=x")
assert not jira.is_valid_issue_key("PROJ-1/comment")
assert not jira.is_valid_issue_key("1-PROJ")
async def test_get_issue_percent_encodes_path(
_jira_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
# Even if a traversal key reaches the util, the path segment is encoded so it
# cannot climb out of /issue/ or inject a query.
calls = _mock_request(monkeypatch, {"key": "x", "fields": {}})
await jira.get_issue("../../../../rest/api/2/permissions")
assert calls[0]["path"] == "/issue/..%2F..%2F..%2F..%2Frest%2Fapi%2F2%2Fpermissions"
async def test_comment_on_issue_encodes_path(
_jira_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
calls = _mock_request(monkeypatch, {"id": "1"})
await jira.comment_on_issue("ABC-1?expand=evil", "hi")
assert calls[0]["path"] == "/issue/ABC-1%3Fexpand%3Devil/comment"
async def test_get_comment_fetches_single_comment(
_jira_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
calls = _mock_request(
monkeypatch,
{
"id": "10050",
"author": {"displayName": "Ada", "emailAddress": "ada@x.com", "accountId": "acc"},
"body": {
"type": "doc",
"content": [{"type": "paragraph", "content": [{"type": "text", "text": "hi"}]}],
},
},
)
result = await jira.get_comment("PROJ-1", "10050")
assert calls[0]["path"] == "/issue/PROJ-1/comment/10050"
assert result["comment"]["author"]["account_id"] == "acc"
assert result["comment"]["body"] == "hi"

View file

@ -0,0 +1,185 @@
"""Tests for Jira webhook PR author linking and repo-mapping cascade."""
from __future__ import annotations
import asyncio
from typing import Any
from unittest.mock import AsyncMock, patch
from agent import webapp
from agent.webhooks import jira as jira_webhook
def _full_issue(*, title: str = "Fix the flaky test") -> dict:
return {
"key": "PROJ-42",
"id": "10001",
"title": title,
"description": "Do the thing",
"url": "https://seahaven.atlassian.net/browse/PROJ-42",
"project_key": "PROJ",
}
def _issue_data(*, account_id: str | None, email: str | None, name: str = "Jane") -> dict:
# jira_webhook resolves account_id -> email once and attaches it to
# comment_author before dispatch (see webapp.jira_webhook).
return {
"key": "PROJ-42",
"project_key": "PROJ",
"triggering_comment": "@openswe fix this",
"triggering_comment_id": "10050",
"comment_author": {"account_id": account_id, "email": email, "name": name},
}
def _run_process(
issue_data: dict,
repo_config: dict[str, str],
*,
active_logins: set[str] | None = None,
) -> tuple[dict, dict, str | None]:
captured: dict[str, Any] = {}
active = {"jane"} if active_logins is None else active_logins
async def fake_dispatch(
thread_id, content, configurable, *, source, metadata=None, client=None
):
captured["configurable"] = configurable
return {"run_id": "run-1"}
async def fake_upsert(
thread_id,
*,
source,
repo_config=None,
github_login="",
user_email="",
title="",
source_context=None,
):
captured["upsert"] = {"github_login": github_login, "user_email": user_email}
return None
async def fake_resolve_login(email):
captured["resolved_email"] = email
return "jane" if email == "jane@example.com" else None
with (
patch.object(
jira_webhook.webapp, "generate_thread_id_from_jira_issue", return_value="thread-1"
),
patch.object(
jira_webhook.webapp,
"fetch_jira_issue_details",
new_callable=AsyncMock,
return_value=_full_issue(),
),
patch.object(
jira_webhook.webapp,
"fetch_jira_issue_comments",
new_callable=AsyncMock,
return_value=[],
),
patch.object(
jira_webhook.webapp, "resolve_login_from_email_async", side_effect=fake_resolve_login
),
patch.object(
jira_webhook.webapp, "is_login_mapped", side_effect=lambda login: login in active
),
patch.object(jira_webhook.webapp, "dispatch_agent_run", side_effect=fake_dispatch),
patch.object(
jira_webhook.webapp, "upsert_agent_thread_owner_metadata", side_effect=fake_upsert
),
patch.object(jira_webhook.webapp, "post_jira_trace_comment", new_callable=AsyncMock),
):
asyncio.run(jira_webhook.process_jira_issue(issue_data, repo_config))
return (
captured.get("configurable", {}),
captured.get("upsert", {}),
captured.get("resolved_email"),
)
def test_jira_configurable_carries_github_login() -> None:
configurable, _upsert, resolved_email = _run_process(
_issue_data(account_id="acc-1", email="jane@example.com"),
{"owner": "langchain-ai", "name": "open-swe"},
)
assert resolved_email == "jane@example.com"
assert configurable["source"] == "jira"
assert configurable["github_login"] == "jane"
assert configurable["user_email"] == "jane@example.com"
assert configurable["jira_issue"]["key"] == "PROJ-42"
assert configurable["jira_issue"]["project_key"] == "PROJ"
assert configurable["jira_issue"]["issue_number"] == "42"
def test_jira_upsert_tags_thread_with_login() -> None:
_configurable, upsert, _email = _run_process(
_issue_data(account_id="acc-1", email="jane@example.com"),
{"owner": "langchain-ai", "name": "open-swe"},
)
assert upsert["github_login"] == "jane"
assert upsert["user_email"] == "jane@example.com"
def test_jira_omits_login_when_unmapped() -> None:
configurable, upsert, resolved_email = _run_process(
_issue_data(account_id="acc-2", email="nobody@example.com"),
{"owner": "langchain-ai", "name": "open-swe"},
)
assert resolved_email == "nobody@example.com"
assert "github_login" not in configurable
assert upsert["github_login"] == ""
def test_jira_omits_login_when_mapping_not_active() -> None:
# A resolvable email whose mapping is pending/inactive must not be attributed.
configurable, upsert, resolved_email = _run_process(
_issue_data(account_id="acc-1", email="jane@example.com"),
{"owner": "langchain-ai", "name": "open-swe"},
active_logins=set(),
)
assert resolved_email == "jane@example.com"
assert "github_login" not in configurable
assert upsert["github_login"] == ""
def test_jira_omits_login_when_no_email_resolved() -> None:
configurable, upsert, resolved_email = _run_process(
_issue_data(account_id=None, email=None),
{"owner": "langchain-ai", "name": "open-swe"},
)
assert resolved_email is None
assert "github_login" not in configurable
assert upsert["github_login"] == ""
def test_repo_cascade_uses_project_mapping(monkeypatch) -> None:
monkeypatch.setattr(
webapp, "JIRA_PROJECT_TO_REPO", {"PROJ": {"owner": "acme", "name": "widgets"}}
)
assert webapp.get_repo_config_from_jira_mapping("PROJ") == {"owner": "acme", "name": "widgets"}
def test_repo_cascade_falls_back_to_default_repo(monkeypatch) -> None:
monkeypatch.setattr(webapp, "JIRA_PROJECT_TO_REPO", {})
monkeypatch.setattr(webapp, "DEFAULT_REPO_OWNER", "langchain-ai")
monkeypatch.setattr(webapp, "DEFAULT_REPO_NAME", "open-swe")
assert webapp.get_repo_config_from_jira_mapping("UNKNOWN") == {
"owner": "langchain-ai",
"name": "open-swe",
}
def test_repo_cascade_empty_without_default(monkeypatch) -> None:
monkeypatch.setattr(webapp, "JIRA_PROJECT_TO_REPO", {})
monkeypatch.setattr(webapp, "DEFAULT_REPO_NAME", "")
assert webapp.get_repo_config_from_jira_mapping("UNKNOWN") == {}

View file

@ -0,0 +1,142 @@
"""Route-level corroboration + input validation for /webhooks/jira.
These cover the hardening from the Phase 2 security review: the unsigned webhook
body is only a pointer (issue_key + comment_id), and the triggering comment's
author and text are re-fetched from Jira server-side. A payload-claimed author
must never be trusted, a malformed issue_key must be rejected, and a comment
that can't be corroborated must be rejected.
"""
from __future__ import annotations
import asyncio
import json
from contextlib import ExitStack
from typing import Any
from unittest.mock import AsyncMock, patch
from agent import webapp
class _FakeRequest:
def __init__(self, body: bytes, headers: dict[str, str] | None = None) -> None:
self.headers = headers or {}
self._body = body
async def body(self) -> bytes:
return self._body
class _FakeBackgroundTasks:
def __init__(self) -> None:
self.tasks: list[tuple[Any, tuple, dict]] = []
def add_task(self, func: Any, *args: Any, **kwargs: Any) -> None:
self.tasks.append((func, args, kwargs))
def _call(
payload: dict[str, Any],
*,
server_comment: dict[str, Any] | None,
email: str | None = "real@example.com",
) -> tuple[dict[str, str], _FakeBackgroundTasks, AsyncMock]:
req = _FakeRequest(json.dumps(payload).encode())
bg = _FakeBackgroundTasks()
get_email = AsyncMock(return_value=email)
with ExitStack() as stack:
stack.enter_context(patch.object(webapp, "verify_jira_secret", return_value=True))
stack.enter_context(
patch.object(webapp, "fetch_jira_comment", new=AsyncMock(return_value=server_comment))
)
stack.enter_context(patch.object(webapp, "get_jira_user_email", new=get_email))
stack.enter_context(
patch.object(webapp, "resolve_login_from_email_async", new=AsyncMock(return_value=None))
)
stack.enter_context(
patch.object(webapp, "get_profile_default_repo", new=AsyncMock(return_value=None))
)
stack.enter_context(
patch.object(
webapp,
"get_repo_config_from_jira_mapping",
return_value={"owner": "langchain-ai", "name": "open-swe"},
)
)
stack.enter_context(patch.object(webapp, "_is_repo_allowed", return_value=True))
result = asyncio.run(webapp.jira_webhook(req, bg))
return result, bg, get_email
def _server_comment(*, account_id: str, name: str, body: str) -> dict[str, Any]:
return {"id": "10050", "body": body, "author": {"account_id": account_id, "name": name}}
def test_malformed_issue_key_rejected() -> None:
result, bg, _ = _call(
{"issue_key": "../../../../rest/api/2/myself", "comment_id": "1"},
server_comment=None,
)
assert result["status"] == "ignored"
assert "issue key" in result["reason"].lower()
assert bg.tasks == []
def test_missing_comment_id_rejected() -> None:
result, bg, _ = _call({"issue_key": "PROJ-42"}, server_comment=None)
assert result["status"] == "ignored"
assert bg.tasks == []
def test_uncorroborated_comment_rejected() -> None:
# fetch_jira_comment returns None (nonexistent / forged) -> hard reject.
result, bg, _ = _call(
{"issue_key": "PROJ-42", "comment_id": "10050", "comment_body": "@openswe do it"},
server_comment=None,
)
assert result["status"] == "ignored"
assert bg.tasks == []
def test_identity_and_body_come_from_server_not_payload() -> None:
# Payload claims a victim's account + benign body; the REAL comment (server)
# has a different author and the actual trigger text. The scheduled task must
# carry the server author, and email lookup must use the server account id.
payload = {
"issue_key": "PROJ-42",
"comment_id": "10050",
"comment_author_account_id": "victim-account-id",
"comment_author_display_name": "Victim",
"comment_body": "totally benign",
}
server = _server_comment(
account_id="real-author-id", name="Real Author", body="@openswe fix the bug"
)
result, bg, get_email = _call(payload, server_comment=server)
assert result["status"] == "accepted"
assert len(bg.tasks) == 1
_func, (issue_data, _repo), _kw = bg.tasks[0]
# Server author wins; payload's victim account is never used.
assert issue_data["comment_author"]["account_id"] == "real-author-id"
assert issue_data["comment_author"]["name"] == "Real Author"
assert issue_data["triggering_comment"] == "@openswe fix the bug"
get_email.assert_awaited_once_with("real-author-id")
def test_project_key_derived_from_issue_key() -> None:
payload = {"issue_key": "OSPROJ-7", "comment_id": "10050", "project_key": "ATTACKER-INJECTED"}
server = _server_comment(account_id="a", name="A", body="@openswe go")
result, bg, _ = _call(payload, server_comment=server)
assert result["status"] == "accepted"
_func, (issue_data, _repo), _kw = bg.tasks[0]
assert issue_data["project_key"] == "OSPROJ"
def test_server_comment_without_mention_ignored() -> None:
# The @openswe check runs on the authoritative server body, not the payload.
payload = {"issue_key": "PROJ-42", "comment_id": "10050", "comment_body": "@openswe do it"}
server = _server_comment(account_id="a", name="A", body="just a normal comment")
result, bg, _ = _call(payload, server_comment=server)
assert result["status"] == "ignored"
assert bg.tasks == []

View file

@ -0,0 +1,133 @@
"""Shared-secret verification for the Jira Automation webhook (AUTHZ)."""
from __future__ import annotations
import hashlib
import hmac
import json
from datetime import UTC, datetime
from types import SimpleNamespace
import pytest
from agent import webapp
_SECRET = "jira-automation-secret"
def _signed_body(secret: str, *, fresh: bool = True) -> tuple[bytes, str]:
ts_ms = datetime.now(UTC).timestamp() * 1000
if not fresh:
ts_ms -= (webapp.JIRA_WEBHOOK_MAX_AGE_SECONDS + 60) * 1000
body = json.dumps({"issue_key": "PROJ-1", "timestamp": ts_ms}).encode()
sig = hmac.new(secret.encode(), body, hashlib.sha256).hexdigest()
return body, sig
def test_valid_secret_accepted(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_SECRET", _SECRET)
headers = {"X-Automation-Webhook-Token": _SECRET}
assert webapp.verify_jira_secret(headers) is True
def test_wrong_secret_rejected(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_SECRET", _SECRET)
headers = {"X-Automation-Webhook-Token": "wrong-token"}
assert webapp.verify_jira_secret(headers) is False
def test_missing_header_rejected(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_SECRET", _SECRET)
assert webapp.verify_jira_secret({}) is False
def test_empty_header_rejected(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_SECRET", _SECRET)
assert webapp.verify_jira_secret({"X-Automation-Webhook-Token": ""}) is False
def test_unset_env_fails_closed(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_SECRET", "")
headers = {"X-Automation-Webhook-Token": _SECRET}
assert webapp.verify_jira_secret(headers) is False
# --- Opt-in HMAC body signature + timestamp (JIRA_WEBHOOK_REQUIRE_SIGNATURE) ---
def test_signature_check_is_noop_when_disabled(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_REQUIRE_SIGNATURE", False)
assert webapp.verify_jira_signature(b"{}", {}) is True
def test_valid_signature_and_fresh_timestamp_accepted(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_REQUIRE_SIGNATURE", True)
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_SECRET", _SECRET)
body, sig = _signed_body(_SECRET)
assert webapp.verify_jira_signature(body, {"X-Openswe-Signature": sig}) is True
def test_missing_signature_rejected_when_required(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_REQUIRE_SIGNATURE", True)
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_SECRET", _SECRET)
body, _sig = _signed_body(_SECRET)
assert webapp.verify_jira_signature(body, {}) is False
def test_wrong_signature_rejected(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_REQUIRE_SIGNATURE", True)
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_SECRET", _SECRET)
body, _sig = _signed_body(_SECRET)
assert webapp.verify_jira_signature(body, {"X-Openswe-Signature": "deadbeef"}) is False
def test_stale_timestamp_rejected(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_REQUIRE_SIGNATURE", True)
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_SECRET", _SECRET)
body, sig = _signed_body(_SECRET, fresh=False)
assert webapp.verify_jira_signature(body, {"X-Openswe-Signature": sig}) is False
# --- Opt-in source-IP allowlist (JIRA_WEBHOOK_IP_ALLOWLIST) ---
def _req(host: str | None) -> object:
client = None if host is None else SimpleNamespace(host=host)
return SimpleNamespace(client=client)
def test_ip_check_is_noop_when_disabled(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_IP_ALLOWLIST", ())
assert webapp.verify_jira_source_ip(_req("9.9.9.9")) is True
def test_ip_in_allowlist_accepted(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_IP_ALLOWLIST", ("10.0.0.0/24",))
assert webapp.verify_jira_source_ip(_req("10.0.0.5")) is True
def test_ip_not_in_allowlist_rejected(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_IP_ALLOWLIST", ("10.0.0.0/24",))
assert webapp.verify_jira_source_ip(_req("192.168.1.1")) is False
def test_ip_missing_client_rejected(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "JIRA_WEBHOOK_IP_ALLOWLIST", ("10.0.0.0/24",))
assert webapp.verify_jira_source_ip(_req(None)) is False
# --- Fail-closed repo allowlist (REQUIRE_REPO_ALLOWLIST) ---
def test_empty_allowlist_allows_all_by_default(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "ALLOWED_GITHUB_ORGS", frozenset())
monkeypatch.setattr(webapp, "ALLOWED_GITHUB_REPOS", frozenset())
monkeypatch.setattr(webapp, "REQUIRE_REPO_ALLOWLIST", False)
assert webapp._is_repo_allowed({"owner": "anyone", "name": "anything"}) is True
def test_empty_allowlist_fails_closed_when_required(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(webapp, "ALLOWED_GITHUB_ORGS", frozenset())
monkeypatch.setattr(webapp, "ALLOWED_GITHUB_REPOS", frozenset())
monkeypatch.setattr(webapp, "REQUIRE_REPO_ALLOWLIST", True)
assert webapp._is_repo_allowed({"owner": "anyone", "name": "anything"}) is False