This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team/agent_team/db/schema.sql
Adam Moussa f916c03818
feat(agent-team): durable GitHub-issue intake de-dup + intake hardening (#32)
* feat(agent-team): durable GitHub-issue intake de-dup (schema v2)

The intake poller de-duped ingested issues in an in-memory set that does
not survive a process restart. A scheduled/cron intake (each run a fresh
process) would therefore re-ingest every still-open labeled issue on every
run and spawn duplicate pipeline tasks. Since the box is read-only (no write
token to remove the intake label), durable de-dup is the only correct guard.

- schema v2: new ingested_issues(source, issue_id, ingested_at) table +
  issue_already_ingested / record_issue_ingested helpers; migrate() adds the
  table to a legacy v1 DB and restamps; init_db creates it.
- github_intake: pluggable IngestStore seam (in-memory default preserved for
  tests/one-off; durable build_ledger_ingest_store for production). Record is
  after start_task succeeds, so a failed intake stays retryable.
- run-team.py intake-github wires the ledger store keyed by github:owner/repo,
  making a scheduled timer idempotent across runs.

978 tests pass (ruff clean).

* fix(agent-team): harden github intake per /sh-security-review (CWE-918, idempotency)

Fixes from the high-recall detector fan-out on the durable-dedup change:

- INTAKE-LOGIC-01 (idempotency): switch the IngestStore seam from
  check-then-record (seen/mark) to claim-then-do (claim/release). The id is
  now reserved BEFORE the non-idempotent start_task side effect, so a crash in
  that window cannot re-spawn a duplicate task on the next run; a raising
  start_task releases the claim so transient failures stay retryable. Adds
  delete_issue_ingested to the schema layer for the release path.
- INTAKE-SSRF-001 / INTAKE-PATHSPLICE-002 (CWE-918) in build_default_issue_client:
  drop the caller-overridable api_root (hardcode GITHUB_API_ROOT) and validate
  owner/repo against an anchored charset before splicing them into the
  token-bearing API URL — mirrors the sibling ci_fetcher BLOCK-3/FIX-3 fixes.

Tests cover cross-process duplicate prevention, release-on-failure retry, and
the owner/repo + api_root rejection. 982 tests pass, ruff clean.

Follow-up (pre-existing, not introduced here): the label-only intake has no
author allowlist (cf. AGENT_TEAM_SLACK_OWNER_IDS on the Slack listener); the
Slack answer gate bounds the blast radius. Track as separate hardening.
2026-06-22 17:20:43 -04:00

87 lines
3.8 KiB
SQL

-- R720 agent-team durable SQLite schema (design §3.3.1, §6.7).
--
-- This file holds the raw DDL statements ONLY. The authoritative copies live
-- as string constants in agent_team/db/schema.py; this companion file mirrors
-- them verbatim for tooling / direct inspection. SQL DDL lives only in these
-- two places.
--
-- The LangGraph SqliteSaver checkpointer creates its OWN tables against this
-- same database file/connection; those are intentionally NOT declared here.
-- pending_questions: the durable human-interaction lifecycle ledger.
-- The LangGraph checkpoint holds graph state; this table holds the question
-- lifecycle (delivery, duplicate/late answers, expiry) and is what delivery,
-- the responder, and restart recovery read. Every race resolves via an atomic
-- compare-and-set against the `status` column under BEGIN IMMEDIATE.
CREATE TABLE IF NOT EXISTS pending_questions (
question_id TEXT PRIMARY KEY,
thread_id TEXT NOT NULL,
turn INTEGER NOT NULL,
status TEXT NOT NULL
CHECK (status IN ('open', 'answered', 'expired', 'superseded')),
transport TEXT NOT NULL,
channel_ref TEXT,
posted_at TEXT,
deadline_at TEXT,
answer_json TEXT,
answered_at TEXT,
answered_via TEXT
);
CREATE INDEX IF NOT EXISTS idx_pending_questions_thread
ON pending_questions (thread_id, turn);
CREATE INDEX IF NOT EXISTS idx_pending_questions_status
ON pending_questions (status);
-- Defense-in-depth: at most one OPEN question may carry a given non-null
-- channel_ref, so a thread reply's thread_ts can never map to two open rows
-- (a partial unique index; NULL channel_refs and closed rows are excluded).
CREATE UNIQUE INDEX IF NOT EXISTS uq_pending_questions_open_channel_ref
ON pending_questions (channel_ref)
WHERE channel_ref IS NOT NULL AND status = 'open';
-- budget_ledger: the persistent shared Claude budget ledger (§6.1, §6.6).
-- One row per accounted spend event; the shared daily cap and the
-- interactive-first reserve are computed by summing over a UTC day. Spend is
-- recorded across ALL R720 Claude work (pipeline + Plane-1 sweeps).
CREATE TABLE IF NOT EXISTS budget_ledger (
entry_id INTEGER PRIMARY KEY AUTOINCREMENT,
thread_id TEXT,
stage TEXT,
model TEXT NOT NULL,
billing_mode TEXT NOT NULL,
input_tokens INTEGER NOT NULL DEFAULT 0,
output_tokens INTEGER NOT NULL DEFAULT 0,
usd_cost REAL NOT NULL DEFAULT 0.0,
recorded_at TEXT NOT NULL,
day_bucket TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_budget_ledger_day
ON budget_ledger (day_bucket);
CREATE INDEX IF NOT EXISTS idx_budget_ledger_thread
ON budget_ledger (thread_id);
-- ingested_issues (schema v2): durable de-dup ledger for the GitHub-issue intake
-- poller. The in-memory poller set does not survive a restart, so a scheduled
-- intake (each run a fresh process) would re-ingest every still-open labeled
-- issue and spawn duplicate tasks. This table records, per (source, issue_id),
-- the issues already turned into tasks so intake is idempotent across restarts.
-- The box is read-only (no write token to remove the intake label), so durable
-- de-dup is the only correct guard. `source` namespaces by repo (e.g.
-- "github:owner/repo") so per-repo issue numbers from different repos cannot collide.
CREATE TABLE IF NOT EXISTS ingested_issues (
source TEXT NOT NULL,
issue_id TEXT NOT NULL,
ingested_at TEXT NOT NULL,
PRIMARY KEY (source, issue_id)
);
-- schema_meta: single-row table recording the applied schema version so
-- migrate() can detect and step forward.
CREATE TABLE IF NOT EXISTS schema_meta (
id INTEGER PRIMARY KEY CHECK (id = 1),
schema_version INTEGER NOT NULL
);