This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/docs/provisioning/OPERATOR-RUNBOOK.md
Adam Moussa a8f00ff676 feat(agent-team): operator dispatch command + runbook fixes
- run-team.py: add the 'dispatch <thread_id>' operator command (P3 option-b).
  The read-only box parks at DISPATCH; this completes it with a just-in-time
  WRITE token: reads candidate_diff + scope from the checkpoint (or --diff/--scope
  files), pushes the head branch + fires workflow_dispatch via dispatch_apply_verify,
  prints the located run_id, and (--write-back) writes it into the task checkpoint
  so VERIFY binds. +2 tests.
- OPERATOR-RUNBOOK: fix the misleading 'systemctl show -p Environment' check (it
  does NOT show EnvironmentFile= vars) -> use /proc/<MainPID>/environ +
  _p3_env_is_configured(); document the operator-initiated dispatch flow + the
  fine-grained-token write-probe caveat.

Suite green, ruff clean. Branch only; not merged.
2026-06-23 20:49:09 -04:00

25 KiB
Raw Blame History

OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling

The on-call runbook for the always-on agent-team-coordinator daemon on the sh-secrev R720 VM. Covers pipeline stalls, stuck/parked tasks, failed human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design §5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement).

CLI used throughout: run-team.py (the entry CLI), run from ~/orchestrator/agent-team with the venv active so it hits the default ledger (state/agent_team.sqlite) and audit log (state/audit.log.jsonl). Do not use agent_team/operator_cli.py — it has divergent verbs (no show, required --db/--audit-log, and a force-resume that supersedes). See DEPLOY-AUDIT.md.

ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team && . .venv/bin/activate

Verb reference (all confirmed in run-team.py:build_parser)

Verb Effect Destructive?
list / list --all / list --parked / list --status <state> read pending questions (default open) no
show <question_id> print one ledger row (JSON) no
redeliver <question_id> clear channel_ref so the reconcile loop re-posts an open question no (audit-logged)
expire <question_id> --confirm force open→expired yes
answer <question_id> --answer <p> [--via <id>] --confirm answer-on-behalf (first-answer-wins CAS) yes
force-resume <question_id> --confirm reopen an expired (parked) question; for answered records resume intent yes
supersede <question_id> --confirm mark a stale open/answered row superseded yes
start --task "..." [--transport ...] [--dry-run] start one task to the human gate no

--operator <name> (global) sets the audit attribution; it defaults to the OS login. Destructive verbs require --confirm and write an attempt-then-outcome record to the audit log before mutating.

First triage for any incident:

systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100
python3 run-team.py list --all          # full ledger snapshot
python3 run-team.py list --parked       # non-open rows (parked-task context)

Incident 1 — Pipeline stall (tasks not advancing)

Symptoms: list shows open (or answered) rows that never progress; the journal shows no tick activity or repeated errors.

Diagnose:

systemctl is-active agent-team-coordinator.service     # "active" expected
journalctl -u agent-team-coordinator.service -e | tail -60
# Daemon dead/looping on restart? Check the unit + deps:
.venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')"

Recover:

  1. If the daemon is dead and Restart=on-failure is flapping, read the journal for the import/auth error. A missing venv dep (D-4/D-5) or a missing CLAUDE_CODE_OAUTH_TOKEN is the usual cause — fix the env / venv, then:
    sudo systemctl restart agent-team-coordinator.service
    
  2. The startup recover() sweep re-drives answered-but-unresumed rows and re-posts open rows that lost their channel_ref, so a clean restart converges from the durable ledger. Confirm with journalctl ... | tail and list --all.
  3. A single task stuck open with a stale/missing post: re-deliver it.
    python3 run-team.py show <qid>
    python3 run-team.py redeliver <qid>     # clears channel_ref; reconcile re-posts
    

Escalate (see the ladder) if a restart does not clear it within one tick cadence and the journal shows a non-transient error.


Incident 2 — Stuck / parked task

A task parks (design §6.6) when its clarifier question expires with no answer, or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked = the durable ledger row is no longer open (it is expired), and an ALARM was raised, not spun on.

Diagnose:

python3 run-team.py list --parked
python3 run-team.py show <qid>          # status, thread_id, deadline_at, answered_via

Recover — depends on why it parked:

  • Expired with no answer (the common case) — un-park by reopening the expired question; it is then re-delivered for an answer:

    python3 run-team.py force-resume <qid> --confirm
    # "force-resume: reopened expired question <qid>; it will be re-delivered"
    python3 run-team.py show <qid>         # status flips back to "open"
    

    Then answer it (Slack or CLI) to drive it forward.

  • Answered but not yet resumed — the recovery sweep handles it; force-resume records intent and reports that (no mutation):

    python3 run-team.py force-resume <qid> --confirm
    # "force-resume: question <qid> is answered and pending resume; the recovery
    #  sweep will resume it (intent recorded)"
    sudo systemctl restart agent-team-coordinator.service   # forces the recover() sweep now
    
  • Stale / wrong question that should be abandoned — supersede it so it stops surfacing as parked context:

    python3 run-team.py supersede <qid> --confirm
    

MAX_PARK FIFO-aging: a task that exceeds the park window escalates (ALARM + a Jira ticket per the ladder) rather than starving silently.


Incident 3 — Failed human-in-the-loop resume

Symptoms: an answer was submitted (Slack or CLI) but the graph did not advance.

Diagnose:

python3 run-team.py show <qid>          # is status "answered"? what answered_via?
journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|<qid>"

Recover:

  • Status is answered but no resume drained — the resume queue is in-process; a daemon restart triggers the recover() sweep that re-drives answered rows via the turn-guarded ResumeWorker (idempotent — an already-advanced thread supersedes-and-skips):
    sudo systemctl restart agent-team-coordinator.service
    python3 run-team.py show <qid>        # confirm it advanced
    
  • Live Slack answer never registered — the listener fails closed. Check:
    journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer"
    # "rejecting Slack answer: owner allowlist is unconfigured ..." -> set
    #   AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart.
    # "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is
    #   not in the allowlist. Add it, or answer via the CLI on their behalf:
    python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
    
  • Answer lost the compare-and-set (not open) — the row was already answered/expired/superseded. Inspect with show; if it parked, go to Incident 2.

Incident 4 — Budget exhaustion mid-pipeline

The shared Claude budget ledger (budget_ledger table) enforces per-call + total nightly caps (design §6.1, §6.6). A stage that would breach the reserve parks the task (deferred, not dropped) and ALARMs — it never loop-drains the pool.

Diagnose:

journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park"
# Inspect today's spend directly (no CLI verb for the budget ledger; read it).
# Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model,
# billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket.
sqlite3 state/agent_team.sqlite \
  "SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok
   FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;"

Recover:

  • Wait for the next budget window — budget-exhausted roles/tasks are deferred via the rotation pointer and picked up next cycle; this is the intended behavior, not a failure. The parked task surfaces in list --parked.
  • Force a specific parked task forward now (e.g. it is urgent and headroom has since freed): un-park it and let the daemon re-run the stage within the remaining cap:
    python3 run-team.py force-resume <qid> --confirm
    
  • Confirm ANTHROPIC_API_KEY is absent — its presence would silently meter to API rates and blow the budget model (the billing seam pops it defensively, but it must not be set):
    grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env   # both must print 0
    

Do not raise the cap to push a task through without Adam's decision — budget realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly parks on budget across nights.


Incident 5 — Transport outage (Slack / GitHub down or misconfigured)

Symptoms: clarifier posts fail; the journal shows transport errors or the inbound listener is not started.

Diagnose:

journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport"
# Did the inbound listener start?
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
#   "started (Socket Mode ...)"  -> inbound up
#   "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off

Recover:

  • Outbound post failed (Slack API / network) — the ledger row stays open with no channel_ref (the responder leaves it for reconcile). After the transport recovers, the recover() sweep (on restart) or a manual redeliver re-posts:
    python3 run-team.py redeliver <qid>
    
  • Inbound listener down / never started — the daemon still posts and expires; it just cannot hear Slack. The CLI answer path is the outage fallback — it runs the identical compare-and-set with no socket:
    python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
    
    Fix SLACK_APP_TOKEN / SLACK_BOT_TOKEN in ~/secrev.env, then sudo systemctl restart agent-team-coordinator.service to re-establish Socket Mode. (D-1: the listener starts only when the transport is live Slack AND SLACK_APP_TOKEN is set; it is optional by design.)
  • Slack listener crash-looping — the listener runs on an isolated daemon thread; a crash is logged (inbound Slack listener thread exited with an error) and takes down only the inbound socket, not the maintenance loop. Restart the service to relaunch the thread once the cause is fixed.

Incident 5b — WS0–WS5 surfaces (HTTP API, /new-task, handbook context)

These were added by the WS-rollout (see agent-team/DEPLOY-R720.md §4b). They layer onto the coordinator; none of them should take down the maintenance loop.

  • HTTP API down / unreachable — the WS1 FastAPI app (agent_team/api.py) is a separate, opt-in process (api.serve(), 127.0.0.1:8765, bearer auth), not started by the coordinator daemon. If /delegate from Claude Code or POST /tasks over HTTP stops working, the coordinator itself is unaffected — check the API process separately:

    curl -sS -o /dev/null -w '%{http_code}\n' \
      -H "Authorization: Bearer $AGENT_TEAM_API_TOKEN" http://127.0.0.1:8765/tasks
    #   405 = API up + authed (GET not allowed on /tasks);  000 = process down;
    #   401 = AGENT_TEAM_API_TOKEN mismatch (client vs ~/secrev.env).
    

    The API refuses to start if AGENT_TEAM_API_TOKEN is unset/empty (logs a RuntimeError). Fix the token, restart the API process. Tasks already in the ledger are unaffected — the API is only an intake/invoke front door; answer via Slack or the CLI as usual.

  • /new-task Slack command not responding — the WS2 slash command is AUTHZ-01 owner-allowlist gated and routes through the same Socket Mode listener as answers. If it silently does nothing, it is almost always the owner allowlist (same failure mode as Incident 3's live-Slack path):

    journalctl -u agent-team-coordinator.service -e | grep -iE "new-task|owner|unauthorized"
    #   unauthorized sender / unconfigured allowlist -> fix AGENT_TEAM_SLACK_OWNER_IDS
    

    Fallback: start the task from the CLI (run-team.py start --task "...") or the HTTP API. If the listener itself is down, see Incident 5 (inbound listener).

  • Handbook dir missing → planner runs without handbook context — the WS5 context_provider (load_handbook_conventions) is fail-safe: if SEA_HAVEN_HANDBOOK_DIR (or ~/.sea-haven/engineering-handbook) is missing or unreadable it returns "" and the planner runs normally, just without handbook conventions injected. This is degraded, not broken — no park, no alarm. Confirm and restore:

    grep '^SEA_HAVEN_HANDBOOK_DIR=' ~/secrev.env
    ls "$(grep '^SEA_HAVEN_HANDBOOK_DIR=' ~/secrev.env | cut -d= -f2)"   # dir present + populated?
    

    Re-sync the handbook (the deploy script does this) and restart the daemon so the planner picks it back up. The WS3 dispatch node is inert (gated) and should never appear in pipeline activity; if it does, treat as an unexpected state and escalate.


Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers)

These come from the nightly checker run, not the coordinator daemon (design §6.4, §6.6):

  • COMPLACENCY ALARM — a checker role missed a planted canary fault. That role is skipped for the night (it never runs silently degraded). Recover: inspect the canary corpus / the role's checker under security-review/checkers/, fix the regression, re-run that checker's dry-run, and confirm the canary passes before re-enabling.
  • COVERAGE ALARM — a role slipped its rotation slot (e.g. budget-deferred). It is deferred via the rotation pointer, never dropped, and picked up next cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from report history) and that the role runs on the next cadence; investigate only if it slips repeatedly.

Both alarms are report + ALARM-only (design D3): nothing posts on a clean state, no auto-Jira/Notion writes from the checker itself. Persistent alarms follow the escalation ladder below.


Incident 7 — P3 rollback (unwind the apply/verify flip)

The P3 apply/verify CI workflow is already LIVE (flipped + provisioned 2026-06-22): the agent-apply environment with its required reviewer, the GitHub App installation (pull-requests: write), the run-name/permissions edits on .github/workflows/agent-team-apply-verify.yml, and enforce_admins branch protection on main. agent-team/scripts/p3_rollback.sh is the inverse — it restores those four privileged surfaces to a recorded baseline and asserts post-restore == baseline (fail-closed: a mismatched restore aborts non-zero).

Run from the operator host (the Mac), NOT the box. No standing box token is used; the script relies on gh being authenticated on the operator host (plus git for the post-merge revert path). The R720 daemon holds no write token by design (P3-PHASE0-DESIGN.md, "Out of scope").

When to trigger:

  • The flip must be unwound (P3 apply/verify is being backed out), OR
  • A premature flip that RAN — an apply/verify run executed before it should have and opened a draft PR / branch that must not exist (use the incident surface, below).

Prerequisites:

  • A baseline JSON recorded BEFORE the flip (the restore target). Default path .security-review/p3-baseline.json or $P3_ROLLBACK_BASELINE, override with --baseline FILE. The script refuses to run if the baseline file is missing (record it BEFORE the flip) — there is no inferred baseline.
  • Target repo from the baseline's repo key, --repo OWNER/NAME, or $AGENT_TEAM_REPO_OWNER/$AGENT_TEAM_REPO_NAME. If both the baseline and --repo are set they must agree (fail-closed on mismatch).

The four privileged surfaces it restores:

  1. workflow — the apply/verify flip itself. Pre-merge: close the flip PR + delete its branch (--flip-pr N [--flip-branch REF]). Post-merge: git revert the flip commit + push + re-run CI (--merged --flip-commit SHA). LIVE: also reverts the workflow YAML to the recorded workflow_baseline_sha so the live file matches baseline.
  2. environment — the agent-apply environment: re-PUTs the recorded required reviewer(s) + deployment branch policy, then asserts the environment exists.
  3. app — the GitHub App: always rotates (revokes) the installation token first (any minted token is now suspect), then reduces permissions to the baseline or uninstalls the installation per app.action.
  4. protection — branch protection on main: re-enables enforce_admins (include_administrators) and asserts it is ON. Refuses to run if the baseline's include_administrators is not true — a rollback must never leave a weaker posture than baseline.

all runs surfaces 1→4 in order.

How to run — always --dry-run first, then --apply:

The script is destructive-safe by default: with no --apply it only PRINTS the plan ([PLAN] ... lines). --apply is the only thing that performs mutations. Inspect the dry-run plan, confirm it targets the right repo + baseline, then re-run with --apply.

cd ~/Documents/repositories/orchestrator/agent-team   # operator host, gh authed

# 1. Dry-run: print the plan only, no mutations.
scripts/p3_rollback.sh all --baseline .security-review/p3-baseline.json

# 2. Apply, once the plan looks right (pre-merge flip example):
scripts/p3_rollback.sh all --apply \
  --baseline .security-review/p3-baseline.json \
  --flip-pr 123 --flip-branch agent-team/apply/flip

# Post-merge workflow revert instead of pre-merge close:
scripts/p3_rollback.sh workflow --apply --merged --flip-commit <SHA>

# A single surface at a time is fine too:
scripts/p3_rollback.sh protection --apply

Premature-flip-that-RAN incident (the incident surface):

Use this when a flip executed prematurely and opened a draft PR/branch. It runs the full incident sequence — rotate, revert, audit (read-only), restore, note:

# Dry-run first (the audit/list steps run read-only either way):
scripts/p3_rollback.sh incident --baseline .security-review/p3-baseline.json

# Apply, with the premature draft PR if known:
scripts/p3_rollback.sh incident --apply \
  --flip-pr <N> --flip-branch agent-team/apply/<task_id>

It performs, in order:

  1. Rotate the App installation token first — anything the premature run minted is suspect.
  2. Revert the draft PR / branch the App opened — close --flip-pr + delete its branch; if no --flip-pr is given it lists open agent-team/apply/* draft PRs to triage.
  3. Audit the Checks trail (read-only — always runs) — recent agent-team-apply-verify.yml runs with conclusions.
  4. Restore the agent-apply environment + branch protection to baseline (surfaces 2 + 4).
  5. File an incident note under .security-review/incidents/p3-premature-flip-<timestamp>.md recording the repo, baseline, flip PR/branch, and actions taken.

After any rollback, confirm the post-restore [OK] asserts printed (the script exits non-zero if any assert failed), and record the action — for the incident path the note is written automatically; otherwise note it on the Jira ticket per the escalation ladder.


Escalation ladder (design §5 / §6.6, resolves Q3)

Anything that does not clear on the first ALARM escalates — but ALARM-only in spirit (nothing posts on a clean state):

  1. Re-alarm on a backoff. A confirmed critical (or a COMPLACENCY / COVERAGE alarm) that persists re-alarms to Slack each night it is still unresolved, on a backoff so it does not spam.
  2. Open a Jira tracking ticket after N nights (default N = 3). If the condition still has not cleared, the coordinator opens an INFRA Jira ticket so it cannot quietly linger. The same ladder applies to a parked task that exceeds MAX_PARK.
  3. Human (Adam) takes it from the Jira ticket. For a parked task, recover via Incidents 2–4 above; for a checker alarm, via Incident 6.

When you resolve an incident, record the action — destructive CLI verbs already write an attributable attempt+outcome record to state/audit.log.jsonl; for non-CLI recoveries note it on the Jira ticket.


Post-incident

  • Confirm the daemon is active and the ledger has no unexpected parked rows (list --parked).
  • If you restored the VM snapshot or wiped the ledger, re-run the relevant PROVISIONING-RUNBOOK steps.
  • Update project_r720_agent_team memory if the incident revealed a durable fact (a new failure mode, a config that must change).

P3 box env wiring (build → dispatch → verify)

The coordinator's environment is loaded from EnvironmentFile=-/home/adam/secrev.env (declared in the unit; the leading - makes it optional so a missing file does not fail the unit). The live P3 path — gated build → dispatch (trigger CI, capture run_id) → verify (read the CI conclusion) — reads four variables from that file at graph-build / dispatch time:

  • AGENT_TEAM_REPO_OWNER — dispatch/verify target owner (fixed at factory time, never read from pipeline state, so model output cannot redirect the target).
  • AGENT_TEAM_REPO_NAME — dispatch/verify target repo (same fail-closed binding).
  • AGENT_TEAM_BASE_BRANCH — PR base branch; optional, defaults to main.
  • AGENT_TEAM_CI_READ_TOKEN — the read-only CI-result token used for the verifier's authenticated conclusion read (falls back to GITHUB_TOKEN).

If owner, repo, or the CI-read token is unset, serve() degrades to the INERT P3 path (one WARNING + a #agent-team notice) rather than crash-looping the daemon (Phase-0 Decision 5). No pull-requests:write / contents:write token and no AGENT_APPLY_APP_ID / AGENT_APPLY_APP_PRIVATE_KEY may live in ~/secrev.env: the apply path mints its write token inside the CI runner from Actions secrets, so the box holds no standing write credential. That invariant is enforced by scripts/assert_no_write_token.py (the A2 audit) at provisioning and in CI — run it before any deploy.

Verifying the vars load

After installing/editing ~/secrev.env and systemctl daemon-reload + systemctl restart agent-team-coordinator.service, confirm the P3 vars reached the running coordinator process.

⚠️ Do NOT use systemctl show -p Environment — it lists only inline Environment= directives and does NOT show vars loaded from EnvironmentFile= (which is how ~/secrev.env is loaded). It comes back empty even when the vars are correctly loaded, so it is misleading here.

Read the actual process environment instead (requires sudo to read another process's environ):

MP=$(systemctl show agent-team-coordinator.service -p MainPID --value)
sudo tr '\0' '\n' < /proc/$MP/environ | grep -E '^AGENT_TEAM_REPO|^AGENT_TEAM_BASE'

The output should list AGENT_TEAM_REPO_OWNER, AGENT_TEAM_REPO_NAME, and AGENT_TEAM_BASE_BRANCH (if set). To confirm the live P3 wiring actually bound (not the INERT path), check the code resolver directly:

cd ~/orchestrator/agent-team && set -a && source ~/secrev.env && set +a \
  && .venv/bin/python -c "from agent_team.coordinator import _p3_env_is_configured; print(_p3_env_is_configured())"

True means the live build+verify path is bound; False means the daemon is on the INERT P3 path (fix ~/secrev.env, reload, restart). The CI-read token is satisfied by AGENT_TEAM_CI_READ_TOKEN or the read-only GITHUB_TOKEN fallback; treat any token value in process output as sensitive. The box must hold NO AGENT_APPLY_APP_ID / AGENT_APPLY_APP_PRIVATE_KEY / write token — verify with python scripts/assert_no_write_token.py (and note its scope-detection caveat in the script header: a fine-grained token's write capability is only definitively confirmed by a live POST /git/refs probe returning 403).

Operator-initiated dispatch (P3 option-b)

The box is read-only, so its in-graph DISPATCH node fail-closes/parks — it never pushes or triggers CI. Completing a dispatch is an explicit operator step with a just-in-time write token (never stored in ~/secrev.env):

# On the box (where the task's candidate_diff lives in the ledger), with a
# WRITE-capable token provided for THIS invocation only:
cd ~/orchestrator/agent-team
GH_TOKEN=<operator pull-requests+contents:write token> \
  .venv/bin/python run-team.py dispatch <thread_id> --write-back

This reads the task's candidate_diff + declared scope from the checkpoint, pushes the head branch, fires the agent-team-apply-verify workflow_dispatch, prints the located run_id, and (--write-back) writes it into the task checkpoint so the box's VERIFY binds to that run. CI then runs guard → build-test → pure-code gate; the privileged gate-and-pr job pauses at the agent-apply environment for your required-reviewer approval before the draft PR opens. Alternatively pass --diff FILE --scope FILE to dispatch a diff without reading the ledger. The token is consumed by gh/git for the one command and never persisted; the box returns to read-only at rest.