This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/docs/provisioning/OPERATOR-RUNBOOK.md
Adam Moussa f46d691e36 docs(agent-team): document the plan-review decision gate (Phase C)
README: two human gates (clarifier + plan-decision), the approve/request-changes/
abandon verbs, free-text-defaults-to-request-changes, MAX_PLAN_GATE_VISITS, and
the planner max_turns reliability fix.
OPERATOR-RUNBOOK: how the gate appears in Slack, the three decision paths
(buttons/modal/free-text), the single-open-gate invariant, ceiling→PARKED, and
the 24h expiry→PARKED→recovery (re-assign / force-resume).
DEPLOY-R720: the pending_questions.kind ledger migration (SCHEMA_VERSION→4,
idempotent additive ALTER on startup) + rollback (restore the ledger backup
before restart if the migration fails).
2026-06-24 11:30:54 -04:00

596 lines
28 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling
The on-call runbook for the always-on `agent-team-coordinator` daemon on the
`sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed
human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and
the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design
§5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement).
> **CLI used throughout: `run-team.py`** (the entry CLI), run from
> `~/orchestrator/agent-team` with the venv active so it hits the default ledger
> (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not**
> use `agent_team/operator_cli.py` — it has divergent verbs (no `show`, required
> `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md.
```bash
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team && . .venv/bin/activate
```
## Verb reference (all confirmed in `run-team.py:build_parser`)
| Verb | Effect | Destructive? |
|---|---|---|
| `list` / `list --all` / `list --parked` / `list --status <state>` | read pending questions (default `open`) | no |
| `show <question_id>` | print one ledger row (JSON) | no |
| `redeliver <question_id>` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) |
| `expire <question_id> --confirm` | force `open`→`expired` | yes |
| `answer <question_id> --answer <p> [--via <id>] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes |
| `force-resume <question_id> --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes |
| `supersede <question_id> --confirm` | mark a stale `open`/`answered` row `superseded` | yes |
| `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no |
`--operator <name>` (global) sets the audit attribution; it defaults to the OS
login. Destructive verbs require `--confirm` and write an attempt-then-outcome
record to the audit log **before** mutating.
First triage for any incident:
```bash
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100
python3 run-team.py list --all # full ledger snapshot
python3 run-team.py list --parked # non-open rows (parked-task context)
```
---
## Incident 1 — Pipeline stall (tasks not advancing)
**Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the
journal shows no `tick` activity or repeated errors.
**Diagnose:**
```bash
systemctl is-active agent-team-coordinator.service # "active" expected
journalctl -u agent-team-coordinator.service -e | tail -60
# Daemon dead/looping on restart? Check the unit + deps:
.venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')"
```
**Recover:**
1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal
for the import/auth error. A missing venv dep (D-4/D-5) or a missing
`CLAUDE_CODE_OAUTH_TOKEN` is the usual cause — fix the env / venv, then:
```bash
sudo systemctl restart agent-team-coordinator.service
```
2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and
re-posts `open` rows that lost their `channel_ref`, so a clean restart
converges from the durable ledger. Confirm with `journalctl ... | tail` and
`list --all`.
3. A single task stuck `open` with a stale/missing post: re-deliver it.
```bash
python3 run-team.py show <qid>
python3 run-team.py redeliver <qid> # clears channel_ref; reconcile re-posts
```
**Escalate** (see the ladder) if a restart does not clear it within one tick
cadence and the journal shows a non-transient error.
---
## Incident 2 — Stuck / parked task
A task parks (design §6.6) when its clarifier question **expires** with no answer,
or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked =
the durable ledger row is no longer `open` (it is `expired`), and an ALARM was
raised, not spun on.
**Diagnose:**
```bash
python3 run-team.py list --parked
python3 run-team.py show <qid> # status, thread_id, deadline_at, answered_via
```
**Recover — depends on why it parked:**
- **Expired with no answer (the common case)** — un-park by reopening the expired
question; it is then re-delivered for an answer:
```bash
python3 run-team.py force-resume <qid> --confirm
# "force-resume: reopened expired question <qid>; it will be re-delivered"
python3 run-team.py show <qid> # status flips back to "open"
```
Then answer it (Slack or CLI) to drive it forward.
- **Answered but not yet resumed** — the recovery sweep handles it; `force-resume`
records intent and reports that (no mutation):
```bash
python3 run-team.py force-resume <qid> --confirm
# "force-resume: question <qid> is answered and pending resume; the recovery
# sweep will resume it (intent recorded)"
sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now
```
- **Stale / wrong question that should be abandoned** — supersede it so it stops
surfacing as parked context:
```bash
python3 run-team.py supersede <qid> --confirm
```
`MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a
Jira ticket per the ladder) rather than starving silently.
---
## Incident 2b — Plan-review decision gate (PLAN ⇄ REVIEW dead-end)
A second human gate opens when the plan↔review loop **cannot auto-converge** (the
review-revision cap is hit) or the planner produced only a partial plan. Instead
of terminally parking, the coordinator suspends on a resumable `interrupt()` and
posts the **plan + reviewer findings** to Slack `#agent-team`, **threaded under
the task root**, with three decision verbs. The ledger row carries
`kind = 'plan_decision'` (a clarifier row is `kind = 'clarify'`); both are
ordinary `pending_questions` rows, so the same `list` / `show` / `force-resume`
verbs apply.
**The three decision paths (all equivalent — pick whichever is handy):**
| Path | How | Effect |
|---|---|---|
| **Buttons** | Block Kit *Approve* / *Request changes* / *Abandon* on the gate message | Approve & Abandon submit immediately; *Request changes* opens a notes modal |
| **Modal** | the *Request changes* button → a one-field "What should change?" modal | submits `request_changes` with your notes |
| **Free-text reply** | reply in the gate thread | parsed kind-aware (below) |
**What each decision does:**
- **Approve** — settles the plan (advances to BUILD / continues the pipeline).
- **Request changes (+ notes)** — loops back to the **planner**, folding the notes
into the review feedback so the re-plan addresses them.
- **Abandon** — fails the task (terminal `FAILED`).
**Free-text mapping (the safe-default rule).** A thread reply is normalized
(lowercase/strip) and matched against small allowlists:
`approve ∈ {approve, approved, yes, ok, lgtm, ship}`;
`abandon ∈ {abandon, reject, cancel, stop, kill}`. **Anything else — any other
prose, including empty/whitespace — maps to *request changes*, carrying the full
reply as the notes.** So typing change notes in the thread (e.g. "use pytest
fixtures instead") requests changes; it can never be misread as an accidental
approve or abandon.
**Single-open-gate invariant.** A thread holds **at most one open question at a
time** — the clarifier row is already answered before the plan stage runs, so a
clarifier-answer and a plan-decision can never be open simultaneously for one
thread. If a second open row for a thread ever appears, treat it as a bug
(the coordinator logs + skips opening it) and inspect with `list --all`.
**Ceiling behavior.** The human request-changes loop is bounded by
`MAX_PLAN_GATE_VISITS` (graph constant, currently **3**, distinct from the
planner's `MAX_PLAN_REVISIONS`). When the gate-visit ceiling is exhausted the
node does **not** re-open the gate — it returns terminal **PARKED** with a
"revision ceiling reached" note, so the loop always terminates.
**Expiry / recovery (gate goes unanswered).** A `plan_decision` row uses the
**same 24h deadline window as the clarifier**. If unanswered, the deadline sweep
**expires** the row → the task **PARKS**, and the coordinator posts a
`⌛ PLAN DECISION EXPIRED` lifecycle notice naming the task + the recovery path,
threaded under the task root. Recover the same way as a parked clarifier:
```bash
python3 run-team.py list --parked
python3 run-team.py show <qid> # kind == "plan_decision", status == expired
python3 run-team.py force-resume <qid> --confirm # reopens the expired gate for re-delivery
```
Then answer it (Slack buttons / modal / free-text reply, or CLI `answer`). If the
task is better restarted from scratch, **re-assign** it instead. There is no
auto-retry — a PARKED plan gate stays PARKED until an operator acts.
---
## Incident 3 — Failed human-in-the-loop resume
**Symptoms:** an answer was submitted (Slack or CLI) but the graph did not
advance.
**Diagnose:**
```bash
python3 run-team.py show <qid> # is status "answered"? what answered_via?
journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|<qid>"
```
**Recover:**
- **Status is `answered` but no resume drained** — the resume queue is in-process;
a daemon restart triggers the `recover()` sweep that re-drives `answered` rows
via the turn-guarded `ResumeWorker` (idempotent — an already-advanced thread
supersedes-and-skips):
```bash
sudo systemctl restart agent-team-coordinator.service
python3 run-team.py show <qid> # confirm it advanced
```
- **Live Slack answer never registered** — the listener fails closed. Check:
```bash
journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer"
# "rejecting Slack answer: owner allowlist is unconfigured ..." -> set
# AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart.
# "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is
# not in the allowlist. Add it, or answer via the CLI on their behalf:
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
```
- **Answer lost the compare-and-set (`not open`)** — the row was already
answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2.
---
## Incident 4 — Budget exhaustion mid-pipeline
The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total
nightly caps (design §6.1, §6.6). A stage that would breach the reserve **parks**
the task (deferred, not dropped) and ALARMs — it never loop-drains the pool.
**Diagnose:**
```bash
journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park"
# Inspect today's spend directly (no CLI verb for the budget ledger; read it).
# Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model,
# billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket.
sqlite3 state/agent_team.sqlite \
"SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok
FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;"
```
**Recover:**
- **Wait for the next budget window** — budget-exhausted roles/tasks are deferred
via the rotation pointer and picked up next cycle; this is the intended
behavior, not a failure. The parked task surfaces in `list --parked`.
- **Force a specific parked task forward now** (e.g. it is urgent and headroom has
since freed): un-park it and let the daemon re-run the stage within the
remaining cap:
```bash
python3 run-team.py force-resume <qid> --confirm
```
- **Confirm `ANTHROPIC_API_KEY` is absent** — its presence would silently meter to
API rates and blow the budget model (the billing seam pops it defensively, but
it must not be set):
```bash
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0
```
Do **not** raise the cap to push a task through without Adam's decision — budget
realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly
parks on budget across nights.
---
## Incident 5 — Transport outage (Slack / GitHub down or misconfigured)
**Symptoms:** clarifier posts fail; the journal shows transport errors or the
inbound listener is not started.
**Diagnose:**
```bash
journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport"
# Did the inbound listener start?
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
# "started (Socket Mode ...)" -> inbound up
# "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off
```
**Recover:**
- **Outbound post failed (Slack API / network)** — the ledger row stays `open`
with no `channel_ref` (the responder leaves it for reconcile). After the
transport recovers, the `recover()` sweep (on restart) or a manual `redeliver`
re-posts:
```bash
python3 run-team.py redeliver <qid>
```
- **Inbound listener down / never started** — the daemon still posts and expires;
it just cannot hear Slack. **The CLI `answer` path is the outage fallback** —
it runs the identical compare-and-set with no socket:
```bash
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
```
Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then
`sudo systemctl restart agent-team-coordinator.service` to re-establish Socket
Mode. (D-1: the listener starts only when the transport is live Slack AND
`SLACK_APP_TOKEN` is set; it is optional by design.)
- **Slack listener crash-looping** — the listener runs on an isolated daemon
thread; a crash is logged (`inbound Slack listener thread exited with an error`)
and takes down only the inbound socket, not the maintenance loop. Restart the
service to relaunch the thread once the cause is fixed.
---
## Incident 5b — WS0–WS5 surfaces (HTTP API, /new-task, handbook context)
These were added by the WS-rollout (see `agent-team/DEPLOY-R720.md` §4b). They
layer onto the coordinator; none of them should take down the maintenance loop.
- **HTTP API down / unreachable** — the WS1 FastAPI app (`agent_team/api.py`) is
a **separate, opt-in process** (`api.serve()`, `127.0.0.1:8765`, bearer auth),
**not** started by the coordinator daemon. If `/delegate` from Claude Code or
`POST /tasks` over HTTP stops working, the coordinator itself is unaffected —
check the API process separately:
```bash
curl -sS -o /dev/null -w '%{http_code}\n' \
-H "Authorization: Bearer $AGENT_TEAM_API_TOKEN" http://127.0.0.1:8765/tasks
# 405 = API up + authed (GET not allowed on /tasks); 000 = process down;
# 401 = AGENT_TEAM_API_TOKEN mismatch (client vs ~/secrev.env).
```
The API refuses to start if `AGENT_TEAM_API_TOKEN` is unset/empty (logs a
`RuntimeError`). Fix the token, restart the API process. Tasks already in the
ledger are unaffected — the API is only an *intake/invoke* front door; answer
via Slack or the CLI as usual.
- **`/new-task` Slack command not responding** — the WS2 slash command is
AUTHZ-01 owner-allowlist gated and routes through the same Socket Mode listener
as answers. If it silently does nothing, it is almost always the owner
allowlist (same failure mode as Incident 3's live-Slack path):
```bash
journalctl -u agent-team-coordinator.service -e | grep -iE "new-task|owner|unauthorized"
# unauthorized sender / unconfigured allowlist -> fix AGENT_TEAM_SLACK_OWNER_IDS
```
Fallback: start the task from the CLI (`run-team.py start --task "..."`) or the
HTTP API. If the listener itself is down, see Incident 5 (inbound listener).
- **Handbook dir missing → planner runs without handbook context** — the WS5
`context_provider` (`load_handbook_conventions`) is **fail-safe**: if
`SEA_HAVEN_HANDBOOK_DIR` (or `~/.sea-haven/engineering-handbook`) is missing or
unreadable it returns `""` and the planner runs normally, just without handbook
conventions injected. This is **degraded, not broken** — no park, no alarm.
Confirm and restore:
```bash
grep '^SEA_HAVEN_HANDBOOK_DIR=' ~/secrev.env
ls "$(grep '^SEA_HAVEN_HANDBOOK_DIR=' ~/secrev.env | cut -d= -f2)" # dir present + populated?
```
Re-sync the handbook (the deploy script does this) and restart the daemon so
the planner picks it back up. The WS3 dispatch node is **inert** (gated) and
should never appear in pipeline activity; if it does, treat as an unexpected
state and escalate.
---
## Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers)
These come from the nightly checker run, not the coordinator daemon (design §6.4,
§6.6):
- **COMPLACENCY ALARM** — a checker role missed a planted canary fault. That role
is **skipped** for the night (it never runs silently degraded). Recover: inspect
the canary corpus / the role's checker under `security-review/checkers/`, fix the
regression, re-run that checker's dry-run, and confirm the canary passes before
re-enabling.
- **COVERAGE ALARM** — a role slipped its rotation slot (e.g. budget-deferred). It
is **deferred via the rotation pointer, never dropped**, and picked up next
cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from
report history) and that the role runs on the next cadence; investigate only if
it slips repeatedly.
Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean
state, no auto-Jira/Notion writes from the checker itself. Persistent alarms
follow the escalation ladder below.
---
## Incident 7 — P3 rollback (unwind the apply/verify flip)
The P3 apply/verify CI workflow is **already LIVE** (flipped + provisioned
2026-06-22): the `agent-apply` environment with its required reviewer, the GitHub
App installation (`pull-requests: write`), the run-name/permissions edits on
`.github/workflows/agent-team-apply-verify.yml`, and `enforce_admins` branch
protection on `main`. `agent-team/scripts/p3_rollback.sh` is the inverse — it
restores those four privileged surfaces to a recorded baseline and asserts
post-restore == baseline (fail-closed: a mismatched restore aborts non-zero).
> **Run from the operator host (the Mac), NOT the box.** No standing box token is
> used; the script relies on `gh` being authenticated on the operator host (plus
> `git` for the post-merge revert path). The R720 daemon holds no write token by
> design (P3-PHASE0-DESIGN.md, "Out of scope").
**When to trigger:**
- The flip must be unwound (P3 apply/verify is being backed out), OR
- A **premature flip that RAN** — an apply/verify run executed before it should
have and opened a draft PR / branch that must not exist (use the `incident`
surface, below).
**Prerequisites:**
- A baseline JSON recorded **BEFORE** the flip (the restore target). Default path
`.security-review/p3-baseline.json` or `$P3_ROLLBACK_BASELINE`, override with
`--baseline FILE`. The script refuses to run if the baseline file is missing
(`record it BEFORE the flip`) — there is no inferred baseline.
- Target repo from the baseline's `repo` key, `--repo OWNER/NAME`, or
`$AGENT_TEAM_REPO_OWNER/$AGENT_TEAM_REPO_NAME`. If both the baseline and `--repo`
are set they must agree (fail-closed on mismatch).
**The four privileged surfaces it restores:**
1. **`workflow`** — the apply/verify flip itself. Pre-merge: close the flip PR +
delete its branch (`--flip-pr N [--flip-branch REF]`). Post-merge: `git revert`
the flip commit + push + re-run CI (`--merged --flip-commit SHA`). LIVE: also
reverts the workflow YAML to the recorded `workflow_baseline_sha` so the live
file matches baseline.
2. **`environment`** — the `agent-apply` environment: re-PUTs the recorded required
reviewer(s) + deployment branch policy, then asserts the environment exists.
3. **`app`** — the GitHub App: **always rotates (revokes) the installation token
first** (any minted token is now suspect), then `reduce`s permissions to the
baseline or `uninstall`s the installation per `app.action`.
4. **`protection`** — branch protection on `main`: re-enables `enforce_admins`
(`include_administrators`) and asserts it is ON. **Refuses to run if the
baseline's `include_administrators` is not `true`** — a rollback must never
leave a weaker posture than baseline.
`all` runs surfaces 1→4 in order.
**How to run — always `--dry-run` first, then `--apply`:**
The script is **destructive-safe by default**: with no `--apply` it only PRINTS
the plan (`[PLAN] ...` lines). `--apply` is the only thing that performs
mutations. Inspect the dry-run plan, confirm it targets the right repo + baseline,
then re-run with `--apply`.
```bash
cd ~/Documents/repositories/orchestrator/agent-team # operator host, gh authed
# 1. Dry-run: print the plan only, no mutations.
scripts/p3_rollback.sh all --baseline .security-review/p3-baseline.json
# 2. Apply, once the plan looks right (pre-merge flip example):
scripts/p3_rollback.sh all --apply \
--baseline .security-review/p3-baseline.json \
--flip-pr 123 --flip-branch agent-team/apply/flip
# Post-merge workflow revert instead of pre-merge close:
scripts/p3_rollback.sh workflow --apply --merged --flip-commit <SHA>
# A single surface at a time is fine too:
scripts/p3_rollback.sh protection --apply
```
**Premature-flip-that-RAN incident (the `incident` surface):**
Use this when a flip executed prematurely and opened a draft PR/branch. It runs
the full incident sequence — rotate, revert, audit (read-only), restore, note:
```bash
# Dry-run first (the audit/list steps run read-only either way):
scripts/p3_rollback.sh incident --baseline .security-review/p3-baseline.json
# Apply, with the premature draft PR if known:
scripts/p3_rollback.sh incident --apply \
--flip-pr <N> --flip-branch agent-team/apply/<task_id>
```
It performs, in order:
1. **Rotate the App installation token first** — anything the premature run minted
is suspect.
2. **Revert the draft PR / branch the App opened** — close `--flip-pr` + delete its
branch; if no `--flip-pr` is given it lists open `agent-team/apply/*` draft PRs
to triage.
3. **Audit the Checks trail** (read-only — always runs) — recent
`agent-team-apply-verify.yml` runs with conclusions.
4. **Restore the `agent-apply` environment + branch protection** to baseline
(surfaces 2 + 4).
5. **File an incident note** under
`.security-review/incidents/p3-premature-flip-<timestamp>.md` recording the
repo, baseline, flip PR/branch, and actions taken.
After any rollback, confirm the post-restore `[OK]` asserts printed (the script
exits non-zero if any assert failed), and record the action — for the `incident`
path the note is written automatically; otherwise note it on the Jira ticket per
the escalation ladder.
---
## Escalation ladder (design §5 / §6.6, resolves Q3)
Anything that does not clear on the first ALARM escalates — but **ALARM-only in
spirit** (nothing posts on a clean state):
1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE
alarm) that persists re-alarms to Slack each night it is still unresolved, on
a backoff so it does not spam.
2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the
condition still has not cleared, the coordinator opens an **INFRA** Jira ticket
so it cannot quietly linger. The same ladder applies to a parked task that
exceeds `MAX_PARK`.
3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via
Incidents 2–4 above; for a checker alarm, via Incident 6.
When you resolve an incident, record the action — destructive CLI verbs already
write an attributable attempt+outcome record to `state/audit.log.jsonl`; for
non-CLI recoveries note it on the Jira ticket.
---
## Post-incident
- Confirm the daemon is `active` and the ledger has no unexpected parked rows
(`list --parked`).
- If you restored the VM snapshot or wiped the ledger, re-run the relevant
PROVISIONING-RUNBOOK steps.
- Update `project_r720_agent_team` memory if the incident revealed a durable
fact (a new failure mode, a config that must change).
## P3 box env wiring (build → dispatch → verify)
The coordinator's environment is loaded from `EnvironmentFile=-/home/adam/secrev.env`
(declared in the unit; the leading `-` makes it optional so a missing file does
not fail the unit). The **live** P3 path — gated build → dispatch (trigger CI,
capture `run_id`) → verify (read the CI conclusion) — reads four variables from
that file at graph-build / dispatch time:
- `AGENT_TEAM_REPO_OWNER` — dispatch/verify target owner (fixed at factory time,
never read from pipeline state, so model output cannot redirect the target).
- `AGENT_TEAM_REPO_NAME` — dispatch/verify target repo (same fail-closed binding).
- `AGENT_TEAM_BASE_BRANCH` — PR base branch; optional, defaults to `main`.
- `AGENT_TEAM_CI_READ_TOKEN` — the **read-only** CI-result token used for the
verifier's authenticated conclusion read (falls back to `GITHUB_TOKEN`).
If owner, repo, or the CI-read token is unset, `serve()` degrades to the INERT P3
path (one WARNING + a `#agent-team` notice) rather than crash-looping the daemon
(Phase-0 Decision 5). No `pull-requests:write` / `contents:write` token and no
`AGENT_APPLY_APP_ID` / `AGENT_APPLY_APP_PRIVATE_KEY` may live in `~/secrev.env`:
the apply path mints its write token **inside** the CI runner from Actions
secrets, so the box holds no standing write credential. That invariant is
enforced by `scripts/assert_no_write_token.py` (the A2 audit) at provisioning and
in CI — run it before any deploy.
### Verifying the vars load
After installing/editing `~/secrev.env` and `systemctl daemon-reload` +
`systemctl restart agent-team-coordinator.service`, confirm the P3 vars reached
the **running coordinator process**.
> ⚠️ Do NOT use `systemctl show -p Environment` — it lists only inline
> `Environment=` directives and does **NOT** show vars loaded from
> `EnvironmentFile=` (which is how `~/secrev.env` is loaded). It comes back empty
> even when the vars are correctly loaded, so it is misleading here.
Read the actual process environment instead (requires sudo to read another
process's `environ`):
```
MP=$(systemctl show agent-team-coordinator.service -p MainPID --value)
sudo tr '\0' '\n' < /proc/$MP/environ | grep -E '^AGENT_TEAM_REPO|^AGENT_TEAM_BASE'
```
The output should list `AGENT_TEAM_REPO_OWNER`, `AGENT_TEAM_REPO_NAME`, and
`AGENT_TEAM_BASE_BRANCH` (if set). To confirm the live P3 wiring actually bound
(not the INERT path), check the code resolver directly:
```
cd ~/orchestrator/agent-team && set -a && source ~/secrev.env && set +a \
&& .venv/bin/python -c "from agent_team.coordinator import _p3_env_is_configured; print(_p3_env_is_configured())"
```
`True` means the live build+verify path is bound; `False` means the daemon is on
the INERT P3 path (fix `~/secrev.env`, reload, restart). The CI-read token is
satisfied by `AGENT_TEAM_CI_READ_TOKEN` or the read-only `GITHUB_TOKEN` fallback;
treat any token value in process output as sensitive. The box must hold NO
`AGENT_APPLY_APP_ID` / `AGENT_APPLY_APP_PRIVATE_KEY` / write token — verify with
`python scripts/assert_no_write_token.py` (and note its scope-detection caveat in
the script header: a fine-grained token's write capability is only definitively
confirmed by a live `POST /git/refs` probe returning `403`).
### Operator-initiated dispatch (P3 option-b)
The box is read-only, so its in-graph DISPATCH node fail-closes/parks — it never
pushes or triggers CI. Completing a dispatch is an explicit operator step with a
**just-in-time** write token (never stored in `~/secrev.env`):
```
# On the box (where the task's candidate_diff lives in the ledger), with a
# WRITE-capable token provided for THIS invocation only:
cd ~/orchestrator/agent-team
GH_TOKEN=<operator pull-requests+contents:write token> \
.venv/bin/python run-team.py dispatch <thread_id> --write-back
```
This reads the task's `candidate_diff` + declared scope from the checkpoint,
pushes the head branch, fires the `agent-team-apply-verify` `workflow_dispatch`,
prints the located `run_id`, and (`--write-back`) writes it into the task
checkpoint so the box's VERIFY binds to that run. CI then runs guard → build-test
→ pure-code gate; the privileged `gate-and-pr` job pauses at the `agent-apply`
environment for your **required-reviewer approval** before the draft PR opens.
Alternatively pass `--diff FILE --scope FILE` to dispatch a diff without reading
the ledger. The token is consumed by `gh`/`git` for the one command and never
persisted; the box returns to read-only at rest.