- run-team.py: add the 'dispatch <thread_id>' operator command (P3 option-b). The read-only box parks at DISPATCH; this completes it with a just-in-time WRITE token: reads candidate_diff + scope from the checkpoint (or --diff/--scope files), pushes the head branch + fires workflow_dispatch via dispatch_apply_verify, prints the located run_id, and (--write-back) writes it into the task checkpoint so VERIFY binds. +2 tests. - OPERATOR-RUNBOOK: fix the misleading 'systemctl show -p Environment' check (it does NOT show EnvironmentFile= vars) -> use /proc/<MainPID>/environ + _p3_env_is_configured(); document the operator-initiated dispatch flow + the fine-grained-token write-probe caveat. Suite green, ruff clean. Branch only; not merged.
532 lines
25 KiB
Markdown
532 lines
25 KiB
Markdown
# OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling
|
||
|
||
The on-call runbook for the always-on `agent-team-coordinator` daemon on the
|
||
`sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed
|
||
human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and
|
||
the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design
|
||
§5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement).
|
||
|
||
> **CLI used throughout: `run-team.py`** (the entry CLI), run from
|
||
> `~/orchestrator/agent-team` with the venv active so it hits the default ledger
|
||
> (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not**
|
||
> use `agent_team/operator_cli.py` — it has divergent verbs (no `show`, required
|
||
> `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md.
|
||
|
||
```bash
|
||
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
|
||
cd ~/orchestrator/agent-team && . .venv/bin/activate
|
||
```
|
||
|
||
## Verb reference (all confirmed in `run-team.py:build_parser`)
|
||
|
||
| Verb | Effect | Destructive? |
|
||
|---|---|---|
|
||
| `list` / `list --all` / `list --parked` / `list --status <state>` | read pending questions (default `open`) | no |
|
||
| `show <question_id>` | print one ledger row (JSON) | no |
|
||
| `redeliver <question_id>` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) |
|
||
| `expire <question_id> --confirm` | force `open`→`expired` | yes |
|
||
| `answer <question_id> --answer <p> [--via <id>] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes |
|
||
| `force-resume <question_id> --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes |
|
||
| `supersede <question_id> --confirm` | mark a stale `open`/`answered` row `superseded` | yes |
|
||
| `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no |
|
||
|
||
`--operator <name>` (global) sets the audit attribution; it defaults to the OS
|
||
login. Destructive verbs require `--confirm` and write an attempt-then-outcome
|
||
record to the audit log **before** mutating.
|
||
|
||
First triage for any incident:
|
||
|
||
```bash
|
||
systemctl status agent-team-coordinator.service
|
||
journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100
|
||
python3 run-team.py list --all # full ledger snapshot
|
||
python3 run-team.py list --parked # non-open rows (parked-task context)
|
||
```
|
||
|
||
---
|
||
|
||
## Incident 1 — Pipeline stall (tasks not advancing)
|
||
|
||
**Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the
|
||
journal shows no `tick` activity or repeated errors.
|
||
|
||
**Diagnose:**
|
||
```bash
|
||
systemctl is-active agent-team-coordinator.service # "active" expected
|
||
journalctl -u agent-team-coordinator.service -e | tail -60
|
||
# Daemon dead/looping on restart? Check the unit + deps:
|
||
.venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')"
|
||
```
|
||
|
||
**Recover:**
|
||
1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal
|
||
for the import/auth error. A missing venv dep (D-4/D-5) or a missing
|
||
`CLAUDE_CODE_OAUTH_TOKEN` is the usual cause — fix the env / venv, then:
|
||
```bash
|
||
sudo systemctl restart agent-team-coordinator.service
|
||
```
|
||
2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and
|
||
re-posts `open` rows that lost their `channel_ref`, so a clean restart
|
||
converges from the durable ledger. Confirm with `journalctl ... | tail` and
|
||
`list --all`.
|
||
3. A single task stuck `open` with a stale/missing post: re-deliver it.
|
||
```bash
|
||
python3 run-team.py show <qid>
|
||
python3 run-team.py redeliver <qid> # clears channel_ref; reconcile re-posts
|
||
```
|
||
|
||
**Escalate** (see the ladder) if a restart does not clear it within one tick
|
||
cadence and the journal shows a non-transient error.
|
||
|
||
---
|
||
|
||
## Incident 2 — Stuck / parked task
|
||
|
||
A task parks (design §6.6) when its clarifier question **expires** with no answer,
|
||
or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked =
|
||
the durable ledger row is no longer `open` (it is `expired`), and an ALARM was
|
||
raised, not spun on.
|
||
|
||
**Diagnose:**
|
||
```bash
|
||
python3 run-team.py list --parked
|
||
python3 run-team.py show <qid> # status, thread_id, deadline_at, answered_via
|
||
```
|
||
|
||
**Recover — depends on why it parked:**
|
||
|
||
- **Expired with no answer (the common case)** — un-park by reopening the expired
|
||
question; it is then re-delivered for an answer:
|
||
```bash
|
||
python3 run-team.py force-resume <qid> --confirm
|
||
# "force-resume: reopened expired question <qid>; it will be re-delivered"
|
||
python3 run-team.py show <qid> # status flips back to "open"
|
||
```
|
||
Then answer it (Slack or CLI) to drive it forward.
|
||
|
||
- **Answered but not yet resumed** — the recovery sweep handles it; `force-resume`
|
||
records intent and reports that (no mutation):
|
||
```bash
|
||
python3 run-team.py force-resume <qid> --confirm
|
||
# "force-resume: question <qid> is answered and pending resume; the recovery
|
||
# sweep will resume it (intent recorded)"
|
||
sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now
|
||
```
|
||
|
||
- **Stale / wrong question that should be abandoned** — supersede it so it stops
|
||
surfacing as parked context:
|
||
```bash
|
||
python3 run-team.py supersede <qid> --confirm
|
||
```
|
||
|
||
`MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a
|
||
Jira ticket per the ladder) rather than starving silently.
|
||
|
||
---
|
||
|
||
## Incident 3 — Failed human-in-the-loop resume
|
||
|
||
**Symptoms:** an answer was submitted (Slack or CLI) but the graph did not
|
||
advance.
|
||
|
||
**Diagnose:**
|
||
```bash
|
||
python3 run-team.py show <qid> # is status "answered"? what answered_via?
|
||
journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|<qid>"
|
||
```
|
||
|
||
**Recover:**
|
||
- **Status is `answered` but no resume drained** — the resume queue is in-process;
|
||
a daemon restart triggers the `recover()` sweep that re-drives `answered` rows
|
||
via the turn-guarded `ResumeWorker` (idempotent — an already-advanced thread
|
||
supersedes-and-skips):
|
||
```bash
|
||
sudo systemctl restart agent-team-coordinator.service
|
||
python3 run-team.py show <qid> # confirm it advanced
|
||
```
|
||
- **Live Slack answer never registered** — the listener fails closed. Check:
|
||
```bash
|
||
journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer"
|
||
# "rejecting Slack answer: owner allowlist is unconfigured ..." -> set
|
||
# AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart.
|
||
# "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is
|
||
# not in the allowlist. Add it, or answer via the CLI on their behalf:
|
||
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
|
||
```
|
||
- **Answer lost the compare-and-set (`not open`)** — the row was already
|
||
answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2.
|
||
|
||
---
|
||
|
||
## Incident 4 — Budget exhaustion mid-pipeline
|
||
|
||
The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total
|
||
nightly caps (design §6.1, §6.6). A stage that would breach the reserve **parks**
|
||
the task (deferred, not dropped) and ALARMs — it never loop-drains the pool.
|
||
|
||
**Diagnose:**
|
||
```bash
|
||
journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park"
|
||
# Inspect today's spend directly (no CLI verb for the budget ledger; read it).
|
||
# Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model,
|
||
# billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket.
|
||
sqlite3 state/agent_team.sqlite \
|
||
"SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok
|
||
FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;"
|
||
```
|
||
|
||
**Recover:**
|
||
- **Wait for the next budget window** — budget-exhausted roles/tasks are deferred
|
||
via the rotation pointer and picked up next cycle; this is the intended
|
||
behavior, not a failure. The parked task surfaces in `list --parked`.
|
||
- **Force a specific parked task forward now** (e.g. it is urgent and headroom has
|
||
since freed): un-park it and let the daemon re-run the stage within the
|
||
remaining cap:
|
||
```bash
|
||
python3 run-team.py force-resume <qid> --confirm
|
||
```
|
||
- **Confirm `ANTHROPIC_API_KEY` is absent** — its presence would silently meter to
|
||
API rates and blow the budget model (the billing seam pops it defensively, but
|
||
it must not be set):
|
||
```bash
|
||
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0
|
||
```
|
||
|
||
Do **not** raise the cap to push a task through without Adam's decision — budget
|
||
realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly
|
||
parks on budget across nights.
|
||
|
||
---
|
||
|
||
## Incident 5 — Transport outage (Slack / GitHub down or misconfigured)
|
||
|
||
**Symptoms:** clarifier posts fail; the journal shows transport errors or the
|
||
inbound listener is not started.
|
||
|
||
**Diagnose:**
|
||
```bash
|
||
journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport"
|
||
# Did the inbound listener start?
|
||
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
|
||
# "started (Socket Mode ...)" -> inbound up
|
||
# "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off
|
||
```
|
||
|
||
**Recover:**
|
||
- **Outbound post failed (Slack API / network)** — the ledger row stays `open`
|
||
with no `channel_ref` (the responder leaves it for reconcile). After the
|
||
transport recovers, the `recover()` sweep (on restart) or a manual `redeliver`
|
||
re-posts:
|
||
```bash
|
||
python3 run-team.py redeliver <qid>
|
||
```
|
||
- **Inbound listener down / never started** — the daemon still posts and expires;
|
||
it just cannot hear Slack. **The CLI `answer` path is the outage fallback** —
|
||
it runs the identical compare-and-set with no socket:
|
||
```bash
|
||
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
|
||
```
|
||
Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then
|
||
`sudo systemctl restart agent-team-coordinator.service` to re-establish Socket
|
||
Mode. (D-1: the listener starts only when the transport is live Slack AND
|
||
`SLACK_APP_TOKEN` is set; it is optional by design.)
|
||
- **Slack listener crash-looping** — the listener runs on an isolated daemon
|
||
thread; a crash is logged (`inbound Slack listener thread exited with an error`)
|
||
and takes down only the inbound socket, not the maintenance loop. Restart the
|
||
service to relaunch the thread once the cause is fixed.
|
||
|
||
---
|
||
|
||
## Incident 5b — WS0–WS5 surfaces (HTTP API, /new-task, handbook context)
|
||
|
||
These were added by the WS-rollout (see `agent-team/DEPLOY-R720.md` §4b). They
|
||
layer onto the coordinator; none of them should take down the maintenance loop.
|
||
|
||
- **HTTP API down / unreachable** — the WS1 FastAPI app (`agent_team/api.py`) is
|
||
a **separate, opt-in process** (`api.serve()`, `127.0.0.1:8765`, bearer auth),
|
||
**not** started by the coordinator daemon. If `/delegate` from Claude Code or
|
||
`POST /tasks` over HTTP stops working, the coordinator itself is unaffected —
|
||
check the API process separately:
|
||
```bash
|
||
curl -sS -o /dev/null -w '%{http_code}\n' \
|
||
-H "Authorization: Bearer $AGENT_TEAM_API_TOKEN" http://127.0.0.1:8765/tasks
|
||
# 405 = API up + authed (GET not allowed on /tasks); 000 = process down;
|
||
# 401 = AGENT_TEAM_API_TOKEN mismatch (client vs ~/secrev.env).
|
||
```
|
||
The API refuses to start if `AGENT_TEAM_API_TOKEN` is unset/empty (logs a
|
||
`RuntimeError`). Fix the token, restart the API process. Tasks already in the
|
||
ledger are unaffected — the API is only an *intake/invoke* front door; answer
|
||
via Slack or the CLI as usual.
|
||
|
||
- **`/new-task` Slack command not responding** — the WS2 slash command is
|
||
AUTHZ-01 owner-allowlist gated and routes through the same Socket Mode listener
|
||
as answers. If it silently does nothing, it is almost always the owner
|
||
allowlist (same failure mode as Incident 3's live-Slack path):
|
||
```bash
|
||
journalctl -u agent-team-coordinator.service -e | grep -iE "new-task|owner|unauthorized"
|
||
# unauthorized sender / unconfigured allowlist -> fix AGENT_TEAM_SLACK_OWNER_IDS
|
||
```
|
||
Fallback: start the task from the CLI (`run-team.py start --task "..."`) or the
|
||
HTTP API. If the listener itself is down, see Incident 5 (inbound listener).
|
||
|
||
- **Handbook dir missing → planner runs without handbook context** — the WS5
|
||
`context_provider` (`load_handbook_conventions`) is **fail-safe**: if
|
||
`SEA_HAVEN_HANDBOOK_DIR` (or `~/.sea-haven/engineering-handbook`) is missing or
|
||
unreadable it returns `""` and the planner runs normally, just without handbook
|
||
conventions injected. This is **degraded, not broken** — no park, no alarm.
|
||
Confirm and restore:
|
||
```bash
|
||
grep '^SEA_HAVEN_HANDBOOK_DIR=' ~/secrev.env
|
||
ls "$(grep '^SEA_HAVEN_HANDBOOK_DIR=' ~/secrev.env | cut -d= -f2)" # dir present + populated?
|
||
```
|
||
Re-sync the handbook (the deploy script does this) and restart the daemon so
|
||
the planner picks it back up. The WS3 dispatch node is **inert** (gated) and
|
||
should never appear in pipeline activity; if it does, treat as an unexpected
|
||
state and escalate.
|
||
|
||
---
|
||
|
||
## Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers)
|
||
|
||
These come from the nightly checker run, not the coordinator daemon (design §6.4,
|
||
§6.6):
|
||
|
||
- **COMPLACENCY ALARM** — a checker role missed a planted canary fault. That role
|
||
is **skipped** for the night (it never runs silently degraded). Recover: inspect
|
||
the canary corpus / the role's checker under `security-review/checkers/`, fix the
|
||
regression, re-run that checker's dry-run, and confirm the canary passes before
|
||
re-enabling.
|
||
- **COVERAGE ALARM** — a role slipped its rotation slot (e.g. budget-deferred). It
|
||
is **deferred via the rotation pointer, never dropped**, and picked up next
|
||
cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from
|
||
report history) and that the role runs on the next cadence; investigate only if
|
||
it slips repeatedly.
|
||
|
||
Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean
|
||
state, no auto-Jira/Notion writes from the checker itself. Persistent alarms
|
||
follow the escalation ladder below.
|
||
|
||
---
|
||
|
||
## Incident 7 — P3 rollback (unwind the apply/verify flip)
|
||
|
||
The P3 apply/verify CI workflow is **already LIVE** (flipped + provisioned
|
||
2026-06-22): the `agent-apply` environment with its required reviewer, the GitHub
|
||
App installation (`pull-requests: write`), the run-name/permissions edits on
|
||
`.github/workflows/agent-team-apply-verify.yml`, and `enforce_admins` branch
|
||
protection on `main`. `agent-team/scripts/p3_rollback.sh` is the inverse — it
|
||
restores those four privileged surfaces to a recorded baseline and asserts
|
||
post-restore == baseline (fail-closed: a mismatched restore aborts non-zero).
|
||
|
||
> **Run from the operator host (the Mac), NOT the box.** No standing box token is
|
||
> used; the script relies on `gh` being authenticated on the operator host (plus
|
||
> `git` for the post-merge revert path). The R720 daemon holds no write token by
|
||
> design (P3-PHASE0-DESIGN.md, "Out of scope").
|
||
|
||
**When to trigger:**
|
||
- The flip must be unwound (P3 apply/verify is being backed out), OR
|
||
- A **premature flip that RAN** — an apply/verify run executed before it should
|
||
have and opened a draft PR / branch that must not exist (use the `incident`
|
||
surface, below).
|
||
|
||
**Prerequisites:**
|
||
- A baseline JSON recorded **BEFORE** the flip (the restore target). Default path
|
||
`.security-review/p3-baseline.json` or `$P3_ROLLBACK_BASELINE`, override with
|
||
`--baseline FILE`. The script refuses to run if the baseline file is missing
|
||
(`record it BEFORE the flip`) — there is no inferred baseline.
|
||
- Target repo from the baseline's `repo` key, `--repo OWNER/NAME`, or
|
||
`$AGENT_TEAM_REPO_OWNER/$AGENT_TEAM_REPO_NAME`. If both the baseline and `--repo`
|
||
are set they must agree (fail-closed on mismatch).
|
||
|
||
**The four privileged surfaces it restores:**
|
||
1. **`workflow`** — the apply/verify flip itself. Pre-merge: close the flip PR +
|
||
delete its branch (`--flip-pr N [--flip-branch REF]`). Post-merge: `git revert`
|
||
the flip commit + push + re-run CI (`--merged --flip-commit SHA`). LIVE: also
|
||
reverts the workflow YAML to the recorded `workflow_baseline_sha` so the live
|
||
file matches baseline.
|
||
2. **`environment`** — the `agent-apply` environment: re-PUTs the recorded required
|
||
reviewer(s) + deployment branch policy, then asserts the environment exists.
|
||
3. **`app`** — the GitHub App: **always rotates (revokes) the installation token
|
||
first** (any minted token is now suspect), then `reduce`s permissions to the
|
||
baseline or `uninstall`s the installation per `app.action`.
|
||
4. **`protection`** — branch protection on `main`: re-enables `enforce_admins`
|
||
(`include_administrators`) and asserts it is ON. **Refuses to run if the
|
||
baseline's `include_administrators` is not `true`** — a rollback must never
|
||
leave a weaker posture than baseline.
|
||
|
||
`all` runs surfaces 1→4 in order.
|
||
|
||
**How to run — always `--dry-run` first, then `--apply`:**
|
||
|
||
The script is **destructive-safe by default**: with no `--apply` it only PRINTS
|
||
the plan (`[PLAN] ...` lines). `--apply` is the only thing that performs
|
||
mutations. Inspect the dry-run plan, confirm it targets the right repo + baseline,
|
||
then re-run with `--apply`.
|
||
|
||
```bash
|
||
cd ~/Documents/repositories/orchestrator/agent-team # operator host, gh authed
|
||
|
||
# 1. Dry-run: print the plan only, no mutations.
|
||
scripts/p3_rollback.sh all --baseline .security-review/p3-baseline.json
|
||
|
||
# 2. Apply, once the plan looks right (pre-merge flip example):
|
||
scripts/p3_rollback.sh all --apply \
|
||
--baseline .security-review/p3-baseline.json \
|
||
--flip-pr 123 --flip-branch agent-team/apply/flip
|
||
|
||
# Post-merge workflow revert instead of pre-merge close:
|
||
scripts/p3_rollback.sh workflow --apply --merged --flip-commit <SHA>
|
||
|
||
# A single surface at a time is fine too:
|
||
scripts/p3_rollback.sh protection --apply
|
||
```
|
||
|
||
**Premature-flip-that-RAN incident (the `incident` surface):**
|
||
|
||
Use this when a flip executed prematurely and opened a draft PR/branch. It runs
|
||
the full incident sequence — rotate, revert, audit (read-only), restore, note:
|
||
|
||
```bash
|
||
# Dry-run first (the audit/list steps run read-only either way):
|
||
scripts/p3_rollback.sh incident --baseline .security-review/p3-baseline.json
|
||
|
||
# Apply, with the premature draft PR if known:
|
||
scripts/p3_rollback.sh incident --apply \
|
||
--flip-pr <N> --flip-branch agent-team/apply/<task_id>
|
||
```
|
||
|
||
It performs, in order:
|
||
1. **Rotate the App installation token first** — anything the premature run minted
|
||
is suspect.
|
||
2. **Revert the draft PR / branch the App opened** — close `--flip-pr` + delete its
|
||
branch; if no `--flip-pr` is given it lists open `agent-team/apply/*` draft PRs
|
||
to triage.
|
||
3. **Audit the Checks trail** (read-only — always runs) — recent
|
||
`agent-team-apply-verify.yml` runs with conclusions.
|
||
4. **Restore the `agent-apply` environment + branch protection** to baseline
|
||
(surfaces 2 + 4).
|
||
5. **File an incident note** under
|
||
`.security-review/incidents/p3-premature-flip-<timestamp>.md` recording the
|
||
repo, baseline, flip PR/branch, and actions taken.
|
||
|
||
After any rollback, confirm the post-restore `[OK]` asserts printed (the script
|
||
exits non-zero if any assert failed), and record the action — for the `incident`
|
||
path the note is written automatically; otherwise note it on the Jira ticket per
|
||
the escalation ladder.
|
||
|
||
---
|
||
|
||
## Escalation ladder (design §5 / §6.6, resolves Q3)
|
||
|
||
Anything that does not clear on the first ALARM escalates — but **ALARM-only in
|
||
spirit** (nothing posts on a clean state):
|
||
|
||
1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE
|
||
alarm) that persists re-alarms to Slack each night it is still unresolved, on
|
||
a backoff so it does not spam.
|
||
2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the
|
||
condition still has not cleared, the coordinator opens an **INFRA** Jira ticket
|
||
so it cannot quietly linger. The same ladder applies to a parked task that
|
||
exceeds `MAX_PARK`.
|
||
3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via
|
||
Incidents 2–4 above; for a checker alarm, via Incident 6.
|
||
|
||
When you resolve an incident, record the action — destructive CLI verbs already
|
||
write an attributable attempt+outcome record to `state/audit.log.jsonl`; for
|
||
non-CLI recoveries note it on the Jira ticket.
|
||
|
||
---
|
||
|
||
## Post-incident
|
||
|
||
- Confirm the daemon is `active` and the ledger has no unexpected parked rows
|
||
(`list --parked`).
|
||
- If you restored the VM snapshot or wiped the ledger, re-run the relevant
|
||
PROVISIONING-RUNBOOK steps.
|
||
- Update `project_r720_agent_team` memory if the incident revealed a durable
|
||
fact (a new failure mode, a config that must change).
|
||
|
||
## P3 box env wiring (build → dispatch → verify)
|
||
|
||
The coordinator's environment is loaded from `EnvironmentFile=-/home/adam/secrev.env`
|
||
(declared in the unit; the leading `-` makes it optional so a missing file does
|
||
not fail the unit). The **live** P3 path — gated build → dispatch (trigger CI,
|
||
capture `run_id`) → verify (read the CI conclusion) — reads four variables from
|
||
that file at graph-build / dispatch time:
|
||
|
||
- `AGENT_TEAM_REPO_OWNER` — dispatch/verify target owner (fixed at factory time,
|
||
never read from pipeline state, so model output cannot redirect the target).
|
||
- `AGENT_TEAM_REPO_NAME` — dispatch/verify target repo (same fail-closed binding).
|
||
- `AGENT_TEAM_BASE_BRANCH` — PR base branch; optional, defaults to `main`.
|
||
- `AGENT_TEAM_CI_READ_TOKEN` — the **read-only** CI-result token used for the
|
||
verifier's authenticated conclusion read (falls back to `GITHUB_TOKEN`).
|
||
|
||
If owner, repo, or the CI-read token is unset, `serve()` degrades to the INERT P3
|
||
path (one WARNING + a `#agent-team` notice) rather than crash-looping the daemon
|
||
(Phase-0 Decision 5). No `pull-requests:write` / `contents:write` token and no
|
||
`AGENT_APPLY_APP_ID` / `AGENT_APPLY_APP_PRIVATE_KEY` may live in `~/secrev.env`:
|
||
the apply path mints its write token **inside** the CI runner from Actions
|
||
secrets, so the box holds no standing write credential. That invariant is
|
||
enforced by `scripts/assert_no_write_token.py` (the A2 audit) at provisioning and
|
||
in CI — run it before any deploy.
|
||
|
||
### Verifying the vars load
|
||
|
||
After installing/editing `~/secrev.env` and `systemctl daemon-reload` +
|
||
`systemctl restart agent-team-coordinator.service`, confirm the P3 vars reached
|
||
the **running coordinator process**.
|
||
|
||
> ⚠️ Do NOT use `systemctl show -p Environment` — it lists only inline
|
||
> `Environment=` directives and does **NOT** show vars loaded from
|
||
> `EnvironmentFile=` (which is how `~/secrev.env` is loaded). It comes back empty
|
||
> even when the vars are correctly loaded, so it is misleading here.
|
||
|
||
Read the actual process environment instead (requires sudo to read another
|
||
process's `environ`):
|
||
|
||
```
|
||
MP=$(systemctl show agent-team-coordinator.service -p MainPID --value)
|
||
sudo tr '\0' '\n' < /proc/$MP/environ | grep -E '^AGENT_TEAM_REPO|^AGENT_TEAM_BASE'
|
||
```
|
||
|
||
The output should list `AGENT_TEAM_REPO_OWNER`, `AGENT_TEAM_REPO_NAME`, and
|
||
`AGENT_TEAM_BASE_BRANCH` (if set). To confirm the live P3 wiring actually bound
|
||
(not the INERT path), check the code resolver directly:
|
||
|
||
```
|
||
cd ~/orchestrator/agent-team && set -a && source ~/secrev.env && set +a \
|
||
&& .venv/bin/python -c "from agent_team.coordinator import _p3_env_is_configured; print(_p3_env_is_configured())"
|
||
```
|
||
|
||
`True` means the live build+verify path is bound; `False` means the daemon is on
|
||
the INERT P3 path (fix `~/secrev.env`, reload, restart). The CI-read token is
|
||
satisfied by `AGENT_TEAM_CI_READ_TOKEN` or the read-only `GITHUB_TOKEN` fallback;
|
||
treat any token value in process output as sensitive. The box must hold NO
|
||
`AGENT_APPLY_APP_ID` / `AGENT_APPLY_APP_PRIVATE_KEY` / write token — verify with
|
||
`python scripts/assert_no_write_token.py` (and note its scope-detection caveat in
|
||
the script header: a fine-grained token's write capability is only definitively
|
||
confirmed by a live `POST /git/refs` probe returning `403`).
|
||
|
||
### Operator-initiated dispatch (P3 option-b)
|
||
|
||
The box is read-only, so its in-graph DISPATCH node fail-closes/parks — it never
|
||
pushes or triggers CI. Completing a dispatch is an explicit operator step with a
|
||
**just-in-time** write token (never stored in `~/secrev.env`):
|
||
|
||
```
|
||
# On the box (where the task's candidate_diff lives in the ledger), with a
|
||
# WRITE-capable token provided for THIS invocation only:
|
||
cd ~/orchestrator/agent-team
|
||
GH_TOKEN=<operator pull-requests+contents:write token> \
|
||
.venv/bin/python run-team.py dispatch <thread_id> --write-back
|
||
```
|
||
|
||
This reads the task's `candidate_diff` + declared scope from the checkpoint,
|
||
pushes the head branch, fires the `agent-team-apply-verify` `workflow_dispatch`,
|
||
prints the located `run_id`, and (`--write-back`) writes it into the task
|
||
checkpoint so the box's VERIFY binds to that run. CI then runs guard → build-test
|
||
→ pure-code gate; the privileged `gate-and-pr` job pauses at the `agent-apply`
|
||
environment for your **required-reviewer approval** before the draft PR opens.
|
||
Alternatively pass `--diff FILE --scope FILE` to dispatch a diff without reading
|
||
the ledger. The token is consumed by `gh`/`git` for the one command and never
|
||
persisted; the box returns to read-only at rest.
|