feat(agent-team): deploy-readiness — serve starts Slack listener + systemd + provisioning docs (#23)
* fix(agent-team): serve() starts the inbound Slack listener (D-1)
Coordinator.serve() now constructs and starts the SlackListener concurrently
with the tick/drain loop on a background daemon thread, but ONLY when the live
transport is a SlackTransport AND SLACK_APP_TOKEN is configured. When Slack is
not the transport or the app token is absent, serve() behaves exactly as before
(tick/recover only) — Slack is never made mandatory.
- New injectable build_listener seam + default_slack_listener_factory sharing
the coordinator's own transport, ledger db_path, and resume_queue put.
- AUTHZ-01 owner-allowlist + open-status CAS untouched: serve() sources
AGENT_TEAM_SLACK_OWNER_IDS in SlackListener.serve, which still fails closed.
- SlackListener.close() added for clean Socket Mode teardown on shutdown;
serve() stops the listener + joins the thread in a finally.
- Tests: start-when-Slack+app-token, no-start otherwise, clean shutdown,
idempotent start, serve start/stop around the loop, listener close().
* fix(agent-team): systemd unit loads ~/orchestrator/.env + uses venv python (D-2/D-7)
D-2: add EnvironmentFile=-/home/adam/orchestrator/.env (optional '-') so the P2
GPT-4.1 review loop's cross_reviewer sub-process can read the non-Claude provider
key once a task reaches REVIEW. Mirrors the sea-haven-secrev unit.
D-7: point ExecStart at the agent-team venv interpreter
(/home/adam/orchestrator/agent-team/.venv/bin/python) instead of
/usr/bin/env python3, which resolved the system interpreter without the
installed deps under systemd's PATH.
All hardening (NoNewPrivileges / ProtectSystem=full / ProtectHome=read-only /
ReadWritePaths) is retained unchanged (locked decision).
* docs(agent-team): land provisioning + operator runbooks under docs/provisioning
- PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5
intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub
App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked
FIXED so the demo can use the live Slack answer path.
- P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the
Slack and operator-CLI answer paths documented for all four exit criteria.
- DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the
operator-CLI divergence kept as provisioning notes.
- OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks,
failed HITL resumes, budget exhaustion, transport outages, and
COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the
re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6).
* fix(agent-team): supervise the Slack listener thread — recurring ALARM + respawn
sh-security-review (logic) MEDIUM: a crashed listener thread was logged once,
then the daemon ran on 'deaf' — posting clarifier questions but receiving no
answers, every gate silently parking, process never exiting so systemd
Restart=on-failure never fired. serve() now calls _supervise_slack_listener()
each pass: when the listener is enabled but its thread is dead, it emits a
recurring ERROR ALARM and respawns via the idempotent starter (self-heal).
No-op when alive or disabled. +3 tests. (authz detector: wiring clean — AUTHZ-01
fail-closed allowlist + open-status CAS intact, dead listener fails SAFE.)
2026-06-18 16:56:21 -04:00
|
|
|
|
# OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling
|
|
|
|
|
|
|
|
|
|
|
|
The on-call runbook for the always-on `agent-team-coordinator` daemon on the
|
|
|
|
|
|
`sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed
|
|
|
|
|
|
human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and
|
|
|
|
|
|
the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design
|
|
|
|
|
|
§5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement).
|
|
|
|
|
|
|
|
|
|
|
|
> **CLI used throughout: `run-team.py`** (the entry CLI), run from
|
|
|
|
|
|
> `~/orchestrator/agent-team` with the venv active so it hits the default ledger
|
|
|
|
|
|
> (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not**
|
|
|
|
|
|
> use `agent_team/operator_cli.py` — it has divergent verbs (no `show`, required
|
|
|
|
|
|
> `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md.
|
|
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
|
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
|
|
|
|
|
|
cd ~/orchestrator/agent-team && . .venv/bin/activate
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
## Verb reference (all confirmed in `run-team.py:build_parser`)
|
|
|
|
|
|
|
|
|
|
|
|
| Verb | Effect | Destructive? |
|
|
|
|
|
|
|---|---|---|
|
|
|
|
|
|
| `list` / `list --all` / `list --parked` / `list --status <state>` | read pending questions (default `open`) | no |
|
|
|
|
|
|
| `show <question_id>` | print one ledger row (JSON) | no |
|
|
|
|
|
|
| `redeliver <question_id>` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) |
|
|
|
|
|
|
| `expire <question_id> --confirm` | force `open`→`expired` | yes |
|
|
|
|
|
|
| `answer <question_id> --answer <p> [--via <id>] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes |
|
|
|
|
|
|
| `force-resume <question_id> --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes |
|
|
|
|
|
|
| `supersede <question_id> --confirm` | mark a stale `open`/`answered` row `superseded` | yes |
|
|
|
|
|
|
| `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no |
|
|
|
|
|
|
|
|
|
|
|
|
`--operator <name>` (global) sets the audit attribution; it defaults to the OS
|
|
|
|
|
|
login. Destructive verbs require `--confirm` and write an attempt-then-outcome
|
|
|
|
|
|
record to the audit log **before** mutating.
|
|
|
|
|
|
|
|
|
|
|
|
First triage for any incident:
|
|
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
|
systemctl status agent-team-coordinator.service
|
|
|
|
|
|
journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100
|
|
|
|
|
|
python3 run-team.py list --all # full ledger snapshot
|
|
|
|
|
|
python3 run-team.py list --parked # non-open rows (parked-task context)
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
## Incident 1 — Pipeline stall (tasks not advancing)
|
|
|
|
|
|
|
|
|
|
|
|
**Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the
|
|
|
|
|
|
journal shows no `tick` activity or repeated errors.
|
|
|
|
|
|
|
|
|
|
|
|
**Diagnose:**
|
|
|
|
|
|
```bash
|
|
|
|
|
|
systemctl is-active agent-team-coordinator.service # "active" expected
|
|
|
|
|
|
journalctl -u agent-team-coordinator.service -e | tail -60
|
|
|
|
|
|
# Daemon dead/looping on restart? Check the unit + deps:
|
|
|
|
|
|
.venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')"
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
**Recover:**
|
|
|
|
|
|
1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal
|
|
|
|
|
|
for the import/auth error. A missing venv dep (D-4/D-5) or a missing
|
|
|
|
|
|
`CLAUDE_CODE_OAUTH_TOKEN` is the usual cause — fix the env / venv, then:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
sudo systemctl restart agent-team-coordinator.service
|
|
|
|
|
|
```
|
|
|
|
|
|
2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and
|
|
|
|
|
|
re-posts `open` rows that lost their `channel_ref`, so a clean restart
|
|
|
|
|
|
converges from the durable ledger. Confirm with `journalctl ... | tail` and
|
|
|
|
|
|
`list --all`.
|
|
|
|
|
|
3. A single task stuck `open` with a stale/missing post: re-deliver it.
|
|
|
|
|
|
```bash
|
|
|
|
|
|
python3 run-team.py show <qid>
|
|
|
|
|
|
python3 run-team.py redeliver <qid> # clears channel_ref; reconcile re-posts
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
**Escalate** (see the ladder) if a restart does not clear it within one tick
|
|
|
|
|
|
cadence and the journal shows a non-transient error.
|
|
|
|
|
|
|
|
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
## Incident 2 — Stuck / parked task
|
|
|
|
|
|
|
|
|
|
|
|
A task parks (design §6.6) when its clarifier question **expires** with no answer,
|
|
|
|
|
|
or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked =
|
|
|
|
|
|
the durable ledger row is no longer `open` (it is `expired`), and an ALARM was
|
|
|
|
|
|
raised, not spun on.
|
|
|
|
|
|
|
|
|
|
|
|
**Diagnose:**
|
|
|
|
|
|
```bash
|
|
|
|
|
|
python3 run-team.py list --parked
|
|
|
|
|
|
python3 run-team.py show <qid> # status, thread_id, deadline_at, answered_via
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
**Recover — depends on why it parked:**
|
|
|
|
|
|
|
|
|
|
|
|
- **Expired with no answer (the common case)** — un-park by reopening the expired
|
|
|
|
|
|
question; it is then re-delivered for an answer:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
python3 run-team.py force-resume <qid> --confirm
|
|
|
|
|
|
# "force-resume: reopened expired question <qid>; it will be re-delivered"
|
|
|
|
|
|
python3 run-team.py show <qid> # status flips back to "open"
|
|
|
|
|
|
```
|
|
|
|
|
|
Then answer it (Slack or CLI) to drive it forward.
|
|
|
|
|
|
|
|
|
|
|
|
- **Answered but not yet resumed** — the recovery sweep handles it; `force-resume`
|
|
|
|
|
|
records intent and reports that (no mutation):
|
|
|
|
|
|
```bash
|
|
|
|
|
|
python3 run-team.py force-resume <qid> --confirm
|
|
|
|
|
|
# "force-resume: question <qid> is answered and pending resume; the recovery
|
|
|
|
|
|
# sweep will resume it (intent recorded)"
|
|
|
|
|
|
sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
- **Stale / wrong question that should be abandoned** — supersede it so it stops
|
|
|
|
|
|
surfacing as parked context:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
python3 run-team.py supersede <qid> --confirm
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
`MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a
|
|
|
|
|
|
Jira ticket per the ladder) rather than starving silently.
|
|
|
|
|
|
|
|
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
## Incident 3 — Failed human-in-the-loop resume
|
|
|
|
|
|
|
|
|
|
|
|
**Symptoms:** an answer was submitted (Slack or CLI) but the graph did not
|
|
|
|
|
|
advance.
|
|
|
|
|
|
|
|
|
|
|
|
**Diagnose:**
|
|
|
|
|
|
```bash
|
|
|
|
|
|
python3 run-team.py show <qid> # is status "answered"? what answered_via?
|
|
|
|
|
|
journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|<qid>"
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
**Recover:**
|
|
|
|
|
|
- **Status is `answered` but no resume drained** — the resume queue is in-process;
|
|
|
|
|
|
a daemon restart triggers the `recover()` sweep that re-drives `answered` rows
|
|
|
|
|
|
via the turn-guarded `ResumeWorker` (idempotent — an already-advanced thread
|
|
|
|
|
|
supersedes-and-skips):
|
|
|
|
|
|
```bash
|
|
|
|
|
|
sudo systemctl restart agent-team-coordinator.service
|
|
|
|
|
|
python3 run-team.py show <qid> # confirm it advanced
|
|
|
|
|
|
```
|
|
|
|
|
|
- **Live Slack answer never registered** — the listener fails closed. Check:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer"
|
|
|
|
|
|
# "rejecting Slack answer: owner allowlist is unconfigured ..." -> set
|
|
|
|
|
|
# AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart.
|
|
|
|
|
|
# "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is
|
|
|
|
|
|
# not in the allowlist. Add it, or answer via the CLI on their behalf:
|
|
|
|
|
|
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
|
|
|
|
|
|
```
|
|
|
|
|
|
- **Answer lost the compare-and-set (`not open`)** — the row was already
|
|
|
|
|
|
answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2.
|
|
|
|
|
|
|
|
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
## Incident 4 — Budget exhaustion mid-pipeline
|
|
|
|
|
|
|
|
|
|
|
|
The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total
|
|
|
|
|
|
nightly caps (design §6.1, §6.6). A stage that would breach the reserve **parks**
|
|
|
|
|
|
the task (deferred, not dropped) and ALARMs — it never loop-drains the pool.
|
|
|
|
|
|
|
|
|
|
|
|
**Diagnose:**
|
|
|
|
|
|
```bash
|
|
|
|
|
|
journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park"
|
|
|
|
|
|
# Inspect today's spend directly (no CLI verb for the budget ledger; read it).
|
|
|
|
|
|
# Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model,
|
|
|
|
|
|
# billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket.
|
|
|
|
|
|
sqlite3 state/agent_team.sqlite \
|
|
|
|
|
|
"SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok
|
|
|
|
|
|
FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;"
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
**Recover:**
|
|
|
|
|
|
- **Wait for the next budget window** — budget-exhausted roles/tasks are deferred
|
|
|
|
|
|
via the rotation pointer and picked up next cycle; this is the intended
|
|
|
|
|
|
behavior, not a failure. The parked task surfaces in `list --parked`.
|
|
|
|
|
|
- **Force a specific parked task forward now** (e.g. it is urgent and headroom has
|
|
|
|
|
|
since freed): un-park it and let the daemon re-run the stage within the
|
|
|
|
|
|
remaining cap:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
python3 run-team.py force-resume <qid> --confirm
|
|
|
|
|
|
```
|
|
|
|
|
|
- **Confirm `ANTHROPIC_API_KEY` is absent** — its presence would silently meter to
|
|
|
|
|
|
API rates and blow the budget model (the billing seam pops it defensively, but
|
|
|
|
|
|
it must not be set):
|
|
|
|
|
|
```bash
|
|
|
|
|
|
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
Do **not** raise the cap to push a task through without Adam's decision — budget
|
|
|
|
|
|
realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly
|
|
|
|
|
|
parks on budget across nights.
|
|
|
|
|
|
|
|
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
## Incident 5 — Transport outage (Slack / GitHub down or misconfigured)
|
|
|
|
|
|
|
|
|
|
|
|
**Symptoms:** clarifier posts fail; the journal shows transport errors or the
|
|
|
|
|
|
inbound listener is not started.
|
|
|
|
|
|
|
|
|
|
|
|
**Diagnose:**
|
|
|
|
|
|
```bash
|
|
|
|
|
|
journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport"
|
|
|
|
|
|
# Did the inbound listener start?
|
|
|
|
|
|
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
|
|
|
|
|
|
# "started (Socket Mode ...)" -> inbound up
|
|
|
|
|
|
# "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
**Recover:**
|
|
|
|
|
|
- **Outbound post failed (Slack API / network)** — the ledger row stays `open`
|
|
|
|
|
|
with no `channel_ref` (the responder leaves it for reconcile). After the
|
|
|
|
|
|
transport recovers, the `recover()` sweep (on restart) or a manual `redeliver`
|
|
|
|
|
|
re-posts:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
python3 run-team.py redeliver <qid>
|
|
|
|
|
|
```
|
|
|
|
|
|
- **Inbound listener down / never started** — the daemon still posts and expires;
|
|
|
|
|
|
it just cannot hear Slack. **The CLI `answer` path is the outage fallback** —
|
|
|
|
|
|
it runs the identical compare-and-set with no socket:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
|
|
|
|
|
|
```
|
|
|
|
|
|
Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then
|
|
|
|
|
|
`sudo systemctl restart agent-team-coordinator.service` to re-establish Socket
|
|
|
|
|
|
Mode. (D-1: the listener starts only when the transport is live Slack AND
|
|
|
|
|
|
`SLACK_APP_TOKEN` is set; it is optional by design.)
|
|
|
|
|
|
- **Slack listener crash-looping** — the listener runs on an isolated daemon
|
|
|
|
|
|
thread; a crash is logged (`inbound Slack listener thread exited with an error`)
|
|
|
|
|
|
and takes down only the inbound socket, not the maintenance loop. Restart the
|
|
|
|
|
|
service to relaunch the thread once the cause is fixed.
|
|
|
|
|
|
|
|
|
|
|
|
---
|
|
|
|
|
|
|
2026-06-23 12:56:08 -04:00
|
|
|
|
## Incident 5b — WS0–WS5 surfaces (HTTP API, /new-task, handbook context)
|
|
|
|
|
|
|
|
|
|
|
|
These were added by the WS-rollout (see `agent-team/DEPLOY-R720.md` §4b). They
|
|
|
|
|
|
layer onto the coordinator; none of them should take down the maintenance loop.
|
|
|
|
|
|
|
|
|
|
|
|
- **HTTP API down / unreachable** — the WS1 FastAPI app (`agent_team/api.py`) is
|
|
|
|
|
|
a **separate, opt-in process** (`api.serve()`, `127.0.0.1:8765`, bearer auth),
|
|
|
|
|
|
**not** started by the coordinator daemon. If `/delegate` from Claude Code or
|
|
|
|
|
|
`POST /tasks` over HTTP stops working, the coordinator itself is unaffected —
|
|
|
|
|
|
check the API process separately:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
curl -sS -o /dev/null -w '%{http_code}\n' \
|
|
|
|
|
|
-H "Authorization: Bearer $AGENT_TEAM_API_TOKEN" http://127.0.0.1:8765/tasks
|
|
|
|
|
|
# 405 = API up + authed (GET not allowed on /tasks); 000 = process down;
|
|
|
|
|
|
# 401 = AGENT_TEAM_API_TOKEN mismatch (client vs ~/secrev.env).
|
|
|
|
|
|
```
|
|
|
|
|
|
The API refuses to start if `AGENT_TEAM_API_TOKEN` is unset/empty (logs a
|
|
|
|
|
|
`RuntimeError`). Fix the token, restart the API process. Tasks already in the
|
|
|
|
|
|
ledger are unaffected — the API is only an *intake/invoke* front door; answer
|
|
|
|
|
|
via Slack or the CLI as usual.
|
|
|
|
|
|
|
|
|
|
|
|
- **`/new-task` Slack command not responding** — the WS2 slash command is
|
|
|
|
|
|
AUTHZ-01 owner-allowlist gated and routes through the same Socket Mode listener
|
|
|
|
|
|
as answers. If it silently does nothing, it is almost always the owner
|
|
|
|
|
|
allowlist (same failure mode as Incident 3's live-Slack path):
|
|
|
|
|
|
```bash
|
|
|
|
|
|
journalctl -u agent-team-coordinator.service -e | grep -iE "new-task|owner|unauthorized"
|
|
|
|
|
|
# unauthorized sender / unconfigured allowlist -> fix AGENT_TEAM_SLACK_OWNER_IDS
|
|
|
|
|
|
```
|
|
|
|
|
|
Fallback: start the task from the CLI (`run-team.py start --task "..."`) or the
|
|
|
|
|
|
HTTP API. If the listener itself is down, see Incident 5 (inbound listener).
|
|
|
|
|
|
|
|
|
|
|
|
- **Handbook dir missing → planner runs without handbook context** — the WS5
|
|
|
|
|
|
`context_provider` (`load_handbook_conventions`) is **fail-safe**: if
|
|
|
|
|
|
`SEA_HAVEN_HANDBOOK_DIR` (or `~/.sea-haven/engineering-handbook`) is missing or
|
|
|
|
|
|
unreadable it returns `""` and the planner runs normally, just without handbook
|
|
|
|
|
|
conventions injected. This is **degraded, not broken** — no park, no alarm.
|
|
|
|
|
|
Confirm and restore:
|
|
|
|
|
|
```bash
|
|
|
|
|
|
grep '^SEA_HAVEN_HANDBOOK_DIR=' ~/secrev.env
|
|
|
|
|
|
ls "$(grep '^SEA_HAVEN_HANDBOOK_DIR=' ~/secrev.env | cut -d= -f2)" # dir present + populated?
|
|
|
|
|
|
```
|
|
|
|
|
|
Re-sync the handbook (the deploy script does this) and restart the daemon so
|
|
|
|
|
|
the planner picks it back up. The WS3 dispatch node is **inert** (gated) and
|
|
|
|
|
|
should never appear in pipeline activity; if it does, treat as an unexpected
|
|
|
|
|
|
state and escalate.
|
|
|
|
|
|
|
|
|
|
|
|
---
|
|
|
|
|
|
|
feat(agent-team): deploy-readiness — serve starts Slack listener + systemd + provisioning docs (#23)
* fix(agent-team): serve() starts the inbound Slack listener (D-1)
Coordinator.serve() now constructs and starts the SlackListener concurrently
with the tick/drain loop on a background daemon thread, but ONLY when the live
transport is a SlackTransport AND SLACK_APP_TOKEN is configured. When Slack is
not the transport or the app token is absent, serve() behaves exactly as before
(tick/recover only) — Slack is never made mandatory.
- New injectable build_listener seam + default_slack_listener_factory sharing
the coordinator's own transport, ledger db_path, and resume_queue put.
- AUTHZ-01 owner-allowlist + open-status CAS untouched: serve() sources
AGENT_TEAM_SLACK_OWNER_IDS in SlackListener.serve, which still fails closed.
- SlackListener.close() added for clean Socket Mode teardown on shutdown;
serve() stops the listener + joins the thread in a finally.
- Tests: start-when-Slack+app-token, no-start otherwise, clean shutdown,
idempotent start, serve start/stop around the loop, listener close().
* fix(agent-team): systemd unit loads ~/orchestrator/.env + uses venv python (D-2/D-7)
D-2: add EnvironmentFile=-/home/adam/orchestrator/.env (optional '-') so the P2
GPT-4.1 review loop's cross_reviewer sub-process can read the non-Claude provider
key once a task reaches REVIEW. Mirrors the sea-haven-secrev unit.
D-7: point ExecStart at the agent-team venv interpreter
(/home/adam/orchestrator/agent-team/.venv/bin/python) instead of
/usr/bin/env python3, which resolved the system interpreter without the
installed deps under systemd's PATH.
All hardening (NoNewPrivileges / ProtectSystem=full / ProtectHome=read-only /
ReadWritePaths) is retained unchanged (locked decision).
* docs(agent-team): land provisioning + operator runbooks under docs/provisioning
- PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5
intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub
App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked
FIXED so the demo can use the live Slack answer path.
- P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the
Slack and operator-CLI answer paths documented for all four exit criteria.
- DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the
operator-CLI divergence kept as provisioning notes.
- OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks,
failed HITL resumes, budget exhaustion, transport outages, and
COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the
re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6).
* fix(agent-team): supervise the Slack listener thread — recurring ALARM + respawn
sh-security-review (logic) MEDIUM: a crashed listener thread was logged once,
then the daemon ran on 'deaf' — posting clarifier questions but receiving no
answers, every gate silently parking, process never exiting so systemd
Restart=on-failure never fired. serve() now calls _supervise_slack_listener()
each pass: when the listener is enabled but its thread is dead, it emits a
recurring ERROR ALARM and respawns via the idempotent starter (self-heal).
No-op when alive or disabled. +3 tests. (authz detector: wiring clean — AUTHZ-01
fail-closed allowlist + open-status CAS intact, dead listener fails SAFE.)
2026-06-18 16:56:21 -04:00
|
|
|
|
## Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers)
|
|
|
|
|
|
|
|
|
|
|
|
These come from the nightly checker run, not the coordinator daemon (design §6.4,
|
|
|
|
|
|
§6.6):
|
|
|
|
|
|
|
|
|
|
|
|
- **COMPLACENCY ALARM** — a checker role missed a planted canary fault. That role
|
|
|
|
|
|
is **skipped** for the night (it never runs silently degraded). Recover: inspect
|
|
|
|
|
|
the canary corpus / the role's checker under `security-review/checkers/`, fix the
|
|
|
|
|
|
regression, re-run that checker's dry-run, and confirm the canary passes before
|
|
|
|
|
|
re-enabling.
|
|
|
|
|
|
- **COVERAGE ALARM** — a role slipped its rotation slot (e.g. budget-deferred). It
|
|
|
|
|
|
is **deferred via the rotation pointer, never dropped**, and picked up next
|
|
|
|
|
|
cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from
|
|
|
|
|
|
report history) and that the role runs on the next cadence; investigate only if
|
|
|
|
|
|
it slips repeatedly.
|
|
|
|
|
|
|
|
|
|
|
|
Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean
|
|
|
|
|
|
state, no auto-Jira/Notion writes from the checker itself. Persistent alarms
|
|
|
|
|
|
follow the escalation ladder below.
|
|
|
|
|
|
|
|
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
## Escalation ladder (design §5 / §6.6, resolves Q3)
|
|
|
|
|
|
|
|
|
|
|
|
Anything that does not clear on the first ALARM escalates — but **ALARM-only in
|
|
|
|
|
|
spirit** (nothing posts on a clean state):
|
|
|
|
|
|
|
|
|
|
|
|
1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE
|
|
|
|
|
|
alarm) that persists re-alarms to Slack each night it is still unresolved, on
|
|
|
|
|
|
a backoff so it does not spam.
|
|
|
|
|
|
2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the
|
|
|
|
|
|
condition still has not cleared, the coordinator opens an **INFRA** Jira ticket
|
|
|
|
|
|
so it cannot quietly linger. The same ladder applies to a parked task that
|
|
|
|
|
|
exceeds `MAX_PARK`.
|
|
|
|
|
|
3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via
|
|
|
|
|
|
Incidents 2–4 above; for a checker alarm, via Incident 6.
|
|
|
|
|
|
|
|
|
|
|
|
When you resolve an incident, record the action — destructive CLI verbs already
|
|
|
|
|
|
write an attributable attempt+outcome record to `state/audit.log.jsonl`; for
|
|
|
|
|
|
non-CLI recoveries note it on the Jira ticket.
|
|
|
|
|
|
|
|
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
## Post-incident
|
|
|
|
|
|
|
|
|
|
|
|
- Confirm the daemon is `active` and the ledger has no unexpected parked rows
|
|
|
|
|
|
(`list --parked`).
|
|
|
|
|
|
- If you restored the VM snapshot or wiped the ledger, re-run the relevant
|
|
|
|
|
|
PROVISIONING-RUNBOOK steps.
|
|
|
|
|
|
- Update `project_r720_agent_team` memory if the incident revealed a durable
|
|
|
|
|
|
fact (a new failure mode, a config that must change).
|