2026-06-18 12:56:42 -04:00
|
|
|
|
# P1 - agent-team Plane-2 coordinator deployment (R720 VM)
|
|
|
|
|
|
|
|
|
|
|
|
Status: **P1 DEPLOY ARTIFACTS** - the systemd unit + this runbook. The coordinator
|
|
|
|
|
|
daemon (`run-team.py serve`) is built by a separate agent; this file is the
|
|
|
|
|
|
operator runbook for standing it up on the always-on R720 VM.
|
|
|
|
|
|
|
|
|
|
|
|
The agent-team coordinator is the long-running Plane-2 brain: it drives the
|
|
|
|
|
|
LangGraph pipeline, owns the durable `pending_questions` ledger, and runs the
|
|
|
|
|
|
Slack Socket Mode inbound listener that receives clarifier answers. It shares the
|
|
|
|
|
|
`sh-secrev` VM and the `~/secrev.env` secrets file with the Path B security sweep
|
|
|
|
|
|
(see `../security-review/DEPLOY-R720.md`), but it is a **service** (always-on),
|
|
|
|
|
|
not a timer-driven oneshot.
|
|
|
|
|
|
|
|
|
|
|
|
## Host
|
|
|
|
|
|
|
|
|
|
|
|
- **Hypervisor:** R720 at `10.10.60.40` (Windows Server 2022, Hyper-V role).
|
|
|
|
|
|
- **VM:** `sh-secrev`, always-on Ubuntu 24.04, Gen2, **4GB / 2 vCPU / 40GB**
|
|
|
|
|
|
dynamic vhdx at `10.10.60.120`.
|
|
|
|
|
|
- **Reach it:** `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only,
|
|
|
|
|
|
NOPASSWD sudo).
|
|
|
|
|
|
|
|
|
|
|
|
Operate on the VM, not from the Mac against the host by hand.
|
|
|
|
|
|
|
|
|
|
|
|
## 1. SNAPSHOT FIRST
|
|
|
|
|
|
|
|
|
|
|
|
**Standing rule: snapshot the VM before any provisioning change.** This box is a
|
|
|
|
|
|
4GB / 2 vCPU / 40GB VM at `10.10.60.120`. Take a Hyper-V checkpoint on the R720
|
|
|
|
|
|
host **before** you install pip deps, the unit, or touch `~/secrev.env`, so the
|
|
|
|
|
|
whole change is one-command reversible (see ROLLBACK). Do not skip this because
|
|
|
|
|
|
"it is only a pip install" - a bad dep set or a wedged service is exactly what
|
|
|
|
|
|
the snapshot exists to undo.
|
|
|
|
|
|
|
|
|
|
|
|
## 2. Prereqs on the VM
|
|
|
|
|
|
|
|
|
|
|
|
Already present from the secrev deploy:
|
|
|
|
|
|
|
|
|
|
|
|
- **Python 3** (3.12) and the `claude` CLI (Node) - the subscription-auth path.
|
|
|
|
|
|
- **Repo:** `~/orchestrator/` (rsync from the Mac, NOT a git clone). The
|
|
|
|
|
|
agent-team package lives at `~/orchestrator/agent-team/`.
|
|
|
|
|
|
|
|
|
|
|
|
New for the coordinator - a dedicated venv under `agent-team/.venv` (excluded
|
|
|
|
|
|
from rsync) with the coordinator/transport deps:
|
|
|
|
|
|
|
|
|
|
|
|
| pip dep | Why |
|
|
|
|
|
|
|---|---|
|
|
|
|
|
|
| `langgraph` | the coordinator pipeline graph |
|
|
|
|
|
|
| `langgraph-checkpoint-sqlite` | `SqliteSaver` checkpointer against the ledger DB |
|
|
|
|
|
|
| `claude-agent-sdk` | subscription-auth Claude invocation seam |
|
|
|
|
|
|
| `slack_sdk` | Slack Web API (post questions, `chat:write`) |
|
|
|
|
|
|
| `slack_bolt` | Socket Mode inbound listener (receive answers) |
|
|
|
|
|
|
|
|
|
|
|
|
The ledger DB defaults to `agent-team/state/agent_team.sqlite`; the audit log to
|
|
|
|
|
|
`agent-team/state/audit.log.jsonl`. Both live under `state/` (gitignored,
|
|
|
|
|
|
never committed).
|
|
|
|
|
|
|
|
|
|
|
|
## 3. Secrets - append to `~/secrev.env` (mode 600, never committed)
|
|
|
|
|
|
|
|
|
|
|
|
The coordinator reads its secrets from the same `~/secrev.env` the secrev sweep
|
|
|
|
|
|
uses. Append these (do not echo them into shell history files; lock the file
|
|
|
|
|
|
down after):
|
|
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
|
echo 'CLAUDE_CODE_OAUTH_TOKEN=...' >> ~/secrev.env # from `claude setup-token`
|
|
|
|
|
|
echo 'SLACK_BOT_TOKEN=xoxb-...' >> ~/secrev.env # bot token, chat:write
|
|
|
|
|
|
echo 'SLACK_APP_TOKEN=xapp-...' >> ~/secrev.env # app-level, connections:write (Socket Mode)
|
|
|
|
|
|
echo 'SLACK_CHANNEL_ID=C0XXXXXXX' >> ~/secrev.env # target clarifier channel
|
|
|
|
|
|
echo 'AGENT_TEAM_SLACK_OWNER_IDS=U0XXXXXXX' >> ~/secrev.env # authorized answerer(s), comma-separated
|
|
|
|
|
|
chmod 600 ~/secrev.env
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
- `CLAUDE_CODE_OAUTH_TOKEN` - subscription OAuth from `claude setup-token`. The
|
|
|
|
|
|
same token type the secrev sweep uses.
|
|
|
|
|
|
- `SLACK_BOT_TOKEN` (`xoxb-`) - bot token with `chat:write`; posts questions.
|
|
|
|
|
|
- `SLACK_APP_TOKEN` (`xapp-`) - app-level token with `connections:write`;
|
|
|
|
|
|
**required for Socket Mode** (opens the inbound WebSocket that hears answers).
|
|
|
|
|
|
- `SLACK_CHANNEL_ID` - the channel id the coordinator posts clarifiers to.
|
|
|
|
|
|
- `AGENT_TEAM_SLACK_OWNER_IDS` - comma-separated Slack **user ids** of the
|
|
|
|
|
|
authorized answerers (e.g. Adam's `U…` id). The inbound listener enforces this
|
|
|
|
|
|
as an owner allowlist (AUTHZ-01): only a sender in this set may answer/steer
|
|
|
|
|
|
the pipeline. **The listener fails closed** - if this is unset/empty it rejects
|
|
|
|
|
|
**every** answer (logs a warning naming `AGENT_TEAM_SLACK_OWNER_IDS`), so it
|
|
|
|
|
|
must be set for the human gate to function. Look up your user id via Slack
|
|
|
|
|
|
profile → "Copy member ID", or the `users.identity` / `auth.test` API.
|
|
|
|
|
|
|
|
|
|
|
|
**CRITICAL:** `ANTHROPIC_API_KEY` must **NOT** be set on this host. A raw API key
|
|
|
|
|
|
would silently win over the subscription OAuth and meter to API rates. The box
|
|
|
|
|
|
runs on subscription OAuth only.
|
|
|
|
|
|
|
fix(agent-team): dispatch via GitHub App so P3 reaches CI (run_id resolves)
The P3 dispatcher's default seams shell out to gh/git, but the R720 box has
no gh and a read-only PAT with no Actions scope — so dispatch_apply_verify
returned no run_id and every task parked at verify ("dispatch unresolved").
Add a GitHub-App auth path: the box mints short-lived (~1h) installation
access tokens from the App private key and uses them for the three dispatch
seams, removing the gh dependency.
- agent_team/github_app.py (new): mint_installation_token (RS256 App JWT,
iss=app_id, iat backdated 60s, exp 9 min; POST /access_tokens) + a lazy
TokenProvider that caches and re-mints near expiry. Secret-safe: the JWT
and token are never logged, never in an exception message, never persisted.
- dispatcher.py: app_branch_pusher / app_workflow_dispatcher / app_run_locator
(additive; gh/git _default_* left untouched). Push auth rides a host-scoped
http.extraHeader via GIT_CONFIG_* env (token never in argv/ps); the REST
run locator maps id->databaseId / created_at->createdAt into select_run_id
and surfaces 4xx promptly instead of silently exhausting the poll window.
- coordinator.py: default_dispatch_node_factory binds the App seams when
AGENT_TEAM_GH_APP_ID / _INSTALLATION_ID / _PRIVATE_KEY are all set; partial
or unreadable config logs one warning and falls back to gh-default (never
raises at serve-start).
- requirements.txt: pin PyJWT, cryptography, requests (App seams + CI fetcher).
- DEPLOY-R720.md / README.md: App dispatch config, permission/scope audit,
env-precedence check, key rotation/revocation + incident response.
Tests: +18 (test_github_app.py new; dispatcher/coordinator additions) covering
JWT claims, cache/re-mint, token-scrub-on-error, REST field mapping + run-name
correlation, and the partial-env inert fallback. Full suite 1523 passing.
2026-06-24 17:07:01 -04:00
|
|
|
|
### 3a. P3 dispatch auth — GitHub App installation tokens (box-side)
|
|
|
|
|
|
|
|
|
|
|
|
The P3 dispatcher (`agent_team/dispatcher.py`) carries an approved diff into org
|
|
|
|
|
|
CI: it pushes the candidate head branch and triggers the
|
|
|
|
|
|
`agent-team-apply-verify.yml` `workflow_dispatch`, then locates the resulting
|
|
|
|
|
|
run id. By default those three seams shell out to `gh`/`git`. **`gh` is not
|
|
|
|
|
|
installed on the box and the box's read-only PAT has no Actions permission**, so
|
|
|
|
|
|
the default path cannot dispatch. Instead the box authenticates with a **GitHub
|
|
|
|
|
|
App**: it holds the App private key and mints short-lived (~1h) **installation
|
|
|
|
|
|
access tokens** on demand (`agent_team/github_app.py`). Set all three of these to
|
|
|
|
|
|
turn the App path on (all-or-nothing — partial config logs one warning and falls
|
|
|
|
|
|
back to the inert gh-default path, it never crashes serve):
|
|
|
|
|
|
|
2026-06-24 17:34:08 -04:00
|
|
|
|
The App is the existing **`agent-team-apply`** App: **app_id `4119505`**,
|
|
|
|
|
|
**installation `141992144`** (on `Sea-Haven-Industries/orchestrator`). These two
|
|
|
|
|
|
IDs are not secrets (only the `.pem` is). Set all three vars (all-or-nothing —
|
|
|
|
|
|
partial config logs one warning and falls back to the inert gh-default path, it
|
|
|
|
|
|
never crashes serve):
|
|
|
|
|
|
|
fix(agent-team): dispatch via GitHub App so P3 reaches CI (run_id resolves)
The P3 dispatcher's default seams shell out to gh/git, but the R720 box has
no gh and a read-only PAT with no Actions scope — so dispatch_apply_verify
returned no run_id and every task parked at verify ("dispatch unresolved").
Add a GitHub-App auth path: the box mints short-lived (~1h) installation
access tokens from the App private key and uses them for the three dispatch
seams, removing the gh dependency.
- agent_team/github_app.py (new): mint_installation_token (RS256 App JWT,
iss=app_id, iat backdated 60s, exp 9 min; POST /access_tokens) + a lazy
TokenProvider that caches and re-mints near expiry. Secret-safe: the JWT
and token are never logged, never in an exception message, never persisted.
- dispatcher.py: app_branch_pusher / app_workflow_dispatcher / app_run_locator
(additive; gh/git _default_* left untouched). Push auth rides a host-scoped
http.extraHeader via GIT_CONFIG_* env (token never in argv/ps); the REST
run locator maps id->databaseId / created_at->createdAt into select_run_id
and surfaces 4xx promptly instead of silently exhausting the poll window.
- coordinator.py: default_dispatch_node_factory binds the App seams when
AGENT_TEAM_GH_APP_ID / _INSTALLATION_ID / _PRIVATE_KEY are all set; partial
or unreadable config logs one warning and falls back to gh-default (never
raises at serve-start).
- requirements.txt: pin PyJWT, cryptography, requests (App seams + CI fetcher).
- DEPLOY-R720.md / README.md: App dispatch config, permission/scope audit,
env-precedence check, key rotation/revocation + incident response.
Tests: +18 (test_github_app.py new; dispatcher/coordinator additions) covering
JWT claims, cache/re-mint, token-scrub-on-error, REST field mapping + run-name
correlation, and the partial-env inert fallback. Full suite 1523 passing.
2026-06-24 17:07:01 -04:00
|
|
|
|
```
|
2026-06-24 17:34:08 -04:00
|
|
|
|
echo 'AGENT_TEAM_GH_APP_ID=4119505' >> ~/secrev.env
|
|
|
|
|
|
echo 'AGENT_TEAM_GH_APP_INSTALLATION_ID=141992144' >> ~/secrev.env
|
2026-06-24 17:37:07 -04:00
|
|
|
|
echo 'AGENT_TEAM_GH_APP_PRIVATE_KEY=/home/adam/.ssh/agent-team-apply.pem' >> ~/secrev.env # PATH to the .pem
|
fix(agent-team): dispatch via GitHub App so P3 reaches CI (run_id resolves)
The P3 dispatcher's default seams shell out to gh/git, but the R720 box has
no gh and a read-only PAT with no Actions scope — so dispatch_apply_verify
returned no run_id and every task parked at verify ("dispatch unresolved").
Add a GitHub-App auth path: the box mints short-lived (~1h) installation
access tokens from the App private key and uses them for the three dispatch
seams, removing the gh dependency.
- agent_team/github_app.py (new): mint_installation_token (RS256 App JWT,
iss=app_id, iat backdated 60s, exp 9 min; POST /access_tokens) + a lazy
TokenProvider that caches and re-mints near expiry. Secret-safe: the JWT
and token are never logged, never in an exception message, never persisted.
- dispatcher.py: app_branch_pusher / app_workflow_dispatcher / app_run_locator
(additive; gh/git _default_* left untouched). Push auth rides a host-scoped
http.extraHeader via GIT_CONFIG_* env (token never in argv/ps); the REST
run locator maps id->databaseId / created_at->createdAt into select_run_id
and surfaces 4xx promptly instead of silently exhausting the poll window.
- coordinator.py: default_dispatch_node_factory binds the App seams when
AGENT_TEAM_GH_APP_ID / _INSTALLATION_ID / _PRIVATE_KEY are all set; partial
or unreadable config logs one warning and falls back to gh-default (never
raises at serve-start).
- requirements.txt: pin PyJWT, cryptography, requests (App seams + CI fetcher).
- DEPLOY-R720.md / README.md: App dispatch config, permission/scope audit,
env-precedence check, key rotation/revocation + incident response.
Tests: +18 (test_github_app.py new; dispatcher/coordinator additions) covering
JWT claims, cache/re-mint, token-scrub-on-error, REST field mapping + run-name
correlation, and the partial-env inert fallback. Full suite 1523 passing.
2026-06-24 17:07:01 -04:00
|
|
|
|
chmod 600 ~/secrev.env
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
Place the App private key on the box and lock it down — it is a write-capable
|
2026-06-24 17:34:08 -04:00
|
|
|
|
credential and must be owner-only. The CI-side Actions secret
|
|
|
|
|
|
`AGENT_APPLY_APP_PRIVATE_KEY` is write-only and cannot be re-exported, so the
|
|
|
|
|
|
box gets its OWN freshly-generated private key for the same App (generate one
|
|
|
|
|
|
under the App settings — the App holds several keys; the CI key keeps working).
|
|
|
|
|
|
Secure-copy it from the Mac (do NOT commit it; it is not in the repo):
|
fix(agent-team): dispatch via GitHub App so P3 reaches CI (run_id resolves)
The P3 dispatcher's default seams shell out to gh/git, but the R720 box has
no gh and a read-only PAT with no Actions scope — so dispatch_apply_verify
returned no run_id and every task parked at verify ("dispatch unresolved").
Add a GitHub-App auth path: the box mints short-lived (~1h) installation
access tokens from the App private key and uses them for the three dispatch
seams, removing the gh dependency.
- agent_team/github_app.py (new): mint_installation_token (RS256 App JWT,
iss=app_id, iat backdated 60s, exp 9 min; POST /access_tokens) + a lazy
TokenProvider that caches and re-mints near expiry. Secret-safe: the JWT
and token are never logged, never in an exception message, never persisted.
- dispatcher.py: app_branch_pusher / app_workflow_dispatcher / app_run_locator
(additive; gh/git _default_* left untouched). Push auth rides a host-scoped
http.extraHeader via GIT_CONFIG_* env (token never in argv/ps); the REST
run locator maps id->databaseId / created_at->createdAt into select_run_id
and surfaces 4xx promptly instead of silently exhausting the poll window.
- coordinator.py: default_dispatch_node_factory binds the App seams when
AGENT_TEAM_GH_APP_ID / _INSTALLATION_ID / _PRIVATE_KEY are all set; partial
or unreadable config logs one warning and falls back to gh-default (never
raises at serve-start).
- requirements.txt: pin PyJWT, cryptography, requests (App seams + CI fetcher).
- DEPLOY-R720.md / README.md: App dispatch config, permission/scope audit,
env-precedence check, key rotation/revocation + incident response.
Tests: +18 (test_github_app.py new; dispatcher/coordinator additions) covering
JWT claims, cache/re-mint, token-scrub-on-error, REST field mapping + run-name
correlation, and the partial-env inert fallback. Full suite 1523 passing.
2026-06-24 17:07:01 -04:00
|
|
|
|
|
|
|
|
|
|
```
|
2026-06-24 17:34:08 -04:00
|
|
|
|
# From the Mac (the key lives in ~/Downloads after generation):
|
2026-06-24 17:37:07 -04:00
|
|
|
|
scp ~/Downloads/agent-team-apply.*.private-key.pem secrev:.ssh/agent-team-apply.pem
|
|
|
|
|
|
# On the box (~/.ssh is already mode 700):
|
|
|
|
|
|
chmod 600 ~/.ssh/agent-team-apply.pem
|
|
|
|
|
|
ls -l ~/.ssh/agent-team-apply.pem # expect -rw-------
|
2026-06-24 17:34:08 -04:00
|
|
|
|
# Then DELETE the Mac copy (the box now holds the only working copy):
|
|
|
|
|
|
# rm ~/Downloads/agent-team-apply.*.private-key.pem
|
fix(agent-team): dispatch via GitHub App so P3 reaches CI (run_id resolves)
The P3 dispatcher's default seams shell out to gh/git, but the R720 box has
no gh and a read-only PAT with no Actions scope — so dispatch_apply_verify
returned no run_id and every task parked at verify ("dispatch unresolved").
Add a GitHub-App auth path: the box mints short-lived (~1h) installation
access tokens from the App private key and uses them for the three dispatch
seams, removing the gh dependency.
- agent_team/github_app.py (new): mint_installation_token (RS256 App JWT,
iss=app_id, iat backdated 60s, exp 9 min; POST /access_tokens) + a lazy
TokenProvider that caches and re-mints near expiry. Secret-safe: the JWT
and token are never logged, never in an exception message, never persisted.
- dispatcher.py: app_branch_pusher / app_workflow_dispatcher / app_run_locator
(additive; gh/git _default_* left untouched). Push auth rides a host-scoped
http.extraHeader via GIT_CONFIG_* env (token never in argv/ps); the REST
run locator maps id->databaseId / created_at->createdAt into select_run_id
and surfaces 4xx promptly instead of silently exhausting the poll window.
- coordinator.py: default_dispatch_node_factory binds the App seams when
AGENT_TEAM_GH_APP_ID / _INSTALLATION_ID / _PRIVATE_KEY are all set; partial
or unreadable config logs one warning and falls back to gh-default (never
raises at serve-start).
- requirements.txt: pin PyJWT, cryptography, requests (App seams + CI fetcher).
- DEPLOY-R720.md / README.md: App dispatch config, permission/scope audit,
env-precedence check, key rotation/revocation + incident response.
Tests: +18 (test_github_app.py new; dispatcher/coordinator additions) covering
JWT claims, cache/re-mint, token-scrub-on-error, REST field mapping + run-name
correlation, and the partial-env inert fallback. Full suite 1523 passing.
2026-06-24 17:07:01 -04:00
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
- `AGENT_TEAM_GH_APP_PRIVATE_KEY` is a **filesystem path** to the `.pem`, not the
|
|
|
|
|
|
key material. The coordinator reads the file at graph-build time; an unreadable
|
|
|
|
|
|
path logs one warning and falls back to gh-default (never crashes).
|
2026-06-24 17:34:08 -04:00
|
|
|
|
- **App permission/scope (verified 2026-06-24).** A test mint of an installation
|
|
|
|
|
|
token for app `4119505` / install `141992144` returned
|
|
|
|
|
|
`{"actions":"write","contents":"write","metadata":"read","pull_requests":"write"}`
|
|
|
|
|
|
with `repository_selection: selected` — i.e. Contents R/W + Actions R/W are in
|
|
|
|
|
|
place and the install is scoped to selected repos (not all). Minting grants
|
|
|
|
|
|
exactly the App's scopes; keep it scoped to `Sea-Haven-Industries/orchestrator`
|
|
|
|
|
|
only. (Follow-up: move to a dedicated least-privilege App to replace
|
|
|
|
|
|
`agent-team-apply`, dropping `pull_requests:write` which the box path does not
|
|
|
|
|
|
need.)
|
fix(agent-team): dispatch via GitHub App so P3 reaches CI (run_id resolves)
The P3 dispatcher's default seams shell out to gh/git, but the R720 box has
no gh and a read-only PAT with no Actions scope — so dispatch_apply_verify
returned no run_id and every task parked at verify ("dispatch unresolved").
Add a GitHub-App auth path: the box mints short-lived (~1h) installation
access tokens from the App private key and uses them for the three dispatch
seams, removing the gh dependency.
- agent_team/github_app.py (new): mint_installation_token (RS256 App JWT,
iss=app_id, iat backdated 60s, exp 9 min; POST /access_tokens) + a lazy
TokenProvider that caches and re-mints near expiry. Secret-safe: the JWT
and token are never logged, never in an exception message, never persisted.
- dispatcher.py: app_branch_pusher / app_workflow_dispatcher / app_run_locator
(additive; gh/git _default_* left untouched). Push auth rides a host-scoped
http.extraHeader via GIT_CONFIG_* env (token never in argv/ps); the REST
run locator maps id->databaseId / created_at->createdAt into select_run_id
and surfaces 4xx promptly instead of silently exhausting the poll window.
- coordinator.py: default_dispatch_node_factory binds the App seams when
AGENT_TEAM_GH_APP_ID / _INSTALLATION_ID / _PRIVATE_KEY are all set; partial
or unreadable config logs one warning and falls back to gh-default (never
raises at serve-start).
- requirements.txt: pin PyJWT, cryptography, requests (App seams + CI fetcher).
- DEPLOY-R720.md / README.md: App dispatch config, permission/scope audit,
env-precedence check, key rotation/revocation + incident response.
Tests: +18 (test_github_app.py new; dispatcher/coordinator additions) covering
JWT claims, cache/re-mint, token-scrub-on-error, REST field mapping + run-name
correlation, and the partial-env inert fallback. Full suite 1523 passing.
2026-06-24 17:07:01 -04:00
|
|
|
|
- The minted token is short-lived (~1h), is **never** logged, never put in an
|
|
|
|
|
|
exception message, and never written to the ledger/graph state; on the
|
|
|
|
|
|
authenticated push it rides in a host-scoped `http.extraHeader` passed via
|
|
|
|
|
|
`GIT_CONFIG_*` env (never in argv/`ps`).
|
|
|
|
|
|
- **Env-file precedence (verify after editing).** The unit loads **both**
|
|
|
|
|
|
`~/secrev.env` and `~/orchestrator/.env` (last-wins). Confirm only the intended
|
|
|
|
|
|
values are present in both so a stale entry can't shadow the App config:
|
|
|
|
|
|
```
|
|
|
|
|
|
grep -nE 'AGENT_TEAM_GH_APP|AGENT_TEAM_REPO_(OWNER|NAME)' ~/secrev.env ~/orchestrator/.env 2>/dev/null
|
|
|
|
|
|
```
|
|
|
|
|
|
|
2026-06-18 12:56:42 -04:00
|
|
|
|
## 4. Deploy steps
|
|
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
|
# 4a. From the Mac - rsync the repo (same pattern/excludes as secrev):
|
|
|
|
|
|
rsync -av --exclude .env --exclude .venv \
|
|
|
|
|
|
~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/
|
|
|
|
|
|
|
|
|
|
|
|
# 4b. On the VM - create + activate the agent-team venv and install deps:
|
|
|
|
|
|
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
|
|
|
|
|
|
cd ~/orchestrator/agent-team
|
|
|
|
|
|
python3 -m venv .venv
|
|
|
|
|
|
. .venv/bin/activate
|
|
|
|
|
|
pip install langgraph langgraph-checkpoint-sqlite claude-agent-sdk slack_sdk slack_bolt
|
|
|
|
|
|
|
|
|
|
|
|
# 4c. Initialize the durable ledger DB (idempotent; creates state/agent_team.sqlite):
|
|
|
|
|
|
python3 run-team.py init-db
|
|
|
|
|
|
|
|
|
|
|
|
# 4d. Install + start the service:
|
|
|
|
|
|
sudo cp systemd/agent-team-coordinator.service /etc/systemd/system/
|
|
|
|
|
|
sudo systemctl daemon-reload
|
|
|
|
|
|
sudo systemctl enable --now agent-team-coordinator.service
|
|
|
|
|
|
|
|
|
|
|
|
# 4e. Verify it is up:
|
|
|
|
|
|
systemctl status agent-team-coordinator.service
|
|
|
|
|
|
journalctl -u agent-team-coordinator.service -e -f
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
The unit runs `python3 run-team.py serve` from
|
|
|
|
|
|
`WorkingDirectory=/home/adam/orchestrator/agent-team` as `User=adam`, loading
|
|
|
|
|
|
secrets from `EnvironmentFile=/home/adam/secrev.env`. `Restart=on-failure` keeps
|
|
|
|
|
|
it up across transient faults; `journalctl -u` is the live log.
|
|
|
|
|
|
|
2026-06-24 11:30:54 -04:00
|
|
|
|
> **⚠️ This deploy carries a ledger migration (SCHEMA_VERSION → 4).** It adds a
|
|
|
|
|
|
> `kind` column to `pending_questions` (values `clarify` | `plan_decision`,
|
|
|
|
|
|
> existing rows default to `clarify`) for the plan-review decision gate. The
|
|
|
|
|
|
> migration is an **additive, idempotent in-place `ALTER TABLE`** run on startup
|
|
|
|
|
|
> (`init_db` / `migrate`, guarded so a second run is a no-op — it never recreates
|
|
|
|
|
|
> the table), so it applies in place against the live box ledger. **Take the
|
|
|
|
|
|
> ledger backup (§1 / the deploy script step) BEFORE restart** — it is the
|
|
|
|
|
|
> migration's safety net (see Rollback, §6). After restart, confirm the column
|
|
|
|
|
|
> landed and the daemon came up clean:
|
|
|
|
|
|
> ```bash
|
|
|
|
|
|
> sqlite3 ~/orchestrator/agent-team/state/agent_team.sqlite \
|
|
|
|
|
|
> "PRAGMA table_info(pending_questions);" | grep kind # expect a 'kind' row
|
|
|
|
|
|
> sqlite3 ~/orchestrator/agent-team/state/agent_team.sqlite \
|
|
|
|
|
|
> "SELECT schema_version FROM schema_meta WHERE id=1;" # expect 4
|
|
|
|
|
|
> journalctl -u agent-team-coordinator.service -e | tail # no migration/import errors
|
|
|
|
|
|
> ```
|
|
|
|
|
|
|
2026-06-23 12:56:08 -04:00
|
|
|
|
## 4b. WS0–WS5 rollout — UPDATE an already-deployed box
|
|
|
|
|
|
|
|
|
|
|
|
The steps above (§1–4) are the **first-time** P1 provision. To bring an
|
|
|
|
|
|
already-deployed coordinator up to the WS0–WS5 rollout, use the attended update
|
|
|
|
|
|
script rather than re-running the manual steps:
|
|
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
|
# From the Mac, after the WS branches have merged to main, snapshot first:
|
|
|
|
|
|
agent-team/scripts/deploy-r720-ws-rollout.sh
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
It is an UPDATE (not a provision): it snapshots-reminds, rsyncs the new code,
|
|
|
|
|
|
rsyncs the engineering handbook to the box, installs the new deps, appends the
|
|
|
|
|
|
new secrets if absent, restarts the coordinator, and smoke-tests. It is
|
|
|
|
|
|
idempotent and fails loudly. What goes **live** after it: WS1 in-process
|
|
|
|
|
|
multi-model invokers (`bind_multi_invoker`, already wired), the WS5 handbook
|
|
|
|
|
|
`context_provider` injected into the planner prompt, and the WS2 Slack
|
|
|
|
|
|
`/new-task` command (AUTHZ-01 owner-allowlist gated). The P3 dispatch/build-verify
|
|
|
|
|
|
path stays **inert** (gated behind the `agent-apply` GitHub Environment approval).
|
|
|
|
|
|
|
|
|
|
|
|
**New venv deps** (the coordinator does not need them; only the optional HTTP
|
|
|
|
|
|
API does) — now pinned in the root `requirements.txt`:
|
|
|
|
|
|
|
|
|
|
|
|
| pip dep | Why |
|
|
|
|
|
|
|---|---|
|
|
|
|
|
|
| `fastapi==0.136.1` | the WS1 HTTP API app (`agent_team/api.py`) |
|
|
|
|
|
|
| `uvicorn==0.46.0` | ASGI server for `api.serve()` |
|
|
|
|
|
|
|
|
|
|
|
|
**New env vars** — append to `~/secrev.env` (mode 600, never committed):
|
|
|
|
|
|
|
|
|
|
|
|
- `SEA_HAVEN_HANDBOOK_DIR` — where `load_handbook_conventions()` reads the
|
|
|
|
|
|
engineering handbook (the script syncs it to `/home/adam/.sea-haven/engineering-handbook`
|
|
|
|
|
|
by default; this var must match). Fail-safe: if the dir is missing the
|
|
|
|
|
|
`context_provider` returns `""` and the planner runs without handbook context.
|
|
|
|
|
|
- `AGENT_TEAM_API_TOKEN` — bearer token for the HTTP API / `/delegate` hook
|
|
|
|
|
|
**only**. Not needed by the coordinator daemon itself. The HTTP API refuses to
|
|
|
|
|
|
start if this is unset/empty.
|
|
|
|
|
|
|
|
|
|
|
|
**The HTTP API is a separate, opt-in process** — it is **not** started by the
|
|
|
|
|
|
coordinator daemon. Run it explicitly (`api.serve()`, binds `127.0.0.1:8765`,
|
|
|
|
|
|
bearer auth) only if you want the `/delegate` Claude Code hook or the
|
|
|
|
|
|
`POST /tasks` / `GET /tasks/{thread_id}` / `POST /orchestrator/invoke` endpoints.
|
|
|
|
|
|
The `/docs` + `/openapi` routes are disabled and it binds loopback by design (do
|
|
|
|
|
|
not change to `0.0.0.0`). See the deploy script's step 6 for how to start it.
|
|
|
|
|
|
|
2026-06-23 14:16:07 -04:00
|
|
|
|
### Status dashboard (optional, LAN/VPN-only, READ-ONLY)
|
|
|
|
|
|
|
2026-06-23 14:39:30 -04:00
|
|
|
|
`agent_team/status_page.py` serves a **live visual pipeline map** of the
|
|
|
|
|
|
coordinator. The top of the page is a hand-rolled inline-SVG diagram of the
|
|
|
|
|
|
agent-team DAG (`INTAKE → CLARIFY ⇄ human gate → PLAN ⇄ REVIEW →
|
|
|
|
|
|
[BUILD → VERIFY → DISPATCH] → DONE`); each stage node is labelled with its model
|
|
|
|
|
|
role (Claude on subscription for CLARIFY/PLAN/VERIFY, GPT-4.1 cross_reviewer for
|
|
|
|
|
|
REVIEW, Gemini for SCAN, DeepSeek fast_coder for BUILD, the Slack owner for the
|
|
|
|
|
|
HUMAN GATE) and is colour-coded by live state — idle / active / awaiting-human
|
|
|
|
|
|
(an **open** pending question) / parked — with a count badge of tasks in that
|
|
|
|
|
|
stage. Hovering (or keyboard-focusing) a node shows what it is working on: the
|
|
|
|
|
|
short `thread_id`, description, status, and waiting age of each task there.
|
|
|
|
|
|
Below the map, the original detail tables remain: all tasks, the human-gate
|
|
|
|
|
|
wait-list, and recent `budget_ledger` spend.
|
|
|
|
|
|
|
|
|
|
|
|
The page **auto-updates without a full reload**: a `GET /api/state` JSON sidecar
|
|
|
|
|
|
returns the same snapshot, and an inline vanilla-JS poller (`fetch()`, no
|
|
|
|
|
|
libraries, no CDN) re-paints node states, counts, the cards, the tooltip data,
|
|
|
|
|
|
and the "last updated" clock every ~4s in place, so hover/scroll/focus survive.
|
|
|
|
|
|
A `<noscript>` 10s meta-refresh is the JS-disabled fallback. Everything
|
|
|
|
|
|
(SVG + CSS + JS) is inline in the served document — nothing is fetched from a
|
|
|
|
|
|
CDN, because the VM is offline/LAN-only. It opens the SQLite ledger
|
|
|
|
|
|
**READ-ONLY** (`mode=ro`), exposes only the two read GETs (`/` and `/api/state`)
|
|
|
|
|
|
and **no mutating endpoints and no auth**.
|
2026-06-23 14:16:07 -04:00
|
|
|
|
|
|
|
|
|
|
It is a **separate, optional process** — `agent-team-status.service` (mirrors the
|
|
|
|
|
|
coordinator unit's hardening; `User=adam`, `EnvironmentFile=-/home/adam/secrev.env`,
|
|
|
|
|
|
venv-python ExecStart, `Restart=on-failure`). Unlike the coordinator it needs **no**
|
|
|
|
|
|
`ReadWritePaths` carve-out (it only reads). It can run side-by-side with the
|
|
|
|
|
|
coordinator (RO SQLite opens coexist with the writer).
|
|
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
|
# install the unit
|
|
|
|
|
|
sudo cp ~/orchestrator/agent-team/systemd/agent-team-status.service /etc/systemd/system/
|
|
|
|
|
|
sudo systemctl daemon-reload
|
|
|
|
|
|
sudo systemctl enable --now agent-team-status.service
|
|
|
|
|
|
systemctl status agent-team-status.service
|
|
|
|
|
|
journalctl -u agent-team-status.service -e -f
|
|
|
|
|
|
|
|
|
|
|
|
# or run it ad hoc from the venv
|
|
|
|
|
|
cd ~/orchestrator/agent-team && . .venv/bin/activate && \
|
|
|
|
|
|
python3 -c "from agent_team.status_page import serve; serve()"
|
|
|
|
|
|
```
|
|
|
|
|
|
|
2026-06-23 14:39:30 -04:00
|
|
|
|
Then browse `http://10.10.60.120:8770/` from the LAN/VPN (the live map polls
|
|
|
|
|
|
`http://10.10.60.120:8770/api/state` itself).
|
2026-06-23 14:16:07 -04:00
|
|
|
|
|
|
|
|
|
|
**Config (env):** `AGENT_TEAM_DB` (default `state/agent_team.sqlite`),
|
|
|
|
|
|
`AGENT_TEAM_STATUS_HOST` (default `0.0.0.0`), `AGENT_TEAM_STATUS_PORT` (default
|
|
|
|
|
|
`8770`).
|
|
|
|
|
|
|
|
|
|
|
|
**Posture:** the sh-secrev VM (`10.10.60.120`, VLAN 60) has no public NIC and sits
|
|
|
|
|
|
behind the UniFi firewall, so `0.0.0.0` reaches the **LAN/VPN only**. Task
|
|
|
|
|
|
descriptions may be sensitive and the page is unauthenticated — **keep it
|
|
|
|
|
|
LAN/VPN-only, never expose it to the public internet.** A missing/locked DB renders
|
|
|
|
|
|
a friendly "no data" page rather than crashing.
|
|
|
|
|
|
|
2026-06-18 12:56:42 -04:00
|
|
|
|
## 5. P1 live exit-criteria demo (§3.3.1)
|
|
|
|
|
|
|
|
|
|
|
|
Demonstrate all four once the service is live. Map each to the operator commands
|
|
|
|
|
|
(`run-team.py list / show / force-resume`, `systemctl`). Run the CLI from the
|
|
|
|
|
|
working dir so it hits the default ledger: `cd ~/orchestrator/agent-team`.
|
|
|
|
|
|
|
|
|
|
|
|
**(a) Crash-safe resume - kill mid-wait, restart, task resumes.**
|
|
|
|
|
|
Start a task, get it to a clarifier wait (`run-team.py list` shows an `open`
|
|
|
|
|
|
question), then:
|
|
|
|
|
|
```
|
|
|
|
|
|
sudo systemctl stop agent-team-coordinator.service
|
|
|
|
|
|
sudo systemctl start agent-team-coordinator.service
|
|
|
|
|
|
journalctl -u agent-team-coordinator.service -e # confirm the task resumes from the ledger/checkpoint
|
|
|
|
|
|
run-team.py show <question_id> # the question is still open, not lost
|
|
|
|
|
|
```
|
|
|
|
|
|
Pass: the task picks up the same waiting question after restart (the LangGraph
|
|
|
|
|
|
`SqliteSaver` checkpoint + the durable ledger survive the kill).
|
|
|
|
|
|
|
|
|
|
|
|
**(b) Duplicate Slack answer is a no-op.**
|
|
|
|
|
|
Answer a question in Slack, then answer the **same** question again.
|
|
|
|
|
|
```
|
|
|
|
|
|
run-team.py show <question_id> # status flipped to answered exactly once; answered_via is the first answer
|
|
|
|
|
|
```
|
|
|
|
|
|
Pass: the first answer wins (`rowcount == 1`); the duplicate hits the
|
|
|
|
|
|
`BEGIN IMMEDIATE` compare-and-set and is ignored (`rowcount == 0`) - no second
|
|
|
|
|
|
resume, no error.
|
|
|
|
|
|
|
|
|
|
|
|
**(c) Past-deadline answer is rejected + the task parks.**
|
|
|
|
|
|
Let a question's `deadline_at` pass with no answer, then answer late.
|
|
|
|
|
|
```
|
|
|
|
|
|
run-team.py show <question_id> # status == expired (auto-expired at deadline)
|
|
|
|
|
|
run-team.py list --parked # the now-parked task surfaces here
|
|
|
|
|
|
```
|
|
|
|
|
|
Pass: the expired question rejects the late answer and the task parks rather than
|
|
|
|
|
|
spins. To un-park it deliberately:
|
|
|
|
|
|
```
|
|
|
|
|
|
run-team.py force-resume <question_id> --confirm # reopens the expired question for re-delivery
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
**(d) Two concurrent tasks resume independently to the correct thread.**
|
|
|
|
|
|
Start two tasks concurrently, each reaching its own clarifier wait.
|
|
|
|
|
|
```
|
|
|
|
|
|
run-team.py list # two distinct open questions, distinct thread_id values
|
|
|
|
|
|
```
|
|
|
|
|
|
Restart the service (as in (a)); answer each in Slack.
|
|
|
|
|
|
Pass: each task resumes to its own `thread_id` / channel - no cross-talk, no
|
|
|
|
|
|
answer routed to the wrong task.
|
|
|
|
|
|
|
|
|
|
|
|
## 6. Rollback
|
|
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
|
# Stop + disable the service and remove the unit:
|
|
|
|
|
|
sudo systemctl disable --now agent-team-coordinator.service
|
|
|
|
|
|
sudo rm /etc/systemd/system/agent-team-coordinator.service
|
|
|
|
|
|
sudo systemctl daemon-reload
|
|
|
|
|
|
|
|
|
|
|
|
# Restore the VM from the pre-provision Hyper-V checkpoint (§1) to undo
|
|
|
|
|
|
# pip deps + any host changes in one step.
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
The ledger is **local state** under `agent-team/state/` (not in git). To reset
|
|
|
|
|
|
it without a full snapshot restore: back it up first, then wipe.
|
|
|
|
|
|
```
|
|
|
|
|
|
cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak} # back up
|
|
|
|
|
|
rm ~/orchestrator/agent-team/state/agent_team.sqlite* # wipe (then re-run init-db)
|
|
|
|
|
|
```
|
2026-06-24 11:30:54 -04:00
|
|
|
|
|
|
|
|
|
|
**Ledger-migration rollback (the schema-v4 `kind` migration).** The deploy takes
|
|
|
|
|
|
a dated ledger backup (`~/agent_team.sqlite.bak-<date>`, written by the
|
|
|
|
|
|
`/sh-deploy-r720` flow / `scripts/deploy-r720.sh`) **before** restart — that
|
|
|
|
|
|
backup is the migration's safety net. The `kind` migration is additive and
|
|
|
|
|
|
idempotent, but **if the `init_db`/`migrate` step fails, or the deploy is rolled
|
|
|
|
|
|
back to pre-v4 code after the migration ran, restore the ledger from that backup
|
|
|
|
|
|
BEFORE restarting the coordinator** (old code does not expect the new column to
|
|
|
|
|
|
matter, but restoring guarantees a clean, pre-migration ledger):
|
|
|
|
|
|
```
|
|
|
|
|
|
sudo systemctl stop agent-team-coordinator.service
|
|
|
|
|
|
cp ~/agent_team.sqlite.bak-<date> ~/orchestrator/agent-team/state/agent_team.sqlite
|
|
|
|
|
|
rm -f ~/orchestrator/agent-team/state/agent_team.sqlite-wal \
|
|
|
|
|
|
~/orchestrator/agent-team/state/agent_team.sqlite-shm # drop stale WAL/SHM
|
|
|
|
|
|
sudo systemctl start agent-team-coordinator.service
|
|
|
|
|
|
```
|
|
|
|
|
|
Restore the ledger backup BEFORE the coordinator restarts — never start the
|
|
|
|
|
|
daemon against a half-migrated or suspect ledger.
|
2026-06-18 12:56:42 -04:00
|
|
|
|
Note the secrets in `~/secrev.env` are NOT removed by rollback - leave them, or
|
|
|
|
|
|
strip the four agent-team keys if you are decommissioning entirely.
|
|
|
|
|
|
|
fix(agent-team): dispatch via GitHub App so P3 reaches CI (run_id resolves)
The P3 dispatcher's default seams shell out to gh/git, but the R720 box has
no gh and a read-only PAT with no Actions scope — so dispatch_apply_verify
returned no run_id and every task parked at verify ("dispatch unresolved").
Add a GitHub-App auth path: the box mints short-lived (~1h) installation
access tokens from the App private key and uses them for the three dispatch
seams, removing the gh dependency.
- agent_team/github_app.py (new): mint_installation_token (RS256 App JWT,
iss=app_id, iat backdated 60s, exp 9 min; POST /access_tokens) + a lazy
TokenProvider that caches and re-mints near expiry. Secret-safe: the JWT
and token are never logged, never in an exception message, never persisted.
- dispatcher.py: app_branch_pusher / app_workflow_dispatcher / app_run_locator
(additive; gh/git _default_* left untouched). Push auth rides a host-scoped
http.extraHeader via GIT_CONFIG_* env (token never in argv/ps); the REST
run locator maps id->databaseId / created_at->createdAt into select_run_id
and surfaces 4xx promptly instead of silently exhausting the poll window.
- coordinator.py: default_dispatch_node_factory binds the App seams when
AGENT_TEAM_GH_APP_ID / _INSTALLATION_ID / _PRIVATE_KEY are all set; partial
or unreadable config logs one warning and falls back to gh-default (never
raises at serve-start).
- requirements.txt: pin PyJWT, cryptography, requests (App seams + CI fetcher).
- DEPLOY-R720.md / README.md: App dispatch config, permission/scope audit,
env-precedence check, key rotation/revocation + incident response.
Tests: +18 (test_github_app.py new; dispatcher/coordinator additions) covering
JWT claims, cache/re-mint, token-scrub-on-error, REST field mapping + run-name
correlation, and the partial-env inert fallback. Full suite 1523 passing.
2026-06-24 17:07:01 -04:00
|
|
|
|
### GitHub App key — compromise / rotation / revocation
|
|
|
|
|
|
|
|
|
|
|
|
The box holds a write-capable App private key, so it needs its own incident path
|
|
|
|
|
|
(separate from the VM snapshot rollback above):
|
|
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
|
# 1. Stop the daemon so no further token mints happen:
|
|
|
|
|
|
sudo systemctl stop agent-team-coordinator.service
|
|
|
|
|
|
# 2. Remove the key + the App env vars from the box (kills the App path -> inert):
|
2026-06-24 17:37:07 -04:00
|
|
|
|
rm -f ~/.ssh/agent-team-apply.pem
|
fix(agent-team): dispatch via GitHub App so P3 reaches CI (run_id resolves)
The P3 dispatcher's default seams shell out to gh/git, but the R720 box has
no gh and a read-only PAT with no Actions scope — so dispatch_apply_verify
returned no run_id and every task parked at verify ("dispatch unresolved").
Add a GitHub-App auth path: the box mints short-lived (~1h) installation
access tokens from the App private key and uses them for the three dispatch
seams, removing the gh dependency.
- agent_team/github_app.py (new): mint_installation_token (RS256 App JWT,
iss=app_id, iat backdated 60s, exp 9 min; POST /access_tokens) + a lazy
TokenProvider that caches and re-mints near expiry. Secret-safe: the JWT
and token are never logged, never in an exception message, never persisted.
- dispatcher.py: app_branch_pusher / app_workflow_dispatcher / app_run_locator
(additive; gh/git _default_* left untouched). Push auth rides a host-scoped
http.extraHeader via GIT_CONFIG_* env (token never in argv/ps); the REST
run locator maps id->databaseId / created_at->createdAt into select_run_id
and surfaces 4xx promptly instead of silently exhausting the poll window.
- coordinator.py: default_dispatch_node_factory binds the App seams when
AGENT_TEAM_GH_APP_ID / _INSTALLATION_ID / _PRIVATE_KEY are all set; partial
or unreadable config logs one warning and falls back to gh-default (never
raises at serve-start).
- requirements.txt: pin PyJWT, cryptography, requests (App seams + CI fetcher).
- DEPLOY-R720.md / README.md: App dispatch config, permission/scope audit,
env-precedence check, key rotation/revocation + incident response.
Tests: +18 (test_github_app.py new; dispatcher/coordinator additions) covering
JWT claims, cache/re-mint, token-scrub-on-error, REST field mapping + run-name
correlation, and the partial-env inert fallback. Full suite 1523 passing.
2026-06-24 17:07:01 -04:00
|
|
|
|
sed -i '/AGENT_TEAM_GH_APP_/d' ~/secrev.env
|
|
|
|
|
|
# 3. In GitHub: rotate (generate a new private key, delete the old one) under the
|
|
|
|
|
|
# App's settings, or uninstall the App from the repo to revoke all access.
|
|
|
|
|
|
# Existing installation tokens are short-lived (~1h) and expire on their own.
|
|
|
|
|
|
# 4. Restart inert (gh-default dispatch) and confirm no App env is loaded:
|
|
|
|
|
|
sudo systemctl start agent-team-coordinator.service
|
|
|
|
|
|
sudo systemctl show agent-team-coordinator.service -p Environment | grep -c AGENT_TEAM_GH_APP # expect 0
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
- **Rotation cadence is Adam's call** — there is no automated rotation. Rotate on
|
|
|
|
|
|
any suspected box compromise, on operator turnover, and on a periodic cadence
|
|
|
|
|
|
Adam sets. Document the rotation event in Confluence.
|
|
|
|
|
|
- Removing the App env vars (or the key file) is a safe partial rollback on its
|
|
|
|
|
|
own: the dispatcher falls back to the inert gh-default path (which is itself
|
|
|
|
|
|
inert without `gh`), so no diff is ever pushed — never a fabricated pass.
|
|
|
|
|
|
|
2026-06-18 12:56:42 -04:00
|
|
|
|
## 7. Security
|
|
|
|
|
|
|
|
|
|
|
|
- **Slack inbound listener (Socket Mode) is the auth + untrusted-input surface.**
|
|
|
|
|
|
It accepts inbound messages over a WebSocket and turns them into ledger
|
|
|
|
|
|
mutations (answering live clarifier questions). It **must pass
|
|
|
|
|
|
`/sh-security-review`** before this is enabled in production - that review is
|
|
|
|
|
|
mandatory for authentication/authorization and untrusted-input handling
|
|
|
|
|
|
changes, and this is both.
|
|
|
|
|
|
- **No IAM / OIDC is involved in P1.** The box runs on subscription OAuth
|
|
|
|
|
|
(`CLAUDE_CODE_OAUTH_TOKEN`) and Slack tokens only; there is no AWS role, no
|
|
|
|
|
|
OIDC trust relationship, no cloud permission surface in this deploy.
|
fix(agent-team): dispatch via GitHub App so P3 reaches CI (run_id resolves)
The P3 dispatcher's default seams shell out to gh/git, but the R720 box has
no gh and a read-only PAT with no Actions scope — so dispatch_apply_verify
returned no run_id and every task parked at verify ("dispatch unresolved").
Add a GitHub-App auth path: the box mints short-lived (~1h) installation
access tokens from the App private key and uses them for the three dispatch
seams, removing the gh dependency.
- agent_team/github_app.py (new): mint_installation_token (RS256 App JWT,
iss=app_id, iat backdated 60s, exp 9 min; POST /access_tokens) + a lazy
TokenProvider that caches and re-mints near expiry. Secret-safe: the JWT
and token are never logged, never in an exception message, never persisted.
- dispatcher.py: app_branch_pusher / app_workflow_dispatcher / app_run_locator
(additive; gh/git _default_* left untouched). Push auth rides a host-scoped
http.extraHeader via GIT_CONFIG_* env (token never in argv/ps); the REST
run locator maps id->databaseId / created_at->createdAt into select_run_id
and surfaces 4xx promptly instead of silently exhausting the poll window.
- coordinator.py: default_dispatch_node_factory binds the App seams when
AGENT_TEAM_GH_APP_ID / _INSTALLATION_ID / _PRIVATE_KEY are all set; partial
or unreadable config logs one warning and falls back to gh-default (never
raises at serve-start).
- requirements.txt: pin PyJWT, cryptography, requests (App seams + CI fetcher).
- DEPLOY-R720.md / README.md: App dispatch config, permission/scope audit,
env-precedence check, key rotation/revocation + incident response.
Tests: +18 (test_github_app.py new; dispatcher/coordinator additions) covering
JWT claims, cache/re-mint, token-scrub-on-error, REST field mapping + run-name
correlation, and the partial-env inert fallback. Full suite 1523 passing.
2026-06-24 17:07:01 -04:00
|
|
|
|
- **GitHub App key on the box (P3 dispatch) — conscious, mitigated trade-off.**
|
|
|
|
|
|
The dispatcher's original design comment said "never a box-held token / operator
|
|
|
|
|
|
host only." Moving dispatch onto the box (so P3 reaches CI without `gh`)
|
|
|
|
|
|
deliberately deviates from that. Mitigations: the App is scoped to **one repo**
|
|
|
|
|
|
with **Contents + Actions only**; minted installation tokens are **short-lived
|
|
|
|
|
|
(~1h)**; the key file is **mode 600**; tokens are **never logged / never in
|
|
|
|
|
|
exception text / never in argv** (push auth rides a host-scoped
|
|
|
|
|
|
`http.extraHeader` via `GIT_CONFIG_*`); CI **re-verifies** the pushed content by
|
|
|
|
|
|
hash and the draft PR is still gated by the `agent-apply` environment's required
|
|
|
|
|
|
reviewer. Any mint/HTTP/git failure **parks** the task (no `run_id` → the verify
|
|
|
|
|
|
gate fails closed) — never a fabricated pass. See §3a for the
|
|
|
|
|
|
permission/scope audit, env-precedence check, and §6 for key rotation/
|
|
|
|
|
|
revocation. **Credential-handling change → `/sh-security-review` is mandatory
|
|
|
|
|
|
before merge** (auth + secret handling surface).
|
2026-06-18 12:56:42 -04:00
|
|
|
|
- `~/secrev.env` stays mode 600 and out of git; `state/` (ledger + audit log) is
|
|
|
|
|
|
gitignored and written 0600.
|