Addresses the confirmed findings from /sh-security-review + the GPT-4.1
cross-review of the Plane-2 scaffold. Full suite: 589 passed; ruff clean.
FIXED (proven-exploitable):
- CI-guard denylist bypass (HIGH): Python fnmatch '**/' is non-recursive, so
root-level template.yaml/*.tf/cdk.json/*.pem/*.key/*-stack.* evaded the
trust-control surface. Replaced fnmatch with a recursive, case-insensitive
glob->regex matcher. (verified: fnmatch('template.yaml','**/template.yaml')==False)
- CI-guard scope bypass (HIGH): a '**' declared_scope made every path in-scope.
Scope is now concrete-prefix confinement (reduces a glob to its leading
metacharacter-free segments; '**' -> empty -> dropped -> unscoped reject).
- Box-side vs CI denylist divergence (MED): builders.py _DENY_PATTERNS now covers
Terraform, *.pem/*.key, CDK stack files, .github/actions, *iam*, bare policy*.json
(case-insensitive), matching the CI surface.
- force-resume was backwards (MED): it superseded the answered row recovery
resumes from, making a stuck task permanently un-resumable while printing
success. Now re-opens an EXPIRED (parked) question via a new reopen_question
CAS helper; never supersedes an answered row; honest exit codes.
- operator attribution (MED): run-team.py --operator defaulted to "" -> now the
OS login, so destructive actions are always attributable.
- audit-log append race (MED): replaced read-modify-rewrite (lost records under
concurrent operators) with an O_APPEND single-line write, mode 600 enforced.
- lstrip("ab/") path-mangling in the symlink error path -> regex prefix strip.
Regression tests added across test_ci_gate_workflow / test_builders / test_run_team
/ test_schema. Design-level findings (resume-worker durability, egress breadth,
answered_at ordering, DB-swap TOCTOU, diff-hash threat-model) are pre-deployment
/ P1-build-proper and recorded with written justification in
agent-team/.security-review/suppressions.json; CI README diff-hash wording made
honest.
|
||
|---|---|---|
| .. | ||
| .security-review | ||
| agent_team | ||
| ci | ||
| tests | ||
| .gitignore | ||
| README.md | ||
| run-team.py | ||
agent-team — R720 Plane-2 FOUNDATION
Pre-deployment scaffolding for the R720 agent-team SDLC pipeline (design:
../docs/r720-agent-team-design.md). This commit ships the Plane-2
FOUNDATION layer only — the durable, transport-agnostic contracts the leaf
builders import verbatim. Nothing here is provisioned, scheduled, or wired to
live infrastructure.
Status: FOUNDATION modules only. No coordinator, no transports' concrete adapters, no CI workflow, no provisioning. Those are later phases (§7).
Layout
agent-team/
agent_team/ # importable package (snake_case)
state_store.py # §6.7 atomic write + integrity-checked read
billing.py # §3.1 claude_invoke billing-mode seam
task_model.py # §3.3 TaskRecord / Phase / PipelineState
db/
schema.py # §3.3.1/§6.7 SQLite DDL + connect/init/migrate
schema.sql # raw DDL, mirrors schema.py verbatim
transport/
base.py # §3.3.1 Transport ABC + QuestionSet/NormalizedAnswer
tests/ # pytest unit tests, one module per source module
The top directory is kebab-case (agent-team/); the importable package is
snake_case (agent_team/), per the engineering handbook.
Modules (contracts)
| Module | Design ref | What it provides |
|---|---|---|
state_store |
§6.7 | atomic_write(path, data) (write-temp → fsync → rename), read_checked(path, *, schema_version) (schema-version + content-hash integrity check, raises IntegrityError), compute_content_hash(data). Pure stdlib; no other agent_team deps. |
db.schema |
§3.3.1, §6.7 | SCHEMA_VERSION, PENDING_QUESTIONS_DDL, BUDGET_LEDGER_DDL, connect() (WAL + foreign_keys + busy_timeout), init_db(), migrate(), and the BEGIN IMMEDIATE compare-and-set helpers (answer_question/expire_question/supersede_question). SQL DDL lives only here. |
billing |
§3.1 | BillingMode{SUBSCRIPTION,API,BEDROCK}, claude_invoke(prompt, *, mode=None, **kw) -> ClaudeResult, resolve_mode(config). Single seam; subscription mode pops any stray ANTHROPIC_API_KEY so OAuth can't be overridden. |
transport.base |
§3.3.1 | Transport ABC (post_question → channel_ref; parse_answer → (question_id, answer, via)), QuestionSet, NormalizedAnswer. Transport-independent; Slack/GitHub/Claude-Code adapters subclass in the leaves. |
task_model |
§3.3 | TaskRecord, TaskStatus, Phase{INTAKE…DONE}, new_thread_id(), PipelineState TypedDict (LangGraph state schema), JSON serialization helpers. Pure model, no I/O. |
Durable human-in-the-loop (§3.3.1)
The pending_questions ledger is the single durable source of truth for the
question lifecycle. Every race (duplicate answers, transport redelivery,
answer-vs-timeout) resolves via one atomic compare-and-set against the status
column, run inside a BEGIN IMMEDIATE transaction so concurrent responders are
serialized — first-answer-wins (rowcount == 1), late/duplicate ignored
(rowcount == 0). The LangGraph SqliteSaver checkpointer creates its own
tables against the same DB file.
Running the tests
cd agent-team
python3 -m pytest tests/ -q
tests/conftest.py puts the package on sys.path, so no install is required.
Not in this commit (later phases)
Coordinator/brain, concrete Slack/GitHub/Claude-Code transport adapters, the
CI apply/verify workflow (§3.3.2), step-ca / Roles Anywhere, scheduling, and
the operator CLI (run-team.py). See ../docs/r720-agent-team-design.md §7
for the phased rollout. Secrets are never committed.