Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.
Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled
KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven
Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 14:17:44 -04:00
|
|
|
"""Unit tests for agent_team.graph (Plane-2 P1 LangGraph wiring; §3.3, §7.1).
|
|
|
|
|
|
|
|
|
|
These exercise the P1 skeleton + human gate wiring:
|
|
|
|
|
|
|
|
|
|
* the pure node functions (intake/clarify-author/plan) in isolation,
|
|
|
|
|
* graph assembly + edge topology,
|
|
|
|
|
* the suspend-on-interrupt / resume-with-Command mechanic end to end,
|
|
|
|
|
* the thread_id-keyed driver seam (start/resume/get_state/pending_question),
|
|
|
|
|
* that the foundation contracts (PipelineState / Phase / TaskStatus /
|
|
|
|
|
QuestionSet) are imported verbatim and round-trip through the wiring.
|
|
|
|
|
|
|
|
|
|
An in-memory checkpointer is injected (the SQLite checkpointer is the
|
|
|
|
|
production store, D9, not constructed in pre-deploy scaffolding).
|
|
|
|
|
"""
|
|
|
|
|
|
|
|
|
|
from __future__ import annotations
|
|
|
|
|
|
|
|
|
|
import pytest
|
|
|
|
|
|
|
|
|
|
# InMemorySaver is the modern name; fall back to MemorySaver on older langgraph.
|
|
|
|
|
try: # pragma: no cover - import shim
|
|
|
|
|
from langgraph.checkpoint.memory import InMemorySaver as _Saver
|
|
|
|
|
except ImportError: # pragma: no cover - import shim
|
|
|
|
|
from langgraph.checkpoint.memory import MemorySaver as _Saver
|
|
|
|
|
|
|
|
|
|
from agent_team import graph as graph_mod
|
|
|
|
|
from agent_team.graph import (
|
|
|
|
|
CLARIFY,
|
|
|
|
|
INTAKE,
|
|
|
|
|
P1_PHASE_SEQUENCE,
|
|
|
|
|
PLAN,
|
2026-06-22 16:14:27 -04:00
|
|
|
build_checkpoint_serde,
|
Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.
Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled
KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven
Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 14:17:44 -04:00
|
|
|
build_graph,
|
|
|
|
|
build_sqlite_checkpointer,
|
|
|
|
|
clarify_node,
|
|
|
|
|
get_pipeline_state,
|
|
|
|
|
intake_node,
|
|
|
|
|
pending_question,
|
|
|
|
|
plan_node,
|
|
|
|
|
plan_phase,
|
|
|
|
|
resume_task,
|
|
|
|
|
start_task,
|
|
|
|
|
thread_config,
|
|
|
|
|
)
|
|
|
|
|
from agent_team.task_model import Phase, PipelineState, TaskStatus
|
|
|
|
|
from agent_team.transport import QuestionSet
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@pytest.fixture()
|
|
|
|
|
def compiled():
|
|
|
|
|
"""A graph compiled with a fresh in-memory checkpointer per test."""
|
|
|
|
|
return build_graph(checkpointer=_Saver())
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# --- Module surface / constants. -------------------------------------------
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_node_name_constants_are_distinct() -> None:
|
|
|
|
|
assert len({INTAKE, CLARIFY, PLAN}) == 3
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_p1_phase_sequence_stops_at_plan() -> None:
|
|
|
|
|
# P1 ends at an approved plan — no BUILD/VERIFY in the wired sequence (§7.1).
|
|
|
|
|
assert P1_PHASE_SEQUENCE == (Phase.INTAKE, Phase.CLARIFY, Phase.PLAN)
|
|
|
|
|
assert Phase.BUILD not in P1_PHASE_SEQUENCE
|
|
|
|
|
assert Phase.VERIFY not in P1_PHASE_SEQUENCE
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# --- Pure node behaviour. ---------------------------------------------------
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_intake_node_activates_and_advances_to_clarify() -> None:
|
|
|
|
|
out = intake_node(PipelineState(thread_id="t", current_phase=Phase.INTAKE.value))
|
|
|
|
|
assert out["status"] == TaskStatus.ACTIVE.value
|
|
|
|
|
assert out["current_phase"] == Phase.CLARIFY.value
|
|
|
|
|
assert out["updated_at"]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_plan_node_lands_approved_plan_and_finishes() -> None:
|
|
|
|
|
out = plan_node(PipelineState(thread_id="t", qa_history=[{"answer": "x"}]))
|
|
|
|
|
assert out["status"] == TaskStatus.DONE.value
|
|
|
|
|
assert out["current_phase"] == Phase.DONE.value
|
|
|
|
|
assert out["plan"]["approved"] is True
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_plan_phase_counts_qa_turns() -> None:
|
|
|
|
|
state = PipelineState(qa_history=[{"answer": "a"}, {"answer": "b"}])
|
|
|
|
|
plan = plan_phase(state)
|
|
|
|
|
assert plan["qa_turns"] == 2
|
|
|
|
|
assert plan["approved"] is True
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_plan_phase_handles_empty_history() -> None:
|
|
|
|
|
assert plan_phase(PipelineState())["qa_turns"] == 0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_clarify_node_suspends_rather_than_falling_through() -> None:
|
|
|
|
|
# Called bare (no running graph), interrupt() refuses to return a value:
|
|
|
|
|
# it raises because there is no runnable context to suspend into. This
|
|
|
|
|
# confirms clarify_node genuinely suspends rather than falling through to
|
|
|
|
|
# its post-interrupt return.
|
|
|
|
|
with pytest.raises(RuntimeError):
|
|
|
|
|
clarify_node(PipelineState(thread_id="t", transport="slack"))
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# --- Graph assembly. --------------------------------------------------------
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_build_graph_without_checkpointer_compiles() -> None:
|
|
|
|
|
# An uncheckpointed graph still compiles (used only for straight-through
|
|
|
|
|
# smoke paths); the driver requires a checkpointer for suspend/resume.
|
|
|
|
|
assert build_graph() is not None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_build_graph_with_checkpointer_compiles(compiled) -> None:
|
|
|
|
|
assert compiled is not None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_graph_nodes_present(compiled) -> None:
|
|
|
|
|
nodes = set(compiled.get_graph().nodes)
|
|
|
|
|
assert {INTAKE, CLARIFY, PLAN} <= nodes
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# --- Suspend / resume end to end. ------------------------------------------
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_start_task_suspends_on_human_gate(compiled) -> None:
|
|
|
|
|
thread_id, state = start_task(compiled, transport="slack")
|
|
|
|
|
# The task ran INTAKE then suspended at CLARIFY's interrupt().
|
|
|
|
|
assert "__interrupt__" in state
|
|
|
|
|
payload = pending_question(compiled, thread_id=thread_id)
|
|
|
|
|
assert payload is not None
|
|
|
|
|
assert payload["thread_id"] == thread_id
|
|
|
|
|
assert payload["transport"] == "slack"
|
|
|
|
|
assert payload["turn"] == 0
|
|
|
|
|
assert payload["deadline"]
|
|
|
|
|
|
|
|
|
|
|
2026-06-23 13:53:39 -04:00
|
|
|
def test_start_task_seeds_task_description_into_state(compiled) -> None:
|
|
|
|
|
# Regression: the intake description (Slack /new-task text, GitHub issue body)
|
|
|
|
|
# must reach the graph state so the clarifier can reason about it. It is
|
|
|
|
|
# seeded into the initial invoke and must persist through INTAKE into the
|
|
|
|
|
# suspended CLARIFY snapshot (intake_node returns only a partial state).
|
|
|
|
|
_thread_id, state = start_task(
|
|
|
|
|
compiled, transport="slack", task="build a login form"
|
|
|
|
|
)
|
|
|
|
|
assert state.get("task") == "build a login form"
|
|
|
|
|
|
|
|
|
|
|
2026-06-23 15:11:19 -04:00
|
|
|
def test_start_task_seeds_slack_thread_ts_into_state(compiled) -> None:
|
|
|
|
|
# One-thread-per-task: the root "📥 Task received" message ts must reach the
|
|
|
|
|
# graph state so every later question/notification threads under it. Like
|
|
|
|
|
# ``task`` it is seeded into the initial invoke and persists through INTAKE
|
|
|
|
|
# into the suspended CLARIFY snapshot.
|
|
|
|
|
_thread_id, state = start_task(
|
|
|
|
|
compiled, transport="slack", slack_thread_ts="1700000000.ROOT"
|
|
|
|
|
)
|
|
|
|
|
assert state.get("slack_thread_ts") == "1700000000.ROOT"
|
|
|
|
|
# It is also surfaced on the pending interrupt payload so the responder can
|
|
|
|
|
# thread the question post.
|
|
|
|
|
payload = pending_question(compiled, thread_id=_thread_id)
|
|
|
|
|
assert payload["slack_thread_ts"] == "1700000000.ROOT"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_start_task_default_slack_thread_ts_is_empty(compiled) -> None:
|
|
|
|
|
# A task with no root post (e.g. a GitHub-issue origin): slack_thread_ts is
|
|
|
|
|
# empty so questions post top-level exactly as before.
|
|
|
|
|
_thread_id, state = start_task(compiled, transport="slack")
|
|
|
|
|
assert state.get("slack_thread_ts") == ""
|
|
|
|
|
|
|
|
|
|
|
Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.
Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled
KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven
Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 14:17:44 -04:00
|
|
|
def test_pending_question_carries_foundation_questionset(compiled) -> None:
|
|
|
|
|
thread_id, _ = start_task(compiled, transport="slack")
|
|
|
|
|
payload = pending_question(compiled, thread_id=thread_id)
|
|
|
|
|
qset = payload["question_set"]
|
|
|
|
|
# Verbatim foundation contract — not a redefinition.
|
|
|
|
|
assert isinstance(qset, QuestionSet)
|
|
|
|
|
assert qset.thread_id == thread_id
|
|
|
|
|
assert qset.question_id == payload["question_id"]
|
|
|
|
|
assert qset.turn == 0
|
|
|
|
|
assert qset.questions # non-empty question-set
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_resume_drives_task_to_done(compiled) -> None:
|
|
|
|
|
thread_id, _ = start_task(compiled, transport="slack")
|
|
|
|
|
final = resume_task(compiled, thread_id=thread_id, answer={"text": "do the thing"})
|
|
|
|
|
assert final["status"] == TaskStatus.DONE.value
|
|
|
|
|
assert final["current_phase"] == Phase.DONE.value
|
|
|
|
|
assert final["plan"]["approved"] is True
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_answer_is_recorded_in_qa_history(compiled) -> None:
|
|
|
|
|
thread_id, _ = start_task(compiled, transport="slack")
|
|
|
|
|
answer = {"text": "scope is X"}
|
|
|
|
|
final = resume_task(compiled, thread_id=thread_id, answer=answer)
|
|
|
|
|
assert len(final["qa_history"]) == 1
|
|
|
|
|
assert final["qa_history"][0]["answer"] == answer
|
|
|
|
|
assert final["qa_history"][0]["turn"] == 0
|
|
|
|
|
|
|
|
|
|
|
2026-06-17 14:47:54 -04:00
|
|
|
def test_question_id_is_stable_across_resume(compiled) -> None:
|
|
|
|
|
# The clarifier node re-executes on resume; the question_id must NOT change
|
|
|
|
|
# between the id delivered at suspend (the ledger key) and the one recorded
|
|
|
|
|
# in qa_history, or the §3.3.1 identity contract breaks.
|
|
|
|
|
thread_id, _ = start_task(compiled, transport="slack")
|
|
|
|
|
delivered = pending_question(compiled, thread_id=thread_id)["question_id"]
|
|
|
|
|
final = resume_task(compiled, thread_id=thread_id, answer="ok")
|
|
|
|
|
assert final["qa_history"][0]["question_id"] == delivered
|
|
|
|
|
|
|
|
|
|
|
Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.
Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled
KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven
Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 14:17:44 -04:00
|
|
|
def test_no_pending_question_after_completion(compiled) -> None:
|
|
|
|
|
thread_id, _ = start_task(compiled, transport="slack")
|
|
|
|
|
resume_task(compiled, thread_id=thread_id, answer="ok")
|
|
|
|
|
assert pending_question(compiled, thread_id=thread_id) is None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_get_pipeline_state_reflects_suspend_then_done(compiled) -> None:
|
|
|
|
|
thread_id, _ = start_task(compiled, transport="slack")
|
|
|
|
|
mid = get_pipeline_state(compiled, thread_id=thread_id)
|
|
|
|
|
# Suspended ON the clarifier gate: INTAKE already advanced the phase to
|
|
|
|
|
# CLARIFY, and the clarifier's post-interrupt write (-> PLAN) has NOT yet
|
|
|
|
|
# committed because the node is paused at interrupt(). Task is mid-flight.
|
|
|
|
|
assert mid["current_phase"] == Phase.CLARIFY.value
|
|
|
|
|
assert mid["status"] == TaskStatus.ACTIVE.value
|
|
|
|
|
resume_task(compiled, thread_id=thread_id, answer="ok")
|
|
|
|
|
done = get_pipeline_state(compiled, thread_id=thread_id)
|
|
|
|
|
assert done["status"] == TaskStatus.DONE.value
|
|
|
|
|
assert done["current_phase"] == Phase.DONE.value
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# --- Thread isolation (§3.3.1 P1 exit criterion (d)). ----------------------
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_two_tasks_suspend_and_resume_independently(compiled) -> None:
|
|
|
|
|
t1, _ = start_task(compiled, transport="slack")
|
|
|
|
|
t2, _ = start_task(compiled, transport="github")
|
|
|
|
|
|
|
|
|
|
assert t1 != t2
|
|
|
|
|
p1 = pending_question(compiled, thread_id=t1)
|
|
|
|
|
p2 = pending_question(compiled, thread_id=t2)
|
|
|
|
|
assert p1["transport"] == "slack"
|
|
|
|
|
assert p2["transport"] == "github"
|
|
|
|
|
assert p1["question_id"] != p2["question_id"]
|
|
|
|
|
|
|
|
|
|
# Resume only t1; t2 must remain suspended on its own gate.
|
|
|
|
|
f1 = resume_task(compiled, thread_id=t1, answer="answer-1")
|
|
|
|
|
assert f1["status"] == TaskStatus.DONE.value
|
|
|
|
|
assert pending_question(compiled, thread_id=t2) is not None
|
|
|
|
|
|
|
|
|
|
f2 = resume_task(compiled, thread_id=t2, answer="answer-2")
|
|
|
|
|
assert f2["status"] == TaskStatus.DONE.value
|
|
|
|
|
assert f2["qa_history"][0]["answer"] == "answer-2"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_explicit_thread_id_is_honoured(compiled) -> None:
|
|
|
|
|
tid, _ = start_task(compiled, thread_id="fixed-thread", transport="slack")
|
|
|
|
|
assert tid == "fixed-thread"
|
|
|
|
|
assert pending_question(compiled, thread_id="fixed-thread") is not None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# --- Durable resume across a fresh graph object (P1 exit criterion (a)). ----
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_resume_works_on_a_new_graph_over_shared_checkpointer() -> None:
|
|
|
|
|
# Simulates a process restart: a NEW compiled graph object built over the
|
|
|
|
|
# SAME checkpointer must resume a task suspended by the first graph object.
|
|
|
|
|
saver = _Saver()
|
|
|
|
|
g1 = build_graph(checkpointer=saver)
|
|
|
|
|
thread_id, _ = start_task(g1, transport="slack")
|
|
|
|
|
|
|
|
|
|
g2 = build_graph(checkpointer=saver) # "after restart"
|
|
|
|
|
assert pending_question(g2, thread_id=thread_id) is not None
|
|
|
|
|
final = resume_task(g2, thread_id=thread_id, answer="post-restart")
|
|
|
|
|
assert final["status"] == TaskStatus.DONE.value
|
|
|
|
|
assert final["qa_history"][0]["answer"] == "post-restart"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# --- Driver-seam helpers. ---------------------------------------------------
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_thread_config_shape() -> None:
|
|
|
|
|
assert thread_config("abc") == {"configurable": {"thread_id": "abc"}}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_start_task_mints_unique_thread_ids(compiled) -> None:
|
|
|
|
|
t1, _ = start_task(compiled, transport="slack")
|
|
|
|
|
t2, _ = start_task(compiled, transport="slack")
|
|
|
|
|
assert t1 != t2
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# --- Production checkpointer factory. --------------------------------------
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_build_sqlite_checkpointer_missing_dep_raises_runtimeerror(
|
|
|
|
|
monkeypatch, tmp_path
|
|
|
|
|
) -> None:
|
|
|
|
|
# When the optional langgraph-checkpoint-sqlite package is absent, the
|
|
|
|
|
# factory must fail loudly with a clear RuntimeError, never silently run
|
|
|
|
|
# uncheckpointed. Force the ImportError path deterministically.
|
|
|
|
|
import builtins
|
|
|
|
|
|
|
|
|
|
real_import = builtins.__import__
|
|
|
|
|
|
|
|
|
|
def _blocking_import(name, *args, **kwargs):
|
|
|
|
|
if name == "langgraph.checkpoint.sqlite":
|
|
|
|
|
raise ImportError("blocked for test")
|
|
|
|
|
return real_import(name, *args, **kwargs)
|
|
|
|
|
|
|
|
|
|
monkeypatch.setattr(builtins, "__import__", _blocking_import)
|
|
|
|
|
with pytest.raises(RuntimeError, match="SQLite checkpointer"):
|
|
|
|
|
build_sqlite_checkpointer(tmp_path / "state.db")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_build_sqlite_checkpointer_builds_when_dep_present(tmp_path) -> None:
|
|
|
|
|
# If the optional package IS installed, the factory returns a checkpointer
|
|
|
|
|
# over the DB path. Skip cleanly where it's absent (pre-deploy scaffolding).
|
|
|
|
|
pytest.importorskip("langgraph.checkpoint.sqlite")
|
2026-06-18 12:56:42 -04:00
|
|
|
cm = build_sqlite_checkpointer(tmp_path / "nested" / "state.db")
|
|
|
|
|
assert cm is not None
|
Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.
Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled
KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven
Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 14:17:44 -04:00
|
|
|
assert (tmp_path / "nested").is_dir()
|
2026-06-18 12:56:42 -04:00
|
|
|
# Contract: the factory returns a CONTEXT MANAGER (SqliteSaver.from_conn_string
|
|
|
|
|
# is a @contextmanager), so callers must enter it. Pin that here so a future
|
|
|
|
|
# change that returns a bare/un-entered object is caught (review FIX).
|
|
|
|
|
assert hasattr(cm, "__enter__") and hasattr(cm, "__exit__")
|
|
|
|
|
with cm as saver:
|
|
|
|
|
# The entered object is the real saver the graph compiles against.
|
|
|
|
|
assert hasattr(saver, "get_next_version")
|
|
|
|
|
|
|
|
|
|
|
2026-06-22 16:14:27 -04:00
|
|
|
# --- Checkpoint serializer / QuestionSet msgpack allowlist (D9). ------------
|
|
|
|
|
# QuestionSet rides in the clarifier interrupt payload and so is msgpack-encoded
|
|
|
|
|
# into every checkpoint. The default serializer deserializes it but logs
|
|
|
|
|
# "Deserializing unregistered type agent_team.transport.base.QuestionSet ... will
|
|
|
|
|
# be blocked in a future version" on each load, and the coming LangGraph default
|
|
|
|
|
# turns that warning into a hard block (breaking durable resume). These pin that
|
|
|
|
|
# build_checkpoint_serde registers the type so it round-trips silently AND is
|
|
|
|
|
# already block-clean.
|
|
|
|
|
|
|
|
|
|
_QSET_MSGPACK_KEY = ("agent_team.transport.base", "QuestionSet")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _capture_serde_warnings():
|
|
|
|
|
"""Attach a capturing handler to the serde logger; return (handler, buffer)."""
|
|
|
|
|
import io
|
|
|
|
|
import logging
|
|
|
|
|
|
|
|
|
|
buf = io.StringIO()
|
|
|
|
|
handler = logging.StreamHandler(buf)
|
|
|
|
|
handler.setLevel(logging.WARNING)
|
|
|
|
|
logger = logging.getLogger("langgraph.checkpoint.serde.jsonplus")
|
|
|
|
|
logger.addHandler(handler)
|
|
|
|
|
logger.setLevel(logging.WARNING)
|
|
|
|
|
return logger, handler, buf
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _reset_serde_warning_dedup() -> None:
|
|
|
|
|
"""Clear the serializer's process-wide warn-once dedup set.
|
|
|
|
|
|
|
|
|
|
jsonplus dedups "unregistered type" warnings across the process lifetime, so
|
|
|
|
|
an earlier test (or the assertion below) could mask a regression. Clearing the
|
|
|
|
|
set makes each assertion observe the live behavior, not a stale dedup.
|
|
|
|
|
"""
|
|
|
|
|
from langgraph.checkpoint.serde import jsonplus as _jp
|
|
|
|
|
|
|
|
|
|
_jp._warned_unregistered_types.clear()
|
|
|
|
|
_jp._warned_blocked_types.clear()
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_build_checkpoint_serde_allowlists_questionset() -> None:
|
|
|
|
|
# The serde must carry QuestionSet on its msgpack allowlist (an explicit
|
|
|
|
|
# collection, NOT the permissive ``True`` sentinel) — that is the registration
|
|
|
|
|
# path the warning recommends and what makes resume forward-compatible.
|
|
|
|
|
pytest.importorskip("langgraph.checkpoint.serde.jsonplus")
|
|
|
|
|
serde = build_checkpoint_serde()
|
|
|
|
|
allowed = serde._allowed_msgpack_modules
|
|
|
|
|
assert allowed is not True # not the warn-on-everything permissive default
|
|
|
|
|
assert _QSET_MSGPACK_KEY in allowed
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_questionset_round_trips_through_serde_without_warning() -> None:
|
|
|
|
|
# The configured serializer must round-trip a QuestionSet AND emit no
|
|
|
|
|
# "unregistered type" warning on deserialize.
|
|
|
|
|
pytest.importorskip("langgraph.checkpoint.serde.jsonplus")
|
|
|
|
|
_reset_serde_warning_dedup()
|
|
|
|
|
logger, handler, buf = _capture_serde_warnings()
|
|
|
|
|
try:
|
|
|
|
|
serde = build_checkpoint_serde()
|
|
|
|
|
qset = QuestionSet(
|
|
|
|
|
thread_id="t-1",
|
|
|
|
|
question_id="q-1",
|
|
|
|
|
turn=0,
|
|
|
|
|
questions=["What is in scope?"],
|
|
|
|
|
context={"phase": Phase.CLARIFY.value},
|
|
|
|
|
)
|
|
|
|
|
encoded = serde.dumps_typed(qset)
|
|
|
|
|
restored = serde.loads_typed(encoded)
|
|
|
|
|
finally:
|
|
|
|
|
logger.removeHandler(handler)
|
|
|
|
|
|
|
|
|
|
assert restored == qset
|
|
|
|
|
assert "unregistered type" not in buf.getvalue()
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_questionset_round_trips_through_sqlite_checkpointer_without_warning(
|
|
|
|
|
tmp_path,
|
|
|
|
|
) -> None:
|
|
|
|
|
# End-to-end against the REAL production saver: drive the graph to the
|
|
|
|
|
# clarifier suspend (which checkpoints a QuestionSet), force a checkpoint
|
|
|
|
|
# load, and resume — asserting no "unregistered type" warning prints and the
|
|
|
|
|
# task still completes. This is the runtime regression guard for the warning.
|
|
|
|
|
pytest.importorskip("langgraph.checkpoint.sqlite")
|
|
|
|
|
_reset_serde_warning_dedup()
|
|
|
|
|
logger, handler, buf = _capture_serde_warnings()
|
|
|
|
|
try:
|
|
|
|
|
cm = build_sqlite_checkpointer(tmp_path / "state.db")
|
|
|
|
|
with cm as saver:
|
|
|
|
|
graph = build_graph(saver)
|
|
|
|
|
thread_id, _ = start_task(graph, transport="slack")
|
|
|
|
|
pending = pending_question(graph, thread_id=thread_id)
|
|
|
|
|
assert isinstance(pending["question_set"], QuestionSet)
|
|
|
|
|
# Force a fresh checkpoint deserialize (the warning's trigger point).
|
|
|
|
|
get_pipeline_state(graph, thread_id=thread_id)
|
|
|
|
|
resumed = resume_task(graph, thread_id=thread_id, answer="scope it")
|
|
|
|
|
finally:
|
|
|
|
|
logger.removeHandler(handler)
|
|
|
|
|
|
|
|
|
|
assert resumed["status"] == TaskStatus.DONE.value
|
|
|
|
|
assert "unregistered type" not in buf.getvalue()
|
|
|
|
|
|
|
|
|
|
|
2026-06-18 12:56:42 -04:00
|
|
|
# --- P2 review-loop wiring. -------------------------------------------------
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _p2_plan_stub(state: PipelineState) -> PipelineState:
|
|
|
|
|
"""Stand-in for the real planner: emit a plan and advance to REVIEW.
|
|
|
|
|
|
|
|
|
|
Mirrors planner.plan_node's contract (sets ``plan`` + phase REVIEW) without a
|
|
|
|
|
model call, so the P2 graph topology + the review loop can be driven in a
|
|
|
|
|
unit test. The revision index tracks prior review rounds.
|
|
|
|
|
"""
|
|
|
|
|
revisions = len(state.get("review_verdicts") or [])
|
|
|
|
|
return PipelineState(
|
|
|
|
|
plan={"phases": ["P1"], "revision": revisions},
|
|
|
|
|
current_phase=Phase.REVIEW.value,
|
|
|
|
|
status=TaskStatus.ACTIVE.value,
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _p2_graph(review_text: str):
|
|
|
|
|
"""Compile a P2 graph whose review invoker returns ``review_text``."""
|
|
|
|
|
from agent_team.nodes import review_loop
|
|
|
|
|
|
|
|
|
|
review_loop.set_review_invoker(lambda prompt, **kw: review_text)
|
|
|
|
|
return build_graph(
|
|
|
|
|
checkpointer=_Saver(),
|
|
|
|
|
live_plan_node=_p2_plan_stub,
|
|
|
|
|
review_node=review_loop.bind_review_node(),
|
|
|
|
|
route_review=review_loop.route_after_review,
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@pytest.fixture
|
|
|
|
|
def restore_review_invoker():
|
|
|
|
|
"""Save/restore the review-loop module-global invoker around a test."""
|
|
|
|
|
from agent_team.nodes import review_loop
|
|
|
|
|
|
|
|
|
|
saved = review_loop._review_invoker
|
|
|
|
|
yield
|
|
|
|
|
review_loop._review_invoker = saved
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_build_graph_review_node_requires_route() -> None:
|
|
|
|
|
from agent_team.nodes import review_loop
|
|
|
|
|
|
|
|
|
|
with pytest.raises(ValueError, match="route_review"):
|
|
|
|
|
build_graph(review_node=review_loop.review_node)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_p2_graph_approve_terminates(restore_review_invoker) -> None:
|
|
|
|
|
# clarify(stub) -> plan(stub->REVIEW) -> review(APPROVE) -> END.
|
|
|
|
|
graph = _p2_graph("VERDICT: APPROVE\nlooks solid")
|
|
|
|
|
thread_id, _ = start_task(graph, transport="slack")
|
|
|
|
|
final = resume_task(graph, thread_id=thread_id, answer="scope is X")
|
|
|
|
|
|
|
|
|
|
# The review node advanced an APPROVED plan toward BUILD and the graph ended.
|
|
|
|
|
assert final["current_phase"] == Phase.BUILD.value
|
|
|
|
|
assert len(final["review_verdicts"]) == 1
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_p2_graph_loops_then_escalates_on_persistent_changes(
|
|
|
|
|
restore_review_invoker,
|
|
|
|
|
) -> None:
|
|
|
|
|
# A reviewer that never approves loops plan<->review until the round cap,
|
|
|
|
|
# then escalates (parks) rather than spinning. Default cap is 3 rounds.
|
|
|
|
|
graph = _p2_graph("VERDICT: REQUEST CHANGES\nstill not ready")
|
|
|
|
|
thread_id, _ = start_task(graph, transport="slack")
|
|
|
|
|
final = resume_task(graph, thread_id=thread_id, answer="scope is X")
|
|
|
|
|
|
|
|
|
|
assert final["current_phase"] == Phase.PARKED.value
|
|
|
|
|
assert final["status"] == TaskStatus.PARKED.value
|
|
|
|
|
assert len(final["review_verdicts"]) == 3 # looped to the cap, then escalated
|
Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.
Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled
KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven
Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 14:17:44 -04:00
|
|
|
|
|
|
|
|
|
2026-06-23 20:11:32 -04:00
|
|
|
# --- Plan-review human decision gate (Phase B2a). ---------------------------
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _plan_gate_graph(review_text: str):
|
|
|
|
|
"""Compile a P2 graph with the plan gate wired (plan_gate=True).
|
|
|
|
|
|
|
|
|
|
A reviewer that never approves drives the plan<->review loop to the review
|
|
|
|
|
cap, which — with the gate wired — suspends on the resumable PLAN_GATE
|
|
|
|
|
interrupt instead of terminally parking.
|
|
|
|
|
"""
|
|
|
|
|
from agent_team.nodes import review_loop
|
|
|
|
|
|
|
|
|
|
review_loop.set_review_invoker(lambda prompt, **kw: review_text)
|
|
|
|
|
return build_graph(
|
|
|
|
|
checkpointer=_Saver(),
|
|
|
|
|
live_plan_node=_p2_plan_stub,
|
|
|
|
|
review_node=review_loop.bind_review_node(),
|
|
|
|
|
route_review=review_loop.route_after_review,
|
|
|
|
|
plan_gate=True,
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _drive_to_plan_gate(graph):
|
|
|
|
|
"""Start a task and resume the clarifier so it lands on the plan gate.
|
|
|
|
|
|
|
|
|
|
Returns ``(thread_id, gate_payload)`` where ``gate_payload`` is the pending
|
|
|
|
|
plan-decision interrupt payload.
|
|
|
|
|
"""
|
|
|
|
|
from agent_team.graph import pending_question
|
|
|
|
|
|
|
|
|
|
thread_id, _ = start_task(graph, transport="slack", slack_thread_ts="ROOT.1")
|
|
|
|
|
resume_task(graph, thread_id=thread_id, answer="scope is X")
|
|
|
|
|
payload = pending_question(graph, thread_id=thread_id)
|
|
|
|
|
return thread_id, payload
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_plan_gate_requires_review_node() -> None:
|
|
|
|
|
with pytest.raises(ValueError, match="plan_gate requires review_node"):
|
|
|
|
|
build_graph(plan_gate=True)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_review_cap_interrupts_at_plan_gate(restore_review_invoker) -> None:
|
|
|
|
|
# Driving to the review-cap dead-end suspends on the plan gate (a pending
|
|
|
|
|
# plan_decision interrupt with the plan + findings) rather than parking.
|
|
|
|
|
from agent_team.graph import PLAN_DECISION_KIND
|
|
|
|
|
|
|
|
|
|
graph = _plan_gate_graph("VERDICT: REQUEST CHANGES\nstill not ready")
|
|
|
|
|
thread_id, payload = _drive_to_plan_gate(graph)
|
|
|
|
|
|
|
|
|
|
assert payload is not None
|
|
|
|
|
assert payload["kind"] == PLAN_DECISION_KIND
|
|
|
|
|
# Mirrors the clarifier contract: thread_id / question_id / turn / transport /
|
|
|
|
|
# deadline / slack_thread_ts all present (so pending_question + ResumeWorker
|
|
|
|
|
# drive it uniformly).
|
|
|
|
|
assert payload["thread_id"] == thread_id
|
|
|
|
|
assert payload["question_id"]
|
|
|
|
|
assert isinstance(payload["turn"], int)
|
|
|
|
|
assert payload["transport"] == "slack"
|
|
|
|
|
assert payload["deadline"]
|
|
|
|
|
assert payload["slack_thread_ts"] == "ROOT.1"
|
|
|
|
|
# Plus the review context the owner decides over.
|
|
|
|
|
assert payload["plan"] is not None
|
|
|
|
|
assert "still not ready" in payload["findings"]
|
|
|
|
|
|
|
|
|
|
# It did NOT terminally park: the task is suspended on the gate interrupt
|
|
|
|
|
# (resumable), not finished. (The status channel still reads the review
|
|
|
|
|
# node's carried-over "parked" until the gate's resume overwrites it; the
|
|
|
|
|
# load-bearing signal is the live pending interrupt.)
|
|
|
|
|
from agent_team.graph import pending_question
|
|
|
|
|
|
|
|
|
|
assert pending_question(graph, thread_id=thread_id) is not None
|
|
|
|
|
snapshot = graph.get_state(thread_config(thread_id))
|
|
|
|
|
assert snapshot.next # graph is suspended, not at a terminal END
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_plan_gate_approve_settles_as_approved_plan(restore_review_invoker) -> None:
|
|
|
|
|
# Command(resume={"decision":"approve"}) -> the same terminal "approved plan"
|
|
|
|
|
# state an auto-approved plan reaches today (phase BUILD, status ACTIVE).
|
|
|
|
|
graph = _plan_gate_graph("VERDICT: REQUEST CHANGES\nnope")
|
|
|
|
|
thread_id, _ = _drive_to_plan_gate(graph)
|
|
|
|
|
final = resume_task(graph, thread_id=thread_id, answer={"decision": "approve"})
|
|
|
|
|
|
|
|
|
|
assert final["current_phase"] == Phase.BUILD.value
|
|
|
|
|
assert final["status"] == TaskStatus.ACTIVE.value
|
|
|
|
|
|
|
|
|
|
|
2026-06-24 14:13:11 -04:00
|
|
|
def _p3_plan_gate_graph(review_text: str, *, ci_result_fetcher):
|
|
|
|
|
"""Compile a graph with BOTH the plan gate AND the P3 build->verify subgraph.
|
|
|
|
|
|
|
|
|
|
The reviewer (``review_text``) never approves, so the plan<->review loop hits
|
|
|
|
|
the cap and suspends on the human PLAN_GATE; a human ``approve`` must then
|
|
|
|
|
route into the build subgraph (BUILD -> VERIFY), exactly as the reviewer's
|
|
|
|
|
own auto-approve "build" route does.
|
|
|
|
|
"""
|
|
|
|
|
from agent_team.nodes import review_loop
|
|
|
|
|
from agent_team.nodes.build_verify_subgraph import (
|
|
|
|
|
make_build_node,
|
|
|
|
|
make_verify_node,
|
|
|
|
|
route_after_verify,
|
|
|
|
|
)
|
|
|
|
|
from agent_team.nodes.verifier import VerifierConfig
|
|
|
|
|
|
|
|
|
|
review_loop.set_review_invoker(lambda prompt, **kw: review_text)
|
|
|
|
|
|
|
|
|
|
def fake_builder(*, plan, config):
|
|
|
|
|
return _p3_diff()
|
|
|
|
|
|
|
|
|
|
build_node = make_build_node(diff_builder=fake_builder)
|
|
|
|
|
verify_node = make_verify_node(
|
|
|
|
|
VerifierConfig(expected_run_id="r1", allowed_scope=["src"]),
|
|
|
|
|
ci_result_fetcher=ci_result_fetcher,
|
|
|
|
|
)
|
|
|
|
|
return build_graph(
|
|
|
|
|
checkpointer=_Saver(),
|
|
|
|
|
live_plan_node=_p3_plan_stub,
|
|
|
|
|
review_node=review_loop.bind_review_node(),
|
|
|
|
|
route_review=review_loop.route_after_review,
|
|
|
|
|
plan_gate=True,
|
|
|
|
|
build_verify=(build_node, verify_node, route_after_verify),
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_plan_gate_approve_enters_build_subgraph_when_p3_wired(
|
|
|
|
|
restore_review_invoker,
|
|
|
|
|
) -> None:
|
|
|
|
|
# Regression for the gate sibling of issue #60: a human approve at the plan
|
|
|
|
|
# gate MUST route into the P3 build subgraph (BUILD -> VERIFY), not dead-end
|
|
|
|
|
# at END leaving the task stuck at phase=build. The reviewer never approves,
|
|
|
|
|
# so the loop caps to the human gate; an authenticated CI pass then carries
|
|
|
|
|
# build -> verify -> DONE — proving the approve edge reached BUILD_NODE.
|
|
|
|
|
def pass_fetcher(state):
|
|
|
|
|
from agent_team.state_store import compute_content_hash
|
|
|
|
|
|
|
|
|
|
diff_hash = compute_content_hash(_p3_diff().encode("utf-8"))
|
|
|
|
|
return {"run_id": "r1", "conclusion": "success", "diff_hash": diff_hash}
|
|
|
|
|
|
|
|
|
|
graph = _p3_plan_gate_graph(
|
|
|
|
|
"VERDICT: REQUEST CHANGES\nnot yet", ci_result_fetcher=pass_fetcher
|
|
|
|
|
)
|
|
|
|
|
thread_id, _ = _drive_to_plan_gate(graph)
|
|
|
|
|
final = resume_task(graph, thread_id=thread_id, answer={"decision": "approve"})
|
|
|
|
|
|
|
|
|
|
# Reached the build->verify PASS terminus, NOT a phase=build dead-end.
|
|
|
|
|
assert final["current_phase"] == Phase.DONE.value
|
|
|
|
|
assert final["status"] == TaskStatus.DONE.value
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_plan_gate_approve_inert_p3_still_settles_at_approved_terminus(
|
|
|
|
|
restore_review_invoker,
|
|
|
|
|
) -> None:
|
|
|
|
|
# With the P3 build subgraph wired but its CI fetcher INERT (no authenticated
|
|
|
|
|
# result), a gate approve enters build->verify and parks at VERIFY (the
|
|
|
|
|
# production-safe default) — it must NOT fabricate a pass. Confirms the
|
|
|
|
|
# approve edge routes through the subgraph, not to END.
|
|
|
|
|
graph = _p3_plan_gate_graph(
|
|
|
|
|
"VERDICT: REQUEST CHANGES\nnot yet", ci_result_fetcher=lambda state: None
|
|
|
|
|
)
|
|
|
|
|
thread_id, _ = _drive_to_plan_gate(graph)
|
|
|
|
|
final = resume_task(graph, thread_id=thread_id, answer={"decision": "approve"})
|
|
|
|
|
|
|
|
|
|
assert final["current_phase"] == Phase.PARKED.value
|
|
|
|
|
assert final["status"] == TaskStatus.PARKED.value
|
|
|
|
|
|
|
|
|
|
|
2026-06-23 20:11:32 -04:00
|
|
|
def test_plan_gate_request_changes_loops_back_with_notes(
|
|
|
|
|
restore_review_invoker,
|
|
|
|
|
) -> None:
|
|
|
|
|
# Command(resume={"decision":"request_changes","notes":"do X"}) loops back to
|
|
|
|
|
# the planner; the notes must reach _format_review_feedback / the planner.
|
|
|
|
|
from agent_team.nodes.planner import _format_review_feedback
|
|
|
|
|
|
|
|
|
|
captured: dict[str, object] = {}
|
|
|
|
|
|
|
|
|
|
def capturing_plan(state):
|
|
|
|
|
captured["feedback"] = _format_review_feedback(
|
|
|
|
|
list(state.get("review_verdicts") or [])
|
|
|
|
|
)
|
|
|
|
|
# After observing the folded-in notes, APPROVE on the re-plan so the
|
|
|
|
|
# graph settles (the review invoker is bound per-test below).
|
|
|
|
|
return _p2_plan_stub(state)
|
|
|
|
|
|
|
|
|
|
from agent_team.nodes import review_loop
|
|
|
|
|
|
|
|
|
|
# First review round REQUEST CHANGES (to reach the gate); after the human
|
|
|
|
|
# request_changes loops back, the next review APPROVES so the task settles.
|
|
|
|
|
texts = iter(
|
|
|
|
|
[
|
|
|
|
|
"VERDICT: REQUEST CHANGES\nnot ready",
|
|
|
|
|
"VERDICT: APPROVE\nnow good",
|
|
|
|
|
]
|
|
|
|
|
)
|
|
|
|
|
last = "VERDICT: APPROVE\nnow good"
|
|
|
|
|
|
|
|
|
|
def invoker(prompt, **kw):
|
|
|
|
|
nonlocal last
|
|
|
|
|
try:
|
|
|
|
|
last = next(texts)
|
|
|
|
|
except StopIteration:
|
|
|
|
|
pass
|
|
|
|
|
return last
|
|
|
|
|
|
|
|
|
|
review_loop.set_review_invoker(invoker)
|
|
|
|
|
graph = build_graph(
|
|
|
|
|
checkpointer=_Saver(),
|
|
|
|
|
live_plan_node=capturing_plan,
|
|
|
|
|
review_node=review_loop.bind_review_node({"max_review_rounds": 1}),
|
|
|
|
|
route_review=review_loop.route_after_review,
|
|
|
|
|
plan_gate=True,
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
thread_id, _ = start_task(graph, transport="slack")
|
|
|
|
|
resume_task(graph, thread_id=thread_id, answer="scope is X")
|
|
|
|
|
final = resume_task(
|
|
|
|
|
graph,
|
|
|
|
|
thread_id=thread_id,
|
|
|
|
|
answer={"decision": "request_changes", "notes": "do X"},
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
# The human notes reached the planner's review-feedback formatter on re-plan.
|
|
|
|
|
assert "do X" in str(captured.get("feedback", ""))
|
|
|
|
|
# And the loop re-entered plan -> review and settled (not stuck at the gate).
|
|
|
|
|
assert final["current_phase"] == Phase.BUILD.value
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_plan_gate_abandon_fails(restore_review_invoker) -> None:
|
|
|
|
|
# Command(resume={"decision":"abandon"}) -> terminal FAILED.
|
|
|
|
|
graph = _plan_gate_graph("VERDICT: REQUEST CHANGES\nnope")
|
|
|
|
|
thread_id, _ = _drive_to_plan_gate(graph)
|
|
|
|
|
final = resume_task(graph, thread_id=thread_id, answer={"decision": "abandon"})
|
|
|
|
|
|
|
|
|
|
assert final["status"] == TaskStatus.FAILED.value
|
|
|
|
|
|
|
|
|
|
|
fix(agent-team): centralize safe decision mapping + remove free-text abandon hair-trigger
Security-review follow-up (LOGIC-01/02/05, all confirmed correctness).
- New transport-neutral `decisions.normalize_decision(raw, *, allow_abandon)` is
the single source of truth: approve-allowlist→approve; abandon-allowlist→abandon
ONLY when allow_abandon; everything else (prose, empty, abandon-verbs when
disallowed) → request_changes with the full reply as notes; idempotent on an
already-formed decision dict. slack_adapter.map_plan_decision is now a thin
wrapper (default allow_abandon=True, no caller churn).
- LOGIC-01/02: graph._parse_decision now delegates to normalize_decision (was:
any unrecognized verb → abandon → FAILED). The graph is now the universal safe
backstop, so EVERY writer that bypassed the listener mapping — operator CLI
answer_on_behalf (raw), Coordinator.submit_answer (raw), the recovery sweep —
loops back on prose instead of silently FAILing the task. Explicit abandon
still abandons (preserves the confirmed-button path).
- LOGIC-05: the Slack FREE-TEXT reply path maps with allow_abandon=False, so a
bare "cancel"/"stop"/"abandon" typed in-thread → request_changes (never
terminal abandon); abandon stays reachable only via the confirm-guarded button.
Tests: graph unrecognized→loops-back (not FAILED), operator raw-prose→request_
changes, free-text destructive verbs→request_changes vs button→abandon,
normalizer idempotency. 1484 passed.
2026-06-24 11:48:04 -04:00
|
|
|
def test_plan_gate_unrecognized_decision_loops_back_not_failed(
|
|
|
|
|
restore_review_invoker,
|
|
|
|
|
) -> None:
|
|
|
|
|
# LOGIC-01/02: an UNRECOGNIZED decision must NOT silently FAIL the task (the
|
|
|
|
|
# old behavior). The graph backstop maps it to request_changes — it loops
|
|
|
|
|
# back to the planner (and, with a reviewer that never approves, re-suspends
|
|
|
|
|
# on the gate), NEVER terminal FAILED, and NEVER an accidental approve.
|
|
|
|
|
from agent_team.graph import pending_question
|
|
|
|
|
|
2026-06-23 20:11:32 -04:00
|
|
|
graph = _plan_gate_graph("VERDICT: REQUEST CHANGES\nnope")
|
|
|
|
|
thread_id, _ = _drive_to_plan_gate(graph)
|
|
|
|
|
final = resume_task(graph, thread_id=thread_id, answer={"decision": "huh?"})
|
|
|
|
|
|
fix(agent-team): centralize safe decision mapping + remove free-text abandon hair-trigger
Security-review follow-up (LOGIC-01/02/05, all confirmed correctness).
- New transport-neutral `decisions.normalize_decision(raw, *, allow_abandon)` is
the single source of truth: approve-allowlist→approve; abandon-allowlist→abandon
ONLY when allow_abandon; everything else (prose, empty, abandon-verbs when
disallowed) → request_changes with the full reply as notes; idempotent on an
already-formed decision dict. slack_adapter.map_plan_decision is now a thin
wrapper (default allow_abandon=True, no caller churn).
- LOGIC-01/02: graph._parse_decision now delegates to normalize_decision (was:
any unrecognized verb → abandon → FAILED). The graph is now the universal safe
backstop, so EVERY writer that bypassed the listener mapping — operator CLI
answer_on_behalf (raw), Coordinator.submit_answer (raw), the recovery sweep —
loops back on prose instead of silently FAILing the task. Explicit abandon
still abandons (preserves the confirmed-button path).
- LOGIC-05: the Slack FREE-TEXT reply path maps with allow_abandon=False, so a
bare "cancel"/"stop"/"abandon" typed in-thread → request_changes (never
terminal abandon); abandon stays reachable only via the confirm-guarded button.
Tests: graph unrecognized→loops-back (not FAILED), operator raw-prose→request_
changes, free-text destructive verbs→request_changes vs button→abandon,
normalizer idempotency. 1484 passed.
2026-06-24 11:48:04 -04:00
|
|
|
assert final["status"] != TaskStatus.FAILED.value
|
|
|
|
|
# request_changes loops plan->review and re-suspends on the gate (reviewer
|
|
|
|
|
# never approves): the task is still pending a human decision, not terminal.
|
|
|
|
|
assert pending_question(graph, thread_id=thread_id) is not None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_plan_gate_operator_raw_prose_loops_back_not_failed(
|
|
|
|
|
restore_review_invoker,
|
|
|
|
|
) -> None:
|
|
|
|
|
# LOGIC-01/02 (operator-CLI-style RAW write): a non-listener writer (operator
|
|
|
|
|
# CLI / Coordinator.submit_answer) stores RAW change-request prose on a
|
|
|
|
|
# plan_decision row — it never passes through the Slack listener's pre-map.
|
|
|
|
|
# The graph backstop (_parse_decision via normalize_decision) maps it to
|
|
|
|
|
# request_changes carrying the full prose as notes, so it loops back to the
|
|
|
|
|
# planner rather than the old silent abandon -> terminal FAILED.
|
|
|
|
|
from agent_team.graph import pending_question
|
|
|
|
|
from agent_team.nodes.planner import _format_review_feedback
|
|
|
|
|
|
|
|
|
|
captured: dict[str, object] = {}
|
|
|
|
|
|
|
|
|
|
def capturing_plan(state):
|
|
|
|
|
captured["feedback"] = _format_review_feedback(
|
|
|
|
|
list(state.get("review_verdicts") or [])
|
|
|
|
|
)
|
|
|
|
|
return _p2_plan_stub(state)
|
|
|
|
|
|
|
|
|
|
from agent_team.nodes import review_loop
|
|
|
|
|
|
|
|
|
|
review_loop.set_review_invoker(lambda prompt, **kw: "VERDICT: REQUEST CHANGES\nno")
|
|
|
|
|
graph = build_graph(
|
|
|
|
|
checkpointer=_Saver(),
|
|
|
|
|
live_plan_node=capturing_plan,
|
|
|
|
|
review_node=review_loop.bind_review_node(),
|
|
|
|
|
route_review=review_loop.route_after_review,
|
|
|
|
|
plan_gate=True,
|
|
|
|
|
)
|
|
|
|
|
thread_id, _ = _drive_to_plan_gate(graph)
|
|
|
|
|
|
|
|
|
|
# A bare RAW string (exactly what answer_question stores for an operator who
|
|
|
|
|
# typed prose) resumes the gate. It must NOT fail the task.
|
|
|
|
|
final = resume_task(
|
|
|
|
|
graph, thread_id=thread_id, answer="please use pytest fixtures instead"
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
assert final["status"] != TaskStatus.FAILED.value
|
|
|
|
|
assert pending_question(graph, thread_id=thread_id) is not None
|
|
|
|
|
# The raw prose was folded into the re-plan as request_changes notes.
|
|
|
|
|
assert "pytest fixtures" in str(captured.get("feedback", ""))
|
2026-06-23 20:11:32 -04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_plan_gate_request_changes_terminates_at_ceiling(
|
|
|
|
|
restore_review_invoker,
|
|
|
|
|
) -> None:
|
|
|
|
|
# Repeated request_changes resumes must eventually hit the gate ceiling and
|
|
|
|
|
# go terminal PARKED ("ceiling reached") — NOT loop unbounded. The reviewer
|
|
|
|
|
# NEVER approves, so every gate visit is request_changes until the ceiling.
|
|
|
|
|
from agent_team.graph import MAX_PLAN_GATE_VISITS, pending_question
|
|
|
|
|
|
|
|
|
|
graph = _plan_gate_graph("VERDICT: REQUEST CHANGES\nstill not ready")
|
|
|
|
|
thread_id, _ = _drive_to_plan_gate(graph)
|
|
|
|
|
|
|
|
|
|
# Each request_changes consumes exactly one gate visit; bound the loop well
|
|
|
|
|
# above the ceiling to prove it terminates on its own, not by our cap.
|
|
|
|
|
for _ in range(MAX_PLAN_GATE_VISITS + 5):
|
|
|
|
|
if pending_question(graph, thread_id=thread_id) is None:
|
|
|
|
|
break
|
|
|
|
|
resume_task(
|
|
|
|
|
graph,
|
|
|
|
|
thread_id=thread_id,
|
|
|
|
|
answer={"decision": "request_changes", "notes": "again"},
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
state = get_pipeline_state(graph, thread_id=thread_id)
|
|
|
|
|
assert pending_question(graph, thread_id=thread_id) is None
|
|
|
|
|
assert state["status"] == TaskStatus.PARKED.value
|
|
|
|
|
assert "ceiling" in (state.get("failure_reason") or "")
|
|
|
|
|
assert state.get("plan_gate_visits") == MAX_PLAN_GATE_VISITS
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_auto_approved_plan_does_not_hit_gate(restore_review_invoker) -> None:
|
|
|
|
|
# No regression: an auto-APPROVED plan settles WITHOUT visiting the gate even
|
|
|
|
|
# when the gate is wired (gate only catches the review-cap dead-end).
|
|
|
|
|
from agent_team.graph import pending_question
|
|
|
|
|
|
|
|
|
|
graph = _plan_gate_graph("VERDICT: APPROVE\nlooks solid")
|
|
|
|
|
thread_id, _ = start_task(graph, transport="slack")
|
|
|
|
|
final = resume_task(graph, thread_id=thread_id, answer="scope is X")
|
|
|
|
|
|
|
|
|
|
assert final["current_phase"] == Phase.BUILD.value
|
|
|
|
|
assert pending_question(graph, thread_id=thread_id) is None
|
|
|
|
|
assert final.get("plan_gate_visits", 0) == 0
|
|
|
|
|
|
|
|
|
|
|
2026-06-18 13:23:03 -04:00
|
|
|
# --- P3 build -> verify subgraph wiring (opt-in). ---------------------------
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _p3_plan_stub(state: PipelineState) -> PipelineState:
|
|
|
|
|
"""P2/P3 planner stub: emit an APPROVED, scoped plan and advance to REVIEW.
|
|
|
|
|
|
|
|
|
|
Like ``_p2_plan_stub`` but carries a ``scope`` so the P3 BUILD node's
|
|
|
|
|
trust-control-surface scan accepts the candidate diff, letting the
|
|
|
|
|
build -> verify topology be driven end to end.
|
|
|
|
|
"""
|
|
|
|
|
revisions = len(state.get("review_verdicts") or [])
|
|
|
|
|
return PipelineState(
|
|
|
|
|
plan={
|
|
|
|
|
"title": "do it",
|
|
|
|
|
"scope": ["src"],
|
|
|
|
|
"phases": ["P1"],
|
|
|
|
|
"revision": revisions,
|
|
|
|
|
},
|
|
|
|
|
current_phase=Phase.REVIEW.value,
|
|
|
|
|
status=TaskStatus.ACTIVE.value,
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _p3_diff() -> str:
|
|
|
|
|
"""A minimal in-scope unified diff the fake builder returns."""
|
|
|
|
|
return "diff --git a/src/foo.py b/src/foo.py\n@@ -1 +1 @@\n-old\n+new\n"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _p3_graph(review_text: str, *, ci_result_fetcher):
|
|
|
|
|
"""Compile a P3 graph: review -> build -> verify with injected seams.
|
|
|
|
|
|
|
|
|
|
The diff builder is a fixed in-scope diff; the CI-result fetcher is injected
|
|
|
|
|
so the test drives the verifier verdict (pass / fail / none) deterministically
|
|
|
|
|
with no live CI.
|
|
|
|
|
"""
|
|
|
|
|
from agent_team.nodes import review_loop
|
|
|
|
|
from agent_team.nodes.build_verify_subgraph import (
|
|
|
|
|
make_build_node,
|
|
|
|
|
make_verify_node,
|
|
|
|
|
route_after_verify,
|
|
|
|
|
)
|
|
|
|
|
from agent_team.nodes.verifier import VerifierConfig
|
|
|
|
|
|
|
|
|
|
review_loop.set_review_invoker(lambda prompt, **kw: review_text)
|
|
|
|
|
|
|
|
|
|
def fake_builder(*, plan, config):
|
|
|
|
|
return _p3_diff()
|
|
|
|
|
|
|
|
|
|
build_node = make_build_node(diff_builder=fake_builder)
|
|
|
|
|
verify_node = make_verify_node(
|
|
|
|
|
VerifierConfig(expected_run_id="r1", allowed_scope=["src"]),
|
|
|
|
|
ci_result_fetcher=ci_result_fetcher,
|
|
|
|
|
)
|
|
|
|
|
return build_graph(
|
|
|
|
|
checkpointer=_Saver(),
|
|
|
|
|
live_plan_node=_p3_plan_stub,
|
|
|
|
|
review_node=review_loop.bind_review_node(),
|
|
|
|
|
route_review=review_loop.route_after_review,
|
|
|
|
|
build_verify=(build_node, verify_node, route_after_verify),
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_build_graph_build_verify_requires_review_node() -> None:
|
|
|
|
|
"""build_verify without review_node is a wiring error (no 'build' route)."""
|
|
|
|
|
from agent_team.nodes.build_verify_subgraph import (
|
|
|
|
|
make_build_node,
|
|
|
|
|
make_verify_node,
|
|
|
|
|
route_after_verify,
|
|
|
|
|
)
|
|
|
|
|
from agent_team.nodes.verifier import VerifierConfig
|
|
|
|
|
|
|
|
|
|
tuple_ = (
|
|
|
|
|
make_build_node(diff_builder=None),
|
|
|
|
|
make_verify_node(VerifierConfig(expected_run_id="")),
|
|
|
|
|
route_after_verify,
|
|
|
|
|
)
|
|
|
|
|
with pytest.raises(ValueError, match="build_verify"):
|
|
|
|
|
build_graph(build_verify=tuple_)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_p3_graph_authenticated_pass_routes_to_done(restore_review_invoker) -> None:
|
|
|
|
|
"""review(APPROVE) -> build -> verify(PASS via fake CI) -> DONE (PR terminus)."""
|
|
|
|
|
|
|
|
|
|
def pass_fetcher(state):
|
|
|
|
|
from agent_team.state_store import compute_content_hash
|
|
|
|
|
|
|
|
|
|
diff_hash = compute_content_hash(_p3_diff().encode("utf-8"))
|
|
|
|
|
return {"run_id": "r1", "conclusion": "success", "diff_hash": diff_hash}
|
|
|
|
|
|
|
|
|
|
graph = _p3_graph("VERDICT: APPROVE\nlooks solid", ci_result_fetcher=pass_fetcher)
|
|
|
|
|
thread_id, _ = start_task(graph, transport="slack")
|
|
|
|
|
final = resume_task(graph, thread_id=thread_id, answer="scope is X")
|
|
|
|
|
|
|
|
|
|
# The authenticated CI pass cleared the gate -> DONE terminus.
|
|
|
|
|
assert final["current_phase"] == Phase.DONE.value
|
|
|
|
|
assert final["status"] == TaskStatus.DONE.value
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_p3_graph_inert_default_parks_at_verify(restore_review_invoker) -> None:
|
|
|
|
|
"""review(APPROVE) -> build -> verify(no CI result) -> BLOCK -> PARKED.
|
|
|
|
|
|
|
|
|
|
With the INERT default (no authenticated CI result) the gate can never
|
|
|
|
|
fabricate a pass, so an approved plan still parks at VERIFY. This is the
|
|
|
|
|
production-safe behavior the opt-in subgraph ships with.
|
|
|
|
|
"""
|
|
|
|
|
graph = _p3_graph(
|
|
|
|
|
"VERDICT: APPROVE\nlooks solid", ci_result_fetcher=lambda state: None
|
|
|
|
|
)
|
|
|
|
|
thread_id, _ = start_task(graph, transport="slack")
|
|
|
|
|
final = resume_task(graph, thread_id=thread_id, answer="scope is X")
|
|
|
|
|
|
|
|
|
|
assert final["current_phase"] == Phase.PARKED.value
|
|
|
|
|
assert final["status"] == TaskStatus.PARKED.value
|
|
|
|
|
|
|
|
|
|
|
2026-06-23 17:51:58 -04:00
|
|
|
def _p3_graph_with_dispatch(review_text: str, *, ci_result_fetcher, dispatch_node):
|
|
|
|
|
"""Compile a P3+ graph: review -> build -> DISPATCH -> verify.
|
|
|
|
|
|
|
|
|
|
Same as ``_p3_graph`` but splices a ``dispatch_node`` between BUILD and
|
|
|
|
|
VERIFY so the reordered topology (design §4 Decision 1) can be driven end to
|
|
|
|
|
end — DISPATCH writes ``state["run_id"]`` before VERIFY reads it.
|
|
|
|
|
"""
|
|
|
|
|
from agent_team.nodes import review_loop
|
|
|
|
|
from agent_team.nodes.build_verify_subgraph import (
|
|
|
|
|
make_build_node,
|
|
|
|
|
make_verify_node,
|
|
|
|
|
route_after_verify,
|
|
|
|
|
)
|
|
|
|
|
from agent_team.nodes.verifier import VerifierConfig
|
|
|
|
|
|
|
|
|
|
review_loop.set_review_invoker(lambda prompt, **kw: review_text)
|
|
|
|
|
|
|
|
|
|
def fake_builder(*, plan, config):
|
|
|
|
|
return _p3_diff()
|
|
|
|
|
|
|
|
|
|
build_node = make_build_node(diff_builder=fake_builder)
|
|
|
|
|
# No static expected_run_id: the gate must bind to the run id DISPATCH wrote.
|
|
|
|
|
verify_node = make_verify_node(
|
|
|
|
|
VerifierConfig(expected_run_id=None, allowed_scope=["src"]),
|
|
|
|
|
ci_result_fetcher=ci_result_fetcher,
|
|
|
|
|
)
|
|
|
|
|
return build_graph(
|
|
|
|
|
checkpointer=_Saver(),
|
|
|
|
|
live_plan_node=_p3_plan_stub,
|
|
|
|
|
review_node=review_loop.bind_review_node(),
|
|
|
|
|
route_review=review_loop.route_after_review,
|
|
|
|
|
build_verify=(build_node, verify_node, route_after_verify),
|
|
|
|
|
dispatch_node=dispatch_node,
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_build_graph_dispatch_node_requires_build_verify() -> None:
|
|
|
|
|
"""dispatch_node without build_verify is a wiring error (nothing to splice)."""
|
|
|
|
|
with pytest.raises(ValueError, match="dispatch_node requires build_verify"):
|
|
|
|
|
build_graph(dispatch_node=lambda state: {})
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_p3_dispatch_runs_before_verify_and_supplies_run_id(
|
|
|
|
|
restore_review_invoker,
|
|
|
|
|
) -> None:
|
|
|
|
|
"""BUILD -> DISPATCH -> VERIFY: DISPATCH captures run_id BEFORE VERIFY reads it.
|
|
|
|
|
|
|
|
|
|
The verify node is wired with NO static expected_run_id, so the only way the
|
|
|
|
|
authenticated-pass gate can bind a verdict is if DISPATCH wrote
|
|
|
|
|
``state["run_id"]`` first. The fetcher keys its conclusion to the dispatched
|
|
|
|
|
run id and asserts it observes that id — proving DISPATCH ran before VERIFY.
|
|
|
|
|
"""
|
|
|
|
|
from agent_team.state_store import compute_content_hash
|
|
|
|
|
|
|
|
|
|
observed: dict[str, object] = {}
|
|
|
|
|
dispatched_run_id = "r-dispatched-007"
|
|
|
|
|
|
|
|
|
|
def dispatch_node(state):
|
|
|
|
|
# Mirror dispatch_invoker's contract: persist the located run identity so
|
|
|
|
|
# the downstream verifier binds the gate to THIS task's dispatched run.
|
|
|
|
|
return {"run_id": dispatched_run_id, "dispatched_at": "2026-06-23T00:00:00Z"}
|
|
|
|
|
|
|
|
|
|
def pass_fetcher(state):
|
|
|
|
|
# The fetcher only sees state["run_id"] if DISPATCH already ran.
|
|
|
|
|
observed["run_id"] = state.get("run_id")
|
|
|
|
|
diff_hash = compute_content_hash(_p3_diff().encode("utf-8"))
|
|
|
|
|
return {
|
|
|
|
|
"run_id": state.get("run_id"),
|
|
|
|
|
"conclusion": "success",
|
|
|
|
|
"diff_hash": diff_hash,
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
graph = _p3_graph_with_dispatch(
|
|
|
|
|
"VERDICT: APPROVE\nlooks solid",
|
|
|
|
|
ci_result_fetcher=pass_fetcher,
|
|
|
|
|
dispatch_node=dispatch_node,
|
|
|
|
|
)
|
|
|
|
|
thread_id, _ = start_task(graph, transport="slack")
|
|
|
|
|
final = resume_task(graph, thread_id=thread_id, answer="scope is X")
|
|
|
|
|
|
|
|
|
|
# VERIFY observed the run id DISPATCH wrote -> DISPATCH ran first.
|
|
|
|
|
assert observed["run_id"] == dispatched_run_id
|
|
|
|
|
# And the per-task-bound authenticated pass cleared the gate -> DONE terminus.
|
|
|
|
|
assert final["current_phase"] == Phase.DONE.value
|
|
|
|
|
assert final["status"] == TaskStatus.DONE.value
|
|
|
|
|
assert final["run_id"] == dispatched_run_id
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_p3_dispatch_node_present_in_graph_topology(restore_review_invoker) -> None:
|
|
|
|
|
"""The DISPATCH vertex is wired between BUILD and VERIFY when injected."""
|
|
|
|
|
from agent_team.graph import BUILD_NODE, DISPATCH_NODE, VERIFY_NODE
|
|
|
|
|
|
|
|
|
|
graph = _p3_graph_with_dispatch(
|
|
|
|
|
"VERDICT: APPROVE\nlooks solid",
|
|
|
|
|
ci_result_fetcher=lambda state: None,
|
|
|
|
|
dispatch_node=lambda state: {"run_id": "r"},
|
|
|
|
|
)
|
|
|
|
|
g = graph.get_graph()
|
|
|
|
|
nodes = set(g.nodes)
|
|
|
|
|
assert {BUILD_NODE, DISPATCH_NODE, VERIFY_NODE} <= nodes
|
|
|
|
|
|
|
|
|
|
# The linear order is BUILD -> DISPATCH -> VERIFY (no direct BUILD -> VERIFY).
|
|
|
|
|
edges = {(e.source, e.target) for e in g.edges}
|
|
|
|
|
assert (BUILD_NODE, DISPATCH_NODE) in edges
|
|
|
|
|
assert (DISPATCH_NODE, VERIFY_NODE) in edges
|
|
|
|
|
assert (BUILD_NODE, VERIFY_NODE) not in edges
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_p3_no_dispatch_falls_back_to_build_then_verify(restore_review_invoker) -> None:
|
|
|
|
|
"""Without a dispatch node the order falls back to BUILD -> VERIFY directly."""
|
|
|
|
|
from agent_team.graph import BUILD_NODE, DISPATCH_NODE, VERIFY_NODE
|
|
|
|
|
|
|
|
|
|
graph = _p3_graph("VERDICT: APPROVE\nlooks solid", ci_result_fetcher=lambda s: None)
|
|
|
|
|
g = graph.get_graph()
|
|
|
|
|
nodes = set(g.nodes)
|
|
|
|
|
assert {BUILD_NODE, VERIFY_NODE} <= nodes
|
|
|
|
|
assert DISPATCH_NODE not in nodes
|
|
|
|
|
|
|
|
|
|
edges = {(e.source, e.target) for e in g.edges}
|
|
|
|
|
assert (BUILD_NODE, VERIFY_NODE) in edges
|
|
|
|
|
|
|
|
|
|
|
2026-06-18 13:23:03 -04:00
|
|
|
def test_p3_graph_route_constants_mirror_subgraph_by_value() -> None:
|
|
|
|
|
"""graph.py's P3 route ids match the subgraph module by value (no cycle)."""
|
|
|
|
|
from agent_team.nodes import build_verify_subgraph as bvs
|
|
|
|
|
|
|
|
|
|
assert graph_mod.APPROVED_ROUTE == bvs.APPROVED_ROUTE
|
|
|
|
|
assert graph_mod.BUILD_ROUTE == bvs.BUILD_ROUTE
|
|
|
|
|
assert graph_mod.PARKED_ROUTE == bvs.PARKED_ROUTE
|
|
|
|
|
|
|
|
|
|
|
Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.
Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled
KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven
Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 14:17:44 -04:00
|
|
|
# --- Module import hygiene. -------------------------------------------------
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_module_imports_without_optional_sqlite_dep() -> None:
|
|
|
|
|
# The module-level import of graph must not pull in the optional SQLite
|
|
|
|
|
# checkpointer (that import is deferred into build_sqlite_checkpointer).
|
|
|
|
|
assert hasattr(graph_mod, "build_graph")
|
|
|
|
|
assert hasattr(graph_mod, "build_sqlite_checkpointer")
|