Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean. Built (pre-deployment scaffold only — nothing provisioned/enabled): - LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable) - nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier - §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder - transports: slack / github / claude_code adapters - ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness - ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up): - builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete) - §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency - operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap - ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.) - P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning, /sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
228 lines
9.4 KiB
Python
228 lines
9.4 KiB
Python
"""Clarifier node — interrupt + 98% confidence loop (design §3.3, §7.1 P1).
|
|
|
|
The clarifier is the first reasoning stage and the **human gate** of the
|
|
Plane-2 pipeline (§3.3). It gathers context, then asks Adam *question-sets
|
|
until it is 98%+ confident*. Each question-set is delivered through a LangGraph
|
|
``interrupt()``: the graph suspends and checkpoints, the question-set is posted
|
|
over the chosen transport (§3.3.1), and the task resumes via
|
|
``Command(resume=...)`` when Adam answers. No progression to planning happens
|
|
until the clarifier clears the confidence bar (§3.3, §7.1 P1).
|
|
|
|
This module is a **leaf** built on the committed foundation contracts, which it
|
|
imports verbatim and never redefines:
|
|
|
|
* :class:`agent_team.task_model.PipelineState` — the LangGraph state schema.
|
|
* :class:`agent_team.task_model.Phase` / :class:`agent_team.task_model.TaskStatus`
|
|
— lifecycle enums written back into the state.
|
|
* :func:`agent_team.task_model.new_thread_id` — thread-id minting (intake).
|
|
* :class:`agent_team.transport.QuestionSet` — the interrupt payload.
|
|
|
|
The node owns only the *loop*: how confidence is assessed and what questions are
|
|
asked are injected as callables so this stays a pure, unit-testable control
|
|
flow with no live Claude call. The real wiring binds Claude through the
|
|
``agent_team.billing.claude_invoke`` seam in a later phase; here the seam is a
|
|
constructor argument so P1 can prove the suspend/resume mechanic
|
|
deterministically.
|
|
|
|
Key design points proven here (the §7.1 P1 "riskiest mechanic"):
|
|
|
|
* **98% loop.** The node calls :func:`langgraph.types.interrupt` repeatedly in a
|
|
``while confidence < threshold`` loop. Each resume replays the node from the
|
|
top; LangGraph returns previously-supplied resume values for already-cleared
|
|
interrupts, so the accumulated Q&A drives confidence upward deterministically.
|
|
* **Turn cap (§7.1).** A clarifier is capped at ``max_turns`` per task; on the
|
|
cap it stops asking, marks the task ``PARKED``/``Phase.PARKED`` (ALARM rather
|
|
than spin), and does not advance to planning.
|
|
* **Human gate.** Only a run that clears the bar writes ``Phase.PLAN`` +
|
|
``TaskStatus.ACTIVE``; nothing else lets the pipeline progress to build.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import uuid
|
|
from collections.abc import Callable, Sequence
|
|
from dataclasses import dataclass
|
|
|
|
from langgraph.types import interrupt
|
|
|
|
from agent_team.task_model import Phase, PipelineState, TaskStatus
|
|
from agent_team.transport import QuestionSet
|
|
|
|
__all__ = [
|
|
"DEFAULT_CONFIDENCE_THRESHOLD",
|
|
"DEFAULT_MAX_TURNS",
|
|
"ClarifierConfig",
|
|
"ConfidenceAssessor",
|
|
"QuestionGenerator",
|
|
"build_question_set",
|
|
"make_clarifier_node",
|
|
]
|
|
|
|
# §3.3 / §7.1 P1: the clarifier must reach "98%+ confident" before the human
|
|
# gate opens. Expressed as a 0..1 fraction; the loop runs while below it.
|
|
DEFAULT_CONFIDENCE_THRESHOLD: float = 0.98
|
|
|
|
# §7.1: "a clarifier is capped at N turns per task, then" parks rather than
|
|
# spinning. A conservative default; callers override per task class.
|
|
DEFAULT_MAX_TURNS: int = 6
|
|
|
|
|
|
# A confidence assessor inspects the running Q&A history (oldest first) and the
|
|
# task state, and returns the current 0..1 confidence that the requirement is
|
|
# understood well enough to plan. Injected so the loop is testable without a
|
|
# live model; the real binding calls Claude through the billing seam.
|
|
ConfidenceAssessor = Callable[[Sequence[object], PipelineState], float]
|
|
|
|
# A question generator produces the next ordered question-set given the Q&A so
|
|
# far and the task state. Injected for the same reason.
|
|
QuestionGenerator = Callable[[Sequence[object], PipelineState], list[str]]
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class ClarifierConfig:
|
|
"""Tunables for :func:`make_clarifier_node` (§3.3, §7.1).
|
|
|
|
``confidence_threshold`` is the 98% bar the loop must clear; ``max_turns``
|
|
is the §7.1 turn cap after which the task parks instead of spinning;
|
|
``transport`` records the channel Adam chose at intake (carried into the
|
|
interrupt payload per §3.3.1) and falls back to the state's ``transport``
|
|
when empty.
|
|
"""
|
|
|
|
confidence_threshold: float = DEFAULT_CONFIDENCE_THRESHOLD
|
|
max_turns: int = DEFAULT_MAX_TURNS
|
|
transport: str = ""
|
|
|
|
def __post_init__(self) -> None:
|
|
if not 0.0 < self.confidence_threshold <= 1.0:
|
|
raise ValueError(
|
|
"confidence_threshold must be in (0, 1]; "
|
|
f"got {self.confidence_threshold!r}"
|
|
)
|
|
if self.max_turns < 1:
|
|
raise ValueError(f"max_turns must be >= 1; got {self.max_turns!r}")
|
|
|
|
|
|
def _new_question_id() -> str:
|
|
"""Mint a fresh ``question_id`` (uuid4 hex) for one question-set (§3.3.1)."""
|
|
return uuid.uuid4().hex
|
|
|
|
|
|
def build_question_set(
|
|
*,
|
|
thread_id: str,
|
|
turn: int,
|
|
questions: list[str],
|
|
context: dict[str, object] | None = None,
|
|
question_id: str | None = None,
|
|
) -> QuestionSet:
|
|
"""Build the :class:`QuestionSet` carried by one ``interrupt()`` (§3.3.1).
|
|
|
|
Mints a ``question_id`` when not supplied. ``turn`` is the monotonic turn
|
|
index within the task; ``questions`` is the ordered prompt list; ``context``
|
|
is optional rendering metadata (repo, summary) the transport adapter may
|
|
surface. The returned payload is exactly the foundation
|
|
:class:`~agent_team.transport.QuestionSet` contract — never a redefinition.
|
|
"""
|
|
return QuestionSet(
|
|
thread_id=thread_id,
|
|
question_id=question_id or _new_question_id(),
|
|
turn=turn,
|
|
questions=list(questions),
|
|
context=dict(context or {}),
|
|
)
|
|
|
|
|
|
def _interrupt_payload(
|
|
question_set: QuestionSet, *, transport: str
|
|
) -> dict[str, object]:
|
|
"""Serialize the interrupt payload (§3.3.1: ``{thread_id, question_id, ...}``).
|
|
|
|
The §3.3.1 interrupt payload carries ``{thread_id, question_id, turn,
|
|
question_set, transport, deadline}``. ``deadline`` is owned by the durable
|
|
ledger/timer seam and is filled in by the responder at delivery time, so it
|
|
is left ``None`` here; the node's contribution is the question-set and its
|
|
identity.
|
|
"""
|
|
return {
|
|
"thread_id": question_set.thread_id,
|
|
"question_id": question_set.question_id,
|
|
"turn": question_set.turn,
|
|
"question_set": question_set,
|
|
"transport": transport,
|
|
"deadline": None,
|
|
}
|
|
|
|
|
|
def make_clarifier_node(
|
|
*,
|
|
assess_confidence: ConfidenceAssessor,
|
|
generate_questions: QuestionGenerator,
|
|
config: ClarifierConfig | None = None,
|
|
) -> Callable[[PipelineState], PipelineState]:
|
|
"""Build the clarifier LangGraph node (§3.3, §7.1 P1).
|
|
|
|
Returns a node callable ``node(state) -> state-delta`` suitable for
|
|
``StateGraph(PipelineState).add_node("clarify", node)``. The node:
|
|
|
|
1. Starts from the task's existing ``qa_history`` (so a resumed run keeps
|
|
prior answers) and assesses confidence.
|
|
2. While confidence is below the threshold **and** the turn cap is not hit,
|
|
generates the next question-set and raises a LangGraph
|
|
:func:`~langgraph.types.interrupt` carrying it. The graph suspends and
|
|
checkpoints; the resume value (Adam's answer) is appended to the running
|
|
Q&A history and confidence is re-assessed. Each resume replays the node
|
|
from the top, so the loop is durable across crashes and restarts
|
|
(§7.1 P1 "resume after restart").
|
|
3. On clearing the bar, advances the task to ``Phase.PLAN`` /
|
|
``TaskStatus.ACTIVE`` — the **human gate opens** (§3.3).
|
|
4. On hitting the turn cap first, parks the task (``Phase.PARKED`` /
|
|
``TaskStatus.PARKED``) instead of spinning (§7.1); it never advances to
|
|
planning, so the pipeline still cannot build.
|
|
|
|
The returned delta writes only the keys it owns (``qa_history``,
|
|
``current_phase``, ``status``) — ``PipelineState`` is ``total=False``, so a
|
|
partial write is the intended per-node checkpoint transition.
|
|
"""
|
|
cfg = config or ClarifierConfig()
|
|
|
|
def clarifier_node(state: PipelineState) -> PipelineState:
|
|
thread_id = state.get("thread_id", "")
|
|
transport = cfg.transport or state.get("transport", "")
|
|
qa_history: list[object] = list(state.get("qa_history", []))
|
|
|
|
# Confidence assessed from whatever Q&A already exists (a fresh task has
|
|
# none, so a context-only assessor may still clear or fall short).
|
|
confidence = assess_confidence(qa_history, state)
|
|
turn = 0
|
|
|
|
while confidence < cfg.confidence_threshold and turn < cfg.max_turns:
|
|
questions = generate_questions(qa_history, state)
|
|
question_set = build_question_set(
|
|
thread_id=thread_id,
|
|
turn=turn,
|
|
questions=questions,
|
|
)
|
|
# Suspend + checkpoint; resume value is Adam's answer for this turn.
|
|
answer = interrupt(_interrupt_payload(question_set, transport=transport))
|
|
qa_history.append(answer)
|
|
confidence = assess_confidence(qa_history, state)
|
|
turn += 1
|
|
|
|
if confidence >= cfg.confidence_threshold:
|
|
# Human gate clears: advance to planning.
|
|
next_phase = Phase.PLAN
|
|
next_status = TaskStatus.ACTIVE
|
|
else:
|
|
# Turn cap hit without clearing the bar: park + ALARM, never spin,
|
|
# never advance to build (§7.1).
|
|
next_phase = Phase.PARKED
|
|
next_status = TaskStatus.PARKED
|
|
|
|
return {
|
|
"qa_history": qa_history,
|
|
"current_phase": next_phase.value,
|
|
"status": next_status.value,
|
|
}
|
|
|
|
return clarifier_node
|