This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team/agent_team/nodes/clarifier.py
Adam Moussa 15a416d31a Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.

Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled

KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven

Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 15:16:12 -04:00

228 lines
9.4 KiB
Python

"""Clarifier node — interrupt + 98% confidence loop (design §3.3, §7.1 P1).
The clarifier is the first reasoning stage and the **human gate** of the
Plane-2 pipeline (§3.3). It gathers context, then asks Adam *question-sets
until it is 98%+ confident*. Each question-set is delivered through a LangGraph
``interrupt()``: the graph suspends and checkpoints, the question-set is posted
over the chosen transport (§3.3.1), and the task resumes via
``Command(resume=...)`` when Adam answers. No progression to planning happens
until the clarifier clears the confidence bar (§3.3, §7.1 P1).
This module is a **leaf** built on the committed foundation contracts, which it
imports verbatim and never redefines:
* :class:`agent_team.task_model.PipelineState` — the LangGraph state schema.
* :class:`agent_team.task_model.Phase` / :class:`agent_team.task_model.TaskStatus`
— lifecycle enums written back into the state.
* :func:`agent_team.task_model.new_thread_id` — thread-id minting (intake).
* :class:`agent_team.transport.QuestionSet` — the interrupt payload.
The node owns only the *loop*: how confidence is assessed and what questions are
asked are injected as callables so this stays a pure, unit-testable control
flow with no live Claude call. The real wiring binds Claude through the
``agent_team.billing.claude_invoke`` seam in a later phase; here the seam is a
constructor argument so P1 can prove the suspend/resume mechanic
deterministically.
Key design points proven here (the §7.1 P1 "riskiest mechanic"):
* **98% loop.** The node calls :func:`langgraph.types.interrupt` repeatedly in a
``while confidence < threshold`` loop. Each resume replays the node from the
top; LangGraph returns previously-supplied resume values for already-cleared
interrupts, so the accumulated Q&A drives confidence upward deterministically.
* **Turn cap (§7.1).** A clarifier is capped at ``max_turns`` per task; on the
cap it stops asking, marks the task ``PARKED``/``Phase.PARKED`` (ALARM rather
than spin), and does not advance to planning.
* **Human gate.** Only a run that clears the bar writes ``Phase.PLAN`` +
``TaskStatus.ACTIVE``; nothing else lets the pipeline progress to build.
"""
from __future__ import annotations
import uuid
from collections.abc import Callable, Sequence
from dataclasses import dataclass
from langgraph.types import interrupt
from agent_team.task_model import Phase, PipelineState, TaskStatus
from agent_team.transport import QuestionSet
__all__ = [
"DEFAULT_CONFIDENCE_THRESHOLD",
"DEFAULT_MAX_TURNS",
"ClarifierConfig",
"ConfidenceAssessor",
"QuestionGenerator",
"build_question_set",
"make_clarifier_node",
]
# §3.3 / §7.1 P1: the clarifier must reach "98%+ confident" before the human
# gate opens. Expressed as a 0..1 fraction; the loop runs while below it.
DEFAULT_CONFIDENCE_THRESHOLD: float = 0.98
# §7.1: "a clarifier is capped at N turns per task, then" parks rather than
# spinning. A conservative default; callers override per task class.
DEFAULT_MAX_TURNS: int = 6
# A confidence assessor inspects the running Q&A history (oldest first) and the
# task state, and returns the current 0..1 confidence that the requirement is
# understood well enough to plan. Injected so the loop is testable without a
# live model; the real binding calls Claude through the billing seam.
ConfidenceAssessor = Callable[[Sequence[object], PipelineState], float]
# A question generator produces the next ordered question-set given the Q&A so
# far and the task state. Injected for the same reason.
QuestionGenerator = Callable[[Sequence[object], PipelineState], list[str]]
@dataclass(frozen=True)
class ClarifierConfig:
"""Tunables for :func:`make_clarifier_node` (§3.3, §7.1).
``confidence_threshold`` is the 98% bar the loop must clear; ``max_turns``
is the §7.1 turn cap after which the task parks instead of spinning;
``transport`` records the channel Adam chose at intake (carried into the
interrupt payload per §3.3.1) and falls back to the state's ``transport``
when empty.
"""
confidence_threshold: float = DEFAULT_CONFIDENCE_THRESHOLD
max_turns: int = DEFAULT_MAX_TURNS
transport: str = ""
def __post_init__(self) -> None:
if not 0.0 < self.confidence_threshold <= 1.0:
raise ValueError(
"confidence_threshold must be in (0, 1]; "
f"got {self.confidence_threshold!r}"
)
if self.max_turns < 1:
raise ValueError(f"max_turns must be >= 1; got {self.max_turns!r}")
def _new_question_id() -> str:
"""Mint a fresh ``question_id`` (uuid4 hex) for one question-set (§3.3.1)."""
return uuid.uuid4().hex
def build_question_set(
*,
thread_id: str,
turn: int,
questions: list[str],
context: dict[str, object] | None = None,
question_id: str | None = None,
) -> QuestionSet:
"""Build the :class:`QuestionSet` carried by one ``interrupt()`` (§3.3.1).
Mints a ``question_id`` when not supplied. ``turn`` is the monotonic turn
index within the task; ``questions`` is the ordered prompt list; ``context``
is optional rendering metadata (repo, summary) the transport adapter may
surface. The returned payload is exactly the foundation
:class:`~agent_team.transport.QuestionSet` contract — never a redefinition.
"""
return QuestionSet(
thread_id=thread_id,
question_id=question_id or _new_question_id(),
turn=turn,
questions=list(questions),
context=dict(context or {}),
)
def _interrupt_payload(
question_set: QuestionSet, *, transport: str
) -> dict[str, object]:
"""Serialize the interrupt payload (§3.3.1: ``{thread_id, question_id, ...}``).
The §3.3.1 interrupt payload carries ``{thread_id, question_id, turn,
question_set, transport, deadline}``. ``deadline`` is owned by the durable
ledger/timer seam and is filled in by the responder at delivery time, so it
is left ``None`` here; the node's contribution is the question-set and its
identity.
"""
return {
"thread_id": question_set.thread_id,
"question_id": question_set.question_id,
"turn": question_set.turn,
"question_set": question_set,
"transport": transport,
"deadline": None,
}
def make_clarifier_node(
*,
assess_confidence: ConfidenceAssessor,
generate_questions: QuestionGenerator,
config: ClarifierConfig | None = None,
) -> Callable[[PipelineState], PipelineState]:
"""Build the clarifier LangGraph node (§3.3, §7.1 P1).
Returns a node callable ``node(state) -> state-delta`` suitable for
``StateGraph(PipelineState).add_node("clarify", node)``. The node:
1. Starts from the task's existing ``qa_history`` (so a resumed run keeps
prior answers) and assesses confidence.
2. While confidence is below the threshold **and** the turn cap is not hit,
generates the next question-set and raises a LangGraph
:func:`~langgraph.types.interrupt` carrying it. The graph suspends and
checkpoints; the resume value (Adam's answer) is appended to the running
Q&A history and confidence is re-assessed. Each resume replays the node
from the top, so the loop is durable across crashes and restarts
(§7.1 P1 "resume after restart").
3. On clearing the bar, advances the task to ``Phase.PLAN`` /
``TaskStatus.ACTIVE`` — the **human gate opens** (§3.3).
4. On hitting the turn cap first, parks the task (``Phase.PARKED`` /
``TaskStatus.PARKED``) instead of spinning (§7.1); it never advances to
planning, so the pipeline still cannot build.
The returned delta writes only the keys it owns (``qa_history``,
``current_phase``, ``status``) — ``PipelineState`` is ``total=False``, so a
partial write is the intended per-node checkpoint transition.
"""
cfg = config or ClarifierConfig()
def clarifier_node(state: PipelineState) -> PipelineState:
thread_id = state.get("thread_id", "")
transport = cfg.transport or state.get("transport", "")
qa_history: list[object] = list(state.get("qa_history", []))
# Confidence assessed from whatever Q&A already exists (a fresh task has
# none, so a context-only assessor may still clear or fall short).
confidence = assess_confidence(qa_history, state)
turn = 0
while confidence < cfg.confidence_threshold and turn < cfg.max_turns:
questions = generate_questions(qa_history, state)
question_set = build_question_set(
thread_id=thread_id,
turn=turn,
questions=questions,
)
# Suspend + checkpoint; resume value is Adam's answer for this turn.
answer = interrupt(_interrupt_payload(question_set, transport=transport))
qa_history.append(answer)
confidence = assess_confidence(qa_history, state)
turn += 1
if confidence >= cfg.confidence_threshold:
# Human gate clears: advance to planning.
next_phase = Phase.PLAN
next_status = TaskStatus.ACTIVE
else:
# Turn cap hit without clearing the bar: park + ALARM, never spin,
# never advance to build (§7.1).
next_phase = Phase.PARKED
next_status = TaskStatus.PARKED
return {
"qa_history": qa_history,
"current_phase": next_phase.value,
"status": next_status.value,
}
return clarifier_node