The clarifier (ClaudeClarifier._turn) called claude_invoke with no max_turns,
inheriting the single-shot default (1). When the model's one turn did not
terminate in a final result the SDK raised 'Reached maximum number of turns (1)'
and, with no salvageable text, the call failed and crashed the clarify node —
leaving the task wedged at clarify with NO question posted to Slack (the human
never sees a clarifier prompt). Observed live on the R720.
Same single-shot flake the planner hit and fixed in PR #58 (_PLANNER_MAX_TURNS=4);
the clarifier never got the headroom. Give it the same: pass max_turns=4 (tools
stay off — still a fast reasoning->JSON completion).
- clarifier_llm.py: _turn passes max_turns=_CLARIFIER_MAX_TURNS (=4).
- tests: clarifier passes max_turns headroom to the invoke seam.
Full suite 1505 passed; ruff clean.
The real #60 root cause: failsafe_production_p3_wiring never bound a diff_builder,
so the build node fell back to builders.default_diff_builder (Claude). The Claude
agentic builder (max_turns + read-only tools) then exhausted its turn cap reading
files before it could emit a diff — 'Reached maximum number of turns (8)' live.
The intended builder already exists: builders_llm.as_diff_builder() routes the
build through DeepSeek fast_coder as a SINGLE model completion (with the
orchestrator's retrieval context + defensive _extract_diff + fail-safe no-op).
A single completion has no agent turn loop, so it CANNOT exhaust turns. Confirmed
get_fast_coder() works in the daemon environment.
- coordinator.py: failsafe_production_p3_wiring (P3-configured branch) now binds
builders_llm.as_diff_builder() as the diff_builder. Without it the build node
silently used the turn-exhausting Claude fallback.
- builders.py: default_diff_builder reverted to a safe SINGLE-SHOT, TOOL-LESS
fallback (max_turns=4, no allowed_tools/budget) — tools are what consumed the
turns; this fallback is no longer the live builder. Keeps _extract_unified_diff.
- tests: failsafe binds a real (callable) diff_builder when P3 is configured;
default_diff_builder is single-shot + tool-less.
Full suite 1504 passed; ruff clean.
The #60 max_turns/tools fix stopped the build-node crash but exposed the next
gap: with tools enabled the agentic builder reads files and its final text is
NARRATION (observed live: candidate_diff = 'Let me read the key source files to
get exact signatures bef…'), not a unified diff. That non-diff dispatched to CI,
could not be applied, and the task parked at verify.
- builders.py: default_diff_builder now extracts the unified diff from the
response via _extract_unified_diff — prefers a fenced ```diff block, else
slices from the first 'diff --git' header, dropping surrounding prose. A reply
with NO diff header raises BuildError so narration fails closed (the node fails
the task with a clear reason) instead of dispatching a bogus diff.
- builders.py: _render_build_prompt now instructs the model that its FINAL
message must be ONLY the unified diff in a single ```diff fenced block, no
narration before/after.
- tests: extract from fenced/bare diff with narration; reject narration-only
(BuildError); default_diff_builder returns the clean diff from a narrated
response and raises on a prose-only reply.
Full suite 1503 passed; ruff clean.
The resume-followup notifier mislabeled build/verify-stage tasks. Two cases,
both surfaced live while smoke-testing the #60 builder fix:
1. AWAITING CI (in-progress): the VERIFY node suspends via interrupt() for the
async CI-wait. Its interrupt payload is not question-shaped, so
pending_question() returns None and the notifier treated the still-suspended
task as a settled PARK -> posted '⚠️ PARKED ... the plan could not be
auto-approved ... re-assign' for a task that had actually approved the plan,
built a diff, and dispatched it to CI. Now detect the suspend (graph
interrupted + dispatched run_id + a built diff) and post an honest
'diff built and dispatched to CI (run X); awaiting verification' notice.
2. SETTLED build/verify PARK: phase inference keyed off review_verdicts and
reported 'review' for a task that reached verify; the copy hardcoded 'the
plan could not be auto-approved'. Now a built diff (candidate_diff/diff_hash)
makes the inferred phase 'verify' and the copy reflect a verification/CI-gate
failure, surfacing the CI conclusion via _summarize_build_blocker.
Plan/review escalation parks (no candidate_diff) are unchanged.
- coordinator.py: _post_resume_followups awaiting-CI branch + build-aware phase
inference + build-aware PARKED copy; add _summarize_build_blocker.
- tests: built-park reports verify + CI conclusion (not auto-approve copy);
awaiting-CI suspend reports in-progress, not parked.
Discovered while smoke-testing the #60 builder fix on the R720: a task that
reaches the human plan-gate and is APPROVED never built. The gate's
GATE_APPROVE_ROUTE was hard-wired to END in build_graph, so plan_gate_node set
phase=BUILD/status=ACTIVE and the graph terminated WITHOUT entering the build
subgraph — the task wedged at phase=build with no build, no error. Only the
reviewer's auto-approve path (review -> build_node) reached the builder; every
human-gate-approved plan silently dead-ended.
The P3 splice only repoints the REVIEW node's build route to BUILD_NODE; the
plan_gate edges are independent and were never updated, so even with P3 fully
wired the gate approve went to END. _apply_plan_decision's 'settle exactly as
an auto-approved plan' intent was broken by the edge map.
- graph.py: when build_verify (P3) is wired, point GATE_APPROVE_ROUTE at
BUILD_NODE (mirroring REVIEW's BUILD_ROUTE: BUILD_NODE); keep END when P3 is
inert (P2 approved-plan terminus). LangGraph resolves the forward reference
to BUILD_NODE at compile().
- tests: gate-approve with P3 wired traverses BUILD -> VERIFY -> DONE on an
authenticated CI pass, and parks at VERIFY (never fabricates a pass) when the
CI fetcher is inert. Existing P2 gate-approve -> END behavior unchanged.
The Plane-2 builder (default_diff_builder) called claude_invoke with no
overrides, inheriting the subscription invoker's single-shot defaults
(max_turns=1, allowed_tools=[]). Diff synthesis is agentic, so the call
died with 'Reached maximum number of turns (1)' and every task failed at
phase=build.
- invoker.py: thread allowed_tools through subscription_invoker and
_collect_subscription_text (default None -> []), so callers can opt in;
single-shot reasoning nodes are unchanged.
- builders.py: default_diff_builder now passes max_turns=8, a read-only
tool allowlist (Read/Grep/Glob), and budget_usd=4.0. No write tools --
the builder returns the diff as data and performs no repo writes (D2/D11).
- Tests: builder agentic-config passthrough; invoker allowed_tools thread +
tool-less default guard (so future nodes must opt in explicitly).
Closes#60