Plane-2 builder calls Claude single-shot (max_turns=1) to generate a code diff — every task fails at the build phase #60

Closed
opened 2026-06-24 16:34:28 +00:00 by amoussa1229 · 0 comments
amoussa1229 commented 2026-06-24 16:34:28 +00:00 (Migrated from github.com)

Summary

The Plane-2 builder node invokes Claude single-shot (max_turns=1, no tools) to generate a unified code diff. Diff generation is an inherently agentic task (it needs multiple turns and file tools), so the call immediately hits Reached maximum number of turns (1) and the task fails at the build phase. This blocks every task from completing end-to-end (plan → review → build → verify → PR).

Discovered live on the R720 (sh-secrev) during a plan-review-gate smoke test on 2026-06-24.

Root cause

agent-team/agent_team/nodes/builders.py → default_diff_builder (~line 188):

prompt = _render_build_prompt(plan)
result: ClaudeResult = claude_invoke(prompt, config=config)   # no max_turns / allowed_tools
return result.text

claude_invoke (via invoker.subscription_invoker) defaults to _DEFAULT_MAX_TURNS = 1 and allowed_tools=[] (the deliberate single-shot config for reasoning→JSON nodes like the clarifier/planner). For the builder that default is wrong: producing a real code diff needs many turns and, almost certainly, file/edit tools — it cannot complete in one tool-less turn.

Evidence (live)

  • Smoke task 6df154edaafa47aeb93782997d5cd27a reached an approved plan (planner + 2 review rounds) and advanced to build, then:
resume of task 6df154ed… raised in-graph; failing the task
(was: Claude Code returned an error result: Reached maximum number of turns (1))
  • Traceback (abridged) — the failing call is the builder, not the planner:
graph.py … build_verify_subgraph.py:165 node
  → builders.py:600 builders_node
  → builders.py:543 build_candidate_diff
  → builders.py:188 default_diff_builder
  → billing.py:143 claude_invoke
  → invoker.py subscription_invoker
  → invoker.py:127 _collect_subscription_text  →  "Reached maximum number of turns (1)"
  • Final task state: status=failed, phase=build, failure_reason="…Reached maximum number of turns".
  • The coordinator's crash-isolation (#53) correctly caught it: the daemon stayed healthy (active, NRestarts=0) and only that one task failed.

Impact

  • No task can finish the pipeline once the build path is wired (plan-approve → build → fails). High severity for the P3 build/verify/draft-PR flow.
  • Surfaces as a ❌ FAILED task at phase=build with the max-turns message.

Why this is its own bug (vs the planner flake)

This is the same class of "single-shot config on a call that needs more" as the planner issue fixed in PR #58, but in a different node and with a different fix: the planner just needed a little turn headroom (max_turns=4, tools still off, because it only emits JSON). The builder must actually write code, so it needs a genuinely agentic config — a higher max_turns and an appropriate allowed_tools set (file read/edit) — not merely max_turns=4. This is a P3 build-node design decision.

Proposed fix

  1. Give default_diff_builder an agentic invoker config: a sufficient max_turns and an allowed_tools allowlist for the diff-generation strategy (or route it through the DeepSeek-edit path the builder docstring references, with appropriate turns/tools).
  2. Decide the diff-generation contract (single unified-diff response vs tool-driven edits) and set max_turns/allowed_tools to match.
  3. Add a hermetic test that the builder invoker is configured non-single-shot (mirrors the planner max_turns passthrough test added in PR #58), so this can't silently regress to the default.
  4. Consider a guard so any future node that needs >1 turn must opt in explicitly, rather than silently inheriting the single-shot default.

Repro

On the R720 (post-#58, with the build path wired): /new-task <anything> in #agent-team, answer the clarifier, let the plan get approved → the task fails at build with Reached maximum number of turns (1).

  • PR #58 (planner reliability + plan-review gate) — fixed the same flake class in the planner; its crash-isolation is what made this fail gracefully instead of crashing the daemon.
  • Secondary observation from the same smoke (file separately if confirmed): the ❌ FAILED lifecycle milestone did not post to Slack for this build-stage failure — worth verifying the _post_resume_followups FAILED path covers failures that originate deep in the build subgraph.
## Summary The Plane-2 **builder** node invokes Claude **single-shot** (`max_turns=1`, no tools) to generate a unified code diff. Diff generation is an inherently agentic task (it needs multiple turns and file tools), so the call immediately hits `Reached maximum number of turns (1)` and the task fails at the `build` phase. This blocks every task from completing end-to-end (plan → review → **build** → verify → PR). Discovered live on the R720 (`sh-secrev`) during a plan-review-gate smoke test on 2026-06-24. ## Root cause `agent-team/agent_team/nodes/builders.py` → `default_diff_builder` (~line 188): ```python prompt = _render_build_prompt(plan) result: ClaudeResult = claude_invoke(prompt, config=config) # no max_turns / allowed_tools return result.text ``` `claude_invoke` (via `invoker.subscription_invoker`) defaults to `_DEFAULT_MAX_TURNS = 1` and `allowed_tools=[]` (the deliberate single-shot config for reasoning→JSON nodes like the clarifier/planner). For the builder that default is wrong: producing a real code diff needs many turns and, almost certainly, file/edit tools — it cannot complete in one tool-less turn. ## Evidence (live) - Smoke task `6df154edaafa47aeb93782997d5cd27a` reached an **approved plan** (planner + 2 review rounds) and advanced to `build`, then: ``` resume of task 6df154ed… raised in-graph; failing the task (was: Claude Code returned an error result: Reached maximum number of turns (1)) ``` - Traceback (abridged) — the failing call is the builder, not the planner: ``` graph.py … build_verify_subgraph.py:165 node → builders.py:600 builders_node → builders.py:543 build_candidate_diff → builders.py:188 default_diff_builder → billing.py:143 claude_invoke → invoker.py subscription_invoker → invoker.py:127 _collect_subscription_text → "Reached maximum number of turns (1)" ``` - Final task state: `status=failed`, `phase=build`, `failure_reason="…Reached maximum number of turns"`. - The coordinator's crash-isolation (#53) correctly caught it: the daemon stayed healthy (`active`, `NRestarts=0`) and only that one task failed. ## Impact - **No task can finish the pipeline** once the build path is wired (plan-approve → build → fails). High severity for the P3 build/verify/draft-PR flow. - Surfaces as a `❌ FAILED` task at `phase=build` with the max-turns message. ## Why this is its own bug (vs the planner flake) This is the **same class** of "single-shot config on a call that needs more" as the planner issue fixed in PR #58, but in a **different node** and with a **different fix**: the planner just needed a little turn headroom (`max_turns=4`, tools still off, because it only emits JSON). The **builder must actually write code**, so it needs a genuinely agentic config — a higher `max_turns` **and** an appropriate `allowed_tools` set (file read/edit) — not merely `max_turns=4`. This is a P3 build-node design decision. ## Proposed fix 1. Give `default_diff_builder` an agentic invoker config: a sufficient `max_turns` and an `allowed_tools` allowlist for the diff-generation strategy (or route it through the DeepSeek-edit path the builder docstring references, with appropriate turns/tools). 2. Decide the diff-generation contract (single unified-diff response vs tool-driven edits) and set `max_turns`/`allowed_tools` to match. 3. Add a hermetic test that the builder invoker is configured non-single-shot (mirrors the planner `max_turns` passthrough test added in PR #58), so this can't silently regress to the default. 4. Consider a guard so any *future* node that needs >1 turn must opt in explicitly, rather than silently inheriting the single-shot default. ## Repro On the R720 (post-#58, with the build path wired): `/new-task <anything>` in `#agent-team`, answer the clarifier, let the plan get approved → the task fails at `build` with `Reached maximum number of turns (1)`. ## Related - PR #58 (planner reliability + plan-review gate) — fixed the same flake class in the planner; its crash-isolation is what made this fail gracefully instead of crashing the daemon. - Secondary observation from the same smoke (file separately if confirmed): the `❌ FAILED` lifecycle milestone did **not** post to Slack for this build-stage failure — worth verifying the `_post_resume_followups` FAILED path covers failures that originate deep in the build subgraph.
This repo is archived. You cannot comment on issues.
No description provided.