fix(agent-team): wire DeepSeek builder (#60) + plan-gate→build routing + accurate build/verify Slack status #61

Merged
amoussa1229 merged 6 commits from fix/agent-team-builder-agentic-invoker into main 2026-06-24 19:52:02 +00:00

6 commits

Author SHA1 Message Date
3ff9a43ca3 fix(agent-team): clarifier turn headroom (max_turns=4) so it can finish its JSON
The clarifier (ClaudeClarifier._turn) called claude_invoke with no max_turns,
inheriting the single-shot default (1). When the model's one turn did not
terminate in a final result the SDK raised 'Reached maximum number of turns (1)'
and, with no salvageable text, the call failed and crashed the clarify node —
leaving the task wedged at clarify with NO question posted to Slack (the human
never sees a clarifier prompt). Observed live on the R720.

Same single-shot flake the planner hit and fixed in PR #58 (_PLANNER_MAX_TURNS=4);
the clarifier never got the headroom. Give it the same: pass max_turns=4 (tools
stay off — still a fast reasoning->JSON completion).

- clarifier_llm.py: _turn passes max_turns=_CLARIFIER_MAX_TURNS (=4).
- tests: clarifier passes max_turns headroom to the invoke seam.

Full suite 1505 passed; ruff clean.
2026-06-24 15:18:28 -04:00
ad6f31c115 fix(agent-team): wire the DeepSeek mechanical-edit builder as the live diff_builder (#60 root fix)
The real #60 root cause: failsafe_production_p3_wiring never bound a diff_builder,
so the build node fell back to builders.default_diff_builder (Claude). The Claude
agentic builder (max_turns + read-only tools) then exhausted its turn cap reading
files before it could emit a diff — 'Reached maximum number of turns (8)' live.

The intended builder already exists: builders_llm.as_diff_builder() routes the
build through DeepSeek fast_coder as a SINGLE model completion (with the
orchestrator's retrieval context + defensive _extract_diff + fail-safe no-op).
A single completion has no agent turn loop, so it CANNOT exhaust turns. Confirmed
get_fast_coder() works in the daemon environment.

- coordinator.py: failsafe_production_p3_wiring (P3-configured branch) now binds
  builders_llm.as_diff_builder() as the diff_builder. Without it the build node
  silently used the turn-exhausting Claude fallback.
- builders.py: default_diff_builder reverted to a safe SINGLE-SHOT, TOOL-LESS
  fallback (max_turns=4, no allowed_tools/budget) — tools are what consumed the
  turns; this fallback is no longer the live builder. Keeps _extract_unified_diff.
- tests: failsafe binds a real (callable) diff_builder when P3 is configured;
  default_diff_builder is single-shot + tool-less.

Full suite 1504 passed; ruff clean.
2026-06-24 14:58:48 -04:00
694161a52c fix(agent-team): builder returns a real unified diff, not tool narration (#60 output contract)
The #60 max_turns/tools fix stopped the build-node crash but exposed the next
gap: with tools enabled the agentic builder reads files and its final text is
NARRATION (observed live: candidate_diff = 'Let me read the key source files to
get exact signatures bef…'), not a unified diff. That non-diff dispatched to CI,
could not be applied, and the task parked at verify.

- builders.py: default_diff_builder now extracts the unified diff from the
  response via _extract_unified_diff — prefers a fenced ```diff block, else
  slices from the first 'diff --git' header, dropping surrounding prose. A reply
  with NO diff header raises BuildError so narration fails closed (the node fails
  the task with a clear reason) instead of dispatching a bogus diff.
- builders.py: _render_build_prompt now instructs the model that its FINAL
  message must be ONLY the unified diff in a single ```diff fenced block, no
  narration before/after.
- tests: extract from fenced/bare diff with narration; reject narration-only
  (BuildError); default_diff_builder returns the clean diff from a narrated
  response and raises on a prose-only reply.

Full suite 1503 passed; ruff clean.
2026-06-24 14:39:50 -04:00
d12880ed09 fix(agent-team): accurate Slack status for build/verify states (no false 'plan could not be auto-approved')
The resume-followup notifier mislabeled build/verify-stage tasks. Two cases,
both surfaced live while smoke-testing the #60 builder fix:

1. AWAITING CI (in-progress): the VERIFY node suspends via interrupt() for the
   async CI-wait. Its interrupt payload is not question-shaped, so
   pending_question() returns None and the notifier treated the still-suspended
   task as a settled PARK -> posted '⚠️ PARKED ... the plan could not be
   auto-approved ... re-assign' for a task that had actually approved the plan,
   built a diff, and dispatched it to CI. Now detect the suspend (graph
   interrupted + dispatched run_id + a built diff) and post an honest
   'diff built and dispatched to CI (run X); awaiting verification' notice.

2. SETTLED build/verify PARK: phase inference keyed off review_verdicts and
   reported 'review' for a task that reached verify; the copy hardcoded 'the
   plan could not be auto-approved'. Now a built diff (candidate_diff/diff_hash)
   makes the inferred phase 'verify' and the copy reflect a verification/CI-gate
   failure, surfacing the CI conclusion via _summarize_build_blocker.

Plan/review escalation parks (no candidate_diff) are unchanged.

- coordinator.py: _post_resume_followups awaiting-CI branch + build-aware phase
  inference + build-aware PARKED copy; add _summarize_build_blocker.
- tests: built-park reports verify + CI conclusion (not auto-approve copy);
  awaiting-CI suspend reports in-progress, not parked.
2026-06-24 14:36:00 -04:00
e42b910a80 fix(agent-team): plan-gate approve routes into the build subgraph (not END)
Discovered while smoke-testing the #60 builder fix on the R720: a task that
reaches the human plan-gate and is APPROVED never built. The gate's
GATE_APPROVE_ROUTE was hard-wired to END in build_graph, so plan_gate_node set
phase=BUILD/status=ACTIVE and the graph terminated WITHOUT entering the build
subgraph — the task wedged at phase=build with no build, no error. Only the
reviewer's auto-approve path (review -> build_node) reached the builder; every
human-gate-approved plan silently dead-ended.

The P3 splice only repoints the REVIEW node's build route to BUILD_NODE; the
plan_gate edges are independent and were never updated, so even with P3 fully
wired the gate approve went to END. _apply_plan_decision's 'settle exactly as
an auto-approved plan' intent was broken by the edge map.

- graph.py: when build_verify (P3) is wired, point GATE_APPROVE_ROUTE at
  BUILD_NODE (mirroring REVIEW's BUILD_ROUTE: BUILD_NODE); keep END when P3 is
  inert (P2 approved-plan terminus). LangGraph resolves the forward reference
  to BUILD_NODE at compile().
- tests: gate-approve with P3 wired traverses BUILD -> VERIFY -> DONE on an
  authenticated CI pass, and parks at VERIFY (never fabricates a pass) when the
  CI fetcher is inert. Existing P2 gate-approve -> END behavior unchanged.
2026-06-24 14:13:11 -04:00
21f2fe54c1 fix(agent-team): builder uses agentic invoker config (read-only tools + turn headroom)
The Plane-2 builder (default_diff_builder) called claude_invoke with no
overrides, inheriting the subscription invoker's single-shot defaults
(max_turns=1, allowed_tools=[]). Diff synthesis is agentic, so the call
died with 'Reached maximum number of turns (1)' and every task failed at
phase=build.

- invoker.py: thread allowed_tools through subscription_invoker and
  _collect_subscription_text (default None -> []), so callers can opt in;
  single-shot reasoning nodes are unchanged.
- builders.py: default_diff_builder now passes max_turns=8, a read-only
  tool allowlist (Read/Grep/Glob), and budget_usd=4.0. No write tools --
  the builder returns the diff as data and performs no repo writes (D2/D11).
- Tests: builder agentic-config passthrough; invoker allowed_tools thread +
  tool-less default guard (so future nodes must opt in explicitly).

Closes #60
2026-06-24 13:13:28 -04:00