fix(agent-team): a crashing pipeline node fails one task, not the whole daemon #53

Merged
amoussa1229 merged 1 commit from fix/coordinator-crash-on-node-exception into main 2026-06-23 20:22:53 +00:00
amoussa1229 commented 2026-06-23 20:21:17 +00:00 (Migrated from github.com)

Summary

A live task on the R720 (d30b697c) crashed the coordinator daemon: the planner's single-shot Claude call raised Reached maximum number of turns (1), the exception propagated out of drain_resumes → the serve loop, and systemd restarted the process — silently, with no Slack notification. A single task's planning failure took down the pipeline for every task, and nobody was told.

Changes (defense in depth)

  1. invoker — _collect_subscription_text tolerates the single-shot turn cap: on a Reached maximum number of turns SDK error it salvages the assistant text already collected (the JSON the planner needs) instead of propagating. Empty salvage or any non-turn error still raises.
  2. coordinator — drain_resumes wraps each resume(); an unhandled node exception now fails that task instead of the daemon: it supersedes the answered question (so the startup recovery sweep can't re-drive it into the same crash on reboot — no crash loop), marks the task FAILED with a short failure_reason, and emits an honest ❌ FAILED line to Slack.
  3. task_model — adds failure_reason to PipelineState + TaskRecord (kept in sync) so the failure detail persists as a real graph channel.
  4. resume_worker — adds ResumeOutcome.FAILED.

Validation

  • Full agent-team suite: 1189 passed; ruff check + ruff format --check clean.
  • The two stuck ledger tasks on the R720 were already closed out separately (marked failed, ledger backed up).

Tests

  • invoker: salvage-on-turn-cap, re-raise-when-no-text, non-turn-error-propagates.
  • coordinator: drain_resumes fails-not-crashes (task FAILED + question superseded + Slack notified), and one failing task does not block others in the same drain.

Notes

The R720 already runs this fix's prior commit (which lowered max_turns to 1 — the change that exposed this crash), so it should be redeployed via /sh-deploy-r720 once this merges. Not yet deployed.

## Summary A live task on the R720 (`d30b697c`) crashed the coordinator daemon: the planner's single-shot Claude call raised `Reached maximum number of turns (1)`, the exception propagated out of `drain_resumes` → the serve loop, and systemd restarted the process — **silently**, with no Slack notification. A single task's planning failure took down the pipeline for every task, and nobody was told. ## Changes (defense in depth) 1. **invoker** — `_collect_subscription_text` tolerates the single-shot turn cap: on a `Reached maximum number of turns` SDK error it salvages the assistant text already collected (the JSON the planner needs) instead of propagating. Empty salvage or any non-turn error still raises. 2. **coordinator** — `drain_resumes` wraps each `resume()`; an unhandled node exception now fails *that* task instead of the daemon: it supersedes the answered question (so the startup recovery sweep can't re-drive it into the same crash on reboot — no crash loop), marks the task `FAILED` with a short `failure_reason`, and emits an honest `❌ FAILED` line to Slack. 3. **task_model** — adds `failure_reason` to `PipelineState` + `TaskRecord` (kept in sync) so the failure detail persists as a real graph channel. 4. **resume_worker** — adds `ResumeOutcome.FAILED`. ## Validation - Full agent-team suite: **1189 passed**; `ruff check` + `ruff format --check` clean. - The two stuck ledger tasks on the R720 were already closed out separately (marked `failed`, ledger backed up). ## Tests - invoker: salvage-on-turn-cap, re-raise-when-no-text, non-turn-error-propagates. - coordinator: `drain_resumes` fails-not-crashes (task FAILED + question superseded + Slack notified), and one failing task does not block others in the same drain. ## Notes The R720 already runs this fix's *prior* commit (which lowered `max_turns` to 1 — the change that exposed this crash), so it should be redeployed via `/sh-deploy-r720` once this merges. Not yet deployed.
This repo is archived. You cannot comment on pull requests.
No description provided.