fix(agent-team): a crashing pipeline node fails one task, not the whole daemon #53
No reviewers
Labels
No labels
app
bug
ci
compliance
content
dependencies
docs
documentation
duplicate
enhancement
github_actions
good first issue
help wanted
infra
invalid
javascript
needs-triage
python
question
tests
wontfix
No milestone
No project
No assignees
1 participant
Due date
No due date set.
Dependencies
No dependencies set.
Reference: adam/orchestrator#53
Loading…
Add table
Reference in a new issue
No description provided.
Delete branch "fix/coordinator-crash-on-node-exception"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Summary
A live task on the R720 (
d30b697c) crashed the coordinator daemon: the planner's single-shot Claude call raisedReached maximum number of turns (1), the exception propagated out ofdrain_resumes→ the serve loop, and systemd restarted the process — silently, with no Slack notification. A single task's planning failure took down the pipeline for every task, and nobody was told.Changes (defense in depth)
_collect_subscription_texttolerates the single-shot turn cap: on aReached maximum number of turnsSDK error it salvages the assistant text already collected (the JSON the planner needs) instead of propagating. Empty salvage or any non-turn error still raises.drain_resumeswraps eachresume(); an unhandled node exception now fails that task instead of the daemon: it supersedes the answered question (so the startup recovery sweep can't re-drive it into the same crash on reboot — no crash loop), marks the taskFAILEDwith a shortfailure_reason, and emits an honest❌ FAILEDline to Slack.failure_reasontoPipelineState+TaskRecord(kept in sync) so the failure detail persists as a real graph channel.ResumeOutcome.FAILED.Validation
ruff check+ruff format --checkclean.failed, ledger backed up).Tests
drain_resumesfails-not-crashes (task FAILED + question superseded + Slack notified), and one failing task does not block others in the same drain.Notes
The R720 already runs this fix's prior commit (which lowered
max_turnsto 1 — the change that exposed this crash), so it should be redeployed via/sh-deploy-r720once this merges. Not yet deployed.