A live task (d30b697c) on the R720 crashed the coordinator: the planner's
single-shot Claude call raised "Reached maximum number of turns (1)", the
exception propagated out of `drain_resumes` through the serve loop, and systemd
restarted the daemon — with no Slack notice, so the failure was silent.
Defense in depth:
1. invoker: `_collect_subscription_text` now tolerates the single-shot turn cap.
When the Agent SDK raises "Reached maximum number of turns" mid-stream it
salvages the assistant text already collected (the JSON the planner needs)
instead of propagating. An empty salvage or any non-turn error still raises.
2. coordinator: `drain_resumes` wraps the per-job `resume()` so an unhandled
node exception fails THAT task instead of the daemon — it supersedes the
answered question (so the startup recovery sweep cannot re-drive it into the
same crash on reboot), marks the task FAILED via `update_state` with a short
`failure_reason`, and surfaces an honest "❌ FAILED" line to Slack.
3. task_model: add `failure_reason` to PipelineState + TaskRecord (kept in sync)
so the terminal-failure detail persists as a real graph channel.
4. resume_worker: add `ResumeOutcome.FAILED`.
Tests: invoker salvage/re-raise/propagate paths; drain_resumes fails-not-crashes,
supersedes the question, notifies, and one failing task does not block others.
1189 passed.
|
||
|---|---|---|
| .. | ||
| db | ||
| nodes | ||
| transport | ||
| __init__.py | ||
| api.py | ||
| billing.py | ||
| ci_fetcher.py | ||
| ci_gate.py | ||
| coordinator.py | ||
| deadline_timer.py | ||
| dispatcher.py | ||
| graph.py | ||
| invoker.py | ||
| invoker_multi.py | ||
| ledger.py | ||
| operator_cli.py | ||
| recovery.py | ||
| responder.py | ||
| resume_worker.py | ||
| state_store.py | ||
| status_page.py | ||
| task_model.py | ||