This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team/agent_team
Adam Moussa f12dfecd95 fix(agent-team): a crashing pipeline node fails one task, not the whole daemon
A live task (d30b697c) on the R720 crashed the coordinator: the planner's
single-shot Claude call raised "Reached maximum number of turns (1)", the
exception propagated out of `drain_resumes` through the serve loop, and systemd
restarted the daemon — with no Slack notice, so the failure was silent.

Defense in depth:

1. invoker: `_collect_subscription_text` now tolerates the single-shot turn cap.
   When the Agent SDK raises "Reached maximum number of turns" mid-stream it
   salvages the assistant text already collected (the JSON the planner needs)
   instead of propagating. An empty salvage or any non-turn error still raises.

2. coordinator: `drain_resumes` wraps the per-job `resume()` so an unhandled
   node exception fails THAT task instead of the daemon — it supersedes the
   answered question (so the startup recovery sweep cannot re-drive it into the
   same crash on reboot), marks the task FAILED via `update_state` with a short
   `failure_reason`, and surfaces an honest "❌ FAILED" line to Slack.

3. task_model: add `failure_reason` to PipelineState + TaskRecord (kept in sync)
   so the terminal-failure detail persists as a real graph channel.

4. resume_worker: add `ResumeOutcome.FAILED`.

Tests: invoker salvage/re-raise/propagate paths; drain_resumes fails-not-crashes,
supersedes the question, notifies, and one failing task does not block others.
1189 passed.
2026-06-23 16:20:45 -04:00
..
db feat(agent-team): durable GitHub-issue intake de-dup + intake hardening (#32) 2026-06-22 17:20:43 -04:00
nodes fix(pipeline): single-shot Claude calls + planner actually reads review findings 2026-06-23 15:49:40 -04:00
transport feat(agent-team): one Slack thread per task — root "Task received" message + threaded questions/milestones 2026-06-23 15:49:40 -04:00
__init__.py Plane 2 foundation: interfaces, SQLite schemas, state-store, billing seam 2026-06-17 15:16:12 -04:00
api.py fix(ws1): harden HTTP API + declare fastapi/uvicorn deps 2026-06-23 12:27:23 -04:00
billing.py Plane 2 foundation: interfaces, SQLite schemas, state-store, billing seam 2026-06-17 15:16:12 -04:00
ci_fetcher.py feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17) 2026-06-18 15:53:26 -04:00
ci_gate.py feat(agent-team): P3-flip Phase 1 — CI trust-boundary hardening (WIP, gated) (#34) 2026-06-22 18:51:52 -04:00
coordinator.py fix(agent-team): a crashing pipeline node fails one task, not the whole daemon 2026-06-23 16:20:45 -04:00
deadline_timer.py Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI) 2026-06-17 15:16:12 -04:00
dispatcher.py feat(agent-team): wire apply/verify into .github/workflows (make it a live GitHub Actions workflow) (#37) 2026-06-22 19:12:17 -04:00
graph.py feat(agent-team): one Slack thread per task — root "Task received" message + threaded questions/milestones 2026-06-23 15:49:40 -04:00
invoker.py fix(agent-team): a crashing pipeline node fails one task, not the whole daemon 2026-06-23 16:20:45 -04:00
invoker_multi.py fix(ws1): skip fastapi TestClient tests when fastapi absent + ruff format 2026-06-23 11:40:46 -04:00
ledger.py Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI) 2026-06-17 15:16:12 -04:00
operator_cli.py Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI) 2026-06-17 15:16:12 -04:00
recovery.py Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI) 2026-06-17 15:16:12 -04:00
responder.py feat(agent-team): one Slack thread per task — root "Task received" message + threaded questions/milestones 2026-06-23 15:49:40 -04:00
resume_worker.py fix(agent-team): a crashing pipeline node fails one task, not the whole daemon 2026-06-23 16:20:45 -04:00
state_store.py Plane 2 foundation: interfaces, SQLite schemas, state-store, billing seam 2026-06-17 15:16:12 -04:00
status_page.py feat(agent-team): live visual pipeline map for the status dashboard 2026-06-23 15:53:04 -04:00
task_model.py fix(agent-team): a crashing pipeline node fails one task, not the whole daemon 2026-06-23 16:20:45 -04:00