shoc-backend/docs/work-orders/phase-7/cloudwatch-alerts.md

2 KiB

Alertas Operacionais — Fase 7

Endpoint de saúde: GET /api/workorders/ops/health (Admin JWT)


Métricas expostas

Campo Fonte Uso
lastWeekRolledRunUtc WorkOrderJobRunState + ledger Job segunda-feira
lastPastDueCacheRunUtc WorkOrderJobRunState Cache opcional PastDue
syncRejectedLast24h WorkOrderAuditLogs Action=SyncRejected Conflito SHOC vs ingest
fieldLockCount WorkOrderFieldLocks Adoção edição manual
syncEnabled Sync:Enabled Estado ponte Dynamo
ingestEnabled WorkOrderIngest:Enabled Ingest direto ativo

Alertas recomendados (Elastic Beanstalk / CloudWatch)

Alerta Condição Severidade Status ops
WeekRolled stale lastWeekRolledRunUtc > 8 dias P1 [ ] criar
WeekRolled job error lastWeekRolledError não nulo P1 [ ] criar
SyncRejected spike syncRejectedLast24h > 50 P2 [ ] criar
API 5xx board ALB/Beanstalk 5xx rate > 1% em /api/workorders/board P1 [ ] criar
Ingest auth failures 401 em /api/workorders/ingest > 10/h P2 [ ] criar
Sync disabled em prod sem cutover syncEnabled=false e ingest não validado P2 [ ] criar

Runbook de criação (ops)

  1. Deploy API com WorkOrderIngest + Sync + jobs habilitados conforme ambiente.
  2. Obter token Admin e validar:
    curl -H "Authorization: Bearer <token>" https://<api-host>/api/workorders/ops/health
    
  3. Criar EventBridge/cron (ex.: a cada 15 min) que chama o health endpoint ou publica métricas customizadas a partir do JSON.
  4. Criar os alarmes da tabela acima no CloudWatch (ALB 5xx + métricas derivadas do health).
  5. Agendar checagem diária manual durante piloto e rollout.
  6. Marcar status ops na tabela quando cada alarme estiver ativo em staging e prod.

Verificação manual (ops)

# Com token Admin
curl -H "Authorization: Bearer <token>" https://<api-host>/api/workorders/ops/health