Add a CloudFormation-managed EC2 status-check alarm for the live syslog
instance, replacing the orphaned EC2-StatusCheck-syslog-server alarm that
still points at the terminated i-0a7470914f0151b97.
The instance is a CFN resource in this stack, so both alarms dimension on
instance.instanceId (Ref) rather than a literal id — they follow the
instance across future replacements (e.g. userDataCausesReplacement).
- EC2-StatusCheck-syslog-server: combined StatusCheckFailed, Maximum >= 1,
300s period, 2 eval periods, treatMissingData=breaching, SNS -> site-alerts.
Mirrors the existing Syslog-NoIncomingLogs SNS reference.
- EC2-StatusCheckSystem-syslog-server-recover: StatusCheckFailed_System with
an EC2 recover action (+ SNS). AWS only allows RECOVER on the _System
metric, not the combined metric, so it is a separate alarm;
treatMissingData=notBreaching per AWS recovery-alarm guidance.
Validated with tsc, cdk synth, cfn-lint (W2001 bootstrap warning only).
The pre-existing orphaned alarm is deleted separately, not here.
The collector's remote-syslog spool had no rotation, so each gateway's
/var/log/remote/<host>/<host>.log grew unbounded. Low risk at the old
~109 events/day, but the gateways now forward ~60k/day. CloudWatch (90d)
is the system of record; the local files are only a CW-agent spool, so
keep a short 7-day compressed window. copytruncate keeps rsyslog's open
dynaFile handles valid (truncate in place).
Applied live already; this codifies it so an instance replacement keeps it
(mirrors the existing netflow-retention timer). Deploying this user-data
change forces an instance replacement (userDataCausesReplacement) — the EIP
re-associates and the forwarding target is unchanged, so do it in a window.
- Build nfdump 1.6.23 from source in user-data (rrdtool-devel for librrd;
not packaged on AL2023) and run nfcapd collectors as systemd units:
nfcapd.service (Ronkonkoma udp/2055), nfcapd-locust.service (Locust udp/2056),
+ netflow-retention.timer (30d sweep). 1 GiB swapfile for build headroom.
- userDataCausesReplacement: true — a user-data change must actually re-run,
so force instance replacement (stateless box, EIP re-associates).
Validated: build + all services active, listeners on 514/2055/2056.
CDK stack for the EC2 syslog collector (rsyslog 514 -> CloudWatch agent ->
unifi-syslog), mirroring the file-share/forgejo pattern. Recreated from the
captured console config; EIP 184.72.154.32 imported + re-associated so the
UniFi forwarding target is unchanged. Deployed + verified 2026-06-09.
Note: deploy role can assume cdk-hnb659fds-* (account-admin via CDK
bootstrap) — same exposure as every org CDK deploy role; per-app qualifier
is a known org-wide follow-up.