mirror of
https://github.com/Sea-Haven-Industries/forgejo.git
synced 2026-09-30 04:13:11 +00:00
Fix daily flapping of backup-verification errors alarm (#23)
The forgejo-backup-verification Lambda runs once per day, so its Errors metric has data for only one hour and is missing for the other ~23h. The errors alarm used TreatMissingData=BREACHING, which treated those 23h of missing data as a breach and flipped the alarm OK->ALARM every day around 11:01 UTC despite zero actual errors. Changes (alarm-only, no instance changes): - ErrorAlarm: TreatMissingData BREACHING -> NOT_BREACHING. No data now means "no errors = healthy" instead of a false breach. - Add forgejo-backup-verification-not-running: Invocations Sum over a 24h period, alarms when < 1 invocation. This is the real "the daily verification never ran" guard that the BREACHING setting was trying (incorrectly) to provide. Both alarms keep the existing action wiring (no SNS/OK actions), per the org convention of never notifying on recovery.
This commit is contained in:
parent
f6b105a9fd
commit
a7536f0bd7
1 changed files with 23 additions and 1 deletions
|
|
@ -99,16 +99,38 @@ export class BackupVerification extends Construct {
|
|||
],
|
||||
});
|
||||
|
||||
// Errors alarm: only fires when the function actually runs and errors.
|
||||
// The function runs once daily, so for ~23h there is no data. Treating
|
||||
// missing data as BREACHING flipped this alarm OK->ALARM every day around
|
||||
// 11:01 UTC even though no error ever occurred. NOT_BREACHING means "no
|
||||
// data = no errors = healthy"; the separate not-running alarm below covers
|
||||
// the "verification never ran" case.
|
||||
new cloudwatch.Alarm(this, "ErrorAlarm", {
|
||||
alarmName: "forgejo-backup-verification-errors",
|
||||
alarmDescription: "Backup verification Lambda is failing — Slack notifications may not be firing",
|
||||
metric: fn.metricErrors({ period: cdk.Duration.hours(1) }),
|
||||
threshold: 1,
|
||||
evaluationPeriods: 1,
|
||||
treatMissingData: cloudwatch.TreatMissingData.BREACHING,
|
||||
treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
|
||||
comparisonOperator: cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
|
||||
});
|
||||
|
||||
// Not-running alarm: fires if the daily verification did not invoke at all
|
||||
// in a 24h window. This is the real "missing run" guard that the errors
|
||||
// alarm's BREACHING setting was previously (and incorrectly) providing.
|
||||
new cloudwatch.Alarm(this, "NotRunningAlarm", {
|
||||
alarmName: "forgejo-backup-verification-not-running",
|
||||
alarmDescription: "Backup verification Lambda has not run in the last 24h — daily verification may be broken",
|
||||
metric: fn.metricInvocations({
|
||||
period: cdk.Duration.hours(24),
|
||||
statistic: cloudwatch.Stats.SUM,
|
||||
}),
|
||||
threshold: 1,
|
||||
evaluationPeriods: 1,
|
||||
treatMissingData: cloudwatch.TreatMissingData.BREACHING,
|
||||
comparisonOperator: cloudwatch.ComparisonOperator.LESS_THAN_THRESHOLD,
|
||||
});
|
||||
|
||||
this.functionArn = fn.functionArn;
|
||||
}
|
||||
}
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue