forgejo/lib/constructs/backup-verification.ts

150 lines
5.5 KiB
TypeScript
Raw Permalink Normal View History

Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
import * as cdk from "aws-cdk-lib";
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. * Replace hardcoded instance ID in README with CloudFormation lookup The instance ID changes on every instance replacement (version upgrades, stack updates). Using a dynamic query prevents stale references and removes a manual update step from the deploy process. * Read backup S3 prefix from SSM parameter at runtime Adds /forgejo/backup-s3-prefix SSM parameter (value: archive) and updates the backup script to fetch it instead of hardcoding the prefix. Eliminates the manual post-deploy sed step. * Address cross-review findings for backup verification - Add size guard before downloading dump in restore test (3.5 GB cap) - Use paginator for list_objects_v2 in S3 checks and restore test - Remove unnecessary overrideLogicalId on GcsTransferCredentials secret - Pass explicit { mode: "daily" } to daily EventBridge rule target - Add fallback for SSM parameter fetch in backup script - Export replica bucket ARN/name from replica stack, consume via props * Add CloudWatch alarm for backup verification Lambda errors Fires on any Lambda error and on missing data (missed schedule). Catches silent failures where the Slack notification never fires. * Add .env to .gitignore Required by org CI conventions check. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
import * as cloudwatch from "aws-cdk-lib/aws-cloudwatch";
import * as cloudwatch_actions from "aws-cdk-lib/aws-cloudwatch-actions";
import * as sns from "aws-cdk-lib/aws-sns";
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
import * as events from "aws-cdk-lib/aws-events";
import * as events_targets from "aws-cdk-lib/aws-events-targets";
import * as iam from "aws-cdk-lib/aws-iam";
import * as lambda from "aws-cdk-lib/aws-lambda";
import * as logs from "aws-cdk-lib/aws-logs";
import * as s3 from "aws-cdk-lib/aws-s3";
import { PythonFunction } from "@aws-cdk/aws-lambda-python-alpha";
import { Construct } from "constructs";
interface BackupVerificationProps {
sourceBucket: s3.IBucket;
replicaBucketName: string;
gcsBucket: string;
gcsSaSecretName: string;
slackWebhookSecretName: string;
}
export class BackupVerification extends Construct {
public readonly functionArn: string;
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
constructor(scope: Construct, id: string, props: BackupVerificationProps) {
super(scope, id);
const fn = new PythonFunction(this, "Function", {
functionName: "forgejo-backup-verification",
entry: "lambda/backup-verification",
runtime: lambda.Runtime.PYTHON_3_12,
architecture: lambda.Architecture.ARM_64,
handler: "handler",
index: "app.py",
memorySize: 512,
ephemeralStorageSize: cdk.Size.gibibytes(4),
timeout: cdk.Duration.minutes(5),
environment: {
SOURCE_BUCKET: props.sourceBucket.bucketName,
REPLICA_BUCKET: props.replicaBucketName,
GCS_BUCKET: props.gcsBucket,
GCS_SA_SECRET_NAME: props.gcsSaSecretName,
SLACK_WEBHOOK_SECRET_NAME: props.slackWebhookSecretName,
},
logRetention: logs.RetentionDays.TWO_MONTHS,
});
props.sourceBucket.grantRead(fn);
fn.addToRolePolicy(
new iam.PolicyStatement({
actions: ["s3:ListBucket", "s3:GetObject"],
resources: [
`arn:aws:s3:::${props.replicaBucketName}`,
`arn:aws:s3:::${props.replicaBucketName}/*`,
],
})
);
const account = cdk.Stack.of(this).account;
const region = cdk.Stack.of(this).region;
fn.addToRolePolicy(
new iam.PolicyStatement({
actions: ["secretsmanager:GetSecretValue"],
resources: [
`arn:aws:secretsmanager:${region}:${account}:secret:${props.gcsSaSecretName}-*`,
`arn:aws:secretsmanager:${region}:${account}:secret:${props.slackWebhookSecretName}-*`,
],
})
);
fn.addToRolePolicy(
new iam.PolicyStatement({
actions: ["ec2:DescribeSnapshots"],
resources: ["*"],
})
);
new events.Rule(this, "DailyCheck", {
ruleName: "forgejo-backup-daily-check",
schedule: events.Schedule.cron({ hour: "8", minute: "0" }),
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. * Replace hardcoded instance ID in README with CloudFormation lookup The instance ID changes on every instance replacement (version upgrades, stack updates). Using a dynamic query prevents stale references and removes a manual update step from the deploy process. * Read backup S3 prefix from SSM parameter at runtime Adds /forgejo/backup-s3-prefix SSM parameter (value: archive) and updates the backup script to fetch it instead of hardcoding the prefix. Eliminates the manual post-deploy sed step. * Address cross-review findings for backup verification - Add size guard before downloading dump in restore test (3.5 GB cap) - Use paginator for list_objects_v2 in S3 checks and restore test - Remove unnecessary overrideLogicalId on GcsTransferCredentials secret - Pass explicit { mode: "daily" } to daily EventBridge rule target - Add fallback for SSM parameter fetch in backup script - Export replica bucket ARN/name from replica stack, consume via props * Add CloudWatch alarm for backup verification Lambda errors Fires on any Lambda error and on missing data (missed schedule). Catches silent failures where the Slack notification never fires. * Add .env to .gitignore Required by org CI conventions check. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
targets: [
new events_targets.LambdaFunction(fn, {
event: events.RuleTargetInput.fromObject({ mode: "daily" }),
}),
],
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
});
new events.Rule(this, "MonthlyRestoreTest", {
ruleName: "forgejo-backup-monthly-restore-test",
schedule: events.Schedule.cron({
hour: "9",
minute: "0",
day: "1",
}),
targets: [
new events_targets.LambdaFunction(fn, {
event: events.RuleTargetInput.fromObject({ mode: "restore-test" }),
}),
],
});
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. * Replace hardcoded instance ID in README with CloudFormation lookup The instance ID changes on every instance replacement (version upgrades, stack updates). Using a dynamic query prevents stale references and removes a manual update step from the deploy process. * Read backup S3 prefix from SSM parameter at runtime Adds /forgejo/backup-s3-prefix SSM parameter (value: archive) and updates the backup script to fetch it instead of hardcoding the prefix. Eliminates the manual post-deploy sed step. * Address cross-review findings for backup verification - Add size guard before downloading dump in restore test (3.5 GB cap) - Use paginator for list_objects_v2 in S3 checks and restore test - Remove unnecessary overrideLogicalId on GcsTransferCredentials secret - Pass explicit { mode: "daily" } to daily EventBridge rule target - Add fallback for SSM parameter fetch in backup script - Export replica bucket ARN/name from replica stack, consume via props * Add CloudWatch alarm for backup verification Lambda errors Fires on any Lambda error and on missing data (missed schedule). Catches silent failures where the Slack notification never fires. * Add .env to .gitignore Required by org CI conventions check. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
// Errors alarm: only fires when the function actually runs and errors.
// The function runs once daily, so for ~23h there is no data. Treating
// missing data as BREACHING flipped this alarm OK->ALARM every day around
// 11:01 UTC even though no error ever occurred. NOT_BREACHING means "no
// data = no errors = healthy"; the separate not-running alarm below covers
// the "verification never ran" case.
// Org operational alarm topic (site-alerts, CMK-encrypted — never
// alias/aws/sns, which CloudWatch cannot publish to). Alarm action only,
// no OK action, per org convention.
const alertTopic = sns.Topic.fromTopicArn(
this,
"AlertTopic",
"arn:aws:sns:us-east-1:328440206208:site-alerts"
);
const errorAlarm = new cloudwatch.Alarm(this, "ErrorAlarm", {
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. * Replace hardcoded instance ID in README with CloudFormation lookup The instance ID changes on every instance replacement (version upgrades, stack updates). Using a dynamic query prevents stale references and removes a manual update step from the deploy process. * Read backup S3 prefix from SSM parameter at runtime Adds /forgejo/backup-s3-prefix SSM parameter (value: archive) and updates the backup script to fetch it instead of hardcoding the prefix. Eliminates the manual post-deploy sed step. * Address cross-review findings for backup verification - Add size guard before downloading dump in restore test (3.5 GB cap) - Use paginator for list_objects_v2 in S3 checks and restore test - Remove unnecessary overrideLogicalId on GcsTransferCredentials secret - Pass explicit { mode: "daily" } to daily EventBridge rule target - Add fallback for SSM parameter fetch in backup script - Export replica bucket ARN/name from replica stack, consume via props * Add CloudWatch alarm for backup verification Lambda errors Fires on any Lambda error and on missing data (missed schedule). Catches silent failures where the Slack notification never fires. * Add .env to .gitignore Required by org CI conventions check. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
alarmName: "forgejo-backup-verification-errors",
alarmDescription: "Backup verification Lambda is failing — Slack notifications may not be firing",
metric: fn.metricErrors({ period: cdk.Duration.hours(1) }),
threshold: 1,
evaluationPeriods: 1,
treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. * Replace hardcoded instance ID in README with CloudFormation lookup The instance ID changes on every instance replacement (version upgrades, stack updates). Using a dynamic query prevents stale references and removes a manual update step from the deploy process. * Read backup S3 prefix from SSM parameter at runtime Adds /forgejo/backup-s3-prefix SSM parameter (value: archive) and updates the backup script to fetch it instead of hardcoding the prefix. Eliminates the manual post-deploy sed step. * Address cross-review findings for backup verification - Add size guard before downloading dump in restore test (3.5 GB cap) - Use paginator for list_objects_v2 in S3 checks and restore test - Remove unnecessary overrideLogicalId on GcsTransferCredentials secret - Pass explicit { mode: "daily" } to daily EventBridge rule target - Add fallback for SSM parameter fetch in backup script - Export replica bucket ARN/name from replica stack, consume via props * Add CloudWatch alarm for backup verification Lambda errors Fires on any Lambda error and on missing data (missed schedule). Catches silent failures where the Slack notification never fires. * Add .env to .gitignore Required by org CI conventions check. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
comparisonOperator: cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
});
errorAlarm.addAlarmAction(new cloudwatch_actions.SnsAction(alertTopic));
// Not-running alarm: fires if the daily verification did not invoke at all
// in a 24h window. This is the real "missing run" guard that the errors
// alarm's BREACHING setting was previously (and incorrectly) providing.
const notRunningAlarm = new cloudwatch.Alarm(this, "NotRunningAlarm", {
alarmName: "forgejo-backup-verification-not-running",
alarmDescription: "Backup verification Lambda has not run in the last 24h — daily verification may be broken",
metric: fn.metricInvocations({
period: cdk.Duration.hours(24),
statistic: cloudwatch.Stats.SUM,
}),
threshold: 1,
evaluationPeriods: 1,
treatMissingData: cloudwatch.TreatMissingData.BREACHING,
comparisonOperator: cloudwatch.ComparisonOperator.LESS_THAN_THRESHOLD,
});
notRunningAlarm.addAlarmAction(new cloudwatch_actions.SnsAction(alertTopic));
this.functionArn = fn.functionArn;
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
}
}