forgejo/lib/forgejo-replica-stack.ts

44 lines
1.3 KiB
TypeScript
Raw Normal View History

Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
import * as cdk from "aws-cdk-lib";
import * as s3 from "aws-cdk-lib/aws-s3";
import { Construct } from "constructs";
export class ForgejoReplicaStack extends cdk.Stack {
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. * Replace hardcoded instance ID in README with CloudFormation lookup The instance ID changes on every instance replacement (version upgrades, stack updates). Using a dynamic query prevents stale references and removes a manual update step from the deploy process. * Read backup S3 prefix from SSM parameter at runtime Adds /forgejo/backup-s3-prefix SSM parameter (value: archive) and updates the backup script to fetch it instead of hardcoding the prefix. Eliminates the manual post-deploy sed step. * Address cross-review findings for backup verification - Add size guard before downloading dump in restore test (3.5 GB cap) - Use paginator for list_objects_v2 in S3 checks and restore test - Remove unnecessary overrideLogicalId on GcsTransferCredentials secret - Pass explicit { mode: "daily" } to daily EventBridge rule target - Add fallback for SSM parameter fetch in backup script - Export replica bucket ARN/name from replica stack, consume via props * Add CloudWatch alarm for backup verification Lambda errors Fires on any Lambda error and on missing data (missed schedule). Catches silent failures where the Slack notification never fires. * Add .env to .gitignore Required by org CI conventions check. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
public readonly replicaBucketArn: string;
public readonly replicaBucketName: string;
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
constructor(scope: Construct, id: string, props?: cdk.StackProps) {
super(scope, id, props);
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. * Replace hardcoded instance ID in README with CloudFormation lookup The instance ID changes on every instance replacement (version upgrades, stack updates). Using a dynamic query prevents stale references and removes a manual update step from the deploy process. * Read backup S3 prefix from SSM parameter at runtime Adds /forgejo/backup-s3-prefix SSM parameter (value: archive) and updates the backup script to fetch it instead of hardcoding the prefix. Eliminates the manual post-deploy sed step. * Address cross-review findings for backup verification - Add size guard before downloading dump in restore test (3.5 GB cap) - Use paginator for list_objects_v2 in S3 checks and restore test - Remove unnecessary overrideLogicalId on GcsTransferCredentials secret - Pass explicit { mode: "daily" } to daily EventBridge rule target - Add fallback for SSM parameter fetch in backup script - Export replica bucket ARN/name from replica stack, consume via props * Add CloudWatch alarm for backup verification Lambda errors Fires on any Lambda error and on missing data (missed schedule). Catches silent failures where the Slack notification never fires. * Add .env to .gitignore Required by org CI conventions check. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
const replicaBucket = new s3.Bucket(this, "ReplicaBucket", {
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
bucketName: "forgejo-backups-replica-328440206208",
encryption: s3.BucketEncryption.S3_MANAGED,
blockPublicAccess: s3.BlockPublicAccess.BLOCK_ALL,
versioned: true,
objectLockEnabled: true,
objectLockDefaultRetention: s3.ObjectLockRetention.governance(
cdk.Duration.days(90)
),
lifecycleRules: [
{
id: "archive-to-glacier",
prefix: "archive/",
transitions: [
{
storageClass: s3.StorageClass.GLACIER,
transitionAfter: cdk.Duration.days(30),
},
],
},
{
id: "cleanup-noncurrent-versions",
noncurrentVersionExpiration: cdk.Duration.days(90),
},
],
removalPolicy: cdk.RemovalPolicy.RETAIN,
});
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. * Replace hardcoded instance ID in README with CloudFormation lookup The instance ID changes on every instance replacement (version upgrades, stack updates). Using a dynamic query prevents stale references and removes a manual update step from the deploy process. * Read backup S3 prefix from SSM parameter at runtime Adds /forgejo/backup-s3-prefix SSM parameter (value: archive) and updates the backup script to fetch it instead of hardcoding the prefix. Eliminates the manual post-deploy sed step. * Address cross-review findings for backup verification - Add size guard before downloading dump in restore test (3.5 GB cap) - Use paginator for list_objects_v2 in S3 checks and restore test - Remove unnecessary overrideLogicalId on GcsTransferCredentials secret - Pass explicit { mode: "daily" } to daily EventBridge rule target - Add fallback for SSM parameter fetch in backup script - Export replica bucket ARN/name from replica stack, consume via props * Add CloudWatch alarm for backup verification Lambda errors Fires on any Lambda error and on missing data (missed schedule). Catches silent failures where the Slack notification never fires. * Add .env to .gitignore Required by org CI conventions check. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
this.replicaBucketArn = replicaBucket.bucketArn;
this.replicaBucketName = replicaBucket.bucketName;
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
}
}