forgejo/cdk.context.json

49 lines
1.5 KiB
JSON
Raw Permalink Normal View History

Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
{
"vpc-provider:account=328440206208:filter.vpc-id=vpc-0d3d4b67bd0cf8a68:region=us-east-1:returnAsymmetricSubnets=true": {
"vpcId": "vpc-0d3d4b67bd0cf8a68",
"vpcCidrBlock": "10.20.0.0/16",
"ownerAccountId": "328440206208",
"availabilityZones": [],
"vpnGatewayId": "vgw-073737d44762dffc2",
"subnetGroups": [
{
"name": "Private",
"type": "Private",
"subnets": [
{
"subnetId": "subnet-04e38c507e96f1926",
"cidr": "10.20.30.0/24",
"availabilityZone": "us-east-1a",
"routeTableId": "rtb-06a2f56f492b9b4de"
},
{
"subnetId": "subnet-0a0b4fc6f296dfba5",
"cidr": "10.20.40.0/24",
"availabilityZone": "us-east-1b",
"routeTableId": "rtb-01e152fe5cabca7d6"
}
]
},
{
"name": "Public",
"type": "Public",
"subnets": [
{
"subnetId": "subnet-0eea820effe1b3ae5",
"cidr": "10.20.10.0/24",
"availabilityZone": "us-east-1a",
"routeTableId": "rtb-0f2232493a5c43fe8"
},
{
"subnetId": "subnet-0012f5895182c1580",
"cidr": "10.20.20.0/24",
"availabilityZone": "us-east-1b",
"routeTableId": "rtb-0f2232493a5c43fe8"
}
]
}
]
},
"ssm:account=328440206208:parameterName=/aws/service/ami-amazon-linux-latest/al2023-ami-kernel-6.1-arm64:region=us-east-1": "ami-0b183bb1259186479"
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
}