- Add size guard before downloading dump in restore test (3.5 GB cap)
- Use paginator for list_objects_v2 in S3 checks and restore test
- Remove unnecessary overrideLogicalId on GcsTransferCredentials secret
- Pass explicit { mode: "daily" } to daily EventBridge rule target
- Add fallback for SSM parameter fetch in backup script
- Export replica bucket ARN/name from replica stack, consume via props
Adds /forgejo/backup-s3-prefix SSM parameter (value: archive)
and updates the backup script to fetch it instead of hardcoding
the prefix. Eliminates the manual post-deploy sed step.
The instance ID changes on every instance replacement (version
upgrades, stack updates). Using a dynamic query prevents stale
references and removes a manual update step from the deploy process.
* Add 3-2-1 backup strategy with cross-region replication and GCS offsite
Implements a fully compliant 3-2-1 backup architecture:
- Copy 1 (live): Harden existing EBS snapshots to 30-day retention
- Copy 2 (near-site): S3 cross-region replication to us-west-2 with
Object Lock (governance 90d) and versioning
- Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project
with 2-year irreversible retention lock
Also adds a verification Lambda that checks all 3 locations daily and
runs monthly restore tests with SQLite integrity checks.
* Enable QEMU in CI for arm64 Lambda Docker builds
* Commit cdk.context.json for CI synth without AWS credentials
Vpc.fromLookup requires cached context to synthesize without
AWS credentials. Required for CI which runs cdk synth without
an OIDC role.
* Fix GCP project ID to sea-haven-backups
* Address code review findings for backup verification
Fix 4 critical issues:
- Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without)
- Add stack dependency so replica deploys before main stack
- Fix DB file extension matching (.sqlite3/.sql instead of .db)
- Replace nonexistent `forgejo restore` command with actual restore steps in README
Fix 4 moderate issues:
- Add timeout=10 to Slack webhook urlopen call
- Add filter='data' to tarfile.extract for PEP 706 compliance
- Add explicit ValueError for unknown handler mode
- Use date-scoped S3/GCS prefix instead of unbounded listing
* Fix backup strategy bug findings
* Handle SQL text dumps separately from binary SQLite in restore test
Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export),
not a binary SQLite file. Opening it directly with sqlite3.connect()
throws DatabaseError. Now imports the SQL dump into a temp DB first.
* Fix GCS backup check: align staleness cutoff and add size validation
GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h),
making the staleness check unreachable. Also added 1MB minimum file
size validation to match the S3 check.
* Rename SECRET_ARN env vars to SECRET_NAME to match actual values
* Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule
* Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt
* Fix restore runbook: trailing-dot cp idiom and Glacier restore step
* Rename GCS service account to match read-only permissions
* Add 4 GiB ephemeral storage to verification Lambda
Monthly restore-test downloads and extracts the full dump tarball
in /tmp. As the dump grows with LFS data, the default 512 MB will
eventually cause ENOSPC failures.
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Monthly restore-test downloads and extracts the full dump tarball
in /tmp. As the dump grows with LFS data, the default 512 MB will
eventually cause ENOSPC failures.
GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h),
making the staleness check unreachable. Also added 1MB minimum file
size validation to match the S3 check.
Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export),
not a binary SQLite file. Opening it directly with sqlite3.connect()
throws DatabaseError. Now imports the SQL dump into a temp DB first.
Fix 4 critical issues:
- Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without)
- Add stack dependency so replica deploys before main stack
- Fix DB file extension matching (.sqlite3/.sql instead of .db)
- Replace nonexistent `forgejo restore` command with actual restore steps in README
Fix 4 moderate issues:
- Add timeout=10 to Slack webhook urlopen call
- Add filter='data' to tarfile.extract for PEP 706 compliance
- Add explicit ValueError for unknown handler mode
- Use date-scoped S3/GCS prefix instead of unbounded listing
Add documentation for S3 backups with Glacier lifecycle, hourly
autodiscovery of new GitHub org repos, daily PAT token refresh,
PAT rotation procedure, and Secrets Manager secret inventory.
Autodiscovery runs hourly — creates Forgejo mirrors for new GitHub
org repos. Token refresh runs daily — propagates the current PAT
from Secrets Manager to all mirror git remotes.