The README covered runtime architecture and operations thoroughly but
never described the infrastructure-as-code layer itself. Readers had no
map of the CDK app: which files define the stacks, what cdk.json is, or
how to synth/diff. Add an Infrastructure as Code section covering the
project layout, the two stacks and their dependency, cdk.json, and the
common CDK commands.
Both alarms changed state but notified nobody (no AlarmActions). Wire
each to the org operational alarm topic (site-alerts, CMK-encrypted)
with an ALARM action only — no OK action per org convention.
The forgejo-backup-verification Lambda runs once per day, so its Errors
metric has data for only one hour and is missing for the other ~23h.
The errors alarm used TreatMissingData=BREACHING, which treated those
23h of missing data as a breach and flipped the alarm OK->ALARM every
day around 11:01 UTC despite zero actual errors.
Changes (alarm-only, no instance changes):
- ErrorAlarm: TreatMissingData BREACHING -> NOT_BREACHING. No data now
means "no errors = healthy" instead of a false breach.
- Add forgejo-backup-verification-not-running: Invocations Sum over a
24h period, alarms when < 1 invocation. This is the real "the daily
verification never ran" guard that the BREACHING setting was trying
(incorrectly) to provide.
Both alarms keep the existing action wiring (no SNS/OK actions), per the
org convention of never notifying on recovery.
Today's instance replacement (uncached AMI lookup resolved a new AL2023
release) destroyed the root volume holding all Forgejo state; restored
manually from the 05:00 S3 dump. This makes replacement harmless:
- New 50 GiB standalone volume (RemovalPolicy.RETAIN) mounted at
/var/lib/forgejo — sqlite db, repositories, and logs all survive
instance replacement and stack deletion. No app.ini path changes.
- Restore-on-boot: if the data volume has no database (first boot or
total volume loss), userdata restores the latest S3 dump
automatically before starting the service. Volume loss self-heals
to <=24h-old state.
- blkid guard: an existing filesystem is mounted, never formatted.
- cachedInContext: true + committed cdk.context.json — AMI changes
(and the instance replacement they force) become deliberate.
- Root volume 50 -> 20 GiB; state no longer lives there.
- forgejo-backup tag on the data volume brings it under the existing
DLM snapshot policy.
Deploy replaces the instance once; restore-on-boot pulls the fresh
17:43 UTC dump.
* Add dependency-review caller workflow
Add a pull_request-triggered caller that invokes the org-level
callable-dependency-review workflow to scan dependency changes and
fail on high-severity advisories.
* chore: retrigger checks
* chore: retrigger dep review (post-fix)
Removing this changed the CloudFormation logical ID from
GcsTransferCredentials to GcsTransferCredentials35DA7E5D,
triggering a replacement that fails because the named secret
already exists.
* Add 3-2-1 backup strategy with cross-region replication and GCS offsite
Implements a fully compliant 3-2-1 backup architecture:
- Copy 1 (live): Harden existing EBS snapshots to 30-day retention
- Copy 2 (near-site): S3 cross-region replication to us-west-2 with
Object Lock (governance 90d) and versioning
- Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project
with 2-year irreversible retention lock
Also adds a verification Lambda that checks all 3 locations daily and
runs monthly restore tests with SQLite integrity checks.
* Enable QEMU in CI for arm64 Lambda Docker builds
* Commit cdk.context.json for CI synth without AWS credentials
Vpc.fromLookup requires cached context to synthesize without
AWS credentials. Required for CI which runs cdk synth without
an OIDC role.
* Fix GCP project ID to sea-haven-backups
* Address code review findings for backup verification
Fix 4 critical issues:
- Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without)
- Add stack dependency so replica deploys before main stack
- Fix DB file extension matching (.sqlite3/.sql instead of .db)
- Replace nonexistent `forgejo restore` command with actual restore steps in README
Fix 4 moderate issues:
- Add timeout=10 to Slack webhook urlopen call
- Add filter='data' to tarfile.extract for PEP 706 compliance
- Add explicit ValueError for unknown handler mode
- Use date-scoped S3/GCS prefix instead of unbounded listing
* Fix backup strategy bug findings
* Handle SQL text dumps separately from binary SQLite in restore test
Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export),
not a binary SQLite file. Opening it directly with sqlite3.connect()
throws DatabaseError. Now imports the SQL dump into a temp DB first.
* Fix GCS backup check: align staleness cutoff and add size validation
GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h),
making the staleness check unreachable. Also added 1MB minimum file
size validation to match the S3 check.
* Rename SECRET_ARN env vars to SECRET_NAME to match actual values
* Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule
* Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt
* Fix restore runbook: trailing-dot cp idiom and Glacier restore step
* Rename GCS service account to match read-only permissions
* Add 4 GiB ephemeral storage to verification Lambda
Monthly restore-test downloads and extracts the full dump tarball
in /tmp. As the dump grows with LFS data, the default 512 MB will
eventually cause ENOSPC failures.
* Replace hardcoded instance ID in README with CloudFormation lookup
The instance ID changes on every instance replacement (version
upgrades, stack updates). Using a dynamic query prevents stale
references and removes a manual update step from the deploy process.
* Read backup S3 prefix from SSM parameter at runtime
Adds /forgejo/backup-s3-prefix SSM parameter (value: archive)
and updates the backup script to fetch it instead of hardcoding
the prefix. Eliminates the manual post-deploy sed step.
* Address cross-review findings for backup verification
- Add size guard before downloading dump in restore test (3.5 GB cap)
- Use paginator for list_objects_v2 in S3 checks and restore test
- Remove unnecessary overrideLogicalId on GcsTransferCredentials secret
- Pass explicit { mode: "daily" } to daily EventBridge rule target
- Add fallback for SSM parameter fetch in backup script
- Export replica bucket ARN/name from replica stack, consume via props
* Add CloudWatch alarm for backup verification Lambda errors
Fires on any Lambda error and on missing data (missed schedule).
Catches silent failures where the Slack notification never fires.
* Add .env to .gitignore
Required by org CI conventions check.
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* Add 3-2-1 backup strategy with cross-region replication and GCS offsite
Implements a fully compliant 3-2-1 backup architecture:
- Copy 1 (live): Harden existing EBS snapshots to 30-day retention
- Copy 2 (near-site): S3 cross-region replication to us-west-2 with
Object Lock (governance 90d) and versioning
- Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project
with 2-year irreversible retention lock
Also adds a verification Lambda that checks all 3 locations daily and
runs monthly restore tests with SQLite integrity checks.
* Enable QEMU in CI for arm64 Lambda Docker builds
* Commit cdk.context.json for CI synth without AWS credentials
Vpc.fromLookup requires cached context to synthesize without
AWS credentials. Required for CI which runs cdk synth without
an OIDC role.
* Fix GCP project ID to sea-haven-backups
* Address code review findings for backup verification
Fix 4 critical issues:
- Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without)
- Add stack dependency so replica deploys before main stack
- Fix DB file extension matching (.sqlite3/.sql instead of .db)
- Replace nonexistent `forgejo restore` command with actual restore steps in README
Fix 4 moderate issues:
- Add timeout=10 to Slack webhook urlopen call
- Add filter='data' to tarfile.extract for PEP 706 compliance
- Add explicit ValueError for unknown handler mode
- Use date-scoped S3/GCS prefix instead of unbounded listing
* Fix backup strategy bug findings
* Handle SQL text dumps separately from binary SQLite in restore test
Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export),
not a binary SQLite file. Opening it directly with sqlite3.connect()
throws DatabaseError. Now imports the SQL dump into a temp DB first.
* Fix GCS backup check: align staleness cutoff and add size validation
GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h),
making the staleness check unreachable. Also added 1MB minimum file
size validation to match the S3 check.
* Rename SECRET_ARN env vars to SECRET_NAME to match actual values
* Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule
* Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt
* Fix restore runbook: trailing-dot cp idiom and Glacier restore step
* Rename GCS service account to match read-only permissions
* Add 4 GiB ephemeral storage to verification Lambda
Monthly restore-test downloads and extracts the full dump tarball
in /tmp. As the dump grows with LFS data, the default 512 MB will
eventually cause ENOSPC failures.
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Add documentation for S3 backups with Glacier lifecycle, hourly
autodiscovery of new GitHub org repos, daily PAT token refresh,
PAT rotation procedure, and Secrets Manager secret inventory.
Autodiscovery runs hourly — creates Forgejo mirrors for new GitHub
org repos. Token refresh runs daily — propagates the current PAT
from Secrets Manager to all mirror git remotes.