* fix(terraform): allow HCP refresh of the log group and parameter
The scoped apply role can create those resources, but CloudWatch and SSM list them on a wildcard ARN. DescribeLogGroups and DescribeParameters need that resource.
* fix(terraform): let the plan role read bucket website config
The S3 provider refreshes GetBucketWebsite. The plan role was denied on the three Forgejo buckets.
* fix(terraform): let the plan role read backup object metadata
HeadObject on the Lambda zip is s3:GetObject. The plan role only had the bucket ARNs.
* fix(terraform): let the plan role read object tags and retention
The S3 provider refreshes tagging, ACL, attributes, and Object Lock on the Lambda zip.
* fix(terraform): scope plan object reads and restore on the data volume
The plan role only needs object reads for the verification zip. Unpacking a dump in /tmp fills the root volume.
* fix(terraform): accept dumps that already contain data/forgejo.db
Today's archive has no gitea-db.sqlite3 at the root. Copy that file only when it is present.
* docs(terraform): keep restore runbook on one bucket and fail closed
Glacier and the download use the prod bucket. GCS unpacks on the data volume. Neither path deletes live repos until the dump has a database.
* docs(terraform): keep optional restore copies from aborting under set -e
if/fi matches user_data.sh. The file comment now says forgejo-services versions also go through the org-account role.
* fix(terraform): restore the dump app.ini with the database
INTERNAL_TOKEN, JWT_SECRET, and LFS_JWT_SECRET live in that file. A restore that keeps the generated file cannot decrypt the dumped secrets.
* feat(terraform): migrate forgejo to HCP Terraform
Move the prod host onto workspace forgejo-prod in the After Hours VPC and freeze CDK push deploys so cutover can happen without applying into seahaven-prod.
* fix(terraform): restore forgejo.db from the dump tarball
Boot restore was copying data/ and repos/ and leaving the sqlite file at the archive root, so a volume restore started Forgejo with no database.
* chore(deps-dev): bump typescript from 6.0.3 to 7.0.2
Bumps [typescript](https://github.com/microsoft/TypeScript) from 6.0.3 to 7.0.2.
- [Release notes](https://github.com/microsoft/TypeScript/releases)
- [Commits](https://github.com/microsoft/TypeScript/commits)
---
updated-dependencies:
- dependency-name: typescript
dependency-version: 7.0.2
dependency-type: direct:development
update-type: version-update:semver-major
...
Signed-off-by: dependabot[bot] <support@github.com>
* fix: regenerate package-lock.json to reflect typescript bump
Update cdk.json to use tsx and regenerate package-lock.json to satisfy CI
* fix: declare and pin tsx runner for CDK app
TypeScript 7's native port is incompatible with ts-node, so cdk.json was
switched to tsx — but it was invoked via `npx`, which fetches an
unpinned copy from the registry at synth/deploy time with no lockfile
entry or integrity check.
Declare tsx as a pinned devDependency (~4.23.0) and invoke the local bin
instead of npx. Update the README's cdk.json section to match. Verified
with `npm run build` (tsc 7.0.2) and `cdk synth` — both green.
---------
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
The README covered runtime architecture and operations thoroughly but
never described the infrastructure-as-code layer itself. Readers had no
map of the CDK app: which files define the stacks, what cdk.json is, or
how to synth/diff. Add an Infrastructure as Code section covering the
project layout, the two stacks and their dependency, cdk.json, and the
common CDK commands.
* Add 3-2-1 backup strategy with cross-region replication and GCS offsite
Implements a fully compliant 3-2-1 backup architecture:
- Copy 1 (live): Harden existing EBS snapshots to 30-day retention
- Copy 2 (near-site): S3 cross-region replication to us-west-2 with
Object Lock (governance 90d) and versioning
- Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project
with 2-year irreversible retention lock
Also adds a verification Lambda that checks all 3 locations daily and
runs monthly restore tests with SQLite integrity checks.
* Enable QEMU in CI for arm64 Lambda Docker builds
* Commit cdk.context.json for CI synth without AWS credentials
Vpc.fromLookup requires cached context to synthesize without
AWS credentials. Required for CI which runs cdk synth without
an OIDC role.
* Fix GCP project ID to sea-haven-backups
* Address code review findings for backup verification
Fix 4 critical issues:
- Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without)
- Add stack dependency so replica deploys before main stack
- Fix DB file extension matching (.sqlite3/.sql instead of .db)
- Replace nonexistent `forgejo restore` command with actual restore steps in README
Fix 4 moderate issues:
- Add timeout=10 to Slack webhook urlopen call
- Add filter='data' to tarfile.extract for PEP 706 compliance
- Add explicit ValueError for unknown handler mode
- Use date-scoped S3/GCS prefix instead of unbounded listing
* Fix backup strategy bug findings
* Handle SQL text dumps separately from binary SQLite in restore test
Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export),
not a binary SQLite file. Opening it directly with sqlite3.connect()
throws DatabaseError. Now imports the SQL dump into a temp DB first.
* Fix GCS backup check: align staleness cutoff and add size validation
GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h),
making the staleness check unreachable. Also added 1MB minimum file
size validation to match the S3 check.
* Rename SECRET_ARN env vars to SECRET_NAME to match actual values
* Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule
* Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt
* Fix restore runbook: trailing-dot cp idiom and Glacier restore step
* Rename GCS service account to match read-only permissions
* Add 4 GiB ephemeral storage to verification Lambda
Monthly restore-test downloads and extracts the full dump tarball
in /tmp. As the dump grows with LFS data, the default 512 MB will
eventually cause ENOSPC failures.
* Replace hardcoded instance ID in README with CloudFormation lookup
The instance ID changes on every instance replacement (version
upgrades, stack updates). Using a dynamic query prevents stale
references and removes a manual update step from the deploy process.
* Read backup S3 prefix from SSM parameter at runtime
Adds /forgejo/backup-s3-prefix SSM parameter (value: archive)
and updates the backup script to fetch it instead of hardcoding
the prefix. Eliminates the manual post-deploy sed step.
* Address cross-review findings for backup verification
- Add size guard before downloading dump in restore test (3.5 GB cap)
- Use paginator for list_objects_v2 in S3 checks and restore test
- Remove unnecessary overrideLogicalId on GcsTransferCredentials secret
- Pass explicit { mode: "daily" } to daily EventBridge rule target
- Add fallback for SSM parameter fetch in backup script
- Export replica bucket ARN/name from replica stack, consume via props
* Add CloudWatch alarm for backup verification Lambda errors
Fires on any Lambda error and on missing data (missed schedule).
Catches silent failures where the Slack notification never fires.
* Add .env to .gitignore
Required by org CI conventions check.
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* Add 3-2-1 backup strategy with cross-region replication and GCS offsite
Implements a fully compliant 3-2-1 backup architecture:
- Copy 1 (live): Harden existing EBS snapshots to 30-day retention
- Copy 2 (near-site): S3 cross-region replication to us-west-2 with
Object Lock (governance 90d) and versioning
- Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project
with 2-year irreversible retention lock
Also adds a verification Lambda that checks all 3 locations daily and
runs monthly restore tests with SQLite integrity checks.
* Enable QEMU in CI for arm64 Lambda Docker builds
* Commit cdk.context.json for CI synth without AWS credentials
Vpc.fromLookup requires cached context to synthesize without
AWS credentials. Required for CI which runs cdk synth without
an OIDC role.
* Fix GCP project ID to sea-haven-backups
* Address code review findings for backup verification
Fix 4 critical issues:
- Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without)
- Add stack dependency so replica deploys before main stack
- Fix DB file extension matching (.sqlite3/.sql instead of .db)
- Replace nonexistent `forgejo restore` command with actual restore steps in README
Fix 4 moderate issues:
- Add timeout=10 to Slack webhook urlopen call
- Add filter='data' to tarfile.extract for PEP 706 compliance
- Add explicit ValueError for unknown handler mode
- Use date-scoped S3/GCS prefix instead of unbounded listing
* Fix backup strategy bug findings
* Handle SQL text dumps separately from binary SQLite in restore test
Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export),
not a binary SQLite file. Opening it directly with sqlite3.connect()
throws DatabaseError. Now imports the SQL dump into a temp DB first.
* Fix GCS backup check: align staleness cutoff and add size validation
GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h),
making the staleness check unreachable. Also added 1MB minimum file
size validation to match the S3 check.
* Rename SECRET_ARN env vars to SECRET_NAME to match actual values
* Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule
* Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt
* Fix restore runbook: trailing-dot cp idiom and Glacier restore step
* Rename GCS service account to match read-only permissions
* Add 4 GiB ephemeral storage to verification Lambda
Monthly restore-test downloads and extracts the full dump tarball
in /tmp. As the dump grows with LFS data, the default 512 MB will
eventually cause ENOSPC failures.
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Add documentation for S3 backups with Glacier lifecycle, hourly
autodiscovery of new GitHub org repos, daily PAT token refresh,
PAT rotation procedure, and Secrets Manager secret inventory.